Method, device and equipment for stack type planning and storage medium

By using machine learning models and policy networks to determine the placement of objects, the problem of insufficient stability and adaptability in traditional hybrid palletizing methods is solved, achieving efficient and stable pallet planning and space utilization.

CN121882489APending Publication Date: 2026-04-17JINGDONG TECH HLDG CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINGDONG TECH HLDG CO LTD
Filing Date
2024-10-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional mixed palletizing methods cannot effectively handle objects of multiple sizes, shapes, and weights, and cannot flexibly adapt to diverse logistics needs. Furthermore, existing algorithms are difficult to meet dynamic logistics and warehousing requirements, resulting in insufficient stability of palletizing solutions in practice or difficulty in meeting specific business needs.

Method used

By using a trained machine learning model, the placement position of the object in the target stacking pattern is determined based on the object's attribute information and the target stacking pattern requirements. Through heuristic algorithms and a policy network of reinforcement learning, candidate positions that meet the stacking constraints are selected to form a stable and efficient stacking pattern.

Benefits of technology

Ensure the stability of stacking and efficient use of storage space, realize stacking planning for specific preferences, improve storage density and adapt to dynamic logistics needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882489A_ABST
    Figure CN121882489A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a stack type planning method and device, equipment and a storage medium. The method comprises the steps that in response to a received stacking request for stacking a plurality of objects, the stacking positions of the plurality of objects in a target stack type are determined based on attribute information of the plurality of objects and stack type requirements of the target stack type by utilizing a trained machine learning model, and determination of the stacking positions comprises the steps that for a given object in the plurality of objects, the stacking positions of the plurality of objects in the target stack type are determined according to the attribute information of the plurality of objects and the stack type requirements of the target stack type; determining at least one stackable space in the target stack type based on the stack type requirement and the determined stacking position of at least one object in the target stack type; based on the attribute information of the given object, determining a candidate position of the given object in at least one stacking space; a candidate position meeting the stacking constraint condition is determined from the at least one stacking space to serve as the stacking position of the given object; and placing the plurality of objects according to the respective stacking positions of the plurality of objects in the target stack type to form the target stack type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computer vision technology, and particularly to methods, apparatus, devices, and storage media for stacking planning. Background Technology

[0002] With the rapid development of robotics technology, it is widely used in automated warehousing. For example, robotic arms are frequently used for stacking objects in automated warehousing. Traditionally, stacking is done according to fixed rules; however, this method becomes more complex when dealing with a mix of objects of different sizes. Therefore, there is a desire to improve the efficiency of stacking mixed objects and reduce labor costs. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for palletizing is provided. The method includes: in response to receiving a palletizing request for multiple objects, using a trained machine learning model, based on attribute information of the multiple objects and palletizing requirements of a target palletizing pattern, determining the palletizing positions of each of the multiple objects within a target palletizing pattern; the determination of the palletizing positions includes: for a given object among the multiple objects, determining at least one stackable space within the target palletizing pattern based on the palletizing requirements and the determined palletizing positions of at least one object in the target palletizing pattern; determining candidate positions of the given object within the at least one stackable space based on attribute information of the given object; determining candidate positions satisfying palletizing constraints from the at least one stackable space as the palletizing positions of the given object; and placing the multiple objects according to their respective palletizing positions within the target palletizing pattern via a control module in the palletizing system to form the target palletizing pattern.

[0004] In a second aspect of this disclosure, an apparatus for pallet planning is provided. The apparatus includes: a placement position determination module configured to, in response to receiving a placement request for multiple objects, determine, using a trained machine learning model, the placement positions of the multiple objects within a target pallet structure based on attribute information of the multiple objects and pallet structure requirements of the target pallet structure. The determination of placement positions includes: for a given object among the multiple objects, determining at least one stackable space within the target pallet structure based on pallet structure requirements and the determined placement positions of at least one object in the target pallet structure; determining candidate positions of the given object within the at least one stackable space based on attribute information of the given object; and determining candidate positions from the at least one stackable space that satisfy placement constraints as the placement positions of the given object; and an object placement module configured to place the multiple objects according to their respective placement positions within the target pallet structure via a control module in a palletizing system to form the target pallet structure.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, which, when executed by a processor, implements the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] It should be understood that the description in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram of an example target environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A schematic block diagram of an example architecture for stacking planning according to some embodiments is shown;

[0012] Figure 3 A flowchart of a process for stacking planning according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic structural block diagram of an apparatus for stacking planning according to some embodiments of the present disclosure is shown; and

[0014] Figure 5 A block diagram of an electronic device suitable for implementing one or more embodiments of the present disclosure is shown. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0017] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0018] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization should be obtained from the relevant users. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.

[0020] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein based on the prompt message.

[0021] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0024] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. The testing phase can sometimes be integrated into the training phase. In the application phase, the trained model can be used to process actual model inputs based on the trained parameter values ​​to determine the corresponding model output.

[0025] As briefly mentioned earlier, with the rapid development of the logistics and warehousing industry, mixed palletizing technology plays a crucial role in improving efficiency and reducing labor costs. Mixed palletizing refers to the task of palletizing unordered and multi-variety boxes. Traditional palletizing methods typically employ fixed rules, applicable to objects of similar size. However, mixed palletizing requires consideration of each object's attributes, including size, shape, and weight, as well as various practical constraints such as stack stability, maximum utilization of loading space, and stacking preferences. Therefore, fixed rules are not effective for mixed palletizing.

[0026] First, traditional mixed palletizing methods typically rely on pre-defined rules and are suitable for simple stacking of similar objects. This approach performs poorly when dealing with objects of varying sizes, shapes, and weights, failing to flexibly adapt to diverse logistics needs. Second, actual palletizing processes involve various constraints, such as pallet stability and stacking preferences. Existing search algorithms struggle to effectively handle these complex constraints, resulting in palletizing schemes that lack stability or fail to meet specific business requirements in practice. Furthermore, existing learning-based methods are poor at adapting to real-time changing conditions and cannot flexibly respond to dynamic logistics and warehousing demands. In real-world applications, the logistics environment and demands are often dynamically changing, while existing algorithms are generally designed based on offline or online settings, making it difficult to adaptively adjust to online and offline palletizing requirements.

[0027] In view of this, an improved scheme for pallet planning is proposed in the embodiments of this disclosure. In this scheme, in response to receiving a palletizing request for multiple objects, a trained machine learning model is used to determine the palletizing positions of each object in the target pallet based on the attribute information of the multiple objects and the palletizing requirements of the target pallet. The determination of the palletizing positions includes: for a given object among the multiple objects, determining at least one stackable space in the target pallet based on the palletizing requirements and the determined palletizing positions of at least one object in the target pallet; determining candidate positions of the given object in the at least one stackable space based on the attribute information of the given object; determining candidate positions that satisfy the palletizing constraints from the at least one stackable space as the palletizing positions of the given object; and placing the multiple objects according to their respective palletizing positions in the target pallet via a control module in the palletizing system to form the target pallet.

[0028] Therefore, this disclosure determines candidate locations for object placement based on object attribute information (e.g., size, shape, and weight), ensuring the stability of the stacking pattern. Furthermore, by selecting the object's placement location from the candidate locations through stacking constraints, stacking pattern planning for specific preferences can be achieved.

[0029] Figure 1 A schematic diagram of an example target environment 100 in which embodiments of the present disclosure can be implemented is shown. This example target environment 100 includes a palletizing system 110 and a machine learning model 155. The palletizing system 110 includes a control module 120, a buffer module (e.g., a buffer wall) 130, and a robotic arm 140. The buffer module 130 can buffer and store a certain number of objects and detect the size and type of the objects using a detection device such as a camera.

[0030] In environment 100, palletizing system 110, based on the palletizing planning method of control module 120 and the object attribute information (e.g., size, category, etc.) provided by cache module 130, designs a palletizing pattern that meets specific stacking requirements. Then, robotic arm 140, according to the palletizing pattern plan of control module 120, stacks the objects onto pallets in the required order and position. In embodiments of this disclosure, the objects to be stacked into a specific palletizing pattern can be various physical objects, including boxes (e.g., logistics boxes).

[0031] In some embodiments, the palletizing system 110 may invoke the machine learning model 155 according to the control module 120 to determine the pallet type planning method. In some examples, the machine learning model 155 may also be deployed in the palletizing system 110, and the palletizing system 110 may use the machine learning model 155 to determine the pallet type planning method.

[0032] Some exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0033] Figure 2 A schematic diagram of an example space 200 for stacking planning according to some embodiments of the present disclosure is shown. The example space 200 can be implemented in... Figure 1 For ease of description, the example embodiment will be described primarily with respect to the palletizing system 110. It should be understood that the actions described with respect to the palletizing system 110 may be performed by the control module 120, the buffer module 130, the robotic arm 140 included in the palletizing system 110, or may be performed by the palletizing system 110 in conjunction with its server (e.g., a server).

[0034] In some embodiments, if the palletizing system 110 receives a request to place multiple objects, it uses a trained machine learning model to determine the placement position of each object within the target pallet type based on the object's attribute information and the pallet type requirements. In some examples, if the palletizing system 110 receives a request to place multiple objects, it invokes a trained machine learning model 155 via the control module 120 to determine the placement position of each object within the pallet type based on the object's attribute information (e.g., size, shape, weight, category, etc.) and the pallet type requirements.

[0035] In some examples, the target stack type has stack type requirements, such as the bottom area of ​​the target stack type not exceeding the bottom area of ​​the pallet, the height to which multiple objects can be placed not exceeding the maximum height of the target stack type, and so on. In some embodiments, the placement positions of multiple objects in the target stack type are stored in a cache module accessible by the palletizing system. In some examples, the palletizing system 110 may store the placement positions of multiple objects in the stack type, determined by calling the trained machine learning model 155 via the control module 120, in the cache module 130.

[0036] In some embodiments, the palletizing system 110, based on the control module 120, places multiple objects according to their stacking positions within a target pallet type to form the target pallet type. In some examples, the palletizing system 110, through the control module 120, controls a robotic arm 140 to place multiple objects onto a pallet according to their respective positions and order within the target pallet type, thereby forming the target pallet type. This allows the algorithm to automatically adapt to pallet type planning in both online and offline palletizing scenarios.

[0037] The following will refer to Figure 2 The palletizing system 110 describes the placement of multiple objects within a target pallet type. For ease of understanding, the following description uses a given object from among the multiple objects as an example.

[0038] In some embodiments, the palletizing system 110 determines at least one stackable space in a target pallet type based on pallet type requirements and the determined stacking position of at least one object in the target pallet type. In some examples, the palletizing system 110 may employ a heuristic Empty Maximal Space (EMS) algorithm to expand at least one stackable space for stackable objects based on the state of already stacked objects.

[0039] In some examples, if the available space E (i.e., the volume of the stack) is determined based on the requirements of the stack type (e.g., the length, width, and height of the stack type), it can be determined through the lower left corner. and spatial dimensions (x) e y e , z e (This indicates that) a dimension of... The object key is placed in the lower left corner of this space. This object will occupy the original space The three largest spaces are:

[0040] To facilitate understanding, projecting three-dimensional space E onto a two-dimensional plane, such as... Figure 2The example spatial structure shown is 200.

[0041] In example space 200, assuming object 210 is placed in the lower left corner of this space, the palletizing system 110 can use the EMS algorithm to determine the stackable spaces 212, 211, and 213 in space E based on the placement position of object 210 and space E. In some examples, stackable space 213 indicates the intersection of stackable spaces 212 and 211. In example space 200, the lower left corner of stackable space 211... The bottom left corner of the stackable space 212

[0042] In some embodiments, the palletizing system 110 determines candidate positions of a given object in at least one stackable space based on attribute information of the given object. In some examples, the palletizing system 110 may utilize a reinforcement learning-based policy network included in the machine learning model 155 to determine candidate positions of a given object in at least one stackable space based on attribute information of the given object.

[0043] In some embodiments, the palletizing system 110 utilizes a policy network to generate at least one candidate action for a given object, each candidate action indicating a candidate position for placing the given object in a stackable space. In some examples, the palletizing system 110 utilizes a policy network for reinforcement learning to match the given object (i.e., the object to be stacked) with at least one stackable space one by one, thereby generating at least one candidate action for the object to be stacked. As shown in Figure 200, each of the at least one candidate action indicates a candidate position for placing the object to be stacked in stackable spaces 211, 212, and 213.

[0044] Understandably, there are multiple locations in stackable spaces 211, 212, and 213 where objects to be stacked can be placed. The stacking system 110 can use a policy network to determine at least one candidate location from the three stackable spaces where the position connected to the vertex of object 210 (e.g., the lower right vertex, the upper left vertex, and the upper right diagonal vertex) is located.

[0045] The palletizing system 110 can generate at least one candidate action for a given object in the following manner.

[0046] In some embodiments, the palletizing system 110 initializes the observation state of the target environment to obtain at least one dynamic node corresponding to at least one of a plurality of objects, each dynamic node indicating the determined placement position of the object. The following description uses an object 210 already placed in the current target pallet type as an example for ease of understanding, but this is merely exemplary and not limited thereto. It is understood that at least one other object may be placed in the current target pallet type. In some examples, the palletizing system 110 initializes the observation state of the target environment and then obtains the dynamic node corresponding to object 210. In some examples, the dynamic node corresponding to object 210 indicates the determined placement position of object 210. In some examples, the palletizing system 110 can generate dynamic nodes based on the observation state using a heuristic algorithm. Dynamic node Γ t The observed state s is obtained by using a node tree. t Transform into a dynamic node Γ t =D nt (s t ).

[0047] In some embodiments, a dynamic node includes internal nodes, at least one leaf node, and node information corresponding to a given object. In some examples, the dynamic node Γ t Internal nodes include internal node I t Leaf node L t and the node information of the object to be stacked n t , Γ t ={I t L t n t To facilitate the use of matrices to represent all nodes, the internal nodes I are... t Leaf node L t and the node information of the object to be stacked n t They are all represented as 9-dimensional vectors.

[0048] Internal nodes include one or more nodes corresponding to one or more objects whose placement positions have been determined among multiple objects. Internal nodes are represented as first vertex information, second vertex information, and flag information indicating whether there is a placement object for each of the one or more objects. The one or more objects include a given object.

[0049] In some examples, internal node I t This includes the nodes corresponding to the placement positions of object 210. Internal node I t It is initialized as an m×9 zero matrix, where m represents the maximum number of objects that can be placed, and 9 is the dimension of each placed object as an internal node attribute. Internal node I t Expressed as an equation As shown, where It is the bottom left vertex of the i-th already stacked object. The top right diagonal vertex of the i-th already stacked object f i This flag indicates whether an object has been placed; it is set to 1 if an object has been placed, and 0 otherwise. (x) i y i , z i The value represents the size of the object represented by the node vector.

[0050] In some embodiments, a leaf node includes a node corresponding to one or more stackable positions in a target stack based on stack type requirements and determined stacking positions of one or more objects. The leaf node is represented by vertex information for each stackable position and flag information indicating whether it is stackable.

[0051] In some examples, leaf node L t This includes the nodes corresponding to the candidate positions where the object to be placed can be placed. Leaf node L t It is initialized as a zero matrix of dimension h×9, where h represents the maximum number of leaf nodes, and 9 is the dimension of each stackable position represented as a leaf node attribute. Leaf node L t Expressed as an equation As shown, where It is the bottom left vertex of the l-th stackable position. f represents the top-right vertex of the object to be stacked in the leaf node. l This flag indicates whether there is a suitable placement position for the object to be placed in the leaf node; it is set to 1 if it exists, and 0 otherwise. (x l y l , z l ) indicates the size of the object to be stacked in the leaf node.

[0052] In some embodiments, node information indicates attribute information of an object that follows a given object among a plurality of objects. In some examples, the node information n of the objects to be stacked... t It is initialized as a k×9 dimensional zero matrix, where k is the number of objects to be placed, and 9 represents each object as n. t Dimensions of the attributes. Node information n of the object to be stacked. t It can be expressed as an equation As shown, where (x k y k , z k ) represents the size of the object to be stacked, f n This flag indicates whether an object to be placed exists. If it exists, set it to 1; otherwise, set it to 0.

[0053] In some embodiments, the palletizing system 110, based on the dynamic nodes corresponding to at least one object, utilizes at least a heuristic algorithm to generate index identifiers for candidate leaf nodes corresponding to a given object from a policy network. The candidate leaf nodes indicate candidate positions for the given object determined based on the already determined placement positions. Then, the palletizing system 110 uses the index identifiers to determine the candidate leaf nodes indexed to the corresponding index identifiers as candidate actions to be performed.

[0054] Subsequently, the palletizing system 110, based on the dynamic nodes corresponding to object 210, including the nodes corresponding to object 210, at least one candidate leaf node, and the attribute information of the object to be palletized, uses the EMS algorithm to generate an index identifier (ID) for at least one candidate leaf node corresponding to the object to be palletized from the policy network. The candidate leaf node indicates a candidate position where the object to be palletized can be placed, determined based on the palletizing position of object 210. Then, the palletizing system 110 uses the index ID of the leaf node to find the corresponding leaf node as a candidate action to be performed.

[0055] The following describes how the palletizing system 110 determines the placement of a given object (i.e., the object to be palletized).

[0056] In some embodiments, the palletizing system 110 determines candidate positions that satisfy palletizing constraints from at least one stackable space as the palletizing positions for a given object. In some embodiments, the palletizing constraints may include stability conforming to a target pallet type. In some examples, by designing appropriate rules, it is determined whether the leaf nodes corresponding to the candidate positions satisfy support stability. Thus, the candidate positions corresponding to the leaf nodes that satisfy support stability can be determined as the palletizing positions for the object to be palletized.

[0057] In some embodiments, the stacking constraints may further include stacking preference information for stacking multiple objects. In some examples, by designing appropriate rules, it is determined whether the leaf nodes corresponding to candidate positions satisfy the stacking preferences. Thus, the candidate positions corresponding to leaf nodes that satisfy the stacking preferences can be determined as the stacking positions for the objects to be stacked. In some embodiments, the stacking constraints may further include predetermined rules determined based on the attribute information of multiple objects. For example, rules may be set such as placing smaller objects on top of larger objects, placing lighter objects on top of heavier objects, etc. Thus, the candidate positions corresponding to leaf nodes that satisfy the stacking rules can be determined as the stacking positions for the objects to be stacked.

[0058] like Figure 2As shown, the palletizing system 110 can determine the lower left corner candidate positions corresponding to stackable spaces 211, 212, and 213 that satisfy the stacking constraints as the stacking positions for a given object. It is understandable that if the object to be stacked is placed at the lower left corner candidate position corresponding to stackable space 213, the object will be in a suspended state. In this case, the palletizing system 110 deletes the node corresponding to the lower left corner candidate position corresponding to stackable space 213 according to the palletizing constraints.

[0059] In some embodiments, if the palletizing system 110 determines multiple candidate positions that satisfy the palletizing constraints from at least one stackable space, it takes the predetermined position of the candidate position located in each of the at least one palletizing space as the palletizing position of a given object. In some examples, the palletizing system 110 matches the object to be palletized against the maximum stackable space one by one; if the object to be palletized can be placed in any of the maximum stackable spaces, a non-zero leaf node can be generated. In some embodiments, the generated non-zero leaf node coordinates It is the bottom left corner of the maximum stackable space.

[0060] In some examples, if there is no maximum available space to place the object to be placed, a zero-leaf node is generated. Understandably, if the object to be placed does not match any available space, it is determined that the object is in an abnormal state.

[0061] The following describes how the palletizing system determines candidate positions that satisfy the palletizing constraints from at least one palletizable space as the palletizing positions for a given object.

[0062] In some embodiments, the palletizing system 110 inputs at least one candidate action into the target environment to obtain the observation state corresponding to each of the at least one candidate action. Then, based on the observation state corresponding to each of the at least one candidate action and the stacking constraints, the palletizing system 110 determines the target action for a given object from the at least one candidate action, wherein the target action indicates the stacking position of the given object.

[0063] In some examples, the palletizing system 110 inputs at least one candidate action into the target environment to obtain an observation state. Then, based on the observation state corresponding to each of the at least one candidate action and the palletizing constraints, the palletizing system 110 determines the target action that indicates the placement position of the object to be palletized from the at least one candidate action.

[0064] In some embodiments, the palletizing system 110 determines the dynamic node corresponding to a given object from the candidate leaf nodes corresponding to at least one candidate action, based on the observation state and placement constraints corresponding to at least one candidate action, and the dynamic node corresponding to the given object indicates the determined placement position of the given object.

[0065] In some examples, after the palletizing system 110 determines the placement position of a given object (i.e., the object to be placed), the internal nodes of the dynamic node corresponding to the placement position of the given object include the node corresponding to the placement position of the object 210 and the node corresponding to the placement position of the given object, so that the palletizing system 110 can determine the placement position of the next object to be placed based at least on the dynamic node of the object 210 and the dynamic node corresponding to the given object.

[0066] By determining candidate placement locations for objects based on their attribute information (e.g., size, shape, and weight), the stability of the stacking pattern can be ensured. Accordingly, by selecting the placement location of objects from the candidate locations through stacking constraints, stacking pattern planning for specific preferences can be achieved, thereby improving the volumetric efficiency of the stacking pattern.

[0067] For ease of understanding, the following describes how to train a machine learning model.

[0068] In some embodiments, an action input is generated into the target environment via a policy network, and the state and reward of the environment feedback are recorded as samples for model training. Specifically, the policy network, based on dynamic nodes Γ... t Output the index IDa of the leaf node t The corresponding leaf node is found using the index ID of the leaf node, which is then used as the final action a′. t Then, the action a′ t Input the simulation environment to obtain the environmental state s t+1 Instant rewards t Is the stack type planning completed? (marked by d) t Repeat this process to collect w sets of sample data, where the information recorded in the sample data includes {(Γ t a t , Γ t+1 r t d t )}.

[0069] Then, the sample data collected through interaction {(Γ t a t , Γ t+1 r t d t Loss function of policy network and the loss function of V-value networks Iteratively update the parameters of the policy network and the V-value network. In some examples, the policy network is a graph attention network, mathematically represented as a t =π θ (a t |Γ t ), whose input dynamic node matrix Γ t Output the index ID of the leaf node a t The V-value network is a multilayer perceptron network, mathematically represented as v = v φ (Γ t ), whose input dynamic node matrix Γ t Output the reward value v for this state. The following are the specific steps for training a machine learning model:

[0070] 1. Input: Preset object size dataset

[0071] 2. Initialize the network parameter φ for V value, and initialize the policy network parameter θ.

[0072] 3.For step=0,1,…,max_step do.

[0073] 4. Reset the simulation environment and initialize the environment observation state s t ;Γ t =D nt (s t The position of the first object.

[0074] 5. For t=0,1,…,n do.

[0075] 6. Generate the optimal leaf node ID a from the policy network. t =π θ (a t |Γ t EMS+ policy network, constraints, determined.

[0076] 7. The corresponding leaf node is indexed by the leaf node ID to determine the action a′ to be executed. t =L s (a t ).

[0077] 8. Perform action a′ t And collect the s feedback from the simulation environment t+1 r t d(new state s and reward r).

[0078] 9. Transform the observed state into node Γ t+1 =D nt (s t+1 The position of the second object.

[0079] 10. Store and record sample data (Γ) t a t r t , Γ t+1 d t ).

[0080] 11.s t =s t+1 , Γ t =D nt (s t ).

[0081] 12. If the termination condition d t If true, then reset the environment and initialize the environment observation state s. t ;Γ t =D nt (s t ).

[0082] 13. End.

[0083] 14. Update the policy network θ according to the loss function of the policy network.

[0084] 15. Update the V-value network φ according to the loss function of the V-value network.

[0085] 16. Output: Policy Network π θ (a t |Γ t ).

[0086] In summary, this disclosure, through an algorithm and the use of a policy network, can ensure the stability of the stacking pattern and achieve stacking pattern planning for specific preferences. Furthermore, by considering the attribute information of different objects and various possible constraints (e.g., stacking pattern stability, loading space, and utilization rate), it can effectively utilize storage space and improve storage density.

[0087] Figure 3 A flowchart of a process 300 for stacking planning according to some embodiments of the present disclosure is shown. Process 300 can be implemented in... Figure 1 The process 300 may be implemented at palletizing system 110, or may be implemented in or included in other equipment used for pallet planning, such as other terminal equipment or service equipment. For ease of description, it is assumed that process 300 may be implemented at palletizing system 110, as described below with reference to... Figure 1 To describe process 300.

[0088] In box 310, in response to receiving a stacking request for multiple objects, the palletizing system 110 uses a trained machine learning model to determine the stacking position of each object in the target pallet type based on the attribute information of the multiple objects and the pallet type requirements of the target pallet type. The determination of the stacking position includes: for a given object among the multiple objects, determining at least one stackable space in the target pallet type based on the pallet type requirements and the determined stacking position of at least one object in the target pallet type; determining candidate positions of the given object in the at least one stackable space based on the attribute information of the given object; and determining candidate positions that satisfy the stacking constraints from the at least one stackable space as the stacking position of the given object.

[0089] In frame 320, the palletizing system 110 places multiple objects according to their respective placement positions in the target pallet type via the control module of the palletizing system to form the target pallet type.

[0090] In some embodiments, determining the placement position of a given object includes: in response to determining a plurality of candidate positions satisfying placement constraints from at least one placement space, taking a predetermined position of the plurality of candidate positions located in each of the at least one placement space as the placement position of the given object.

[0091] In some embodiments, the stacking constraints include at least one of the following: stability of the target stack type, stacking preference information for stacking multiple objects, and / or predetermined rules determined based on the attribute information of multiple objects.

[0092] In some embodiments, the placement positions of multiple objects in the target pallet type are stored in a cache module accessible by the palletizing system.

[0093] In some embodiments, the machine learning model includes a reinforcement learning-based policy network, and determining whether a given object can be placed at a candidate position in at least one stackable space includes generating at least one candidate action for the given object using the policy network, each candidate action indicating a candidate position for placing the given object in the stackable space; and determining a candidate position satisfying stacking constraints from at least one stackable space as a stacking position for the given object includes: inputting at least one candidate action into a target environment to obtain an observation state corresponding to each of the at least one candidate action; and determining a target action for the given object from at least one candidate action based on the observation state corresponding to each of the at least one candidate action and the stacking constraints, the target action indicating a stacking position for the given object.

[0094] In some embodiments, generating at least one candidate action includes: initializing the observation state of the target environment to obtain at least one dynamic node corresponding to at least one of a plurality of objects, each dynamic node indicating a determined placement position of the object; based on the dynamic node corresponding to the at least one object, generating an index identifier for a candidate leaf node corresponding to a given object by a policy network using at least a heuristic algorithm, the candidate leaf node indicating a candidate position of the given object determined based on the determined placement position; and determining the candidate leaf node indexed to the index identifier as a candidate action to be executed via the index identifier.

[0095] In some embodiments, determining the target action for a given object from at least one candidate action based on the observation state and placement constraints corresponding to at least one candidate action includes: determining the dynamic node corresponding to the given object from the candidate leaf nodes corresponding to at least one candidate action based on the observation state and placement constraints corresponding to at least one candidate action, wherein the dynamic node corresponding to the given object indicates the determined placement position of the given object.

[0096] In some embodiments, a dynamic node includes an internal node, at least one leaf node, and node information corresponding to a given object. The internal node includes one or more nodes corresponding to one or more objects whose stacking positions have been determined among a plurality of objects. The internal node is represented as first vertex information, second vertex information, and flag information indicating whether a stacking object exists for each of the one or more objects. The one or more objects include the given object. The at least one leaf node includes nodes corresponding to one or more stackable positions in a target stack based on the stacking requirements and the determined stacking positions of the one or more objects. The leaf node is represented as vertex information for each stackable position and flag information indicating whether it is stackable. The node information indicates attribute information of the object following the given object among the plurality of objects.

[0097] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an apparatus 400 for palletizing according to some embodiments of the present disclosure is shown. The apparatus 400 may be implemented in or included in the palletizing system 110, or implemented in or included in other devices for palletizing, such as other terminal devices or service devices. The various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0098] As shown in the figure, the device 400 includes a stacking position determination module 410, configured to, in response to receiving a stacking request for stacking multiple objects, use a trained machine learning model to determine the stacking position of each of the multiple objects in a target stacking pattern based on the attribute information of the multiple objects and the stacking pattern requirements of the target stacking pattern. The determination of the stacking position includes: for a given object among the multiple objects, determining at least one stackable space in the target stacking pattern based on the stacking pattern requirements and the determined stacking position of at least one object in the target stacking pattern; determining candidate positions of the given object in the at least one stackable space based on the attribute information of the given object; and determining candidate positions from the at least one stackable space that satisfy the stacking constraints as the stacking position of the given object. The device 400 also includes an object placement module 420, configured to place multiple objects according to their respective stacking positions in the target stacking pattern via a control module in the stacking system to form the target stacking pattern.

[0099] In some embodiments, the placement position determination module 410 is further configured to, in response to determining a plurality of candidate positions satisfying placement constraints from at least one placement space, use a predetermined position of the plurality of candidate positions located in each of the at least one placement space as the placement position of the given object.

[0100] In some embodiments, the stacking constraints include at least one of the following: stability of the target stack type, stacking preference information for stacking multiple objects, and / or predetermined rules determined based on the attribute information of multiple objects.

[0101] In some embodiments, the placement positions of multiple objects in the target pallet type are stored in a cache module accessible by the palletizing system.

[0102] In some embodiments, the machine learning model includes a reinforcement learning-based policy network, and determining whether a given object can be placed at at least one candidate position in a stackable space includes generating at least one candidate action for the given object using the policy network, each candidate action indicating a candidate position for placing the given object in the stackable space; and the stacking position determination module 410 is further configured to input the at least one candidate action into a target environment to obtain an observation state corresponding to each of the at least one candidate action; and to determine a target action for the given object from the at least one candidate action based on the observation state corresponding to each of the at least one candidate action and stacking constraints, the target action indicating a stacking position for the given object.

[0103] In some embodiments, the placement position determination module 410 is further configured to initialize the observation state of the target environment to obtain at least one dynamic node corresponding to at least one of a plurality of objects, each dynamic node indicating the determined placement position of the object; based on the dynamic node corresponding to at least one object, using at least a heuristic algorithm, generate an index identifier for a candidate leaf node corresponding to a given object by a policy network, the candidate leaf node indicating a candidate position of the given object determined based on the determined placement position; and determine the candidate leaf node indexed to the index identifier as a candidate action to be performed via the index identifier.

[0104] In some embodiments, based on the observation state and placement constraints corresponding to at least one candidate action, the placement position determination module 410 is further configured to determine the dynamic node corresponding to the given object from the candidate leaf nodes corresponding to at least one candidate action, based on the observation state and placement constraints corresponding to at least one candidate action, wherein the dynamic node corresponding to the given object indicates the determined placement position of the given object.

[0105] In some embodiments, a dynamic node includes an internal node, at least one leaf node, and node information corresponding to a given object. The internal node includes one or more nodes corresponding to one or more objects whose stacking positions have been determined among a plurality of objects. The internal node is represented as first vertex information, second vertex information, and flag information indicating whether a stacking object exists for each of the one or more objects. The one or more objects include the given object. The at least one leaf node includes nodes corresponding to one or more stackable positions in a target stack based on the stacking requirements and the determined stacking positions of the one or more objects. The leaf node is represented as vertex information for each stackable position and flag information indicating whether it is stackable. The node information indicates attribute information of the object following the given object among the plurality of objects.

[0106] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 Stack type monitoring system 110 or Figure 4 The device 400 shown.

[0107] like Figure 5As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0108] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0109] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0110] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0111] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0112] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0113] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0114] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0115] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0117] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A stacking type planning method, comprising: In response to receiving a stacking request for multiple objects, a trained machine learning model is used to determine the stacking position of each object within the target stacking pattern based on the attribute information of the multiple objects and the stacking pattern requirements of the target stacking pattern. The determination of the stacking position includes: for a given object among the multiple objects, Based on the stack type requirements and the determined stacking position of at least one object in the target stack type, at least one stackable space in the target stack type is determined; Based on the attribute information of the given object, determine the candidate position of the given object in the at least one stackable space; From the at least one stackable space, candidate positions that satisfy the stacking constraints are determined as the stacking positions of the given object; and The multiple objects are placed according to their respective placement positions in the target stack type via the control module of the palletizing system to form the target stack type.

2. The method according to claim 1, wherein determining the placement position of the given object comprises: In response to determining a plurality of candidate positions satisfying the stacking constraints from the at least one stackable space, a predetermined position of the plurality of candidate positions located in each of the at least one stacking space is taken as the stacking position of the given object.

3. The method according to claim 1, wherein the stacking constraint includes at least one of the following: The stability meets the target stack type requirements. Stacking preference information for stacking the multiple objects, and / or A predetermined rule determined based on the attribute information of the multiple objects.

4. The method of claim 1, wherein the stacking position of each of the plurality of objects in the target stack type is stored in a cache module accessible by the palletizing system.

5. The method of claim 1, wherein the machine learning model comprises a policy network based on reinforcement learning, and wherein determining whether the given object can be placed at a candidate position in the at least one stackable space comprises, The policy network is used to generate at least one candidate action for the given object, each candidate action indicating a candidate position to place the given object in the stackable space; And determining candidate positions satisfying the stacking constraints from the at least one stackable space as the stacking positions of the given object includes: The at least one candidate action is input into the target environment to obtain the observation state corresponding to each of the at least one candidate action; as well as Based on the observation state corresponding to each of the at least one candidate action and the stacking constraint, a target action for the given object is determined from the at least one candidate action, the target action indicating the stacking position of the given object.

6. The method of claim 5, wherein generating the at least one candidate action comprises: Initialize the observation state of the target environment to obtain at least one dynamic node corresponding to at least one of the plurality of objects, each dynamic node indicating the determined placement position of the object; Based on the dynamic nodes corresponding to the at least one object, at least using a heuristic algorithm, the policy network generates index identifiers for candidate leaf nodes corresponding to the given object, wherein the candidate leaf nodes indicate candidate positions of the given object determined based on the determined placement positions. as well as The candidate leaf node indexed to the index identifier is determined as the candidate action to be executed.

7. The method according to claim 5, wherein determining the target action for the given object from the at least one candidate action based on the observation state corresponding to each of the at least one candidate action and the stacking constraint includes: Based on the observation state corresponding to each of the at least one candidate action and the stacking constraints, the dynamic node corresponding to the given object is determined from the candidate leaf nodes corresponding to each of the at least one candidate action, and the dynamic node corresponding to the given object indicates the determined stacking position of the given object.

8. The method according to claim 7, wherein the dynamic node includes internal nodes, the at least one leaf node, and node information corresponding to the given object. The internal nodes include one or more nodes corresponding to one or more objects whose placement positions have been determined among the plurality of objects. Each internal node is represented by first vertex information, second vertex information, and flag information indicating the presence of a placement object for each of the one or more objects. The one or more objects include the given object. The at least one leaf node includes a node corresponding to one or more stackable positions in the target stack type based on the stack type requirements and the determined stacking positions of the one or more objects. The leaf node is represented as vertex information for each stackable position and flag information indicating whether it is stackable. The node information indicates the attribute information of the object following the given object among the plurality of objects.

9. An apparatus for stacking planning, comprising: A stacking position determination module is configured to, in response to receiving a stacking request for multiple objects, utilize a trained machine learning model to determine the stacking position of each object within the target stacking pattern based on the attribute information of the multiple objects and the stacking pattern requirements of the target stacking pattern. The determination of the stacking position includes: for a given object among the multiple objects... Based on the stack type requirements and the determined stacking position of at least one object in the target stack type, at least one stackable space in the target stack type is determined; Based on the attribute information of the given object, determine the candidate position of the given object in the at least one stackable space; From the at least one stackable space, candidate positions that satisfy the stacking constraints are determined as the stacking positions of the given object; and The object placement module is configured to place the plurality of objects according to their respective placement positions in the target stack type via a control module in the palletizing system, so as to form the target stack type.

10. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 8.

12. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 8.