Palletizing method and device based on multi-modal embodied intelligence, equipment and medium

CN122607723APending Publication Date: 2026-08-21BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610729662.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]这种高度依赖预设条件和人工经验的工作模式,限制了码垛机器人在动态且非结构化物流场景中的适应能力与部署效率,暴露出其灵活性差且智能化水平低的问题

Benefits of technology

[0040]This application provides a palletizing method, apparatus, device, and medium based on multimodal embodied intelligence. The method includes: acquiring a multimodal task instruction and inputting the multimodal task instruction into a multimodal model to obtain multiple preset shapes output by the multimodal model, and a preset stacking method for each preset shape; acquiring a pallet image and inputting the pallet image into a pallet recognition model to obtain the shape and size of the target pallet output by the pallet recognition model; obtaining the coordinate position of the target pallet in the pallet based on the shape, size, and stacking method of the target pallet, and palletizing the target pallet according to the coordinate position and the stacking method of the target pallet. The following technical effects were achieved: Based on multimodal task instructions, multiple preset shapes and corresponding preset stacking methods were obtained through a preset multimodal model, enabling the determination of stacking methods for different shapes of stacked items and improving the flexibility and intelligence level of the palletizing robot; Based on the stacked item image, the shape of the target stacked item was determined from multiple preset shapes and the size of the target stacked item was determined from multiple preset sizes through a preset stacked item recognition model, enabling the determination of stacked items of different shapes and sizes and improving the flexibility and intelligence level of the palletizing robot; Based on the shape, size, and stacking method of the target stacked item, the coordinate position of the target stacked item in the stack was obtained, and combined with the stacking method of the target stacked item, the palletizing of the target stacked item was achieved, improving the flexibility and intelligence level of the palletizing robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122607723A_ABST
    Figure CN122607723A_ABST
Patent Text Reader

Abstract

The application provides a palletizing method and device based on multi-modal embodiment intelligence, equipment and medium, relating to the technical field of robots. The method comprises: acquiring a multi-modal task instruction, inputting the multi-modal task instruction into a multi-modal model, obtaining a plurality of preset shapes output by the multi-modal model, and a preset stacking mode of a stack corresponding to each preset shape; acquiring a stack image and inputting the stack image into a stack recognition model to obtain the shape and size of a target stack output by the stack recognition model; obtaining the coordinate position of the target stack in the stack according to the shape, size and stacking mode of the target stack, and palletizing the target stack according to the stacking mode of the target stack according to the coordinate position. The method of the application improves the adaptability and deployment efficiency of the palletizing robot in dynamic and unstructured logistics scenarios, thereby improving the flexibility and intelligent level of the palletizing robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, and in particular to a palletizing method, apparatus, device and medium based on multimodal embodied intelligence. Background Technology

[0002] In current logistics scenarios, palletizing robots generally adopt an automated control method based on fixed programs. Their operating logic mainly relies on pre-set motion trajectories and grasping schemes. In practical applications, palletizing robots typically use a single vision sensor and simple image processing algorithms to identify and locate pallets, while the design of the pallet shape requires manual pre-setting of stacking parameters and offline programming to complete path planning and motion choreography.

[0003] Lacking the ability to perceive complex environments and make autonomous decisions, palletizing robots can only recognize standard stacks of limited sizes, making it difficult to handle diverse material forms. When faced with stacks of varying sizes or mixed categories, palletizing robots cannot autonomously generate suitable gripping and palletizing plans, requiring manual intervention for parameter adjustments and task replanning.

[0004] This working mode, which relies heavily on preset conditions and human experience, limits the adaptability and deployment efficiency of palletizing robots in dynamic and unstructured logistics scenarios, exposing their poor flexibility and low level of intelligence. Summary of the Invention

[0005] This application provides a palletizing method, apparatus, equipment, and medium based on multimodal embodied intelligence, which improves the adaptability and deployment efficiency of palletizing robots in dynamic and unstructured logistics scenarios, thereby enhancing the flexibility and intelligence level of palletizing robots.

[0006] The first aspect of this application provides a palletizing method based on multimodal embodied intelligence, the method comprising:

[0007] The system acquires multimodal task instructions and inputs them into a preset multimodal model to obtain multiple preset shapes output by the multimodal model, as well as preset stacking methods for each preset shape; where multimodal task instructions refer to task instructions in text or voice form.

[0008] The system acquires an image of a stack of objects displaying the target stack, and inputs the image into a preset stack recognition model to obtain the shape and size of the target stack output by the stack recognition model. The stack recognition model stores multiple preset sizes corresponding to each preset shape. The shape of the target stack is one of multiple preset shapes, and the size of the target stack is one of multiple preset sizes corresponding to the shape of the target stack.

[0009] Based on the shape, size, and stacking method of the target stack, the coordinate position of the target stack in the stack is obtained, and the target stack is stacked according to the coordinate position and the stacking method of the target stack; wherein, the stacking method of the target stack is the preset stacking method of the stack corresponding to the shape of the target stack.

[0010] In one possible design, the target stack is stacked according to its coordinate position and the stacking method, including:

[0011] Obtain the grasping pose of the target stack of objects;

[0012] Based on the coordinate position, the target stack is stacked in the simulation space according to the grasping posture and stacking method of the target stack.

[0013] When the simulation results indicate successful palletizing, the target stack is palletized.

[0014] In one possible design, if the image of the stack is a depth image, then obtaining the grasping pose of the target stack includes:

[0015] The image of the stack is input into a preset grasping algorithm model to obtain multiple candidate grasping poses of the target stack output by the grasping algorithm model, as well as the confidence level corresponding to each candidate grasping pose of the target stack.

[0016] Based on the confidence level corresponding to each candidate grasp pose of the target stack, the grasp pose of the target stack is selected from multiple candidate grasp poses of the target stack.

[0017] In one possible design, the stack image shows multiple stacks of objects. Based on the confidence level corresponding to each candidate grasp pose of the target stack, the grasp pose of the target stack is selected from the multiple candidate grasp poses, including:

[0018] The DBSCAN clustering algorithm is used to perform the first clustering of multiple candidate grasping poses, resulting in multiple first clusters.

[0019] The candidate grasping pose with the highest confidence in each first cluster is determined as the first grasping pose of the corresponding first cluster, resulting in multiple first grasping poses.

[0020] The K-Means clustering algorithm is used to perform a second clustering on multiple first grasp poses to obtain multiple second clusters; each of the multiple second clusters corresponds to a stack in the stack image;

[0021] The first grasping pose with the highest confidence in each second cluster is determined as the second grasping pose of the corresponding second cluster, resulting in multiple second grasping poses.

[0022] Determine the grasping pose of the target stack from multiple second grasping poses.

[0023] In one possible design, before performing the first clustering of multiple candidate grasping poses using the DBSCAN clustering algorithm to obtain multiple first clusters, the method further includes:

[0024] Based on the confidence level corresponding to each candidate grasp pose, multiple candidate grasp poses are filtered out; among them, the confidence level corresponding to the candidate grasp poses that are not filtered out is greater than the confidence level corresponding to the candidate grasp poses that are filtered out.

[0025] In one possible design, before performing a second clustering of multiple first grasping poses using the K-Means clustering algorithm to obtain multiple second clusters, the method further includes:

[0026] Input the stack image into the preset quantity recognition model to obtain the stack quantity output by the quantity recognition model; where the stack quantity refers to the number of multiple stacks in the stack image;

[0027] The number of stacks is determined as the number of clusters in the K-Means clustering algorithm.

[0028] In one possible design, each preset shape corresponds to at least one preset stacking method of the stack. Then, based on the shape, size, and stacking method of the target stack, the coordinate position of the target stack within the stack is obtained, including:

[0029] Based on the shape and stacking method of the target stack, a coordinate position preset algorithm corresponding to the target stack is determined from multiple coordinate position preset algorithms; wherein, each of the multiple coordinate position preset algorithms is applied to a combination of a preset shape and a preset stacking method;

[0030] Based on the size of the target stack, the coordinate position of the target stack in the stack is obtained through a preset algorithm corresponding to the coordinate position of the target stack.

[0031] A second aspect of this application provides a palletizing device based on multimodal embodied intelligence, the device comprising:

[0032] The multimodal instruction analysis module is used to acquire multimodal task instructions and input them into a preset multimodal model to obtain multiple preset shapes output by the multimodal model, as well as the preset stacking method of each preset shape; wherein, the multimodal task instructions refer to task instructions in text or voice form;

[0033] The stack image analysis module is used to acquire stack images displaying target stacks, input the stack images into a preset stack recognition model, and obtain the shape and size of the target stack output by the stack recognition model. The stack recognition model stores multiple preset sizes corresponding to each preset shape. The shape of the target stack is one of multiple preset shapes, and the size of the target stack is one of multiple preset sizes corresponding to the shape of the target stack.

[0034] The stacking planning module is used to obtain the coordinate position of the target stack in the stack according to the shape, size and stacking method of the target stack, and stack the target stack according to the coordinate position and the stacking method of the target stack; wherein, the stacking method of the target stack is the preset stacking method of the stack corresponding to the shape of the target stack.

[0035] A third aspect of this application provides an electronic device, including: a memory, and a memory communicatively connected to a processor;

[0036] The memory stores the instructions that the computer executes;

[0037] When the processor executes computer execution instructions stored in memory, it implements the palletizing method based on multimodal embodied intelligence for any of the first aspects.

[0038] The fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the palletizing method based on multimodal embodied intelligence according to any one of the first aspects.

[0039] The fifth aspect of this application provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the palletizing method based on multimodal embodied intelligence according to any one of the first aspects.

[0040] This application provides a palletizing method, apparatus, device, and medium based on multimodal embodied intelligence. The method includes: acquiring a multimodal task instruction and inputting the multimodal task instruction into a multimodal model to obtain multiple preset shapes output by the multimodal model, and a preset stacking method for each preset shape; acquiring a pallet image and inputting the pallet image into a pallet recognition model to obtain the shape and size of the target pallet output by the pallet recognition model; obtaining the coordinate position of the target pallet in the pallet based on the shape, size, and stacking method of the target pallet, and palletizing the target pallet according to the coordinate position and the stacking method of the target pallet. The following technical effects were achieved: Based on multimodal task instructions, multiple preset shapes and corresponding preset stacking methods were obtained through a preset multimodal model, enabling the determination of stacking methods for different shapes of stacked items and improving the flexibility and intelligence level of the palletizing robot; Based on the stacked item image, the shape of the target stacked item was determined from multiple preset shapes and the size of the target stacked item was determined from multiple preset sizes through a preset stacked item recognition model, enabling the determination of stacked items of different shapes and sizes and improving the flexibility and intelligence level of the palletizing robot; Based on the shape, size, and stacking method of the target stacked item, the coordinate position of the target stacked item in the stack was obtained, and combined with the stacking method of the target stacked item, the palletizing of the target stacked item was achieved, improving the flexibility and intelligence level of the palletizing robot. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 A flowchart illustrating the palletizing method based on multimodal embodied intelligence provided in this application embodiment. Figure 1 ;

[0043] Figure 2 A flowchart illustrating the palletizing method based on multimodal embodied intelligence provided in this application embodiment. Figure 2 ;

[0044] Figure 3 A schematic diagram illustrating the principle of overlapping stacking provided in an embodiment of this application;

[0045] Figure 4 A schematic diagram illustrating the principle of staggered stacking provided in an embodiment of this application;

[0046] Figure 5 A schematic diagram illustrating the principle of a tower-type stack provided in an embodiment of this application;

[0047] Figure 6 A flowchart illustrating the palletizing method based on multimodal embodied intelligence provided in this application embodiment. Figure 3 ;

[0048] Figure 7 A schematic diagram of the structure of a palletizing device based on multimodal embodied intelligence provided in an embodiment of this application;

[0049] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0050] Figure label:

[0051] 710 - Multimodal instruction analysis module; 720 - Stacking image analysis module; 730 - Stacking planning module;

[0052] 810 - Processor; 820 - Memory; 830 - Communication components; 840 - Bus. Detailed Implementation

[0053] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0054] In this application, the terms "first" and "second" are used to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, nor do they necessarily imply difference. It should be noted that in this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner. In this application, "at least one" means one or more, and "more than one" means two or more.

[0055] It should be noted that the phrase "at the moment" in this application can refer to the instant a certain situation occurs, or to a period of time after the occurrence of a certain situation; this application does not impose a specific limitation on this. Furthermore, the palletizing method based on multimodal embodied intelligence provided in this application is merely an example; palletizing methods based on multimodal embodied intelligence may also include more or less content. The user information (including but not limited to user device information and user personal information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in one or more embodiments of this application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of related data must comply with relevant laws, regulations, and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0056] To facilitate a clear description of the technical solution of this application, some of the terms and technologies involved in this application are briefly introduced below:

[0057] Stacked goods: refers to a single item or product with a specific shape and size. In logistics scenarios, stacked goods are the basic operational unit. Depending on their shape and size, different stacking methods are required to ensure the stability of the stack and the efficiency of space utilization.

[0058] A stack is a structure or whole formed by stacking multiple items of the same shape and size in a certain way. The stacking method depends on the shape and size of the items and the stacking requirements.

[0059] Palletizing is the process of placing items into designated locations within a stack based on their shape, size, and stacking method. Palletizing requires precise control over the gripping posture, position, and orientation of the items.

[0060] To clearly understand the technical solution of this application, the solutions of the prior art will be described in detail first.

[0061] In current logistics scenarios, palletizing robots generally adopt an automated control method based on fixed programs, that is, automated control is carried out through the Robot Operating System (ROS), and its operating logic mainly depends on the pre-set motion trajectory and grasping scheme.

[0062] ROS is a robot control platform designed specifically for robot software development. It is an open-source meta-operating system, or post-operating system, providing services similar to an operating system, including hardware abstraction description, low-level driver management, execution of common functions, inter-program message passing, and program distribution package management. It also provides tools and libraries for acquiring, building, writing, and executing multi-machine fusion programs. For palletizing robots, a digital twin simulation system for the robotic arm was built based on the ROS framework. Through a cross-platform control interface, intelligent palletizing operations of the robotic arm in a real production environment were achieved, realizing dual-end execution of palletizing tasks.

[0063] In practical applications, palletizing robots typically use a single vision sensor and simple image processing algorithms to identify and locate palletized items, while the design of the pallet shape requires manual pre-setting of stacking parameters and offline programming to complete path planning and motion choreography.

[0064] Lacking the ability to perceive complex environments and make autonomous decisions, palletizing robots can only recognize standard stacks of limited sizes, making it difficult to handle diverse material forms. When faced with stacks of varying sizes or mixed categories, palletizing robots cannot autonomously generate suitable gripping and palletizing plans, requiring manual intervention for parameter adjustments and task replanning.

[0065] This working mode, which relies heavily on preset conditions and human experience, limits the adaptability and deployment efficiency of palletizing robots in dynamic and unstructured logistics scenarios, exposing their poor flexibility and low level of intelligence.

[0066] Therefore, to address the aforementioned technical issues, the research found that by deeply integrating multimodal inputs, embodied intelligence models, and palletizing robots, a complete closed loop from environmental perception to autonomous execution can be constructed, enabling palletizing robots to better perform task-oriented palletizing. Among these, the embodied intelligence model allows the palletizing robot to make better decisions, thus improving its intelligence level.

[0067] Specifically, to address the issue that palletizing robots cannot determine the stacking method for different shaped pallets in various pallet scenarios, an embodied intelligence model used for multimodal task instruction recognition is applied to the palletizing robot, and combined with multimodal input, the stacking method for different shaped pallets is determined.

[0068] To address the issue that palletizing robots cannot determine the shape and size of palletized items in various scenarios, an embodied intelligent model for image recognition is applied to the palletizing robot, and combined with images of the palletized items, the determination of palletized items with different shapes and sizes is achieved.

[0069] To address the issue that palletizing robots cannot palletize items according to different stacking methods in various scenarios, an embodied intelligent model for coordinate position calculation and grasping pose determination is applied to the palletizing robot. Combined with the stacking method of the items, it enables the palletizing of items of different shapes and sizes.

[0070] Based on the above-mentioned inventive discovery, the technical solution of this application is proposed.

[0071] The following section introduces the application scenarios of the palletizing method based on multimodal embodied intelligence provided in this application.

[0072] In one embodiment of this application, a common palletizing scenario is provided for a single palletized item. Specifically, most palletizing systems typically operate in conjunction with conveyor belt systems, where palletized items are sequentially transported to the palletizing robot's work area via conveyor belts or other conveying devices. When the palletizing robot detects a single palletized item arriving at a designated position, it identifies and grasps the item using a palletizing method based on multimodal embodied intelligence, and then palletizes it to the target coordinate position of the pallet. After all palletized items are completed, a common pallet is obtained.

[0073] In another embodiment of this application, a scenario for mixed stacking of multiple stacks of items is provided. Specifically, multiple stacks of items with different shapes and / or sizes are mixed and stacked within the working area of ​​a palletizing robot. The palletizing robot identifies and grasps each stack using a palletizing method based on multimodal embodied intelligence and stacks it to the target coordinate position of the stack. After all stacks are stacked, a mixed stack is obtained.

[0074] In both application scenarios described above, the palletizing robot stacks of items with the same shape and size into the same pile. For example, stacks of cubes measuring 12cm × 12cm × 12cm are stacked into the first pile; stacks of cubes measuring 16cm × 16cm × 16cm are stacked into the second pile; and stacks of cuboids measuring 12cm × 42cm × 12cm are stacked into the third pile. It is important to note that each of the first, second, and third piles includes at least one stack.

[0075] The technical solutions of this application will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0076] Figure 1 A flowchart illustrating the palletizing method based on multimodal embodied intelligence provided in this application embodiment. Figure 1 .like Figure 1 As shown in the embodiments of this application, the executing entity can be a palletizing device based on multimodal embodied intelligence. This device can be located in an electronic device, and the device can be a palletizing robot. Therefore, the palletizing method based on multimodal embodied intelligence provided in the embodiments of this application includes the following steps:

[0077] S101. Obtain the multimodal task instruction and input the multimodal task instruction into the preset multimodal model to obtain multiple preset shapes output by the multimodal model, as well as the preset stacking method of the stack corresponding to each preset shape.

[0078] Specifically, multimodal refers to an interaction method or technological path that integrates multiple information forms such as text, speech, and images for understanding and expression. Multimodal task instructions refer to task instructions in text or speech form. A multimodal model is a large-scale model, a powerful artificial intelligence model with a huge number of parameters, capable of handling complex tasks and possessing versatility, widely used in fields such as natural language processing and multimodal understanding. A multimodal model is an artificial intelligence model capable of simultaneously understanding and generating multiple information forms such as text, speech, and images to achieve cross-modal perception and intelligent interaction. Multimodal models are used to extract preset shapes and preset stacking methods from multimodal task instructions through cue word engineering.

[0079] After the user inputs multimodal task commands into the multimodal embodied intelligence-based palletizing device, the device analyzes these commands using a pre-set multimodal model. The model, through prompt word engineering, organizes a task list in a fixed format based on the commands, extracting multiple preset shapes and their corresponding preset stacking methods from the complex commands. It's important to note that the deployment of the multimodal model on the palletizing robot primarily includes local deployment and cloud deployment; this embodiment does not limit the specific deployment method.

[0080] This application leverages the superior capabilities of multimodal models in text understanding, speech understanding, and visual perception understanding to comprehend user-inputted multimodal task instructions and make intelligent decisions. For example, if a user inputs a textual task instruction: "Stack cube-shaped objects in an overlapping manner, stack rectangular objects in a staggered manner, and stack cylindrical objects in a tower manner," the multimodal model will output: "When the object shape is cube, the stacking method is overlapping; when the object shape is rectangular, the stacking method is staggered; when the object shape is cylindrical, the stacking method is tower."

[0081] In one possible design, the user inputs task instructions in the form of voice. Voice input is superior to text input in terms of efficiency and convenience. Therefore, to cope with scenarios with more modal inputs, multimodal models are also used to convert real-time user-inputted speech into standard natural language text through voice acquisition, format conversion, speech model transcription, and transcription result parsing, providing a voice input channel for multimodal task instructions.

[0082] S102. Obtain an image of the stacked object displaying the target stacked object, and input the stacked object image into a preset stacked object recognition model to obtain the shape and size of the target stacked object output by the stacked object recognition model.

[0083] Specifically, in ordinary palletizing applications, after inputting multimodal task commands into the multimodal model, the camera is activated. When a single pallet is detected reaching a designated position, the camera captures an image of that pallet. In hybrid palletizing applications, after inputting multimodal task commands into the multimodal model, the camera is activated and captures images of the work area to obtain pallet images. When capturing images of the pallet, the camera can be fixedly mounted outside the end effector of the palletizing robot's arm, for example, directly above the gripping point of the arm. It should be noted that the deployment methods of the pallet recognition model and other subsequent embodied intelligence models on the palletizing robot mainly include local deployment and cloud deployment; this application embodiment does not limit the specific deployment methods.

[0084] In one possible design, the multimodal task instruction also includes a stack image. The user then takes a picture of the stack and uploads it to the palletizing device based on multimodal embodied intelligence. Step S101 includes: acquiring the multimodal task instruction, inputting the multimodal task instruction into a preset multimodal model, and obtaining multiple preset shapes output by the multimodal model, preset stacking methods for each preset shape, and a stack image.

[0085] The stack recognition model is used to extract key information from stack images to obtain the shape and size of the target stack. For this model, a scalable "stack library" is first constructed to improve its accuracy and adaptability in stack recognition tasks. Under a fixed camera viewpoint, for stacks with significant size differences, the stack recognition model can reliably identify the size of the target stack using the stack library. To this end, multiple preset stack shapes can be pre-defined, and each preset shape can have multiple preset sizes. Therefore, the stack recognition model stores multiple preset sizes corresponding to each preset shape in the stack library.

[0086] After acquiring the image of the stacked item, the palletizing device based on multimodal embodied intelligence inputs the image into the stacking item recognition model. The stacking item recognition model performs discrimination and recognition in the stacking item database to obtain the shape and size of the target stacked item. The shape of the target stacked item is one of a variety of preset shapes, and the size of the target stacked item is one of a variety of preset sizes corresponding to the shape of the target stacked item.

[0087] In one possible design, the stack recognition model needs to be fine-tuned based on the various preset shapes output by the multimodal model to ensure the size of the target stack is output.

[0088] In one possible design, the stacking image may show multiple stacks, and the stacking recognition model outputs the shape and size of each stack. It should be noted that the target stack is any one of the multiple stacks in the stacking image. The method for stacking these stacks is similar to the method for stacking the target stack, and will not be described again in the embodiments of this application.

[0089] S103. Based on the shape, size and stacking method of the target stack, obtain the coordinate position of the target stack in the stack, and stack the target stack according to the coordinate position and the stacking method of the target stack.

[0090] Specifically, the multimodal embodied intelligence-based palletizing device determines which stack a target stack belongs to, its coordinate position within that stack, and its stacking pattern based on the target stack's shape, size, and stacking method. The stacking pattern is a preset pattern for the stack corresponding to the target stack's shape; to ensure stability and efficiency during palletizing, only stacks of the same shape and size are placed within a single stack.

[0091] Coordinate position refers to the location of a stacked item in three-dimensional space, defined by a mathematical coordinate system. It precisely describes the specific position of the target stacked item within the stack, including its horizontal position and vertical height. For example, with a corner of the stack as the origin, the X and Y axes are defined as the horizontal directions, and the Z axis as the vertical height. Based on the target stacked item's coordinate position and the stacking method, assuming no overlap between stacked items and that they do not exceed the stack boundaries, the specific position of the target stacked item within the stack is calculated. Then, based on this coordinate position, the palletizing robot's robotic arm is guided to grasp the target stacked item and place it in the designated position, completing the palletizing process. During this process, the palletizing device based on multimodal embodied intelligence can optimize the robotic arm's movement path, avoiding collisions with the surrounding environment and already stacked items.

[0092] This application provides a palletizing method based on multimodal embodied intelligence. The method includes: acquiring a multimodal task instruction and inputting the multimodal task instruction into a multimodal model to obtain multiple preset shapes output by the multimodal model, and a preset stacking method for each preset shape; acquiring a pallet image and inputting the pallet image into a pallet recognition model to obtain the shape and size of the target pallet output by the pallet recognition model; obtaining the coordinate position of the target pallet in the pallet based on the shape, size, and stacking method of the target pallet, and palletizing the target pallet according to the coordinate position and the stacking method of the target pallet. The following technical effects were achieved: Based on multimodal task instructions, multiple preset shapes and corresponding preset stacking methods were obtained through a preset multimodal model, enabling the determination of stacking methods for different shapes of stacked items and improving the flexibility and intelligence level of the palletizing robot; Based on the stacked item image, the shape of the target stacked item was determined from multiple preset shapes and the size of the target stacked item was determined from multiple preset sizes through a preset stacked item recognition model, enabling the determination of stacked items of different shapes and sizes and improving the flexibility and intelligence level of the palletizing robot; Based on the shape, size, and stacking method of the target stacked item, the coordinate position of the target stacked item in the stack was obtained, and combined with the stacking method of the target stacked item, the palletizing of the target stacked item was achieved, improving the flexibility and intelligence level of the palletizing robot.

[0093] Figure 2 A flowchart illustrating the palletizing method based on multimodal embodied intelligence provided in this application embodiment. Figure 2 ,like Figure 2 As shown, the palletizing method based on multimodal embodied intelligence provided in this application embodiment is... Figure 1 The palletizing method based on multimodal embodied intelligence provided in this embodiment is further refined. In one possible design, each preset shape corresponds to at least one preset stacking method for the stack. For example, when the shape of the stack is a cube, the stacking method can be an overlapping method indicated by the multimodal task instruction, or other forms, such as honeycomb and mesh. The palletizing method based on multimodal embodied intelligence provided in this embodiment includes the following steps.

[0094] S201. Obtain the multimodal task instruction and input the multimodal task instruction into the preset multimodal model to obtain multiple preset shapes output by the multimodal model, as well as the preset stacking method of the stack corresponding to each preset shape.

[0095] S202. Obtain an image of the stacked object displaying the target stack, and input the stacked object image into a preset stacked object recognition model to obtain the shape and size of the target stacked object output by the stacked object recognition model.

[0096] S203. Based on the shape and stacking method of the target stack, determine the preset coordinate position algorithm corresponding to the target stack from multiple preset coordinate position algorithms.

[0097] S204. Based on the dimensions of the target stack, the coordinate position of the target stack in the stack is obtained through a preset algorithm for the coordinate position of the target stack.

[0098] Specifically, the shape of the stacked items can include cubes, cuboids, and cylinders; in order to achieve stacking of various preset shapes, multiple preset stacking methods are designed for each of the cube, cuboid, and cylinder shapes.

[0099] When the shape of the stack is a cube, the preset stacking method of the corresponding stack can be overlapping, honeycomb, or grid, etc.; preferably, taking into account the stability and efficiency during stacking, the stacking method of the target stack is determined to be overlapping according to the multimodal task instruction.

[0100] When the shape of the stack is a cuboid, the preset stacking method of the corresponding stack can be staggered, layered, or jigsaw puzzle, etc.; preferably, the stability of stacking is improved by placing multiple layers in staggered directions, and the stacking failure caused by slippage and center of gravity shift is reduced while maintaining space utilization. Therefore, the stacking method of the target stack is determined to be staggered according to the multimodal task instruction.

[0101] When the shape of the stack is cylindrical, the preset stacking method of the corresponding stack can be tower, plum blossom, spiral, etc.; preferably, the tower arrangement increases the contact area between the stacks to ensure that the cylindrical stacks do not roll and avoid the collapse of the stack shape. Therefore, the stacking method of the target stack is determined to be tower according to the multimodal task instruction.

[0102] Taking these three preset shapes and the various preset stacking methods corresponding to each preset shape as examples, multiple coordinate position preset algorithms were designed. Each of these algorithms is applied to a combination of a preset shape and a preset stacking method, which can satisfy the stacking of the corresponding preset shape within a certain area. When calculating the coordinate position, the coordinate position preset algorithm corresponding to the target stack is determined from the multiple coordinate position preset algorithms to calculate the coordinate position of the target stack in the corresponding stack.

[0103] Figure 3 This is a schematic diagram illustrating the principle of overlapping stacking provided in an embodiment of this application. Figure 3As shown, when the target stacking method is overlapping, this algorithm is designed around the intelligent palletizing scenario of a palletizing robot. It is suitable for the standardized arrangement of cubic stacks of the same size within a regular stacking area, that is, stacking cubic stacks of the same size layer by layer without offset between adjacent stacks. The shape of the stack is simplified to a cube. Let A be the area of ​​one surface of the stack, then the length of one side of the stack is determined by integer division calculation, which is L. The calculation formula is as follows: .

[0104] Given the side dimensions (l, w, h) of each stack, representing its length, width, and height respectively, and assuming the number of stacks to be stacked is N, then the number of stacks per layer is... The formulas for calculating the number of stack layers are as follows:

[0105]

[0106]

[0107]

[0108] Additionally, let the global coordinate system offset be... , serving as a three-dimensional offset correction term for the origin of the stacking area. Considering the grasping requirements of the palletizing robot's robotic arm, an operation interval g is introduced in each direction to ensure sufficient spatial freedom for the robotic arm during the palletizing process. For the stack in the i-th row and j-th column of the k-th stack, the coordinates of the center point of its upper surface are:

[0109]

[0110]

[0111]

[0112] The orientation of each palletizing point is set to no rotation by default (i.e., r=q=yaw=0 in quaternions or Euler angles), suitable for symmetrical objects or fixed-orientation grasping scenarios. Finally, the coordinate positions of each item in each row and column of each layer are obtained. These calculated coordinate positions are visualized to check the validity and correctness of the palletizing. The visualization results are as follows: Figure 3 As shown.

[0113] Figure 4 This is a schematic diagram illustrating the principle of interleaved stacking provided in an embodiment of this application. Figure 4 As shown, when the target stack is staggered, the stack area is simplified to a cube. Let A be the area of ​​one surface of the stack. Then, the length L of one side of the stack is determined by integer division. The calculation formula is as follows: .

[0114] Given the side dimensions (l, w, h) of each stack, representing its length, width, and height respectively, determine the number of stacks per layer. The calculation formula for the number of layers in a stack is similar to that when the stacking method of the target stack is overlapping, and will not be repeated in the embodiments of this application.

[0115] The target stack is staggered, meaning that even-numbered stacks maintain their orientation, while odd-numbered stacks are rotated 90° around the z-axis. This design significantly enhances interlayer bonding and reduces the risk of collapse in engineering practice, but also introduces geometric variations that present more complex route planning challenges.

[0116] Let the global coordinate system offset be... An operation interval g is introduced in each direction. For the stack in the i-th row and j-th column of the k-th stack, when k is even, the coordinates of the center point of its upper surface are:

[0117]

[0118]

[0119]

[0120]

[0121] When k is odd, the coordinates of the center point of its upper surface are:

[0122]

[0123]

[0124]

[0125]

[0126] Finally, the coordinates of each item in each row and column of each layer are obtained. These calculated coordinates are then visualized to check the validity and correctness of the palletizing. The visualization results are as follows: Figure 4 As shown.

[0127] Figure 5 This is a schematic diagram illustrating the principle of a tower-type stack provided in an embodiment of this application. Figure 5 As shown, when the target stack is stacked in a tower configuration, the shape of the stack is simplified to a cuboid. Let A be the area of ​​one surface of the stack. Then, the length of one side of the stack is determined by integer division, and the formula is as follows: .

[0128] Given the side dimensions (l, w, h) of each stack, representing its length, width, and height respectively, determine the number of stacks per layer. The calculation formula for the number of layers in a stack is similar to that when the stacking method of the target stack is overlapping, and will not be repeated in the embodiments of this application.

[0129] Let the global coordinate system offset be... An operation interval g is introduced in each direction. For the stack in the i-th row and j-th column of the k-th stack layer, the coordinates of the center point of its upper surface are:

[0130]

[0131]

[0132]

[0133]

[0134] Finally, the coordinates of each item in each row and column of each layer are obtained. These calculated coordinates are then visualized to check the validity and correctness of the palletizing. The visualization results are as follows: Figure 5 As shown.

[0135] S205. Obtain the grasping pose of the target stack.

[0136] Specifically, in the application scenario of ordinary stacking, when the target stack is transported to the working area of ​​the palletizing robot, and its relative position and rotation angle with other stacks are the same, the palletizing device based on multimodal embodied intelligence extracts the preset grasping posture of the target stack from the preset program.

[0137] In conventional palletizing applications, when a target pallet is transported to the palletizing robot's work area and its relative position and / or rotation angle differs from other pallets, or in hybrid palletizing applications, a palletizing device based on multimodal embodied intelligence calculates the preset grasping pose of the target pallet using a certain algorithm. This algorithm can be a pose estimation algorithm based on 3D vision, a localization algorithm based on feature points, or a pose estimation algorithm based on a deep learning grasping model.

[0138] S206. Based on the coordinate position, stack the target stack in the simulation space according to the grasping posture and stacking method of the target stack.

[0139] S207. When the simulation results indicate that the palletizing is successful, palletize the target stack.

[0140] Specifically, after obtaining the grasping pose and coordinate position, the palletizing device based on multimodal embodied intelligence generates a task list. This task list clearly outlines the core content required for palletizing, including the dimensions of the items to be palletized, the stacking method of each item, and the coordinate position of each item within its corresponding stack. The task list is then combined with the aforementioned grasping pose to form a complete palletizing task execution instruction. This instruction provides a precise basis for subsequent robotic arm path calculation and control, realizing full-process control from understanding the multimodal task instruction to palletizing the items.

[0141] Before stacking the target stack, the stacking of the target stack is simulated in the simulation space.

[0142] First, based on the multimodal embodied intelligence palletizing device, a digital twin simulation system for the robotic arm is built within the robot operating system framework. The parameters of the robotic arm joints and various controllers are configured to achieve the maximum simulation of the palletizing robot in the simulation space. At the same time, it is connected to the real robotic arm to achieve joint communication with the real robotic arm.

[0143] Secondly, the palletizing device based on multimodal embodied intelligence is used to build a scene of the physical object to avoid physical collisions and unreasonable movement trajectories in the experiment.

[0144] Finally, the palletizing device based on multimodal embodied intelligence is used to simulate the palletizing of the target stack. When the simulation result indicates that the palletizing has failed, the coordinate position and / or grasping pose of the target stack are reacquired, and the palletizing of the target stack is simulated again in the simulation space; when the simulation result indicates that the palletizing has succeeded, the target stack is palletized in the experiment.

[0145] The technical effects of this application embodiment are: the stacking of the target stack is simulated in the simulation space, and the target stack is stacked when the simulation result indicates that the stacking is successful, which improves the safety and reliability of stacking; the coordinate position of the target stack in the stacking is calculated according to the stacking method of the target stack, as well as the number of stacks in each layer and the number of stack layers, realizing the stacking of target stacks of different shapes and sizes.

[0146] In one possible design, the image of the stack is a depth image, then S205 obtains the grasping pose of the target stack, including:

[0147] S2051. Input the image of the stack into the preset grasping algorithm model to obtain multiple candidate grasping poses of the target stack output by the grasping algorithm model, and the confidence level corresponding to each candidate grasping pose of the target stack.

[0148] Specifically, in typical palletizing scenarios, when a target pallet is transported to the palletizing robot's working area, it is often impossible to ensure that its relative position and rotation angle are the same as other pallets. Alternatively, in mixed palletizing scenarios, multiple pallets often exist in an unstructured, mixed-stacking manner, posing a significant challenge to intelligent grasping and stacking tasks. Traditional methods based on rule-based stacking or pure visual perception have significant limitations in dealing with occlusion between pallets and pallets of various shapes and sizes. To improve the grasping capabilities of palletizing robots in both typical and mixed palletizing scenarios, a grasping algorithm model is integrated into a palletizing device based on multimodal embodied intelligence.

[0149] The grasping algorithm model is used for grasping objects in cluttered scenes. It can be a GraspNet model (an open-source benchmark model for grasping algorithms). Using a depth image dataset, it designs a grasping pose prediction network based on point cloud input, significantly improving the robustness and evaluation efficiency of grasping prediction. The grasping algorithm model takes a depth image as input, performs network prediction calculations, and outputs multiple candidate grasping poses of the target object in the depth image, along with the confidence score for each candidate pose. The output includes (x, y, z, r, p, yaw, s), where x, y, and z represent the coordinates of a candidate grasping pose, r, p, and yaw represent the orientation of the candidate grasping pose, and s represents the confidence score of the candidate grasping pose.

[0150] It should be noted that the stack image may show multiple stacks. In this case, the grasping algorithm model outputs multiple candidate grasping poses for each stack, as well as the confidence level for each candidate grasping pose.

[0151] S2052. Based on the confidence level corresponding to each candidate grasping pose of the target stack, the grasping pose of the target stack is obtained from multiple candidate grasping poses of the target stack.

[0152] Specifically, in actual operation, it is impossible to respond to all candidate grasping poses simultaneously. Therefore, it is necessary to select the grasping pose of the target stack from multiple candidate grasping poses. This can be done by determining the candidate grasping pose with the highest confidence among the multiple candidate grasping poses of the target stack as the grasping pose of the target stack; or by using a certain algorithm to determine one of the multiple candidate grasping poses of the target stack as the grasping pose of the target stack.

[0153] The technical effect of this application embodiment is that by inputting the image of the stacked goods into the preset grasping algorithm model, the grasping pose of the target stacked goods in the image of the stacked goods is determined, thereby realizing the intelligent grasping and stacking of the target stacked goods.

[0154] Figure 6A flowchart illustrating the palletizing method based on multimodal embodied intelligence provided in this application embodiment. Figure 3 In one possible design, in a hybrid stacking application scenario, the stack image shows multiple stacks. Then, in step S2052, based on the confidence level corresponding to each candidate grasping pose of the target stack, the grasping pose of the target stack is selected from the multiple candidate grasping poses, including:

[0155] S601. Filter multiple candidate grasp poses based on the confidence level corresponding to each candidate grasp pose.

[0156] Specifically, the grasping algorithm model converts the stack image O into a 3D point cloud image and generates M candidate grasping poses located on the 3D point cloud image. and the m-th candidate grasping pose Corresponding confidence level ,in, ∈[0,1]. The set of M candidate grasp poses is denoted as G.

[0157] The confidence score corresponding to each candidate grasping pose directly represents its reliability or success rate. However, the number of candidate grasping poses is large. To support subsequent cluster analysis, these candidate grasping poses undergo preliminary screening to reduce computational load. Specifically, for the M candidate grasping poses... A preliminary screening is performed, retaining only multiple high-confidence grasping poses with a confidence level greater than or equal to a confidence threshold, where the confidence threshold can be 0.9; or, M candidate grasping poses are sorted in ascending order of confidence level. The data is sorted sequentially, retaining only the highest-confidence crawl poses that fall within the ranking threshold. The ranking threshold can be 90%. High-confidence crawl poses are those that were not filtered out; other candidate crawl poses are those that were filtered out. The confidence level of the candidate crawl poses that were not filtered out is higher than the confidence level of the candidate crawl poses that were filtered out.

[0158] For example, after filtering multiple candidate grasping poses based on a confidence threshold, the set of candidate grasping poses that are not filtered out is represented as:

[0159]

[0160] in, This refers to the confidence threshold.

[0161] S602. Using the DBSCAN clustering algorithm, multiple candidate grasping poses are clustered for the first time to obtain multiple first clusters.

[0162] S603. The candidate grasping pose with the highest confidence in each first cluster is determined as the first grasping pose of the corresponding first cluster, resulting in multiple first grasping poses.

[0163] Specifically, the number of remaining candidate grasping poses is still relatively large. To further reduce the computational load, a density-based spatial clustering of applications with noise (DBSCAN) algorithm is used to identify high-density regions of candidate grasping poses with relatively lenient clustering conditions. It's important to note that the clustering results do not correspond to the number of stacks N, which is usually greater than or equal to N; the number of stacks refers to the number of stacks in the stack image. Assuming there are k first clusters, the set of the k first clusters is represented as:

[0164]

[0165] Multiple candidate grasp poses are clustered in the first stage to obtain multiple first clusters. In each first cluster, the candidate grasp pose with the highest confidence is selected and determined as the first grasp pose of the corresponding first cluster, so as to further reduce the number of candidate grasp poses.

[0166] The first grasping pose of the first cluster is represented as:

[0167]

[0168] The set of k first grasp poses is represented as:

[0169]

[0170] S604. Input the stack image into the preset quantity recognition model to obtain the stack quantity output by the quantity recognition model.

[0171] Specifically, the quantity recognition model is used for visual perception and scene understanding, analyzing images of stacked items to automatically identify and output the quantity of items, replacing traditional manual techniques. A palletizing device based on multimodal embodied intelligence invokes the quantity recognition model, taking an image of a stack containing multiple items as input, and uses the model to infer the quantity of items.

[0172] S605. Determine the number of stacked items as the number of clusters for the K-Means clustering algorithm.

[0173] S606. Using the K-Means clustering algorithm, multiple first grasping poses are clustered a second time to obtain multiple second clusters.

[0174] Each of the multiple second clusters corresponds to one stack in the stack image.

[0175] S607. The first grasping pose with the highest confidence in each second cluster is determined as the second grasping pose of the corresponding second cluster, thus obtaining multiple second grasping poses.

[0176] Specifically, the number of stacked items is determined as the number of clusters. A K-means clustering algorithm is used to perform a second clustering of multiple first grasping poses, clearly defining the cluster divisions corresponding to the stacked items. Finally, the first grasping pose with the highest confidence in each second cluster is selected as the second grasping pose for the corresponding second cluster, achieving efficient and accurate stacked item grasping pose determination.

[0177] Assuming there are N categories, consider the set of k first grasp poses. As input, execute the K-Means clustering algorithm, with N clusters. Multiple second clusters are represented as follows:

[0178]

[0179] The j-th second cluster The second grasping pose is represented as:

[0180]

[0181] The set of N second grasping poses is represented as:

[0182]

[0183] S608. Determine the grabbing pose of the target stack from multiple second grabbing poses.

[0184] Specifically, since multiple second grasping positions each correspond to one stacked item in the stack image, resulting in multiple second grasping poses, the palletizing device based on multimodal embodied intelligence can only grasp multiple stacked items in order of confidence. Because it cannot identify the correspondence between each stacked item and its grasping pose, it cannot yet achieve grasping of a target stacked item. Therefore, the stacked item image and multiple second grasping poses are input into a preset cross-modal representation fusion unit. By setting some preset conditions, alignment between the grasping poses and the stacked items is achieved. Finally, the alignment result is stored in a dictionary, enabling intelligent grasping of target stacked items in mixed stacking scenarios.

[0185] The technical effect of this application embodiment is that: through cluster analysis, the target stack's grasping pose is selected from multiple candidate grasping poses in the stack image; and the target stack indicated by the multimodal task instruction is stacked using a pre-set multimodal model, stack recognition model, grasping algorithm model, and quantity recognition model, thereby realizing the multimodal embodied intelligence of the palletizing robot.

[0186] Figure 7 A schematic diagram of the structure of the palletizing device based on multimodal embodied intelligence provided in the embodiments of this application is shown below. Figure 7 As shown in the embodiments of this application, the palletizing device based on multimodal embodied intelligence can be located in an electronic device. This palletizing device based on multimodal embodied intelligence includes:

[0187] The multimodal instruction analysis module 710 is used to acquire multimodal task instructions and input the multimodal task instructions into a preset multimodal large model to obtain multiple preset shapes output by the multimodal large model, as well as the preset stacking method of the stack corresponding to each preset shape; wherein, the multimodal task instructions refer to task instructions in the form of text or voice.

[0188] The stack image analysis module 720 is used to acquire stack images displaying target stacks, input the stack images into a preset stack recognition model, and obtain the shape and size of the target stack output by the stack recognition model; wherein, the stack recognition model stores multiple preset sizes corresponding to each preset shape; the shape of the target stack is one of multiple preset shapes, and the size of the target stack is one of multiple preset sizes corresponding to the shape of the target stack;

[0189] The stacking planning module 730 is used to obtain the coordinate position of the target stack in the stack according to the shape, size and stacking method of the target stack, and stack the target stack according to the coordinate position and the stacking method of the target stack; wherein, the stacking method of the target stack is the preset stacking method of the stack corresponding to the shape of the target stack.

[0190] The palletizing device based on multimodal embodied intelligence provided in this application embodiment can perform... Figure 1 The technical solution of the method embodiment shown has the same implementation principle and technical effect as... Figure 1 The method embodiments shown are similar, and will not be described again in the embodiments of this application.

[0191] Meanwhile, the palletizing device based on multimodal embodied intelligence provided in this application embodiment is a further refinement based on the palletizing device based on multimodal embodied intelligence provided in the previous application embodiment.

[0192] In one possible design, the stacking planning module 730 includes:

[0193] The grasping pose acquisition module is used to acquire the grasping pose of the target stack of objects;

[0194] The stacking simulation module is used to stack target stacks in simulation space according to their coordinate positions, gripping postures, and stacking methods.

[0195] The stacking control module is used to stack the target stack when the simulation results indicate that the stacking is successful.

[0196] In one possible design, if the image of the stack is a depth image, then the grasping pose acquisition module includes:

[0197] The candidate grasping pose determination module is used to input the image of the stack into a preset grasping algorithm model to obtain multiple candidate grasping poses of the target stack output by the grasping algorithm model, as well as the confidence level corresponding to each candidate grasping pose of the target stack.

[0198] Candidate grasping pose selection and determination is used to select the grasping pose of the target stack from multiple candidate grasping poses based on the confidence level corresponding to each candidate grasping pose of the target stack.

[0199] In one possible design, if the stack image shows multiple stacks, then the candidate grab pose selection includes:

[0200] The first clustering module is used to perform the first clustering of multiple candidate grasping poses using the DBSCAN clustering algorithm to obtain multiple first clusters;

[0201] The first extraction module is used to determine the candidate grasping pose with the highest confidence in each first cluster as the first grasping pose of the corresponding first cluster, thereby obtaining multiple first grasping poses.

[0202] The second clustering module is used to perform a second clustering on multiple first grasping poses using the K-Means clustering algorithm to obtain multiple second clusters; wherein each of the multiple second clusters corresponds to a stack in the stack image;

[0203] The second extraction module is used to determine the first grasping pose with the highest confidence in each second cluster as the second grasping pose of the corresponding second cluster, thereby obtaining multiple second grasping poses.

[0204] The image mapping module is used to determine the grasping pose of the target stack from multiple second grasping poses.

[0205] In one possible design, the palletizing device based on multimodal embodied intelligence also includes:

[0206] The third extraction module is used to filter multiple candidate grasping poses based on the confidence level corresponding to each candidate grasping pose; wherein, the confidence level corresponding to the candidate grasping poses that are not filtered out is greater than the confidence level corresponding to the candidate grasping poses that are filtered out.

[0207] In one possible design, the palletizing device based on multimodal embodied intelligence also includes:

[0208] The stack quantity determination module is used to input the stack image into a preset quantity recognition model and obtain the stack quantity output by the quantity recognition model; where the stack quantity refers to the number of multiple stacks in the stack image;

[0209] The cluster number determination module is used to determine the number of stacks as the number of clusters for the K-Means clustering algorithm.

[0210] In one possible design, each preset shape corresponds to at least one preset stacking method for the stack, then the stacking planning module 730 includes:

[0211] The coordinate position algorithm determination module is used to determine the corresponding preset coordinate position algorithm for the target stack from multiple preset coordinate position algorithms based on the shape and stacking method of the target stack; wherein, each of the multiple preset coordinate position algorithms is applied to a combination of a preset shape and a preset stacking method;

[0212] The coordinate position calculation module is used to obtain the coordinate position of the target stack in the stack according to the size of the target stack and through a preset algorithm for the coordinate position of the target stack.

[0213] The palletizing device based on multimodal embodied intelligence provided in this application embodiment can perform... Figures 1 to 6 The technical solution of the method embodiment shown has the same implementation principle and technical effect as... Figures 1 to 6 The method embodiments shown are similar, and will not be described again in the embodiments of this application.

[0214] This application also provides an electronic device. Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device includes at least one processor 810 and a memory 820. The electronic device also includes a communication component 830. The processor 810, memory 820, and communication component 830 are connected via a bus 840.

[0215] In a specific implementation, at least one processor 810 executes computer execution instructions stored in memory 820, causing at least one processor 810 to implement the palletizing method based on multimodal embodied intelligence described in the above embodiments.

[0216] The specific implementation process of processor 810 can be found in the above method embodiments, and its implementation principle and technical effect are similar. The embodiments of this application will not be repeated here.

[0217] In the above embodiments, it should be understood that the processor 810 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0218] The memory 820 may include high-speed RAM memory, and may also include non-volatile memory NVM, such as at least one disk storage.

[0219] Bus 840 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 840 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 840 in the accompanying drawings of this application is not limited to only one bus or one type of bus.

[0220] The above description addresses the functions implemented by electronic devices and main control devices, and introduces the solutions provided in the embodiments of this application. It is understood that, in order to achieve the above functions, the electronic device or main control device includes hardware structures and / or software modules corresponding to the execution of each function. By combining the units and algorithm steps of the various examples described in the embodiments disclosed in this application, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of this application.

[0221] This application also provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, these instructions are used to implement the palletizing method based on multimodal embodied intelligence described in the above embodiments. In the specific implementation of the aforementioned palletizing method based on multimodal embodied intelligence, each module can be implemented as a processor.

[0222] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0223] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, the processor and the readable storage medium can exist as discrete components in an electronic device or a host device.

[0224] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the palletizing method based on multimodal embodied intelligence described in the above embodiments.

[0225] The computer program is stored in a readable storage medium, and at least one processor can read the computer program from the readable storage medium and execute the computer program to perform the scheme provided in any of the above embodiments.

[0226] Those skilled in the art will understand that all or part of the steps in the above-described embodiments can be implemented using hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps included in the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0227] The technical solutions of this application have been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A palletizing method based on multimodal embodied intelligence, characterized in that, The method includes: Obtain multimodal task instructions and input the multimodal task instructions into a preset multimodal model to obtain multiple preset shapes output by the multimodal model, as well as preset stacking methods for each preset shape; wherein, the multimodal task instructions refer to task instructions in text or voice form; A stack image displaying the target stack is acquired, and the stack image is input into a preset stack recognition model to obtain the shape and size of the target stack output by the stack recognition model; wherein, the stack recognition model stores multiple preset sizes corresponding to each preset shape; the shape of the target stack is one of the multiple preset shapes, and the size of the target stack is one of the multiple preset sizes corresponding to the shape of the target stack; Based on the shape, size, and stacking method of the target stack, the coordinate position of the target stack in the stack is obtained, and the target stack is stacked according to the coordinate position and the stacking method of the target stack; wherein, the stacking method of the target stack is a preset stacking method of the stack corresponding to the shape of the target stack.

2. The palletizing method based on multimodal embodied intelligence according to claim 1, characterized in that, The step of stacking the target stack according to the coordinate position and the stacking method of the target stack includes: Obtain the grasping pose of the target stack; Based on the coordinate positions, the target stack is stacked in the simulation space according to the grasping posture and stacking method of the target stack. When the simulation results indicate successful palletizing, the target stack is palletized.

3. The palletizing method based on multimodal embodied intelligence according to claim 2, characterized in that, If the image of the stack is a depth image, then obtaining the grasping pose of the target stack includes: The image of the stack is input into a preset grasping algorithm model to obtain multiple candidate grasping poses of the target stack output by the grasping algorithm model, and the confidence level corresponding to each candidate grasping pose of the target stack. Based on the confidence level corresponding to each candidate grasping pose of the target stack, the grasping pose of the target stack is obtained from multiple candidate grasping poses of the target stack.

4. The palletizing method based on multimodal embodied intelligence according to claim 3, characterized in that, The image of the stacked objects shows multiple stacked objects. The step of selecting the grasping pose of the target stacked object from the multiple candidate grasping poses based on the confidence level corresponding to each candidate grasping pose of the target stacked object includes: The DBSCAN clustering algorithm is used to perform the first clustering of multiple candidate grasping poses to obtain multiple first clusters. The candidate grasping pose with the highest confidence in each of the first clusters is determined as the first grasping pose of the corresponding first cluster, thus obtaining multiple first grasping poses. The K-Means clustering algorithm is used to perform a second clustering on multiple first grasping poses to obtain multiple second clusters; wherein each of the multiple second clusters corresponds to a stack of objects in the stack image; The first grasping pose with the highest confidence in each second cluster is determined as the second grasping pose of the corresponding second cluster, resulting in multiple second grasping poses; The grasping pose of the target stack is determined from a plurality of second grasping poses.

5. The palletizing method based on multimodal embodied intelligence according to claim 4, characterized in that, Before performing the first clustering of multiple candidate grasping poses using the DBSCAN clustering algorithm to obtain multiple first clusters, the method further includes: Based on the confidence level corresponding to each candidate grasping pose, multiple candidate grasping poses are filtered; wherein, the confidence level corresponding to the candidate grasping poses that are not filtered out is greater than the confidence level corresponding to the candidate grasping poses that are filtered out.

6. The palletizing method based on multimodal embodied intelligence according to claim 4, characterized in that, Before performing a second clustering of multiple first grasping poses using the K-Means clustering algorithm to obtain multiple second clusters, the method further includes: The stack image is input into a preset quantity recognition model to obtain the stack quantity output by the quantity recognition model; wherein, the stack quantity refers to the number of multiple stacks in the stack image; The number of stacks is determined as the number of clusters in the K-Means clustering algorithm.

7. The palletizing method based on multimodal embodied intelligence according to any one of claims 1 to 6, characterized in that, Each of the preset shapes corresponds to at least one preset stacking method of the stack. Therefore, obtaining the coordinate position of the target stack within the stack based on its shape, size, and stacking method includes: Based on the shape and stacking method of the target stack, a coordinate position preset algorithm corresponding to the target stack is determined from multiple coordinate position preset algorithms; wherein, each of the multiple coordinate position preset algorithms is applied to a combination of a preset shape and a preset stacking method; Based on the dimensions of the target stack, the coordinate position of the target stack in the stack is obtained through a preset algorithm for the coordinate position of the target stack.

8. A palletizing device based on multimodal embodied intelligence, characterized in that, The device includes: A multimodal instruction analysis module is used to acquire multimodal task instructions and input the multimodal task instructions into a preset multimodal model to obtain multiple preset shapes output by the multimodal model, as well as preset stacking methods for each preset shape; wherein, the multimodal task instructions refer to task instructions in text or voice form; The stack image analysis module is used to acquire a stack image displaying a target stack, and input the stack image into a preset stack recognition model to obtain the shape and size of the target stack output by the stack recognition model; wherein, the stack recognition model stores multiple preset sizes corresponding to each preset shape; the shape of the target stack is one of the multiple preset shapes, and the size of the target stack is one of the multiple preset sizes corresponding to the shape of the target stack; The stacking planning module is used to obtain the coordinate position of the target stack in the stack according to the shape, size and stacking method of the target stack, and stack the target stack according to the coordinate position and the stacking method of the target stack; wherein, the stacking method of the target stack is a preset stacking method of the stack corresponding to the shape of the target stack.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; When the processor executes the computer execution instructions stored in the memory, it is used to implement the palletizing method based on multimodal embodied intelligence as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the palletizing method based on multimodal embodied intelligence as described in any one of claims 1 to 7.