Robot sorting control system

By constructing a layered architecture consisting of a perception layer, a planning layer, and an execution layer, and utilizing embodied robots for flat part sorting, the problem of existing systems being unable to cope with dynamic changes is solved, achieving efficient and flexible sorting of flat parts and reducing damage.

CN121553659APending Publication Date: 2026-02-24SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511587507.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing robot sorting control systems are unable to flexibly respond to dynamically changing sorting tasks, which makes flat items prone to wrinkles, damage and dirt during logistics sorting, and requires additional manpower.

Method used

A layered architecture consisting of a perception layer, a planning layer, and an execution layer is constructed. The perception layer acquires environmental data through multiple sensors, the planning layer performs task planning, and the execution layer calls upon the upper and lower limbs and navigation strategies for motion control, thus employing an embodied robot to adapt to complex environments.

Benefits of technology

It improves the flexibility and adaptability of the robot sorting control system in complex and ever-changing sorting scenarios, reduces the risk of damage to flat parts, and lowers the manpower requirement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121553659A_ABST
    Figure CN121553659A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a robot sorting control system, and belongs to the technical field of logistics. The system comprises a sensing layer which is used for obtaining current environment sensing data of a robot in a logistics site; the planning layer is in communication connection with the sensing layer, and the planning layer is used for receiving a flat part sorting instruction, carrying out sorting task planning according to the flat part sorting instruction and the current environment sensing data to obtain a sub-task planning sequence, and determining a to-be-executed sub-task from the sub-task planning sequence; and the execution layer is in communication connection with the planning layer and the sensing layer, and the execution layer is used for calling at least one of a preset upper limb strategy, a preset lower limb strategy and a preset navigation strategy according to the to-be-executed sub-task and the current environment sensing data to perform task execution control on the upper limb and / or the lower limb of the robot. According to the embodiment of the invention, the flexibility of executing the sorting task by the robot can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of logistics technology, and in particular to a robot sorting control system. Background Technology

[0002] Flat parcels refer to express shipments containing documents, tickets, certificates, and coupons. In logistics sorting, flat parcels are usually bundled and transported together with other small items. This method easily leads to wrinkles, damage, and dirt on flat parcels. While individually packaging flat parcels throughout the entire process can solve these problems, this method requires more manpower in the sorting stage.

[0003] With the development of robotics technology, robot-based flat-piece sorting solutions have emerged. However, the robot sorting control systems in these technologies typically rely on predefined programs or fixed workstation deployments, making it difficult to flexibly respond to dynamically changing sorting tasks. Summary of the Invention

[0004] The main objective of this application is to propose a robot sorting control system that can improve the flexibility of robots in performing sorting tasks.

[0005] To achieve the above objectives, a first aspect of this application provides a robot sorting control system, the system comprising: The perception layer is used to acquire the robot's current environmental perception data in the logistics site; wherein the robot is a body-bound robot, and the robot includes upper limbs and lower limbs; The planning layer is communicatively connected to the perception layer. The planning layer is used to receive flat part sorting instructions, perform sorting task planning based on the flat part sorting instructions and the current environmental perception data, obtain a sub-task planning sequence, and determine the sub-tasks to be executed from the sub-task planning sequence. An execution layer, which is communicatively connected to the planning layer and the perception layer, is used to invoke at least one of preset upper limb strategies, lower limb strategies and navigation strategies to perform task execution control on the robot's upper limbs and / or lower limbs based on the sub-task to be executed and the current environmental perception data.

[0006] The robot sorting control system proposed in this application improves its flexibility in dealing with complex and ever-changing sorting scenarios by constructing a layered architecture of perception, planning, and execution layers. Specifically, the perception layer focuses on the real-time acquisition and fusion of environmental data, providing the system with dynamic environmental perception capabilities. The planning layer enables task parsing and sequence planning, allowing the system to understand and decompose task instructions and dynamically adjust the execution plan based on environmental feedback. The execution layer integrates multiple action strategies, enabling it to flexibly call and combine different action strategies to complete actual operations based on corresponding sub-task requirements and real-time environmental data. Furthermore, by employing an embodied robot as the execution end for sorting flat items, and utilizing the mobility of the embodied robot's lower limbs, the method proposed in this application can be adapted to existing, unmodified manual sorting sites and environmental facilities, without requiring site modifications to adapt to automated equipment. Simultaneously, by utilizing the humanoid upper limb structure of the embodied robot, the robot can adapt to the operational actions required by different sorting tasks, thus achieving near-human perceptual flexibility and operational adaptability. Attached Figure Description

[0007] Figure 1 This is a schematic diagram of a robot sorting control system provided in an embodiment of this application; Figure 2 This is a communication diagram of the robot sorting control system provided in the embodiments of this application. Detailed Implementation

[0008] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0009] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., used in the specification, claims, and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0010] First, the terms used in the embodiments of this application will be explained: Humanoid robots (also known as embodied robots) are intelligent robots that mimic the structure of the human body. They achieve efficient interaction with the physical world by integrating the coordinated operation of the upper limbs, lower limbs, and torso. These robots typically possess human-like upper limb structures, including multi-degree-of-freedom shoulder, elbow, and wrist joints, as well as dexterous hands, enabling them to perform fine tasks such as grasping, carrying, and manipulating tools. Simultaneously, the lower limbs often mimic the structure of legs and feet, equipped with joint drive and balance control systems to achieve mobility capabilities such as walking, climbing, and obstacle avoidance. Through the deep integration of sensors, processors, and mechanical structures, humanoid robots can adapt to various complex environments.

[0011] Flat parcels refer to express shipments containing documents, tickets, certificates, and coupons. In logistics sorting, flat parcels are usually bundled and transported together with other small items. This method easily leads to wrinkles, damage, and dirt on flat parcels. While individually packaging flat parcels throughout the entire process can solve these problems, this method requires more manpower in the sorting stage.

[0012] With the development of robotics technology, robot-based flat-piece sorting solutions have emerged. However, the robot sorting control systems in these technologies typically rely on predefined programs or fixed workstation deployments, making it difficult to flexibly respond to dynamically changing sorting tasks.

[0013] Based on this, embodiments of this application provide a robot sorting control system that can improve the flexibility of robots in performing sorting tasks.

[0014] Reference Figure 1 The robot sorting control system provided in this application embodiment includes a perception layer 110, a planning layer 120, and an execution layer 130. The perception layer is used to acquire current environmental perception data of the robot in the logistics area. The planning layer is communicatively connected to the perception layer. The planning layer receives flat-piece sorting instructions, performs sorting task planning based on the flat-piece sorting instructions and current environmental perception data, obtains a sub-task planning sequence, and determines the sub-tasks to be executed from the sub-task planning sequence. The execution layer is communicatively connected to the planning layer and the perception layer. The execution layer is used to invoke at least one of preset upper limb strategies, lower limb strategies, and navigation strategies to control the robot's task execution based on the sub-tasks to be executed and the current environmental perception data.

[0015] In this embodiment, the robot can refer to the execution end of the flat-piece sorting task. In this embodiment, the robot can be a humanoid robot, including upper and lower limbs. The perception layer serves as the data acquisition foundation for the robot sorting control system, used to acquire current environmental perception data of the logistics site through multiple sensors integrated within the robot body. The planning layer serves as the intelligent decision-making center of the robot sorting control system, establishing a communication connection with the perception layer to convert flat-piece sorting instructions into executable specific operation sequences. Specifically, the planning layer can be deployed on a cloud server to utilize the powerful computing resources of the cloud to run intelligent planning strategies. For example, the planning layer can adopt a Visual Language Model (VLM) architecture, which is an artificial intelligence model capable of simultaneously understanding visual scene information and natural language instructions. Thus, the planning layer can use VLM to perform deep semantic understanding and logical reasoning on the current environmental perception data provided by the perception layer and the received flat-piece sorting instructions (such as "start sorting" operation instructions), thereby realizing sorting task planning and obtaining a sub-task planning sequence. A subtask planning sequence can contain multiple subtasks with logical order and dependencies. For example, a subtask planning sequence might include: move to the temporary storage area → pick up the flat item → scan the barcode → move to the sorting cabinet → deliver the flat item. It's understandable that the "scan the barcode" subtask is used to obtain information about the flat item delivery slot.

[0016] Furthermore, the planning layer also has dynamic adjustment capabilities, which can re-plan the subsequent task sequence in real time based on environmental changes and sub-task completion status during task execution, and determine the highest priority sub-task to be executed from the current sequence, thereby improving the adaptability and robustness of the entire sorting process.

[0017] The execution layer serves as the motion control center of the robot's sorting control system, communicating with the planning and perception layers. It combines the sub-tasks to be executed from the planning layer with real-time environmental data (i.e., current environmental perception data) provided by the perception layer, converting them into specific motion control for the robot. Specifically, the execution layer can be deployed within the robot body, coordinating three control strategies (upper limb strategy, lower limb strategy, and navigation strategy) to control the robot's upper and / or lower limbs to execute corresponding sub-tasks. The upper limb strategy can employ a Visual Language Action Model (VLA), an artificial intelligence method that maps visual perception to verbal commands and translates them into specific actions. The upper limb strategy enables the robot to dynamically adjust the motion trajectories of its robotic arm and end effector based on real-time environmental perception data, thereby completing precise operations such as grasping and delivering flat objects, and reducing wrinkles or damage to objects during grasping. The lower limb strategy can be implemented using reinforcement learning (RL), a machine learning method that learns optimal decisions autonomously through interaction with the environment and based on a reward mechanism. Lower limb strategies enable robots to autonomously adjust their gait and center of gravity using real-time perceived environmental information and obstacle distribution. This allows them to maintain stable balance and avoid dynamic obstacles during movement, ensuring safe movement within logistics environments. Navigation strategies can employ Simultaneous Localization and Mapping (SLAM) technology. These strategies allow the robot to continuously update its environmental map and accurately locate itself during movement. Combined with real-time obstacle information, global path replanning is performed, guiding the robot precisely to its target location (such as sorting cabinets or temporary storage areas).

[0018] This application's embodiments improve the flexibility of the robot sorting control system in handling complex and ever-changing sorting scenarios by constructing a layered architecture of perception, planning, and execution layers. Specifically, the perception layer focuses on the real-time acquisition and fusion of environmental data, providing the system with dynamic environmental perception capabilities. The planning layer enables task parsing and sequence planning, allowing the system to understand and decompose task instructions and dynamically adjust the execution plan based on environmental feedback. The execution layer integrates multiple action strategies, enabling it to flexibly invoke and combine different action strategies to complete actual operations based on corresponding sub-task requirements and real-time environmental data.

[0019] In this embodiment, the entire robot sorting control system can be constructed using the ROS (Robot Operating System) architecture. Based on ROS's distributed node communication mechanism, modular design and decoupling of the system are achieved. In the ROS architecture, each functional layer (such as the perception layer, planning layer, and execution layer) is encapsulated as an independent node. These nodes exchange data through communication modes such as topics, services, and actions. Specifically, topic communication uses a publish-subscribe model to achieve asynchronous data stream transmission between nodes. The topic communication between different nodes is described below.

[0020] Reference Figure 2 In some embodiments, the system further includes an interaction interface, which is used for: The interactive interface retrieves input data from the object. The interactive interface publishes task topics based on the object input data, and the task topics are used to output flat part sorting instructions.

[0021] In this embodiment, an interactive interface service can be implemented based on an interactive interface. The interactive interface can refer to the entry point for external interaction and is used to receive and convert user operation instructions. Specifically, the interactive interface can be deployed in the cloud and acquire object input data through various human-computer interaction methods. Object input data can refer to control instructions initiated by the operating object or business system, including but not limited to control signals in the form of voice commands, graphical interface operation instructions, etc. For example, control instructions can include sorting instructions and movement action instructions such as "start sorting," "start changing boxes," and "move to sorting cabinet." After receiving the object input data, the interactive interface can return the receiving result and convert the object input data into a unified task topic within the system for publication. A task topic is a communication mechanism used in the system to transmit structured information. The task topic carries the flat-piece sorting instructions obtained after converting the object input data. In actual deployment, a message middleware can be used to implement a publish-subscribe mechanism for topics to ensure that flat-piece sorting instructions can be reliably transmitted to the corresponding functional level (such as the planning layer).

[0022] The interactive interface service implemented in this application based on the interactive interface improves the human-machine interaction efficiency and ease of operation of the robot sorting control system by providing an external command access mechanism.

[0023] In some embodiments, the execution layer is used to publish subtask topics, which output the execution status of the subtasks currently being executed by the robot. The planning layer receives flat part sorting instructions, performs sorting task planning based on the instructions and current environmental perception data, obtains a subtask planning sequence, and determines the subtasks to be executed. This can be achieved based on the following communication mechanism: The planning layer subscribes to task topics published by the interactive interface and obtains flat component sorting instructions based on the task topics. The planning layer subscribes to subtask topics published by the execution layer and obtains the subtask execution status based on the subtask topics. The planning layer subscribes to environmental topics published by the perception layer and obtains the robot's current environmental perception data based on these topics. The planning layer publishes planning topics based on flat part sorting instructions, subtask execution status, and current environmental perception data. The planning topics output subtask planning sequences and subtasks to be executed.

[0024] In this embodiment, a task planning service can be implemented based on the planning layer. This service, through a multi-source information subscription mechanism and dynamic planning, converts flat part sorting instructions into an executable task sequence. Specifically, the planning layer can subscribe to task topics, subtask topics, and environmental topics. Task topics are information carriers containing user operation instructions published by the interaction interface. Subtask topics are status information streams reflecting the latest completion status of subtasks published by the execution layer. Environmental topics can be comprehensive data streams containing the robot's real-time pose and surrounding environment perception data published by the perception layer. By simultaneously subscribing to these three topics, the planning layer can continuously acquire flat part sorting instructions (such as "start sorting"), subtask execution status (such as "moving," "grabbing completed," etc.), and current environmental perception data (i.e., real-time perception data reflecting the dynamic changes in the robot's environment). Based on this multi-dimensional input data, the planning layer can invoke the VLM agent to perform task parsing and planning, obtain a subtask planning sequence, and determine the highest-priority subtask to be executed from the sequence (i.e., determining the next subtask to be executed based on environmental perception data and task execution status). Finally, the planning layer can publish the planning results through planning topics. These topics can include a complete task sequence and detailed information about the currently pending subtasks (such as subtask instructions and content). This provides operational guidance for subsequent control operations in the execution layer.

[0025] This application's embodiment implements a task planning service based on the planning layer. By integrating real-time information such as instructions, environmental data, and task execution progress, the system can make intelligent decisions based on complete contextual information, generating task planning schemes that both conform to the intent of the instructions and adapt to actual environmental conditions. This approach not only ensures that sorting tasks can be executed in an orderly manner but also enables the system to dynamically adjust to environmental changes and task interruptions, thus improving the system's flexibility.

[0026] In some embodiments, the perception layer acquires the robot's current environmental perception data in the logistics site, which can be achieved based on the following communication mechanism: The perception layer acquires the robot's current LiDAR point cloud, inertial measurement data, and depth image; The perception layer publishes environmental topics based on lidar point clouds, inertial measurement data, and depth images. These environmental topics are used to output current environmental perception data, which includes 3D maps, pose data, and depth data.

[0027] In this embodiment, perception and fusion services can be implemented based on the perception layer. These services, through multi-sensor data acquisition and fusion processing, can construct a complete representation of the robot's operating environment. Specifically, the perception layer can acquire, in real time, three types of sensor data, including but not limited to: LiDAR point clouds (a large set of three-dimensional spatial coordinate points obtained by emitting laser beams and receiving reflected signals from the LiDAR; these point clouds can depict the geometric contours of the robot's surrounding environment and the spatial distribution of obstacles), inertial measurement data (object motion parameters provided by the inertial measurement unit (IMU) including three-axis acceleration and three-axis angular velocity, which can be used to track the robot's own motion state and attitude changes), and depth images (image data acquired by a depth camera containing distance information corresponding to each pixel, typically presented in RGB-D format, i.e., containing both color and depth information). After preprocessing such as synchronization timestamps, these multi-source sensor data can be integrated using a pre-defined fusion algorithm. For example, the LIO (LiDAR-Inertial Odometry) algorithm can be invoked. LIO (LiDAR Point Cloud) is a technology that tightly couples and fuses LiDAR point clouds with inertial measurement data. By jointly optimizing point cloud matching constraints and inertial measurement constraints, it can generate real-time 3D environmental maps and robot pose data. Finally, the processing results can be published through an environmental topic. The current environmental perception data output by the environmental topic can include a 3D map (a 3D environmental map generated through real-time mapping technology, identifying feasible areas, obstacle locations, and other spatial structure information), pose data (i.e., the robot's real-time position and orientation information relative to the 3D environmental map, including parameters such as coordinates and orientation), and depth data (i.e., RGB-D images, which can provide visual support for subsequent object recognition and manipulation). It is understandable that... Figure 2 The multimedia in the middle perception layer 110 can refer to image streams, audio streams, etc., obtained based on sensors, and these data can also be used as environmental perception data.

[0028] The embodiments of this application implement a perception and fusion service based on the perception layer. By collaboratively utilizing multiple source sensors, it achieves complete acquisition and fusion processing of the three-dimensional geometric structure of the environment, the robot's own motion state, and the scene's visual features. This enables the robot to accurately understand its own position in the environment and clearly perceive its spatial relationship with surrounding objects, thereby providing reliable environmental support for subsequent task execution and improving the operational reliability of the robot sorting control system in complex dynamic environments.

[0029] In some embodiments, the execution layer invokes at least one of preset upper limb strategies, lower limb strategies, and navigation strategies to perform task execution control on the robot's upper limbs and / or lower limbs based on the sub-task to be executed and the current environmental perception data. This can be achieved based on the following communication mechanism: The execution layer subscribes to environmental topics published by the perception layer and obtains current environmental perception data based on the environmental topics. The execution layer subscribes to planning topics published by the planning layer and obtains sub-tasks to be executed based on the planning topics; The execution layer subscribes to sorting plan topics and obtains target compartment information based on the sorting plan topics; The execution layer publishes control topics based on the current environmental perception data, the sub-tasks to be executed, and the target grid information. The control topics are used to output a robot action sequence determined according to at least one of the upper limb strategy, lower limb strategy, and navigation strategy. The robot action sequence is used to control the robot's upper limbs and / or lower limbs to execute the sub-tasks to be executed.

[0030] In this embodiment, decision-making and control services can be implemented based on the execution layer. These services achieve precise conversion from task instructions to robot actions by establishing a multi-source information subscription mechanism and an intelligent decision-making system. Specifically, the execution layer can adopt a topic-based publish-subscribe communication mechanism, establishing three data subscription channels, including but not limited to: environment topics, planning topics, and sorting plan topics. Environment topics and planning topics have already been described and will not be repeated here. Sorting plan topics are communication content carrying business logic data, which can be parsed to obtain target slot information. Target slot information may include parameters such as the identifier and location coordinates of the specific sorting slot to which the flat part needs to be delivered. After obtaining the above multi-source information, the execution layer can perform fusion analysis. For example, based on the type of the sub-task to be executed and the environmental state, it can invoke one or more combinations of preset upper limb strategies, lower limb strategies, and navigation strategies to obtain a decision result. Finally, the execution layer can publish the decision result (i.e., the robot action sequence) through the control topic. The robot action sequence can be a set of control instructions arranged in chronological order, detailing the motion parameters of each joint of the robot. The robot's motion sequence can be received and executed by the robot's underlying drive system, thereby driving the robot's upper and / or lower limbs to accurately complete the corresponding sub-tasks and realize robot motion control.

[0031] Understandably, after driving the robot to execute the corresponding subtask, the execution layer can also publish subtask topics to output the execution status of the current subtask.

[0032] The implementation of this application is based on the decision and control service implemented at the execution layer. By integrating real-time information on environmental status, task instructions, and business data (such as target grid information), the robot can intelligently select and combine corresponding control strategies based on complete contextual information. This effectively coordinates the fine operation of the upper limbs, the stable movement of the lower limbs, and the overall precise navigation, ensuring that it can reliably complete operations such as grasping, handling, and delivery in various complex logistics sorting scenarios.

[0033] In some embodiments, the robot sorting control system further includes a business interface, which is used for: The business interface obtains sorting data from the logistics site; The business interface publishes a sorting plan topic based on the sorting data, and the sorting plan topic is used to output target compartment information.

[0034] In this embodiment, business data services can be implemented based on a business interface. These services act as a data bridge between the robot sorting control system and the logistics business system, handling data interaction tasks related to sorting business logic. Specifically, the business interface can be deployed in the cloud, establishing a connection with the logistics business system to obtain sorting data from the logistics site. Sorting data can refer to a structured dataset containing business information such as waybill numbers, destination codes, sorting path planning, and grid allocation methods. In practice, sorting data can originate from various business systems, such as warehouse management systems (WMS), sorting scheduling systems, or order management platforms. After obtaining the sorting data, data cleaning, format conversion, and business logic processing can be performed, extracting target grid information. Then, a sorting plan topic is published based on the target grid information. This topic ensures that the execution layer's decision-making and control services can obtain accurate sorting business guidance in a timely manner, thereby guiding the robot to deliver different flat items to the correct target grids.

[0035] This application embodiment implements a business data service based on a business interface. By acquiring and processing sorting data from the logistics site in real time, it generates accurate target compartment information and publishes the target compartment information through a topic mechanism. This enables the robot to obtain the correct delivery target for each sorted item in a timely manner, ensuring the accuracy and efficiency of the sorting operation.

[0036] Reference Figure 1 In some embodiments, the planning layer triggers sorting task planning according to a first frequency, and the execution layer triggers task execution control of the robot's upper and / or lower limbs according to a second frequency, wherein the second frequency is greater than the first frequency.

[0037] In this embodiment, a differentiated triggering frequency mechanism can be used to achieve reasonable allocation of computing resources and optimization of system performance. Specifically, the sorting task planning function of the planning layer can be triggered at a low frequency, such as periodically planning and adjusting tasks at a first frequency (e.g., 5Hz, i.e., once every 200 milliseconds, with no specific limitation on the value of the first frequency). The reason for setting the planning layer to be triggered at a low frequency is that when the planning layer calls the VLM model for task planning, it needs to perform complex semantic understanding, environmental analysis, and multi-step reasoning, which involves a large number of parameter calculations. A lower execution frequency can ensure the comprehensiveness and accuracy of planning decisions and reduce unnecessary consumption of computing resources. Secondly, the sorting task planning processed by the planning layer is a high-level decision, and the decision results remain valid for a period of time, so high-frequency updates are not required.

[0038] The robot control at the execution layer can be triggered at a high frequency, such as generating and issuing upper limb and / or lower limb control commands at a second frequency (e.g., 20Hz, i.e., executing once every 50 milliseconds, with no specific limitation on the value of the second frequency). The reason for setting the execution layer to be triggered at a high frequency is based on the real-time requirements of robot motion control. The upper limb strategy (VLA) processed by the execution layer needs to quickly adjust the grasping action based on real-time visual feedback (i.e., real-time adjustment of upper limb joint space), the lower limb strategy (RL) needs to respond promptly to balance maintenance during upper limb movement or lower limb movement (i.e., real-time adjustment of lower limb joint space), and the navigation strategy also needs to correct the motion trajectory in real time based on sensor data. These control operations are sensitive to timeliness; a high execution frequency ensures that the system can respond promptly to environmental changes and achieve smooth and precise motion control.

[0039] In some embodiments, the planning layer is deployed in the cloud, while the perception and execution layers are deployed within the robot body. Specifically, the robot sorting control system of this application embodiment can adopt a layered deployment architecture that coordinates the cloud and the robot body. The planning layer, serving as the system's intelligent decision-making center, is deployed in the cloud (i.e., a remote server cluster providing scalable computing resources via a network). This deployment method fully utilizes the computing power and elastic resources of the cloud to support complex task planning and decision-making reasoning by artificial intelligence models such as the VLM model. The perception and execution layers, on the other hand, are directly deployed within the robot body (i.e., local computing units, related sensors, and actuators installed on the robot body). This deployment method ensures the real-time performance and reliability of environmental perception and motion control.

[0040] In some embodiments, the execution layer includes an upper limb control module, which is used to control the robot's upper limbs to perform task execution by invoking a preset upper limb strategy according to the corresponding sub-task to be executed. The upper limb strategy may include a hierarchical route generation strategy utilizing a visual-speech model.

[0041] In this embodiment, the upper limb control module can serve as a control module for the robot's fine-grained operations. It is used to invoke upper limb strategies to guide the robot's upper limbs in completing various sorting actions based on the sub-task to be executed (such as "picking up flat parts" or "delivering flat parts"). Specifically, the upper limb control module can be a control unit deployed within the robot body. By parsing the sub-task to be executed, it invokes a preset upper limb strategy (i.e., an intelligent decision-making method specifically for robot upper limb motion planning) to generate corresponding control commands. The upper limb strategy can employ a hierarchical route generation strategy based on a visual language model (VLA). This strategy can include upper-level routes and lower-level routes. The upper-level route can convert abstract task instructions into specific action sequences; for example, "picking up flat parts" can be decomposed into action sequences such as "locating and moving to the workbench → adjusting pose → moving the gripper." The lower-level route can further convert these action sequences into specific joint motion trajectories and force control parameters.

[0042] Understandably, the upper limb control module can make decisions at a frequency of 5Hz, performing complex reasoning based on the VLA model. Simultaneously, it operates the underlying control at a frequency of 20Hz, adjusting the motion state of each joint in the upper limb in real time. This setup allows the robot to dynamically adjust its motion parameters based on real-time visual feedback. For example, during grasping, it can fine-tune the robotic arm's trajectory based on the object's actual position and posture; during delivery, it can optimize the delivery angle based on the spatial constraints of the grid, thus ensuring stable and precise operation of flat parts of various sizes.

[0043] In some embodiments, the execution layer further includes a lower limb control module, which is used to invoke a preset lower limb strategy to perform task execution control on the robot's lower limbs according to the corresponding sub-task to be executed. The lower limb strategy includes a balance maintenance strategy and a walking obstacle avoidance strategy.

[0044] In this embodiment, the lower limb control module serves as the robot's movement and stabilization control module. It is used to control the robot's lower limbs to perform movements and maintain balance by invoking lower limb strategies based on the sub-task to be executed (such as "moving to the sorting cabinet" or "driving to the temporary storage area"). Specifically, the lower limb control module can be a control unit deployed within the robot body. It generates corresponding control commands by parsing the sub-task to be executed and invoking preset lower limb strategies (i.e., decision-making methods specifically for the robot's lower body movement and balance control). Lower limb strategies can include balance maintenance strategies and obstacle avoidance strategies. The balance maintenance strategy refers to a control method that maintains the robot's overall stability by adjusting its center of gravity and gait in real time. The obstacle avoidance strategy refers to a method that dynamically plans the movement trajectory and avoids obstacles during movement. Both strategies can be trained and optimized based on reinforcement learning (RL).

[0045] Understandably, the lower limb control module can operate at a high frequency of 20Hz to ensure timely response to environmental changes and dynamic obstacles.

[0046] The robot sorting control system provided in this application improves the overall flexibility of the system by constructing a hierarchical control architecture of perception, planning, and execution layers, combined with the modular communication mechanism of ROS. Specifically, in terms of control flow, the hierarchical design decouples environmental perception, task planning, and action execution, allowing each layer to be optimized independently (e.g., the perception layer focuses on real-time acquisition and fusion of multi-sensor data, the planning layer relies on cloud computing resources to call complex VLM models for dynamic task sequence planning, and the execution layer flexibly calls and combines upper limb strategies, lower limb strategies, and navigation strategies according to specific sub-task requirements), thereby realizing the transformation from overall task to specific actions and adaptive execution. In terms of control method, the ROS-based distributed communication framework refines the entire control flow into multiple loosely coupled functional nodes, and data interaction between modules is achieved through a standardized topic publication and subscription mechanism. This approach gives the system high configurability and scalability, enabling it to quickly adapt to different sorting process changes, site layout adjustments, and changes in task requirements. Furthermore, by employing an embodied robot as the execution end for sorting flat items, and utilizing the mobility of the robot's lower limbs, the method proposed in this application can be adapted to existing, unmodified manual sorting sites and environmental facilities, eliminating the need for site modifications to accommodate automated equipment. Simultaneously, by utilizing the humanoid upper limb structure of the embodied robot, the robot can adapt to the operational movements required for different sorting tasks, thus achieving near-human perceptual flexibility and operational adaptability.

[0047] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0048] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0049] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0050] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0051] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0052] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0053] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0054] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0055] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0056] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0057] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A robot sorting control system, characterized in that, The system includes: A perception layer is used to acquire the robot's current environmental perception data in the logistics site; wherein, the robot is a body-bound robot, and the robot includes upper limbs and lower limbs; The planning layer is communicatively connected to the perception layer. The planning layer is used to receive flat part sorting instructions, perform sorting task planning based on the flat part sorting instructions and the current environmental perception data, obtain a sub-task planning sequence, and determine the sub-tasks to be executed from the sub-task planning sequence. An execution layer, which is communicatively connected to the planning layer and the perception layer, is used to invoke at least one of preset upper limb strategies, lower limb strategies and navigation strategies to perform task execution control on the robot's upper limbs and / or lower limbs based on the sub-task to be executed and the current environmental perception data.

2. The system according to claim 1, characterized in that, The execution layer, based on the sub-task to be executed and current environmental perception data, invokes at least one of preset upper limb strategies, lower limb strategies, and navigation strategies to perform task execution control on the robot's upper limbs and / or lower limbs, including: The execution layer subscribes to the environmental topics published by the perception layer and obtains the current environmental perception data based on the environmental topics. The execution layer subscribes to the planning topics published by the planning layer and obtains the sub-tasks to be executed based on the planning topics; The execution layer subscribes to sorting plan topics and obtains target compartment information based on the sorting plan topics; The execution layer publishes control topics based on the current environmental perception data, the sub-tasks to be executed, and the target grid information. The control topics are used to output a robot action sequence determined according to at least one of the upper limb strategy, the lower limb strategy, and the navigation strategy. The robot action sequence is used to control the robot's upper limbs and / or lower limbs to execute the sub-tasks to be executed.

3. The system according to claim 2, characterized in that, The execution layer is also used to publish subtask topics, which are used to output the execution status of the subtasks currently being executed by the robot; the system also includes an interaction interface. The planning layer receives flat-piece sorting instructions, performs sorting task planning based on the instructions and current environmental perception data, obtains a sub-task planning sequence, and determines the sub-tasks to be executed, including: The planning layer subscribes to the task topics published by the interactive interface and obtains flat part sorting instructions based on the task topics. The planning layer subscribes to the subtask topics published by the execution layer and obtains the subtask execution status based on the subtask topics; The planning layer subscribes to environmental topics published by the perception layer and obtains the robot's current environmental perception data based on the environmental topics. The planning layer publishes planning topics based on the flat part sorting instructions, subtask execution status, and current environmental perception data. The planning topics output subtask planning sequences and subtasks to be executed.

4. The system according to claim 2, characterized in that, The perception layer acquires the robot's current environmental perception data in the logistics site, including: The perception layer acquires the robot's current LiDAR point cloud, inertial measurement data, and depth image; The perception layer publishes environmental topics based on the lidar point cloud, the inertial measurement data, and the depth image. The environmental topics are used to output current environmental perception data, which includes a 3D map, pose data, and depth data.

5. The system according to claim 2, characterized in that, The system also includes a business interface, which is used for: The business interface obtains sorting data from the logistics site; The business interface publishes a sorting plan topic based on the sorting data, and the sorting plan topic is used to output target compartment information.

6. The system according to claim 3, characterized in that, The interactive interface is used for: The interactive interface retrieves input data from the object. The interactive interface publishes task topics based on the input data of the object, and the task topics are used to output flat part sorting instructions.

7. The system according to any one of claims 1 to 6, characterized in that, The planning layer triggers sorting task planning based on a first frequency, and the execution layer triggers task execution control of the robot's upper and / or lower limbs based on a second frequency, wherein the second frequency is greater than the first frequency.

8. The system according to claim 7, characterized in that, The planning layer is deployed in the cloud, while the perception layer and the execution layer are deployed inside the robot itself.

9. The system according to claim 1, characterized in that, The execution layer includes an upper limb control module, which is used to call a preset upper limb strategy to control the robot's upper limbs to perform tasks according to the corresponding sub-task to be executed. The upper limb strategy includes a hierarchical route generation strategy using a visual language model.

10. The system according to claim 2, characterized in that, The execution layer also includes a lower limb control module, which is used to call a preset lower limb strategy to control the robot's lower limbs to perform tasks according to the corresponding sub-task to be executed. The lower limb strategy includes a balance maintenance strategy and a walking obstacle avoidance strategy.