Flat part sorting method and device, electronic equipment and storage medium
By using humanoid robots for task breakdown and vision-driven motion sequence control, the problems of damage and efficiency of flat parts in logistics operations have been solved, achieving efficient and low-damage sorting operations.
Patent Information
- Application Number
- CN202511587977.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-24
AI Technical Summary
In existing technologies, flat parts are prone to abnormalities such as wrinkles, punctures, tears, moisture damage, and dirt during logistics operations. Furthermore, the separate containerization and handling of all parts requires additional manpower and reduces sorting efficiency.
Humanoid robots are used for sorting flat parts. By using task breakdown and motion sequence control driven by visual observation data, the sorting of flat parts is automated.
It improves the efficiency and quality of flat part sorting, reduces labor costs, minimizes damage and operational errors, and enhances the system's flexibility and dynamic adaptability.
Smart Images

Figure CN121553660A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of logistics technology, and in particular to a method and apparatus for sorting flat parts, electronic equipment and storage medium. Background Technology
[0002] Flat mail refers to express mail containing documents, certificates, coupons, or similar items, or mail packaged in document envelopes.
[0003] Currently, in the operational model of logistics companies, flat items are usually packaged and transported together with other small items. This method easily leads to abnormalities such as wrinkles, punctures, tears, moisture damage, and dirt on flat items. To solve these problems, related technologies implement separate containerization processing for flat items at every stage. However, this method requires additional manpower, which reduces sorting efficiency to some extent. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, electronic device, and storage medium for sorting flat parts, which can improve the sorting efficiency of flat parts.
[0005] To achieve the above objectives, a first aspect of this application proposes a method for sorting flat parts, the method comprising: The flat parts sorting task is broken down into multiple sub-tasks; The task status of each subtask is determined. When the task status indicates that it is to be executed, the humanoid robot is controlled to move to the task position of the subtask to be executed, and the first visual observation data after the humanoid robot moves to the task position is obtained. A sorting action sequence is generated based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and the humanoid robot is controlled according to the sorting action sequence.
[0006] To achieve the above objectives, a second aspect of this application provides a flat parts sorting device, the device comprising: The task splitting unit is used to split the flat part sorting task into multiple sub-tasks; The observation data acquisition unit is used to determine the task status of each sub-task. When the task status indicates that it is to be executed, it controls the humanoid robot to run to the task position of the sub-task to be executed and acquires the first visual observation data after the humanoid robot runs to the task position. The robot control unit is used to generate a sorting action sequence based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and to control the humanoid robot according to the sorting action sequence.
[0007] To achieve the above objectives, a third aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in any one of the embodiments of the first aspect.
[0008] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any one of the embodiments of the first aspect.
[0009] The flat-piece sorting method, apparatus, electronic device, and storage medium proposed in this application significantly improve the automation level and overall efficiency of flat-piece sorting operations by introducing a humanoid robot as the executor and adopting a task-splitting control strategy. Specifically, based on the humanoid robot's anthropomorphic morphology and multi-degree-of-freedom operation capabilities, the humanoid robot can adapt to various sorting environments, not only reducing labor costs but also minimizing damage to flat-pieces and operational errors during sorting, thereby achieving simultaneous improvement in sorting efficiency and quality. Secondly, by breaking down the complex flat-piece sorting task into multiple sub-tasks that can be executed in parallel or sequentially, and planning targeted action sequences for the humanoid robot for different sub-tasks, task scheduling becomes more flexible and system response more agile. This reduces the complexity of single-time decisions and increases the system's adaptability to dynamic environments and its ability to handle anomalies. Attached Figure Description
[0010] Figure 1 This is a flowchart of an embodiment of the flat part sorting method provided in this application; Figure 2 This is a flowchart of an embodiment of the flat part sorting task provided in this application; Figure 3 This is a flowchart of an embodiment of the action model training method provided in this application; Figure 4 This is a schematic diagram of the system structure provided in the embodiments of this application; Figure 5 This is a flowchart of an embodiment of the data acquisition model provided in this application; Figure 6 This is a flowchart of one embodiment of the training mode provided in this application; Figure 7 This is a flowchart of one embodiment of the task execution mode provided in this application; Figure 8 This is a schematic diagram of a flat part sorting device provided in an embodiment of this application; Figure 9This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0012] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., used in the specification, claims, and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0013] Flat mail refers to express mail containing documents, certificates, coupons, or similar items, or mail packaged in document envelopes.
[0014] Currently, in logistics company operations, flat items are typically packaged and transported together with other small items. This method easily leads to abnormalities such as wrinkles, punctures, tears, moisture damage, and dirt on flat items. To address these issues, related technologies implement separate container processing for flat items throughout the entire process. However, this method requires additional manpower, which reduces sorting efficiency to some extent. Furthermore, although there are methods using robotic arms for sorting, traditional robotic arms struggle to meet the requirements of flexible production because sorting involves operations such as grasping, scanning, delivering, transporting, and exchanging boxes for flat items.
[0015] Based on this, embodiments of this application provide a flat part sorting method and apparatus, electronic device and storage medium, which can improve the sorting efficiency of flat parts.
[0016] The flat-piece sorting method provided in this application relates to the field of logistics technology. The logistics transfer method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms; the software can be an application implementing the flat-piece sorting method, but is not limited to the above forms.
[0017] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network personal computers (PCs), minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0018] Please see Figure 1 , Figure 1 This is an optional flowchart of the flat component sorting method provided in the embodiments of this application. In some embodiments of this application, Figure 1 The specific methods may include, but are not limited to, steps S110 to S130.
[0019] Step S110: The flat part sorting task is split into multiple sub-tasks; Step S120: Determine the task status of each subtask. When the task status indicates that it is to be executed, control the humanoid robot to move to the task position of the subtask to be executed, and obtain the first visual observation data after the humanoid robot moves to the task position. Step S130: Generate a sorting action sequence based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and control the humanoid robot according to the sorting action sequence.
[0020] In step S110 of some embodiments, the flat-piece sorting task can refer to the entire process task related to flat-piece sorting. It is understood that, in the embodiments of this application, the flat-piece sorting task can be completed independently by at least one humanoid robot. Figure 2 As shown, taking a single humanoid robot as an example, the flat part sorting task may include: 1. Control the humanoid robot to move to the sorting cabinet. The sorting cabinet can be a storage device with multiple compartments, each compartment holding one bin. The bin is a container for storing flat parts. Different compartments can correspond to different material flow directions. Sensors and other equipment are used to detect whether each compartment has a bin. When a compartment without a bin is detected, control the humanoid robot to move to the empty bin acquisition point to obtain an empty bin. After obtaining an empty bin, control the humanoid robot to move to the empty bin position in the sorting cabinet (i.e., the compartment without a bin) and place the obtained empty bin into that empty bin position.
[0021] 2. When all compartments are filled with material bins, control the humanoid robot to move to the supply area (the supply area can be a centralized storage area for flat parts to be sorted, which may include flat parts destined for different flows). Use sensors and other equipment to determine if there are any unprocessed flat parts in the supply area. Once it is determined that all flat parts have been processed, control the humanoid robot to move to the preset waiting area to stand by.
[0022] 3. When it is determined that there are unprocessed flat parts, control the humanoid robot to grab the flat parts and use scanning equipment to obtain the surface information of the flat parts (such as QR codes or barcodes) to obtain the target flow direction of the grabbed flat parts.
[0023] 4. Based on the target flow of the flat piece being grabbed, control the humanoid robot to move to the corresponding slot in the sorting cabinet, and control the humanoid robot to deliver the flat piece being grabbed into the material box in that slot.
[0024] 5. After delivery, sensors and other equipment determine if the bin is full. When a bin is full, the humanoid robot moves to the empty bin retrieval area to obtain an empty bin. After obtaining an empty bin, the humanoid robot moves to the full bin position in the sorting cabinet (i.e., the slot where the full bin is located) to perform bin exchange (i.e., remove the full bin and place the empty bin in the full bin position). Alternatively, the robot can be controlled to transport the full bin to a corresponding area for temporary storage. This allows other automated equipment (such as AGVs) to transport the full bin from the temporary storage area to the shipping area, thus achieving seamless integration of the sorting and shipping processes.
[0025] 6. When the bin is not full, control the humanoid robot to move back to the feeding area to continue processing other flat parts to be sorted.
[0026] In this way, the flat part sorting task can be broken down into multiple sub-tasks based on timing logic or function. For example, sub-tasks such as flat part delivery task and bin replacement task can be obtained, and different sub-tasks can be performed by the same or different humanoid robots.
[0027] In step S120 of some embodiments, the task status of each subtask is obtained. The task status can refer to the stage of the subtask in its lifecycle, such as pending execution, completed, interrupted, etc. Taking the flat part delivery task as an example, the task status of the subtask can be determined according to the number of flat parts in the supply area. If the number of flat parts is greater than a preset threshold (such as 0), the task status of the flat part delivery task is determined to be pending execution. For subtasks with a pending execution status, the task location corresponding to the subtask can be obtained. The task location can be a physical space identifier such as the coordinates of the sorting cabinet slot, the shelf location, or the work area. In this way, based on technologies such as laser SLAM (simultaneous localization and mapping), the prior map of the sorting environment and real-time sensor data can be fused to control the humanoid robot executing the corresponding subtask to run to the task location. After the humanoid robot runs to the task location, the first visual observation data can be obtained according to the vision sensor on the humanoid robot. The first visual observation data is used to characterize the visual state of the target object (such as a bin or flat part) and the surrounding environment from the current subtask perspective. For example, the first visual observation data can include scene color images, depth point clouds, three-dimensional structural data, etc.
[0028] In step S130 of some embodiments, the task instruction can refer to natural language or structured commands (which can be in text, voice, or other formats). The task instruction describes the target corresponding to the sub-task to be executed. For example, the task instruction can be "go to change the box", "go to deliver the flat part", "deliver the flat part to the XX compartment", etc. In this way, the sorting action sequence can be determined based on the first visual observation data and the task instruction. The sorting action sequence can refer to a set of actions that need to be performed from the current state of the humanoid robot to enable it to complete the corresponding sub-task. For example, it can include actions such as robotic arm movement, end effector movement (such as the movement trajectory of the gripping end used to grip the flat part), and navigation movement.
[0029] In some embodiments, a sorting action sequence can be generated based on a pre-trained action model. For example... Figure 3 As shown, the training method for the action model may include, but is not limited to, steps S310 to S320.
[0030] Step S310: Obtain the position of the first sample task of any sample task in the sample task set, and obtain the first sample visual observation data and the first sample joint state after the humanoid robot runs to the first sample task position. Call the action model to generate the first predicted action sequence based on the first sample visual observation data, the first sample joint state and the task instructions corresponding to the sample task. Step S320: Obtain the preset teaching trajectory sequence corresponding to the sample task, and train the action model based on the preset teaching trajectory sequence and the first predicted action sequence.
[0031] In step S310 of some embodiments, the sample task set may refer to a collection of multiple sub-task examples pre-built for training the action model. The sample task set can be constructed in a simulation environment. The sample task set may contain multiple sample tasks, each corresponding to a specific task objective and environmental configuration, thus allowing the humanoid robot to acquire different first sample visual observation data. The first sample visual observation data may refer to the information collected after controlling the humanoid robot to run to the first sample task position corresponding to the sample task. Specifically, the first sample visual observation data is used to characterize the visual state of the sample object (such as a bin, flat part) and its surrounding environment. For example, the first sample visual observation data may include scene color images, depth point clouds, three-dimensional structural data, etc. The first sample task position may refer to the operational position that the humanoid robot needs to reach in the sample task. The first sample joint state may refer to the state of the torque, rotation, and other parameters of all motion joints after the humanoid robot runs to the first sample task position.
[0032] After acquiring the first sample visual observation data and the first sample joint state, the action model to be trained (which can be a Vision-Language-Motion Model (VLA) or a multimodal neural network model, etc.) can be invoked. The first sample visual observation data, the first sample joint state, and the task instructions corresponding to the sample task are used as the input data for the action model. Based on the input data, the action model can infer the actions required for the humanoid robot to perform the sample task, and obtain the first predicted action sequence. The first predicted action sequence may include the humanoid robot's joint angles, the trajectory of the end effector, etc.
[0033] In step S320 of some embodiments, the motion model can be trained based on the difference between the first predicted motion sequence and the preset teaching trajectory sequence. The teaching trajectory sequence can refer to a standard motion sequence for completing a corresponding sample task, demonstrated by an expert or recorded through high-precision manual operation. For example, a humanoid robot can be guided to record joint space or task space paths using a teaching pendant, or the teaching trajectory sequence can be obtained in a simulation environment using demonstration recording tools. The data structure of the teaching trajectory sequence is consistent with the first predicted motion sequence to ensure comparability.
[0034] This application embodiment collects multimodal information (visual observation data and joint state) after the robot reaches the task position and combines it with task instructions to drive the motion model to generate a first predicted motion sequence. Then, the teaching trajectory is used as a supervision signal to train the model, thereby effectively realizing the automated learning and optimization of humanoid robot sorting skills.
[0035] In some embodiments, step S320 may include, but is not limited to, the following steps: Determine the action similarity between the preset teaching trajectory sequence and the first predicted action sequence, and pre-train the action model based on the action similarity; Obtain the position of the second sample task for any sample task in the sample task set, and obtain the second sample visual observation data and the second sample joint state after the humanoid robot runs to the second sample task position. Call the pre-trained action model to generate the second predicted action sequence based on the second sample visual observation data, the second sample joint state and the task instructions corresponding to the sample task. The reward score of the second predicted action sequence is determined according to the preset reward function, and the parameters of the action model are adjusted according to the reward score.
[0036] In this embodiment, the difference between the taught trajectory sequence and the first predicted action sequence can be quantified using action similarity. For example, when the taught trajectory sequence and the first predicted action sequence are continuous actions, the mean squared error can be used to directly compare the action differences at corresponding time steps, or the Dynamic Time Warping (DTW) algorithm can be used to align and calculate the overall difference between the two sequences. When the taught trajectory sequence and the first predicted action sequence are discrete actions, cross-entropy loss can be used to measure the similarity in action classification. After determining the action similarity, the model parameters of the action model are optimized and updated based on the action similarity measurement results and the backpropagation algorithm to achieve pre-training of the action model. In this way, a high-performance initial strategy can be provided for subsequent fine-tuning training, reducing the exploration time required for subsequent fine-tuning and improving the stability of training.
[0037] After pre-training the action model, it can be fine-tuned based on other sample tasks in the sample task set, such as using reinforcement learning, to further optimize the model. Specifically, firstly, similar to the steps described above, the position of a second sample task (which may be the same as or different from the sample task used to determine the first predicted action sequence) is obtained from any sample task in the sample task set. Then, the visual observation data and joint states of the second sample task after the humanoid robot reaches the second sample task position are obtained. In this way, the pre-trained action model can be invoked, using the second sample visual observation data, the second sample joint states, and the task instructions corresponding to the current sample task as input data to obtain the second predicted action sequence inferred by the action model.
[0038] Then, the performance of the motion model can be further fine-tuned and improved using the feedback reward score. Specifically, the reward function can be a pre-set function used to quantify the effect of task execution. The reward function maps the state changes of the humanoid robot after executing an action sequence (such as a second predicted action sequence) to a specific reward score. For example, in a flat object delivery task, the reward function can give a positive reward for the humanoid robot successfully grasping the flat object, a negative reward for a collision, and a small negative reward for a long grasping time. After the humanoid robot executes the relevant actions based on the second predicted action sequence, the environment (such as a simulator) can calculate a cumulative reward value (i.e., a reward score) based on the action result and the reward function. Subsequently, reinforcement learning algorithms can be used to adjust the parameters of the motion model based on this reward score.
[0039] Unlike the pre-training phase, which minimizes the difference between the training and demonstration trajectories, the fine-tuning phase aims to maximize the expected cumulative reward. The algorithm adjusts the model's policy distribution by calculating the gradient of the reward score with respect to the model parameters and updating these parameters along the gradient direction. This makes the action model more likely to generate action sequences that yield high rewards in the future, rather than simply mimicking the training trajectories. Through numerous such iterative interactions, the action model can continuously explore and discover behaviors that are superior or more robust than the pre-trained policy. In other words, it can optimize its decision-making policy (i.e., infer action sequences) based on the actual performance of the task, thereby acquiring the ability to handle unseen scenarios.
[0040] The following sections explain the specific control strategies for the humanoid robot when the subtasks are flat part delivery and hopper replacement.
[0041] First, let's explain the flat-piece delivery task.
[0042] In some embodiments, when the subtask to be executed is a flat part delivery task, the task position can refer to the target delivery slot. Step S120, "controlling the humanoid robot to move to the task position of the subtask to be executed and acquiring the first visual observation data after the humanoid robot moves to the task position," may include, but is not limited to, the following steps: Control the humanoid robot to move to the preset flat part temporary storage area to grasp flat parts; After the flat part is grasped, a movement trajectory sequence is generated based on the second visual observation data of the humanoid robot, the preset scanning instructions and the state of the first joint. The first gripping end of the humanoid robot is then controlled to move the grasped flat part to the scanning area of the scanning device for scanning operation according to the movement trajectory sequence. The target delivery slot for the flat part is obtained based on the scanning operation. The humanoid robot is controlled to move to the target delivery slot, and the first visual observation data after the humanoid robot moves to the target delivery slot is obtained.
[0043] In this embodiment, the flat parts storage area is the aforementioned supply area (i.e., the physical work area for centrally storing flat parts to be sorted). The flat parts storage area can be located at a specific workbench, the end of a conveyor belt, or a specific shelf level, etc., without specific limitations. The optimal path for a humanoid robot to reach the flat parts storage area is planned using laser SLAM navigation technology, and the humanoid robot is controlled to move to the storage area based on the planned path. After arriving at the flat parts storage area, the humanoid robot can be controlled to perform flat parts grasping operations. For example, the end of the humanoid robot's robotic arm is equipped with a gripper, and one side of the gripper can be equipped with a suction cup. Thus, the gripping of flat parts can be achieved based on this suction cup.
[0044] After the humanoid robot grasps the flat object, its second visual observation data can be acquired. This second visual observation data refers to a new round of environmental images and depth information collected by the humanoid robot's vision system (such as a vision sensor) after grasping. Based on this second visual observation data, the relative pose between the humanoid robot and the scanning device and its corresponding scanning area can be determined. The preset scanning command can refer to a predefined, structured command or natural language instruction (text or spoken language format) that requires the execution of a scanning operation. The content of the preset scanning command can include the device identifier of the scanning device (e.g., scanning device A), the scanning action type (e.g., barcode recognition), or a simplified "de-scan" command. The first joint state refers to the state of the torque, rotation, and other parameters of each joint after the humanoid robot has grasped the flat object. These parameters collectively define the overall posture of the humanoid robot and the current position of the first gripping end (i.e., the end of the robotic arm). Thus, the second visual observation data, the preset scanning command, and the first joint state can be used as input data for the previously trained motion model (or other methods, such as other network structures), to obtain a sequence of movement trajectories. A movement trajectory sequence can refer to an optimal robotic arm movement trajectory calculated from the current holding position to the scanning area, using the first joint state as the starting point, the scanning area perceived by the second visual observation data as the ending point, and a preset scanning command as the task objective. The movement trajectory sequence can include a series of path points and their corresponding joint states or end-effector poses. Thus, the first gripping end of the humanoid robot can be controlled to move carrying the flat piece based on the movement trajectory sequence, ensuring that it eventually enters the effective scanning area of the scanning device, thereby triggering the scanning device (such as a scanner fixed on a workbench or a vision scanning module integrated into the humanoid robot) to successfully read the delivery information of the flat piece. It is understood that if the delivery information is not recognized on one side of the flat piece, the first gripping end can be controlled to flip the flat piece for re-recognition. It is also understood that when using the motion model to determine the movement trajectory sequence, the motion model can be trained based on a sample task set before step S120.
[0045] The delivery information can include the flow direction of the flat piece, allowing the target delivery slot to be determined based on this flow direction. The target delivery slot can refer to the slot on the sorting cabinet corresponding to the flow direction of the flat piece. Navigation technology is used to control the humanoid robot to move to the target delivery slot. After the humanoid robot arrives at the target delivery slot, its first visual observation data can be acquired. This data reflects the state of the target delivery slot and the material bins within it, providing data support for subsequent accurate delivery. It is understandable that a motion model can be used to determine the movement trajectory of the robotic arm's end effector (such as the first gripper end or a robotic arm with a dexterous hand) based on the first visual observation data, the task instruction for the flat piece delivery task (such as "deliver"), and the current joint state of the humanoid robot. This allows control of the robotic arm's end effector to move based on the determined trajectory, thereby completing the flat piece delivery task. It is also understandable that when using a robotic arm with a dexterous hand for flat piece delivery, a motion model can be used to generate a sequence of actions describing the first gripper end transferring the flat piece to the dexterous hand; this will not be elaborated further.
[0046] The embodiments of this application realize the fully automated operation of humanoid robots from grasping and scanning to precise delivery, improving the intelligence level and work efficiency of flat part sorting.
[0047] In some embodiments, "controlling the humanoid robot to move to the preset flat part temporary storage area to grasp the flat part" may include, but is not limited to, the following steps: Control the humanoid robot to move to the preset flat component temporary storage area, obtain the first pose data of the temporary storage device in the preset flat component temporary storage area relative to the humanoid robot, and adjust the pose of the humanoid robot according to the first pose data; The system acquires the third visual observation data of the humanoid robot after pose adjustment and the state of the second joint of the humanoid robot. Based on the third visual observation data, the state of the second joint and the preset grasping instructions, it generates a first end effector sequence and controls the first gripping end of the humanoid robot to move to the temporary storage device to grasp flat parts according to the first end effector sequence.
[0048] In this embodiment, the temporary storage device can refer to a flat device used to support or accommodate it, such as a file basket, material bin, or tray. After the humanoid robot reaches the flat storage area, it can identify and locate the temporary storage device. Specifically, the first pose data of the temporary storage device relative to the humanoid robot can be calculated by applying a deep learning-based object detection algorithm or a point cloud registration algorithm. The first pose data describes the position and orientation of the temporary storage device in three-dimensional space (with reference to the robot coordinate system). Thus, the humanoid robot's pose can be adjusted based on the first pose data. The pose adjustment can be achieved through the movement of the humanoid robot's overall base and collaborative changes in upper body posture. Through pose adjustment, the humanoid robot can eventually stop in the optimal operating pose relative to the temporary storage device, such as having the humanoid robot face the temporary storage device directly at a suitable distance, laying the foundation for subsequent precise grasping. It is understood that the humanoid robot can undergo gradual and multiple pose adjustments to achieve the optimal operating pose, and no specific limitation is made thereto.
[0049] After pose adjustment, third-vision observation data can be acquired based on the humanoid robot's vision system. This data provides detailed visual information about the interior of the temporary storage device and the flat components it contains, including their stacking state, surface texture, and precise 3D geometric information. Simultaneously, the second joint state of the humanoid robot can be obtained. Then, a trained motion model can be invoked, using the third-vision observation data, second joint state, and preset grasping instructions (such as "grab the top-level file") as input data to obtain a first end-effector motion sequence (or another method, such as using a different network structure, can be employed to obtain the first end-effector motion sequence). This first end-effector motion sequence can be a planned trajectory for the robotic arm's end effector (such as the first gripper end) to move from its current position to the target gripping point and complete the grasping process. Thus, based on the first end-effector motion sequence, the first gripper end can be controlled to enter the temporary storage device to grasp the flat component.
[0050] Understandably, the robotic arm at the first gripping end can be controlled to rise vertically, and the dexterous hand can be controlled to move to a fixed position to grasp the flat parts, so as to remove the excess flat parts through the dexterous hand.
[0051] Secondly, the task of changing the material bin is explained.
[0052] In some embodiments, when the subtask to be executed is a bin replacement task, the task location may refer to the target bin compartment, and the task instructions corresponding to the subtask to be executed may include replacement instructions and placement instructions. Step S130 may include, but is not limited to, the following steps: The state of the third joint of the humanoid robot is obtained. Based on the state of the third joint, the first visual observation data and the replacement command, a second end effector sequence is generated. Based on the second end effector sequence, the first gripping end of the humanoid robot is controlled to move to the target material changing grid to grip the corresponding material box. After completing the grabbing of the material box, control the humanoid robot to move to the preset placement position to place the grabbed material box; After the material bins are placed, the humanoid robot is controlled to move to the preset material picking area to grab a new material bin, and then the humanoid robot is controlled to move to the target material changing grid. The third end effector sequence is generated based on the fourth visual observation data of the humanoid robot at the target material changing grid, the placement command, and the fourth joint state. The first gripping end of the humanoid robot is then controlled to place the new material box into the target material changing grid according to the third end effector sequence.
[0053] In this embodiment, the third joint state refers to the state of the torque, rotation, and other parameters of each joint after the humanoid robot moves to the target material bin slot. The target material bin slot refers to a slot where the placed material bin (such as the target material bin) is full. The first visual observation data can be a color image or depth point cloud data of the area where the target material bin slot is located, collected by the humanoid robot's vision system. The replacement command can refer to a predefined control command, such as "replace material bin". Thus, a trained motion model can be called, and the third joint state, the first visual observation data, and the replacement command can be used as input data for the motion model to obtain a second end effector sequence. The second end effector sequence can refer to a motion trajectory planned for the end effector of the robotic arm (such as the first gripping end) from its current position to the target gripping point and complete the material bin gripping. Thus, the first gripping end can be controlled to enter the target material bin slot to grip the material bin based on the second end effector sequence.
[0054] After successfully grabbing the bin, the humanoid robot can be controlled to move to a preset placement location and place the grabbed bin. It's understood that the preset placement location can be a pre-defined temporary storage area, conveyor belt interface, or temporary storage rack, etc. Similarly, the humanoid robot can be controlled to perform bin placement operations based on a motion model, which will not be elaborated further.
[0055] After placing a full bin, a humanoid robot needs to be controlled to retrieve a new bin for replenishment. Specifically, the humanoid robot can be controlled to move to a preset retrieval area to grab a new bin and then return to the target material exchange slot carrying the new bin. The preset retrieval area can refer to a supply area for storing empty bins, such as a specific shelf, an automatic replenishment station, or a material cart. Similarly, the humanoid robot can be controlled to grab a new bin and return based on a motion model, which will not be elaborated further.
[0056] After the humanoid robot returns to the target material changing slot, its fourth visual observation data and fourth joint state can be acquired. The fourth visual observation data can be color images and depth point cloud data of the area where the target material changing slot is located, collected through the humanoid robot's vision system. The fourth joint state refers to the state of the torque, rotation, and other parameters of each joint after the humanoid robot returns to the target material changing slot. The fourth joint state reflects the posture of the humanoid robot while carrying the material box. The placement command can refer to predefined control commands, such as "place material box" or "place empty material box to slot XX," etc. Thus, the motion model can be invoked, using the fourth visual data, placement command, and fourth joint state as input data to obtain the third end-effector motion sequence. The third end-effector motion sequence can refer to a planned trajectory for the robotic arm end effector (such as the first gripping end) to move from its current position into the target material changing slot to place the new material box. For example, the other side of the gripper is flat to ensure the material box is placed stably on the first gripping end. The third end action sequence can refer to the set of actions that describe the first gripping end carrying the material box and gradually pushing the material box into the target material changing grid.
[0057] It is understood that the sorting action sequence described above may include a second end action sequence and a third end action sequence.
[0058] The embodiments of this application realize the automated replacement of sorting bins by humanoid robots, improve the continuity and intelligence of the flat parts sorting process, and thus improve work efficiency.
[0059] In some embodiments, a second end effector sequence is generated based on the third joint state, first visual observation data, and replacement instructions. The first gripping end of the humanoid robot is then controlled to move to the target material exchange slot to grip the corresponding material box based on the second end effector sequence. This may include, but is not limited to, the following steps: The second end effector sequence is generated based on the state of the third joint, the first visual observation data, and the replacement command. The first gripping end of the humanoid robot is then moved to the bottom of the corresponding material box in the target material changing compartment based on the second end effector sequence. Acquire the second pose data of the corresponding material box relative to the second gripping end of the humanoid robot, and control the second gripping end to move to the top of the corresponding material box based on the second pose data; According to the preset pull teaching trajectory, the second gripping end is controlled to pull the corresponding material box away from the target material changing grid, and the first gripping end is controlled to move synchronously so that the first gripping end can grip the material box pulled out of the target material changing grid.
[0060] In this embodiment, after generating the second end-effector action sequence, the first gripping end (such as a clamping plane) can be controlled to move to the bottom of the target material box in the target material exchange port according to the second end-effector action sequence. It is understood that the bottom of each compartment in the sorting cabinet can be partially hollowed out to facilitate the first gripping end to perform related operations while the material box is placed.
[0061] After the first gripping end is in place, the second pose data can be acquired. The second pose data refers to the spatial position and orientation information of the target bin relative to the second gripping end (with reference to the humanoid robot coordinate system). Based on the second pose data, the pose of the second gripping end can be adjusted to a standard pose, laying the foundation for subsequent pulling control. Furthermore, trajectory planning can be performed using motion models or other methods. This allows the second gripping end (usually configured as an actuator with gripping or hooking functions, such as a dexterous hand or hook-like gripper) to move to the top of the target bin according to the planned trajectory. At this point, the first and second gripping ends grip the target bin from two directions, thus enabling the target bin to be removed from the target material exchange slot through coordinated control of the first and second gripping ends. Specifically, the pulling teaching trajectory can refer to a standard action sequence, demonstrated by an expert or recorded through high-precision manual operation, that completes the bin pulling action starting from a standard pose. The pulling teaching trajectory can include parameters such as the pose and velocity of the second gripping end at various time points during the pulling process. In this way, the second gripping end can be controlled to pull the target material box away from the target material changing grid (i.e., the exit direction of the target material changing grid) according to the pull teaching trajectory. At the same time, the first gripping end and the second gripping end are controlled to move in coordination, so that after the second gripping end pulls the target material box outward, the first gripping end is still placed at the bottom of the target material box, which facilitates the subsequent operation of carrying the material box.
[0062] It is understandable that during the process of placing the target bin into the preset position and grabbing a new bin, the first and second gripping ends can be controlled to operate in coordination, and no specific limitations are made therein. In addition, similar to the control operations described above, sub-tasks may also include placing empty bins in the empty bin positions of the sorting cabinet, etc., which will not be elaborated further.
[0063] The flat-piece sorting method provided in this application significantly improves the automation level and overall efficiency of flat-piece sorting operations by introducing a humanoid robot as the executor and adopting a task-splitting control strategy, meeting the flexible and precise sorting needs of flat-pieces in logistics operations. Specifically, based on the humanoid robot's anthropomorphic morphology and multi-degree-of-freedom operation capabilities, the humanoid robot can adapt to various sorting environments, replacing manual labor in repetitive grasping, handling, delivery, and box-changing operations. This not only reduces labor costs but also minimizes damage to flat-pieces and operational errors during sorting, thereby achieving simultaneous improvement in sorting efficiency and quality. Secondly, by breaking down the complex flat-piece sorting task into multiple sub-tasks that can be executed in parallel or sequentially, and planning targeted action sequences for the humanoid robot for different sub-tasks, task scheduling becomes more flexible and system response more agile. This reduces the complexity of single-time decisions and increases the system's adaptability to dynamic environments and its ability to handle anomalies. Furthermore, by integrating multimodal perception data, teaching trajectories, and reinforcement learning optimization strategies, the action model of this application embodiment possesses autonomous learning and generalization capabilities. This enables the humanoid robot to perform relevant tasks normally when facing bins of different sizes, changing environmental conditions, and temporary obstacles.
[0064] In one specific embodiment, the flat component sorting method of this application can be based on, as follows: Figure 4 The system architecture shown is implemented below. The following sections describe each module within the system architecture.
[0065] 1. Humanoid robot.
[0066] The humanoid robot is the terminal that performs the flat-piece sorting task in the embodiments of this application. All algorithms and strategies are ultimately translated into the actual actions of the humanoid robot to enable it to perform tasks such as flat-piece delivery and box swapping in actual logistics scenarios. The humanoid robot can be equipped with necessary sensors (such as cameras, LiDAR, etc.) and actuators (such as gripping ends, etc.).
[0067] Robot teleoperation can refer to the demonstration by experts or the simulation of a robot's actions during actual tasks through high-precision manual operation. Teaching trajectories can refer to the optimal sequence of actions recorded through teleoperation to complete different tasks.
[0068] 2. On-body control and intelligent services.
[0069] The body control and intelligent service module is used to receive policy instructions and convert them into low-level control signals, while managing the basic state and services of the humanoid robot.
[0070] The humanoid robot control includes: Control Mode: Defines the operating mode of the humanoid robot, which may include remote operation, recording mode, calibration mode, etc. Remote operation refers to the operating mode of remotely controlling the humanoid robot to perform tasks. Recording mode refers to the operating mode of collecting and storing multimodal data of the humanoid robot during task execution. Calibration mode refers to the operating mode of performing initial calibration on the humanoid robot.
[0071] Control interface: Provides an interface for controlling the humanoid robot, which may include motion control (sending specific motion commands), motion acquisition (acquiring motion feedback from the humanoid robot), and perception control (managing sensor data streams), etc.
[0072] Policy control includes: Interaction planning: This is used for high-level decision-making and semantic understanding of the humanoid robot's ontology. For example, a Vision-Language Model (VLM) can be used to understand images and instructions, while a Large Language Model (LLM) can be used to understand task context and perform logical reasoning.
[0073] Action policy: This is responsible for generating specific actions. For example, VLA can output action instructions (i.e., action models) based on vision and task commands. RL (Reinforcement Learning) strategies optimize decisions through continuous trial and error.
[0074] Navigation strategy: Used for movement planning of humanoid robots. SLAM is used for real-time localization and mapping, while VLN (Visual-Language Navigation) enables humanoid robots to navigate based on natural language commands.
[0075] 3. Data and simulation services.
[0076] The data and simulation service module serves as the "infrastructure" supporting system development and training. This module can interact with the "body control and intelligence service" for data acquisition, policy updates, and policy verification. Data acquisition refers to collecting data (such as observation data and joint states) from actual operations and teleoperations of the humanoid robot. Policy updates involve using the collected data to continuously train and update policy models such as VLA through methods like imitation learning (based on taught trajectories) and reinforcement learning (optimization based on rewards), thereby improving model performance. Policy verification involves testing and validating the updated policy in a simulation environment before deploying it in a real humanoid robot, ensuring the policy's safety and effectiveness.
[0077] The data and simulation service module may include: Dataset management: This feature enables data processing, data visualization, data annotation, and fragment management. Data processing, visualization, and annotation refer to cleaning, organizing, visualizing, and labeling the collected raw data to transform it into a high-quality dataset suitable for training. Fragment management refers to managing key segments of the data (such as a successful grabbing action) for easy reuse and analysis.
[0078] Simulation environment: Simulation platforms such as Nvidia ISAAC Sim and MuJoCo can be set up. Nvidia ISAAC Sim provides realistic visual rendering and physical simulation, suitable for training vision-based VLA models. MuJoCo can be used to train and validate motion control strategies.
[0079] Furthermore, based on the above system architecture, the following robot operation modes can be implemented. The relationships between the modules in the system structure under different operation modes are illustrated below in flowchart form.
[0080] First, the data acquisition mode will be explained. For example... Figure 5 As shown, the system first determines whether to enable the simulation environment. If not, it connects to the humanoid robot or teleoperation device via the control interface. If so, it enters the simulation environment module to perform scene environment setup and model import operations. Then, the system enters control mode and performs humanoid robot initialization calibration to ensure the accuracy of data acquisition. After calibration, the system switches to remote operation control mode, where it can control the humanoid robot's movement and continuously acquire and record its multimodal data (such as observation data and joint states). This data acquisition mode enables the collection of high-quality task execution data in real or virtual environments through teleoperation or robot movement control, providing data support for imitation learning and strategy training, or for actual task execution.
[0081] Secondly, the training mode will be explained. For example... Figure 6As shown, the system first determines whether to enable the simulation environment. If not, it connects to the humanoid robot via the control interface; if so, it enters the simulation environment module to perform scene environment setup and model import. Then, the system enters control mode and performs initial calibration of the humanoid robot to ensure the accuracy of data acquisition. After calibration, the system enters a policy update loop: first, it acquires the humanoid robot's observation state via the control interface; then, the policy system generates or updates the control policy based on current observations, historical data, and reward scores (e.g., adjusting model parameters through reinforcement learning algorithms) to continuously adapt the policy to the environment and improve decision-making performance. The updated policy is converted into specific actions (e.g., action sequences) via the control interface and executed. The reward signal from the environment after execution is scored by the policy system to quantitatively evaluate the policy's merits and guide optimization, thus forming a closed-loop learning process of "perception-decision-execution-evaluation," driving the policy to continuously evolve iteratively, ultimately achieving intelligent behavior capable of efficiently and autonomously completing complex tasks.
[0082] Finally, the task execution mode is explained. For example... Figure 7 As shown, firstly, the system determines whether to enable the simulation environment. If not, it connects to the humanoid robot body through the control interface; if so, it enters the simulation environment module to perform operations such as scene environment construction and model import. Then, the system enters control mode and performs humanoid robot initialization calibration to ensure the accuracy of data acquisition. After calibration, the system enters a control loop: first, it acquires real-time observation data of the humanoid robot through the control interface (joint status and other information can also be acquired); then, the strategy system performs overall task planning (such as determining the sub-task sequence) based on the current observation data and task objectives, thus decomposing abstract instructions into specific executable steps. Further, the strategy system calls the corresponding sub-task strategy (such as the action model corresponding to flat part delivery and bin replacement) based on the planning results to generate a specific action sequence, thus converting decisions into low-level executable instructions. Finally, the generated action sequence is sent to the humanoid robot for execution through the control interface, thereby completing the corresponding sub-task. Therefore, this embodiment can form a closed-loop automatic control process of "perception-planning-decision-execution," enabling the humanoid robot to intelligently complete the flat part sorting task.
[0083] Reference Figure 8 This application also provides a flat part sorting device, the device 800 comprising: The task splitting unit 810 is used to split the flat part sorting task into multiple sub-tasks; The observation data acquisition unit 820 is used to determine the task status of each sub-task. When the task status indicates that it is to be executed, it controls the humanoid robot to run to the task position of the sub-task to be executed and acquires the first visual observation data after the humanoid robot runs to the task position. The robot control unit 830 is used to generate a sorting action sequence based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and to control the humanoid robot according to the sorting action sequence.
[0084] It should be noted that the flat part sorting device provided in this application embodiment is used to implement the flat part sorting method provided in the above embodiment, and the specific implementation process corresponds to the flat part sorting method in the above embodiment. It can be referred to the aforementioned flat part sorting method, and will not be repeated here.
[0085] This application also provides an electronic device (i.e., a computer device), which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it can implement any of the flat component sorting methods described in the above embodiments. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0086] Reference Figure 9 , Figure 9 This illustration shows the hardware structure of an electronic device according to another embodiment, the electronic device comprising: The processor 910 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 920 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 920 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 920 and called and executed by the processor 910 using the flat component sorting method of the embodiments of this application. The input / output interface 930 is used to implement information input and output; The communication interface 940 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 950 transmits information between various components of the device (e.g., processor 910, memory 920, input / output interface 930, and communication interface 940); The processor 910, memory 920, input / output interface 930 and communication interface 940 are connected to each other within the device via bus 950.
[0087] This application also provides a computer-readable storage medium storing a computer program for causing a computer to execute the flat part sorting method described in the above embodiments.
[0088] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0089] This invention also provides a computer program product that stores program instructions, which, when executed by a computer, cause the computer to implement the flat part sorting method described in any of the above embodiments.
[0090] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0091] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0094] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0095] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0097] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0100] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for sorting flat parts, characterized in that, The method includes: The flat parts sorting task is broken down into multiple sub-tasks; The task status of each subtask is determined. When the task status indicates that it is to be executed, the humanoid robot is controlled to move to the task position of the subtask to be executed, and the first visual observation data after the humanoid robot moves to the task position is obtained. A sorting action sequence is generated based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and the humanoid robot is controlled according to the sorting action sequence.
2. The method according to claim 1, characterized in that, The pre-trained action model is invoked to generate a sorting action sequence based on the visual observation data and the task instructions corresponding to the sub-tasks to be executed. The training method of the action model includes: Obtain the position of the first sample task of any sample task in the sample task set, and obtain the first sample visual observation data and the first sample joint state after the humanoid robot runs to the first sample task position. Call the action model to generate the first predicted action sequence based on the first sample visual observation data, the first sample joint state and the task instruction corresponding to the sample task. Obtain a preset teaching trajectory sequence corresponding to the sample task, and train the action model based on the preset teaching trajectory sequence and the first predicted action sequence.
3. The method according to claim 2, characterized in that, The step of training the action model based on the preset teaching trajectory sequence and the first predicted action sequence includes: Determine the action similarity between the preset teaching trajectory sequence and the first predicted action sequence, and pre-train the action model based on the action similarity; Obtain the second sample task position of any sample task in the sample task set, and obtain the second sample visual observation data and the second sample joint state after the humanoid robot runs to the second sample task position. Call the pre-trained action model to generate a second predicted action sequence based on the second sample visual observation data, the second sample joint state and the task instruction corresponding to the sample task. The reward score of the second predicted action sequence is determined according to a preset reward function, and the parameters of the action model are adjusted according to the reward score.
4. The method according to claim 1, characterized in that, The sub-task to be executed includes a flat component delivery task, the task location includes a target delivery slot, the humanoid robot is controlled to move to the task location of the sub-task to be executed, and the first visual observation data after the humanoid robot moves to the task location is acquired includes: The humanoid robot is controlled to move to a preset flat component storage area to grasp flat components. After the flat part is grasped, a movement trajectory sequence is generated based on the second visual observation data of the humanoid robot, the preset scanning command and the state of the first joint. The first gripping end of the humanoid robot is then controlled to move the grasped flat part to the scanning area of the scanning device for scanning operation according to the movement trajectory sequence. The target delivery slot for the flat part is obtained based on the scanning operation, the humanoid robot is controlled to run to the target delivery slot, and the first visual observation data after the humanoid robot runs to the target delivery slot is obtained.
5. The method according to claim 4, characterized in that, The step of controlling the humanoid robot to move to the preset flat component temporary storage area to grasp flat components includes: Control the humanoid robot to run to the preset flat component temporary storage area, obtain the first pose data of the temporary storage device in the preset flat component temporary storage area relative to the humanoid robot, and adjust the pose of the humanoid robot according to the first pose data; The system acquires the third visual observation data of the humanoid robot after pose adjustment and the second joint state of the humanoid robot. Based on the third visual observation data, the second joint state and the preset grasping command, it generates a first end effector sequence and controls the first gripping end of the humanoid robot to move to the temporary storage device to grasp flat parts according to the first end effector sequence.
6. The method according to claim 1, characterized in that, The sub-task to be executed includes a bin replacement task, the task location includes a target replacement compartment, and the task instructions corresponding to the sub-task to be executed include a replacement instruction and a placement instruction. The step of generating a sorting action sequence based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and controlling the humanoid robot according to the sorting action sequence, includes: The state of the third joint of the humanoid robot is obtained, and a second end effector sequence is generated based on the state of the third joint, the first visual observation data and the replacement instruction. The first gripping end of the humanoid robot is then controlled to move to the target material replacement grid to grip the corresponding material box based on the second end effector sequence. After completing the grabbing of the bin, the humanoid robot is controlled to move to the preset placement position to place the grabbed bin. After the material bin is placed, the humanoid robot is controlled to move to the preset material picking area to grab a new material bin, and then the humanoid robot is controlled to move to the target material changing grid. Based on the fourth visual observation data of the humanoid robot at the target material changing slot, the placement command, and the fourth joint state, a third end effector sequence is generated, and the first gripping end of the humanoid robot is controlled to place the new material box into the target material changing slot according to the third end effector sequence.
7. The method according to claim 6, characterized in that, The step of generating a second end effector sequence based on the third joint state, the first visual observation data, and the replacement command, and controlling the first gripping end of the humanoid robot to move to the target material changing grid to grip the corresponding material box according to the second end effector sequence, includes: A second end effector sequence is generated based on the state of the third joint, the first visual observation data, and the replacement command. The first gripping end of the humanoid robot is then moved to the bottom of the corresponding material box in the target material replacement compartment based on the second end effector sequence. Acquire the second pose data of the corresponding material box relative to the second gripping end of the humanoid robot, and control the second gripping end to move to the top of the corresponding material box according to the second pose data; According to the preset pull teaching trajectory, the second gripping end is controlled to pull the corresponding material box away from the target material changing grid, and the first gripping end is controlled to move synchronously so that the first gripping end can grip the material box pulled out of the target material changing grid.
8. A flat parts sorting device, characterized in that, The device includes: The task splitting unit is used to split the flat part sorting task into multiple sub-tasks; The observation data acquisition unit is used to determine the task status of each sub-task. When the task status indicates that it is to be executed, it controls the humanoid robot to run to the task position of the sub-task to be executed and acquires the first visual observation data after the humanoid robot runs to the task position. The robot control unit is used to generate a sorting action sequence based on the first visual observation data and the task instructions corresponding to the sub-task to be executed, and to control the humanoid robot according to the sorting action sequence.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Article sorting method and related equipment
CN108357886A
A warehouse logistics quick picking system and method
CN109711770A
Sorting system and method
CN110404830A
Unmanned intensive intelligent storeroom
CN112722672A
Certificate sorting method, device and system and storage medium
CN116116736A