Pickup abnormality detection method and pickup abnormality detection device

CN122807940APending Publication Date: 2026-09-25BEIJING JIQING JUYUAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611272576.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]在仓储拣选场景中,机械臂可以从库存容器中拣选出订单所需的目标物品,在拣选过程中,可能会出现双拣异常情况,即机械臂在拾取目标物品时,会由于粘连等情况将其他物品也一起拾取出来,导致拣选任务异常,从而影响仓储系统的整体拣选效率

Benefits of technology

[0025]本申请实施例提供的拾取异常检测方法和拾取异常检测装置,首先,通过多个采集装置从不同角度同步采集取放装置的拣选图像,克服单一视角下因遮挡导致的部分物品不可见问题,为后续物品计数提供更全面的数据基础。其次,基于第一预设模型(即实例分割模型)输出图像的掩码信息和逐像素概率图,以对取放装置本体进行精准过滤,有效避免取放装置上的部件被误识别为物品所导致的虚警。之后,可以通过两条并列的分支一和分支二分别快速判决取放装置拾取的物品数量,并在两个分支判决物品数量一致时,基于该物品数量确定取放装置是否发生双拣异常。本申请通过双分支确定拾取物品数量,可以避免误检情况,提升了机器人拣选作业的准确性和可靠性,与相关技术中的双拣异常检测方式相比,本申请的检测方法具有更快的响应速度,适用于高速拣选场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807940A_ABST
    Figure CN122807940A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of warehouse logistics, and discloses a picking abnormality detection method and a picking abnormality detection device. The method comprises the following steps: after a taking and placing device picks target articles from a target container, a plurality of acquisition devices are used to acquire a plurality of image information corresponding to the taking and placing device; based on the plurality of image information and a first preset model, a plurality of candidate pixel regions corresponding to each image information are determined; based on the plurality of candidate pixel regions corresponding to each image information, a first quantity of articles picked by the taking and placing device is determined; based on the plurality of candidate pixel regions corresponding to each image information and a second preset model, a second quantity of articles picked by the taking and placing device is determined; in the case that the first quantity and the second quantity are consistent, a picking result of the taking and placing device is determined, and the picking result is used to control the taking and placing device to perform a picking operation. The application can improve the accuracy of double-picking abnormality detection of the taking and placing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of warehousing and logistics technology, and in particular to a method and device for detecting picking anomalies. Background Technology

[0002] In warehouse picking scenarios, robotic arms can pick the target items required for an order from inventory containers. During the picking process, double picking anomalies may occur, that is, when the robotic arm picks up the target item, other items may also be picked up due to adhesion or other reasons, resulting in picking task abnormalities and affecting the overall picking efficiency of the warehouse system. Summary of the Invention

[0003] To address the aforementioned problems, embodiments of this application provide a picking anomaly detection method and a picking anomaly detection device, which can improve the accuracy of dual-picking anomaly detection and increase the picking efficiency of the warehousing system. Specifically, embodiments of this application disclose the following technical solutions: The first aspect of this application provides a method for detecting picking anomalies, comprising: after a picking device picks up a target item from a target container, acquiring multiple image information corresponding to the picking device through multiple acquisition devices; determining multiple candidate pixel regions corresponding to each image information based on the multiple image information and a first preset model; wherein the multiple candidate pixel regions include a first pixel region corresponding to the item picked up by the picking device and a second pixel region corresponding to the body of the picking device; determining a first number of items picked up by the picking device based on the multiple candidate pixel regions corresponding to each image information; determining a second number of items picked up by the picking device based on the multiple candidate pixel regions corresponding to each image information and a second preset model; and determining the picking result of the picking device when the first number and the second number are consistent, and controlling the picking device to perform a picking operation based on the picking result.

[0004] In some embodiments, when the first quantity and the second quantity are the same, determining the picking result of the pick-up and placement device and controlling the pick-up and placement device to perform a picking operation based on the picking result includes: when both the first quantity and the second quantity indicate that the number of items picked up by the pick-up and placement device is one, determining that the picking result of the pick-up and placement device is a normal picking, and controlling the pick-up and placement device to transfer the picked-up item to the target location; or, when both the first quantity and the second quantity indicate that the number of items picked up by the pick-up and placement device is multiple, determining that the picking result of the pick-up and placement device is a picking anomaly, controlling the pick-up and placement device to transfer the picked-up item to the target container, and picking up the target item again from the target container.

[0005] In some embodiments, after determining the plurality of candidate pixel regions corresponding to each image information, the method further includes: performing dilation processing on the first pixel region and the second pixel region corresponding to each image information to obtain a plurality of effective pixel regions corresponding to each image information; wherein the plurality of effective pixel regions include the dilated first pixel region and the dilated second pixel region; determining at least one target first pixel region that is directly or indirectly adjacent to the dilated second pixel region based on the pixel overlap relationship between each effective pixel region; and generating a gripping area corresponding to the picking and placing device in each image information based on the at least one target first pixel region.

[0006] In some embodiments, the method further includes: if the first quantity and the second quantity are inconsistent, determining a third quantity of items picked up by the picking and placing device based on the gripping area corresponding to the picking and placing device in each image information, a preset prompt word, and a third preset model; determining the picking result of the picking and placing device based on the third quantity, and controlling the picking and placing device to perform a picking operation based on the picking result.

[0007] In some embodiments, determining the third quantity of items picked up by the picking and placing device based on the gripping area corresponding to the picking and placing device in each image information, a preset prompt word, and a third preset model includes: determining a target image in multiple image information; cropping each target image based on the gripping area in the target image to obtain a partial image of the gripping area corresponding to the target image; and inputting the partial image of the gripping area and the preset prompt word into the third preset model to obtain the third quantity output by the third preset model.

[0008] In some embodiments, determining the picking result of the pick-up and placement device based on the third quantity and controlling the pick-up and placement device to perform a picking operation based on the picking result includes: when the third quantity is one, determining that the picking result of the pick-up and placement device is a normal picking, and controlling the pick-up and placement device to transfer the picked item to the target location; or, when the third quantity is greater than one, determining that the picking result of the pick-up and placement device is a picking anomaly, controlling the pick-up and placement device to transfer the picked item to the target container, and picking up the target item again from the target container.

[0009] In some embodiments, determining the first number of items picked up by the pick-up and place device based on the multiple candidate pixel regions corresponding to each image information includes: determining the first number of items in each image information based on the number of at least one target first pixel region; and determining the first number of items picked up by the pick-up and place device based on the first number of items in each image information.

[0010] In some embodiments, determining the first number of items picked up by the picking and placing device based on the first number of items in each image information includes: comparing the size of the first number of items in each image information and determining the largest first number of items as the first number.

[0011] In some embodiments, determining the second quantity of items picked up by the picking and placing device based on multiple candidate pixel regions corresponding to each image information and a second preset model includes: cropping each image information based on the gripping area corresponding to the picking and placing device in each image information to obtain a cropped image; inputting each cropped image into a sub-classification network corresponding to each acquisition device to output a probability distribution of the quantity of items in each image information at multiple preset quantity levels through each sub-classification network; wherein, the second preset model includes multiple sub-classification networks; determining the second quantity of items in each image information based on each probability distribution; and determining the second quantity based on the second quantity of items in each image information.

[0012] In some embodiments, determining the quantity of the second item in each image information based on the probability distribution of the quantity of items in each image information across multiple preset quantity levels includes: determining the probability distribution of the quantity of items in the first image acquired by the first acquisition device across multiple preset quantity levels; wherein the first acquisition device is any one of multiple acquisition devices; determining the preset quantity level with the largest probability distribution as the corresponding target quantity level in the first image; and determining the quantity of the second item in the first image based on the target quantity level; wherein the multiple preset quantity levels include a first level, a second level, and a third level, the first level indicating that the quantity of the second item is zero, the second level indicating that the quantity of the second item is one, and the second quantity level indicating that the quantity of the second item is multiple.

[0013] In some embodiments, determining the second quantity based on the number of second items in each image information includes: comparing the size of the number of second items in each image information and determining the largest number of second items as the second quantity.

[0014] In some embodiments, the method further includes: when the second acquisition device is a depth image acquisition device, acquiring a second image acquired by the second acquisition device and determining the depth point cloud information corresponding to the second image; wherein, the plurality of acquisition devices includes the second acquisition device; based on the gripping area of ​​the pick-and-place device in the second image, performing spatial range extraction on the depth point cloud information, and determining the extracted point cloud as the point cloud within the gripping area of ​​the pick-and-place device; performing three-dimensional spatial clustering processing on the point cloud within the gripping area to determine the number of point cloud clusters within the gripping area. The above-mentioned determination of the number of second items in each image information based on various probability distributions includes: determining the number of second items in the second image based on the probability distribution of the number of items in the second image across multiple preset quantity levels and the number of point cloud clusters within the gripping area.

[0015] In some embodiments, determining multiple candidate pixel regions corresponding to each image information based on multiple image information and a first preset model includes: acquiring a preset region of interest corresponding to each of the multiple acquisition devices; inputting the image information corresponding to each acquisition device and the preset region of interest corresponding to each acquisition device into a sub-segmentation network corresponding to each acquisition device, so as to output mask information and pixel-wise probability map corresponding to each image information through each sub-segmentation network; wherein, the first preset model includes multiple sub-segmentation networks, the mask information indicates the pixel region corresponding to the candidate object in the image information, and the pixel-wise probability map indicates the probability that each pixel in the image information belongs to the body of the picking and placing device. Based on the mask information and pixel-wise probability map corresponding to each image information, multiple candidate pixel regions corresponding to each image information are determined.

[0016] In some embodiments, determining multiple candidate pixel regions corresponding to each image information based on the mask information and pixel-wise probability map corresponding to each image information includes: determining candidate regions corresponding to at least one candidate object in the first image based on the mask information corresponding to the first image acquired by the first acquisition device; wherein the first acquisition device is any one of multiple acquisition devices; determining the pixel probability of each pixel in each candidate region belonging to the body of the picking and placing device based on each candidate region and the pixel-wise probability map corresponding to the first image; and filtering out the first pixel region and the second pixel region corresponding to the first image from at least one candidate region based on the pixel probability of each pixel in each candidate region and a preset probability threshold.

[0017] In some embodiments, the above-mentioned selection of a first pixel region and a second pixel region corresponding to a first image from at least one candidate region based on the pixel probability of each pixel in each candidate region and a preset probability threshold includes: determining the average pixel probability of the candidate region based on the pixel probability of each pixel in each candidate region, deleting candidate regions whose average pixel probability exceeds the preset probability threshold, and obtaining the first pixel region corresponding to the first image; in the first image, determining the pixels whose pixel probability exceeds the preset probability threshold as the main pixels of the pick-and-place device, and determining the second pixel region based on the main pixels of the pick-and-place device.

[0018] In some embodiments, determining the second pixel region based on the body pixels of the pick-and-place device includes: performing connected component processing on the body pixels of the pick-and-place device to determine at least one connected component corresponding to the pick-and-place device; determining the pixel area of ​​each connected component; and determining the connected component with the largest pixel area as the second pixel region in the first image.

[0019] In some embodiments, the above-mentioned acquisition of multiple image information corresponding to the pick-up and place device through multiple acquisition devices includes: when the end mechanism of the pick-up and place device picks up the target item from the target container and transfers the target item to a preset acquisition area, acquiring multiple image information corresponding to the end mechanism through multiple acquisition devices; wherein, the multiple acquisition devices include multiple acquisition devices set in the working area and / or at least one acquisition device on the pick-up and place device.

[0020] A second aspect of this application provides a picking anomaly detection device, including an acquisition module, a determination module, and a control module. The acquisition module is configured to: acquire multiple image information corresponding to the picking and placing device via multiple acquisition devices after the picking and placing device picks up a target item from a target container. The determination module is configured to: determine multiple candidate pixel regions corresponding to each image information based on the multiple image information and a first preset model; wherein the multiple candidate pixel regions include a first pixel region corresponding to the item picked up by the picking and placing device and a second pixel region corresponding to the body of the picking and placing device; determine a first number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each image information; determine a second number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each image information and a second preset model; and determine the picking result of the picking and placing device if the first number and the second number are consistent. The control module is configured to: control the picking and placing device to perform a picking operation based on the picking result.

[0021] A third aspect of this application provides an electronic device, including: a processor and a memory, wherein the memory is used to store computer-executable instructions; and the processor is used to read the instructions from the memory and execute the instructions to implement the picking anomaly detection method described in the first aspect above.

[0022] A fourth aspect of this application provides a computer-readable storage medium storing computer program instructions, which, when read by a computer, execute the picking anomaly detection method described in the first aspect.

[0023] A fifth aspect of this application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the picking anomaly detection method described in the first aspect.

[0024] The sixth aspect of this application provides a computer program that, when executed by a processor, can implement the picking anomaly detection method described in the first aspect.

[0025] The picking anomaly detection method and device provided in this application firstly acquire picking images of the picking device from different angles simultaneously using multiple acquisition devices, overcoming the problem of some items being invisible due to occlusion from a single viewpoint, and providing a more comprehensive data foundation for subsequent item counting. Secondly, based on the mask information and pixel-by-pixel probability map of the output image from the first preset model (i.e., instance segmentation model), the picking device itself is accurately filtered, effectively avoiding false alarms caused by components on the picking device being misidentified as items. Then, the number of items picked by the picking device can be quickly determined through two parallel branches, branch one and branch two. When the number of items determined by the two branches is consistent, it is determined whether a double-picking anomaly has occurred based on the number of items. This application determines the number of picked items through a dual-branch method, which avoids false detections and improves the accuracy and reliability of robot picking operations. Compared with double-picking anomaly detection methods in related technologies, the detection method of this application has a faster response speed and is suitable for high-speed picking scenarios. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A schematic diagram of a warehousing system provided in an embodiment of this application; Figure 2 This is a schematic diagram of a pick-and-place device provided in an embodiment of this application; Figure 3 A schematic diagram of an anomaly detection method provided in an embodiment of this application; Figure 4 A schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application; Figure 5 A schematic diagram illustrating yet another method for picking up anomalies provided in an embodiment of this application; Figure 6 A schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application; Figure 7 A schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application; Figure 8 A schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application; Figure 9 A schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application; Figure 10A schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application; Figure 11 This is a schematic diagram of an anomaly detection device provided in an embodiment of this application; Figure 12 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, and to make the above-mentioned objectives, features and advantages of the embodiments of the present invention more apparent and understandable, the technical solutions in the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0029] Figure 1 This is a schematic diagram of a warehousing system provided as an embodiment of this application. Figure 1 As shown, the warehousing system 100 includes multiple carriers 10, robots 20, work areas 30, and pick-and-place devices 40.

[0030] Exemplarily, the carrier 10 may include multiple storage locations, which may be used to place containers, bins, or goods directly, or even the original packaging of the goods. This application does not limit this; the following embodiments use the placement of containers as an example for illustrative purposes. For example, the carrier 10 may be a shelf, a trolley, a cage cart, etc.

[0031] For example, such as Figure 1 As shown, the robots 20 in the warehousing system 100 may include various types of robots, such as the first robot 21 and the second robot 22. The application scenarios and functions of different types of robots may differ.

[0032] In some examples, the first robot 21 may be referred to as a handling robot, which is primarily used to handle the carrier 10 or containers placed on the carrier 10. For example, the first robot 21 can move inventory objects (such as carriers or containers) from the inventory area to the work area 30 to perform corresponding operations.

[0033] In some examples, the second robot 22 may be referred to as a picking robot. The second robot 22 can perform movement operations or picking operations. For example, the second robot 22 can pick up a container or pick up items from a container.

[0034] In some examples, the second robot 22 can be a robot with a robotic arm, such as a unibody robot.

[0035] For example, the warehousing system 100 may include multiple work areas 30, which may be areas for performing operations such as sorting, picking or packing.

[0036] In some examples, a first robot 21 and / or conveyor line can transport containers (such as inventory containers) from the inventory area to a picking station in the work area 30, where a second robot 22 located in the work area 30 performs the picking operation on the inventory containers at the picking station. The second robot 22 can pick the items required for the order from the inventory containers and deliver them to the corresponding order container.

[0037] In some embodiments, the warehousing system 100 further includes at least one picking and placing device 40. For example, the picking and placing device 40 may be a robotic arm device, and a robotic arm used to perform picking operations may also be called a picking robotic arm.

[0038] In some examples, the pick-and-place device 40 may be a robotic arm device independently set in the work area for performing the picking / placing of items; or, the pick-and-place device 40 may also be a robotic arm device set on the robot 20 for performing the picking / placing of items. For example, when the robot 20 is a embodied robot, the pick-and-place device 40 may be a robotic arm on the embodied robot.

[0039] Figure 2 This is a schematic diagram of a pick-and-place device provided in an embodiment of this application. Figure 2 As shown, the pick-and-place device 40 may include a base 41, an articulated arm 42, an end effector 43, and a plurality of replaceable end tools 44.

[0040] In some examples, the base 41 serves as the bottom support structure for the pick-and-place device 40, used to secure the device to the workstation floor, control table, or robot body. The base 41 provides a stable mounting foundation for the articulated arm 42, ensuring the stability of the pick-and-place device during movement and operation.

[0041] In some examples, the articulated arm 42 is the motion mechanism of the pick-and-place device 40. One end of the articulated arm 42 can be connected to the base 41, and the other end can be connected to the end effector 43. For example, the articulated arm 42 includes multiple joints and connecting arms. Each joint is provided with a driving component. Through the coordinated driving of each driving component, the articulated arm 42 can achieve multi-degree-of-freedom position and posture adjustment in three-dimensional space. It should be noted that the specific form of the articulated arm 42 is not limited in the embodiments of this application. The articulated arm 42 can be a six-axis articulated arm structure, a four-axis articulated arm structure, or other multi-degree-of-freedom robotic arm structures.

[0042] In some examples, the end effector 43 can be the end actuation component of the articulated arm 42, used to connect to and drive the end tool 44. The end effector 43 can be precisely moved to the target position and adjusted to the corresponding posture under the drive of the articulated arm 42. The end tool 44 is the actuation component of the pick-and-place device 40 that performs specific pick-up operations, and is detachably mounted on the end effector 43.

[0043] In some examples, the end effector 44 may include multiple types, and different types of end effectors 44 can be used to pick up different items. For example, the end effector 44 may include two main categories: suction cup tools and gripper tools. It should be noted that the end effector 44 may also include more types, but this embodiment of the application does not limit this.

[0044] It should be noted that, Figure 2 This is only one implementation of the picking and placing device 40. The picking and placing device 40 can also be implemented in other ways. This application does not limit the specific form of the picking and placing device, as long as the picking and placing device can pick items from the inventory container to the order container.

[0045] For example, the warehousing system 100 may also include a control device, which may be a server or a terminal device. The terminal device may include at least one of a personal computer, a laptop computer, a smartphone, a tablet computer, and a portable wearable device; the server may include a standalone server or a server cluster consisting of multiple servers, which is not limited in this embodiment.

[0046] In some examples, the control device and the pick-and-place device 40 are communicatively connected for data communication. For example, the control device can communicate with the pick-and-place device 40 via a local area network (LAN), wireless local area network (WLAN), or other networks.

[0047] In some examples, during robotic arm picking scenarios, a double-pickup anomaly may occur when the robotic arm picks the items required for an order from the inventory container. This double-pickup anomaly indicates that the robotic arm picks up multiple items at once, instead of just one. For example, when picking up the items required for an order, the robotic arm may pick up other items due to adhesion or other reasons, leading to a picking anomaly and impacting the overall picking efficiency of the warehousing system.

[0048] To promptly detect double-picking anomalies and prevent subsequent picking failures, two common methods are used to detect them. One method uses weight sensors to measure the weight difference between the inventory containers before and after picking, thus determining the number of items picked. This method has a significant response delay and a high false positive rate for items with similar weights. The other method uses a single camera combined with image recognition to count the number of items picked by the robotic arm. However, this method is susceptible to limitations in viewing angle, leading to misjudgments, unnecessary return actions, and repeated picking, thus reducing the picking efficiency of the warehouse system.

[0049] To address the aforementioned issues, this application employs multiple vision cameras to simultaneously acquire picking images of the pick-and-place device from different angles. It then combines this with pixel-wise probability maps from an instance segmentation model (i.e., a first preset model) for on-device filtering of the pick-and-place device, and utilizes parallel computation and mutual verification of two independent fast decision branches. This effectively solves the problems of high false positive rates and missed detections in related technologies regarding double-picking anomalies. Furthermore, this application, through its fast and slow dual-channel architecture, can complete decisions with extremely low latency in normal scenarios, triggering a multimodal visual language model for semantic verification only in a few divergent scenarios. This balances real-time performance and accuracy, making it widely applicable to robotic picking scenarios in the field of warehouse logistics automation.

[0050] The following is combined Figure 3 The implementation process of the picking anomaly detection method provided in the embodiments of this application will be described. It should be noted that the picking anomaly detection method provided in the embodiments of this application can be implemented by the above-mentioned warehousing system 100, for example, by the control device in the warehousing system 100.

[0051] For example, when the pick-and-place device 40 is set independently (such as in the work area), the control device can be the robot management system (RMS) in the warehouse system 100; when the pick-and-place device 40 is set on the robot, the control device can be the robot management system (RMS) in the warehouse system 100, or it can be the controller in the robot body. This application embodiment does not limit this.

[0052] Figure 3 This is a schematic diagram of an anomaly detection method provided in an embodiment of this application. Figure 3 As shown, the picking anomaly detection method includes steps 310 to 350 as shown below.

[0053] Step 310: After the pick-up and place device picks up the target item from the target container, multiple image information corresponding to the pick-up and place device is acquired through multiple acquisition devices.

[0054] In some examples, the target container can be any inventory container in the warehousing system that holds items. For example, the target container can be located at a picking station in the work area, or it can be placed in other locations, which is not limited in this application embodiment. The target container can contain multiple items. The control device issues picking instructions to the picking and placing device to instruct the picking and placing device to pick up the required items (such as target items) from the target container and remove the target items from the target container to transfer them to the target location.

[0055] For example, after the pick-up and place device retrieves the target item from the target container, multiple image information corresponding to the pick-up and place device can be acquired through multiple acquisition devices in the warehousing system. Each image information may include an image of the item picked up by the pick-up and place device.

[0056] In some examples, the acquisition device can be a vision camera, such as at least one of an RGB-D depth camera, a structured light camera, a time-of-flight (TOF) camera, and a traditional RGB industrial camera. This application does not limit this, and the following embodiments can be illustrated by taking a camera as an example.

[0057] In some embodiments, when the end mechanism of the pick-and-place device picks up a target item from a target container and transfers the target item to a preset collection area, multiple image information corresponding to the end mechanism is acquired by multiple collection devices. These multiple collection devices include multiple collection devices located in the work area and / or at least one collection device on the pick-and-place device.

[0058] For example, after receiving a picking instruction from the control device, the picking and placing device controls its end effector to perform a picking operation in the target container to remove the target item from the target container. Once the end effector completes the picking action and lifts the target item to a preset collection area, multiple collection devices can be triggered to capture images of the end effector's grip from their respective perspectives.

[0059] In some examples, the end effector can stop in a preset acquisition area for a preset time to allow each acquisition device to take pictures, or the end effector can pass through the preset acquisition area at a lower speed to allow each acquisition device to take pictures. This application embodiment does not limit this. The preset acquisition area can be located at a preset distance above the target container, and the preset acquisition area can ensure that multiple acquisition devices can capture the holding status of the end effector.

[0060] For example, the multiple acquisition devices include multiple acquisition devices (such as referred to as the first acquisition device) disposed in the work area, and / or at least one acquisition device (such as referred to as the second acquisition device) on the pick-and-place device. The first acquisition device may be fixedly installed in the work area, such as on a support frame or workbench in the work area, and the first acquisition device can capture images of the gripping status of the end mechanism of the pick-and-place device from an external perspective. The second acquisition device may be installed on the pick-and-place device, such as near the end mechanism, and the second acquisition device can capture images of the gripping status of the end mechanism at close range.

[0061] In some examples, all of the multiple acquisition devices may be the first acquisition device, or the multiple acquisition devices may include both the first acquisition device and the second acquisition device; this application embodiment does not limit this. To improve the accuracy of detection, the multiple acquisition devices in this application embodiment may include at least one first acquisition device and at least one second acquisition device. By combining and arranging them together, they can complement each other. The first acquisition device provides the global field of view, while the second acquisition device provides local details, thereby effectively overcoming the structural occlusion problem of a single viewpoint.

[0062] It should be noted that the embodiments of this application do not limit the number of the first acquisition device and the second acquisition device. For example, there can be multiple first acquisition devices and only one second acquisition device. The embodiments of this application also do not limit the type of the first acquisition device and the second acquisition device. The first acquisition device and the second acquisition device can be the same type of camera or different types of cameras. Alternatively, multiple first acquisition devices can be the same type of camera or different types of cameras.

[0063] For example, at least one first acquisition device includes, but is not limited to, a bird's-eye view camera positioned directly above the target container, and a side-view camera positioned on the left and / or right side of the work area. A second acquisition device may include a wrist camera mounted on the end mechanism of the pick-and-place device.

[0064] In some examples, each acquisition device can simultaneously capture images after receiving the shooting command from the control device, so as to ensure that each acquisition device obtains an image of the gripping status of the end mechanism of the picking and placing device at the same moment, and avoids misalignment of multi-view images in time due to slight vibrations of the end mechanism.

[0065] Step 320: Based on multiple image information and the first preset model, determine multiple candidate pixel regions corresponding to each image information.

[0066] For example, after completing image acquisition, each acquisition device can send the acquired image information to the control device. The control device performs image processing using a pre-trained first preset model to obtain multiple candidate pixel regions corresponding to each image information. These multiple candidate pixel regions may include a first pixel region corresponding to the item picked up by the picking and placing device and a second pixel region corresponding to the picking and placing device itself.

[0067] In some examples, the first preset model can be a pre-trained network model, which can process multiple images separately to obtain the first pixel region and the second pixel region corresponding to each image information.

[0068] In some examples, the first pixel region may be a pixel region of at least one candidate object picked up by the pick-and-place device, as identified by the first preset model from the image. The candidate object may be an item picked up by the end-effector, and at least one candidate object may include the target item and other items; alternatively, at least one candidate object may only include the target item, and this embodiment of the application does not limit this.

[0069] In some examples, the second pixel region may be a pixel region of the picking and placing device body identified from the image by the first preset model. For example, the second pixel region may include the region corresponding to the end effector and / or other components on the picking and placing device (such as an articulated arm).

[0070] The following is combined Figure 4 The process of determining multiple candidate pixel regions corresponding to each image information using a first preset model is explained.

[0071] Figure 4 This is a schematic diagram of another anomaly detection method provided in an embodiment of this application. Figure 4 As shown, step 320 above includes steps 321 to 323 as shown below.

[0072] Step 321: Obtain the preset interest area corresponding to each of the multiple acquisition devices.

[0073] For example, a corresponding region of interest (such as a preset region of interest) can be pre-configured for each acquisition device. The preset region of interest can be the projection area of ​​the spatial range that each acquisition device needs to focus on onto the image plane. The preset regions of interest for each acquisition device are pre-set and stored, and the control device can obtain the preset region of interest for that acquisition device after acquiring the image information acquired by each acquisition device.

[0074] In some examples, the preset regions of interest for each acquisition device can be different. For instance, for a bird's-eye view camera positioned above a target container, its preset region of interest can include the pixel area inside the target container, thus masking the border and outer area of ​​the target container in the image. For a wrist camera positioned on a pick-and-place device, its preset region of interest can be the pixel area surrounding the end effector.

[0075] Step 322: Input the image information corresponding to each of the multiple acquisition devices and the preset interest region corresponding to each acquisition device into the sub-segmentation network corresponding to each acquisition device, so as to output the mask information and pixel-by-pixel probability map corresponding to each image information through each sub-segmentation network.

[0076] For example, the first preset model may include multiple sub-segmentation networks. The number of sub-segmentation networks may be related to the number of acquisition devices (i.e., multiple acquisition devices) performing the picking detection. Each segmentation sub-network may process the image information acquired by each acquisition device independently.

[0077] For example, taking a total of 4 acquisition devices as an example, the first preset model may include 4 parallel sub-segmentation networks. Each sub-segmentation network processes the image information acquired by one acquisition device to output the mask information and pixel-wise probability map corresponding to the image information.

[0078] In some examples, the control unit can create an independent inference namespace for each acquisition device, with a sub-segmentation network running in each namespace. Each sub-segmentation network can share the same model structure and weight parameters, and each sub-segmentation network can independently load the preset region of interest for its corresponding acquisition device.

[0079] In some examples, for each acquisition device, the control device uses the image information it acquires as input (hereinafter referred to as the input image). The sub-segmentation network can spatially crop the input image based on the preset region of interest corresponding to the acquisition device to retain the image data within the preset region of interest, thereby reducing the computational load of the sub-segmentation network and focusing on the region of interest.

[0080] For example, mask information can indicate a mask corresponding to at least one candidate object in the image information. The sub-segmentation network can aggregate pixels belonging to the same entity into a single mask and assign a unique identifier to each mask. For instance, mask information can include object masks of at least one candidate object picked up by the end-effector.

[0081] For example, a pixel-wise probability map can indicate the probability that each pixel in an image belongs to the body of the pick-and-place device. That is, the sub-segmentation network can calculate a probability value for each pixel in the input image, which can reflect the likelihood that the pixel belongs to the pixel region of the pick-and-place device body.

[0082] In some examples, the sub-segmentation network can distinguish between pixels on the device itself and pixels not on the device itself in the input image. For instance, assuming the device is mounted on a robot, pixels on the device itself can also be called robot pixels, and the pixel-by-pixel probability map can also be called the pixel-by-pixel robot probability map.

[0083] It should be noted that the pixel-by-pixel probability map makes independent judgments for each pixel. Even if a pixel is identified as an object mask by the sub-segmentation network, that pixel can still be assigned a probability value belonging to the object itself. For example, the pixel probability of a pixel belonging to a candidate object in the image can be lower than the pixel probability of a pixel belonging to the object itself.

[0084] Step 323: Based on the mask information and pixel-by-pixel probability map corresponding to each image information, determine multiple candidate pixel regions corresponding to each image information.

[0085] For example, the control device can spatially align each object mask in the mask information of each image information with the corresponding pixel-wise probability map, thereby determining the pixel region (i.e., the first pixel region) that belongs to the candidate object and the pixel region (i.e., the second pixel region) that belongs to the pick-and-place device body from the image information.

[0086] In some examples, after outputting the mask information and pixel-by-pixel probability map corresponding to each image information through the first preset model, candidate pixel regions can be filtered out from the image data (i.e., mask information and pixel-by-pixel probability map) through filtering processing, that is, the first pixel region corresponding to the item and the second pixel region corresponding to the picking and placing device body can be filtered out.

[0087] For example, since the filtering process for each image information is similar, the following embodiment uses the first acquisition device among multiple acquisition devices and the first image acquired by the first acquisition device as an example to illustrate the implementation process of step 323 above. It should be noted that the processing methods for other image information are similar, and to avoid repetition, they will not be described again here.

[0088] Figure 5 This is a schematic diagram illustrating another method for detecting picking anomalies provided in an embodiment of this application. Figure 5 As shown, step 323 above includes steps 3231 to 3233 as shown below.

[0089] Step 3231: Based on the mask information corresponding to the first image acquired by the first acquisition device, determine the candidate region corresponding to at least one candidate object in the first image.

[0090] For example, for any one of the multiple acquisition devices (such as the first acquisition device), the control device can extract the segmentation result corresponding to each candidate object from the mask information corresponding to the first image. The mask information contains all candidate objects identified by the sub-segmentation network in the first image, and each candidate object corresponds to a binary instance mask. The control device traverses all instance masks in the mask information and determines the corresponding pixel region (called the candidate region) for each instance mask. Each candidate region may include multiple pixels.

[0091] For example, in the first image acquired by the first acquisition device, three independent candidate objects can be detected by the sub-segmentation network, corresponding to item A, item B and item C respectively. The pixel regions corresponding to the three items are candidate region a, candidate region b and candidate region c respectively.

[0092] Step 3232: Based on the pixel-by-pixel probability map corresponding to each candidate region and the first image, determine the pixel probability that each pixel in each candidate region belongs to the main body of the picking and placing device.

[0093] For example, the control device uses each candidate region (i.e. each instance mask) as a spatial template and maps it onto the pixel-by-pixel probability map corresponding to the first acquisition device.

[0094] In some examples, for each candidate region in the first image, the control device can extract the probability value corresponding to each pixel within the coverage area of ​​the candidate region on the pixel-by-pixel probability map, thereby determining the pixel probability that each pixel in the candidate region belongs to the main body of the picking and placing device. For example, the pixel probability corresponding to each pixel can be a value between 0 and 1.

[0095] For example, in candidate region 'a' corresponding to item A, if item A covers p1 pixels, the probability values ​​corresponding to these p1 pixels can be extracted from the pixel-by-pixel probability map. Similarly, in candidate region 'b' corresponding to item B, if item B covers p2 pixels, the pixel probabilities corresponding to these p2 pixels can be extracted from the pixel-by-pixel probability map. Through this pixel mapping, the pixel probability corresponding to each pixel in each candidate region can be obtained, that is, the probability that each pixel belongs to the main body of the picking and placing device.

[0096] Step 3233: Based on the pixel probability of each pixel in each candidate region and a preset probability threshold, select the first pixel region and the second pixel region corresponding to the first image from at least one candidate region.

[0097] In some examples, after determining the pixel probability of each pixel in each candidate region, at least one candidate region can be filtered based on the pixel probability and a preset probability threshold to select the first pixel region corresponding to the item and the second pixel region corresponding to the picking and placing device body from at least one candidate region.

[0098] In some embodiments, step 3233 includes: determining the average pixel probability of a candidate region based on the pixel probability of each pixel in each candidate region, deleting candidate regions whose average pixel probability exceeds a preset probability threshold, and obtaining the first pixel region corresponding to the first image.

[0099] For example, after determining the pixel probability of each pixel in each candidate region corresponding to the first image, the average pixel probability of all pixels in the candidate region can be calculated to obtain the average pixel probability of the candidate region. The average pixel probability reflects the probability distribution of the candidate region as a whole belonging to the pick-and-place device body.

[0100] In some examples, a high average pixel probability indicates that most pixels in the candidate region are judged by the first preset model to likely belong to the pick-and-place device itself; that is, the candidate region may be the pixel area covered by the end effector (such as a suction cup or gripper) on the pick-and-place device, and not the object area. Conversely, a low average pixel probability indicates that most pixels in the candidate region are judged by the first preset model to be unlikely to belong to the pick-and-place device itself; that is, the candidate region is likely the area covered by the object picked up by the end effector.

[0101] In some examples, after determining the average pixel probability of the candidate region, the control device compares the average pixel probability with a preset probability threshold, and further filters the first and second pixel regions based on the comparison result. The preset probability threshold can be set according to requirements, such as 0.7 or 0.8, etc., and this embodiment does not limit this setting.

[0102] For example, if the average pixel probability exceeds a preset probability threshold, it can be determined that the candidate area belongs to the pick-and-place device body. Therefore, the candidate area can be removed from the item candidate list, that is, the candidate area cannot be used as the first pixel area.

[0103] For example, if the average pixel probability does not exceed a preset probability threshold, it can be determined that the candidate region belongs to an item. Therefore, the candidate region can be marked as the first pixel region, that is, the pixel region corresponding to the item, and the candidate region can be kept in the item candidate list.

[0104] In some examples, after filtering by the preset probability threshold, candidate regions belonging to the device body can be excluded from at least one candidate region. The remaining candidate regions (i.e., candidate regions in the item candidate list) are the pixel regions corresponding to the items in the first image, thus obtaining the first pixel region corresponding to the first image. The number of first pixel regions can be one or more.

[0105] In some embodiments, in the first image, pixels with a pixel probability exceeding a preset probability threshold are identified as the main pixels of the pick-and-place device, and a second pixel region is determined based on the main pixels of the pick-and-place device.

[0106] For example, after obtaining the first pixel region corresponding to the first image, the pixels belonging to the body of the picking and placing device can be further filtered out based on the pixels in the first image whose pixel probability exceeds a preset probability threshold, and a second pixel region can be obtained based on the body pixels. The second pixel region is the region that covers all body pixels.

[0107] In some examples, among all the pixels of a pixel whose probability exceeds a preset probability threshold (hereinafter referred to as the body pixels), some body pixels may be distributed in discontinuous regions. In this case, it indicates that there are noise pixels in the body pixels. Therefore, connectivity processing can be performed on the body pixels to filter out noise pixels from the body pixels in order to obtain the second pixel region.

[0108] In some embodiments, the anomaly detection method may further include: performing connected component processing on the main pixels of the picking and placing device to determine at least one connected component corresponding to the picking and placing device; determining the pixel area of ​​each connected component; and determining the connected component with the largest pixel area as the second pixel region in the first image.

[0109] For example, after determining the main pixel of the pick-and-place device, the main pixel can be further processed as a connected component.

[0110] In some examples, connected component processing is performed on the ontology pixels to divide interconnected ontology pixels into connected components, and a unique identifier is assigned to each connected component. After determining at least one connected component using multiple ontology pixels, the pixel area of ​​each connected component can be determined. The pixel area of ​​a connected component can be related to the number of pixels contained within that component. For example, pixel area is positively correlated with the number of pixels; the more pixels, the larger the pixel area of ​​the connected component.

[0111] In some examples, after obtaining the pixel area of ​​each connected component, the connected component with the largest remaining pixel area can be identified as the second pixel region. This connectivity processing ensures that the second pixel region contains only the main body of the pick-and-place device, eliminating glitches and isolated noise caused by the prediction noise of the first preset model, making subsequent analysis more accurate and reliable.

[0112] In some examples, after the processing in step 320 above, at least one first pixel region and a second pixel region corresponding to each image information can be obtained. After determining the first pixel region and the second pixel region corresponding to each image information, the gripping area of ​​the pick-and-place device can be further determined. The gripping area of ​​the pick-and-place device can be a pixel region connected to the end of the pick-and-place device. That is, it is necessary to select the first pixel region connected to the second pixel region from at least one first pixel region.

[0113] The following is combined Figure 6 The process of determining the gripping area of ​​the pick-up and drop-off device is explained.

[0114] Figure 6 This is a schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application. For example... Figure 6 As shown, after step 320 above, the method further includes steps 610 to 630 as shown below.

[0115] Step 610: Dilation processing is performed on the first pixel region and the second pixel region corresponding to each image information to obtain multiple effective pixel regions corresponding to each image information.

[0116] For example, the control device performs dilation processing on all first pixel regions and second pixel regions corresponding to each image information to expand the range of each first pixel region and second pixel region, thereby obtaining multiple effective pixel regions. The multiple effective pixel regions include the dilated first pixel regions and the dilated second pixel regions.

[0117] In some examples, the dilation process can expand the edges of each effective pixel region outward by a preset number of pixels based on a preset size. The preset size can be set according to requirements, and this embodiment does not limit it. After the dilation process, the first pixel region corresponding to each image information is expanded outward by the preset number of pixels, resulting in a dilated first pixel region; the second pixel region is also expanded outward by the preset number of pixels, resulting in a dilated second pixel region. The dilated first pixel region and the dilated second pixel region are collectively referred to as the multiple effective pixel regions corresponding to the image information.

[0118] It should be noted that due to factors such as lighting, shadows, or surface reflections, objects that are physically close together may have tiny gaps of a few pixels in the image. Dilation processing can create pixel overlap between adjacent but non-overlapping areas, thus enabling the correct determination of the number of items adjacent to the pick-up and drop-off device in subsequent processing.

[0119] Step 620: Based on the pixel overlap relationship between each effective pixel region, determine at least one target first pixel region that is directly or indirectly adjacent to the dilated second pixel region.

[0120] For example, after determining multiple valid pixel regions, the pixel overlap relationship between any two valid pixel regions can be determined. The pixel overlap relationship indicates whether there are overlapping pixels between any two valid pixel regions; if overlapping pixels exist, it indicates that pixel overlap occurs between the two valid pixel regions.

[0121] In some examples, based on the pixel overlap relationship between each effective pixel region, a first pixel region that is directly or indirectly adjacent to the dilated second pixel region can be selected from at least one first pixel region and determined as the target first pixel region. That is, when the dilated first pixel region is directly or indirectly adjacent to the dilated second pixel region, the first pixel region can be determined as the target first pixel region.

[0122] In some examples, where there is pixel overlap between the dilated first pixel region and the dilated second pixel region, it indicates that the dilated first pixel region and the dilated second pixel region are directly adjacent.

[0123] For example, if there is at least one pixel overlap between the first pixel region a1 corresponding to item A and the second pixel region b corresponding to the loading and unloading device body, the first pixel region a1 corresponding to item A can be referred to as the target first pixel region directly adjacent to the second pixel region b.

[0124] In some examples, the first pixel region after dilation is not directly adjacent to the second pixel region after dilation, but there is pixel overlap between the first pixel region after dilation and another first pixel region after dilation, that is, the first pixel region after dilation is adjacent to the other first pixel region after dilation. In the case that the other first pixel region after dilation is directly adjacent to the second pixel region after dilation, it indicates that the first pixel region after dilation and the second pixel region after dilation are indirectly adjacent.

[0125] For example, if the first pixel region a1 corresponding to item A is directly adjacent to the second pixel region, there is no pixel overlap between the first pixel region a2 corresponding to item B and the second pixel region b corresponding to the loading / unloading device body, and there is at least one pixel overlap between the first pixel region a2 corresponding to item B and the first pixel region a1 corresponding to item A. Therefore, the first pixel region a2 is directly adjacent to the first pixel region a1. In this case, the first pixel region a2 is indirectly adjacent to the second pixel region b.

[0126] In some examples, each valid pixel region can be treated as a node to construct an object adjacency matrix. For any two valid pixel regions, check if they have a pixel intersection, i.e., determine if a pixel belongs to both valid pixel regions simultaneously. If the pixel intersection area is greater than 0, the two valid pixel regions are spatially adjacent, and the corresponding position in the adjacency matrix is ​​marked as 1. If the pixel intersection area is 0, the two valid pixel regions are spatially non-adjacent, and the corresponding position in the adjacency matrix is ​​marked as 0.

[0127] In some examples, after constructing the adjacency matrix, the transitive closure of the adjacency matrix can be calculated. The solution for the transitive closure can be implemented using a disjoint-set data structure or other algorithms; this application does not limit the specific implementation. For example, the disjoint-set data structure can be initialized first, making each valid pixel region an independent set; then, all elements marked as 1 in the adjacency matrix are traversed, and the sets containing the corresponding two nodes are merged; after the traversal, all the dilated first pixel regions that are in the same set as the dilated second pixel region are the target first pixel regions directly or indirectly adjacent to the main body of the pick-and-place device.

[0128] It should be noted that after the above expansion process, even if an item does not directly contact the end of the pick-and-place device—for example, if item A is gripped by the gripper and item B adheres to item A, preventing direct contact with the gripper—the item will still be identified as indirectly adjacent to the pick-and-place device. Item B forms a connection path with the pick-and-place device through item A. This system can accurately count all items currently held by the pick-and-place device, regardless of whether these items directly contact the robot or indirectly through other items.

[0129] Step 630: Based on at least one target first pixel region, generate the gripping area corresponding to the picking and placing device in each image information.

[0130] For example, after determining at least one target pixel region, all target first pixel regions can be merged to obtain a complete binary mask image, which is the image of the gripping area corresponding to the picking and placing device in the image information.

[0131] In some examples, other first pixel regions that are not connected to the end of the pick-and-place device have been filtered out through step 620 above. Therefore, the resulting gripping area can cover the overall outline of all items actually held by the end of the pick-and-place device. That is, the gripping area includes the items actually held by the pick-and-place device and does not include distracting objects in the background.

[0132] It should be noted that step 320 above yields multiple candidate pixel regions corresponding to each image information, and step 630 above yields the gripping area corresponding to the pick-and-place device. Subsequently, based on the candidate pixel regions and gripping areas corresponding to each image information, the number of items picked up by the pick-and-place device can be further determined.

[0133] In some examples, embodiments of this application may provide two parallel methods for determining the number of items picked up by the pick-up and place device. These two parallel methods may be referred to as Branch 1 and Branch 2, respectively. That is, Branch 1 and Branch 2 may be two implementations of determining the number of items picked up by the pick-up and place device. For example, step 330 below may be an implementation of Branch 1, and step 340 below may be an implementation of Branch 2.

[0134] Step 330: Based on multiple candidate pixel regions corresponding to each image information, determine the first number of items picked up by the picking and placing device.

[0135] In some examples, step 330 is the first branch for determining the number of items picked up by the pick-up and place-down device. This first branch can also be called fast decision branch one. Branch one can determine the number of items picked up by the pick-up and place-down device based on the target first pixel area determined in step 620 above. The number of items determined by branch one is called the first quantity.

[0136] In some embodiments, step 330 may include: determining the number of first items in each image information based on the number of at least one target first pixel region; and determining the first number of items picked up by the pick-and-place device based on the number of first items in each image information.

[0137] For example, for the image information acquired by each acquisition device, the number of at least one target first pixel region can be determined through step 620 above. Based on the number of target first pixel regions, the number of items determined from the perspective of the acquisition device can be determined, referred to as the first item quantity. The first item quantity can also be referred to as the local counting result from the perspective of each acquisition device.

[0138] For example, the number of items picked up by the picking and placing device from the perspective of the first acquisition device can be determined based on the number of at least one target first pixel region in the first image.

[0139] In some examples, the number of target first pixel regions corresponding to each image information can be the number of first items identified in that image information.

[0140] In some examples, the number of first items detected by each acquisition device may be the same or different due to occlusion or other reasons. For example, the target first pixel area determined by the first acquisition device is 3, that is, the number of first items in the first acquisition device's view can be 3; the target first pixel area determined by the second acquisition device's view is 2, that is, the number of first items in the second acquisition device's view can be 2.

[0141] For example, after obtaining the first item quantity from the perspective of each acquisition device, the first item quantities from all acquisition device perspectives can be fused across perspectives to determine the first item quantity picked up by the pick-up and drop-off device. Here, the first quantity is the global calculation result determined by the first branch.

[0142] In some embodiments, the picking anomaly detection method may further include: comparing the size of the number of first items in each image information, and determining the largest number of first items as the first quantity.

[0143] For example, after determining the number of first items calculated by each of the acquisition devices, the control device can compare the number of first items and determine the largest number of first items as the first quantity, that is, determine the largest number of first items as the number of items determined by the first branch.

[0144] In some examples, besides determining the maximum quantity of the first item as the first quantity, embodiments of this application may also determine the first quantity in other ways. For example, other fusion strategies such as weighted voting, majority voting, and probability distribution weighted averaging may be used to fuse multiple quantities of the first item to obtain the first quantity, and embodiments of this application do not limit this.

[0145] Step 340: Based on multiple candidate pixel regions corresponding to each image information and a second preset model, determine the second number of items picked up by the picking and placing device.

[0146] It should be noted that the execution order of steps 330 and 340 is not limited in the embodiments of this application. Steps 330 and 340 are two steps that are executed in parallel. Step 340 can be executed at the same time as step 330, or step 340 can be executed after step 330, or step 340 can be executed before step 330.

[0147] In some examples, step 340 is a second branch for determining the number of items picked up by the pick-up and place-down device. This second branch can also be called a quick decision branch two. Branch two can determine the number of items picked up by the pick-up and place-down device based on the gripping area of ​​the pick-up and place-down device determined in step 630 above and the second preset model. This can be called the second quantity.

[0148] In some examples, the second preset model can be a pre-trained network model, such as a classification model used to probabilistically classify the number of items in an image.

[0149] The following is combined Figure 7 The implementation process of branch two is explained.

[0150] Figure 7 This is a schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application. For example... Figure 7 As shown, step 340 above includes steps 341 to 344 as shown below.

[0151] Step 341: Based on the gripping area corresponding to the picking and placing device in each image information, cropping is performed on each image information to obtain a cropped image.

[0152] For example, the control device can perform cropping processing on each image information based on the gripping area corresponding to the picking and placing device in each image information to obtain the cropped image corresponding to each image information.

[0153] In some examples, the control device can determine the minimum bounding rectangle of the gripping area and, based on this minimum bounding rectangle, crop a partial image containing only the gripping area from the image information (i.e., the original image). The minimum bounding rectangle can be the smallest rectangle capable of completely encompassing all pixels within the gripping area.

[0154] It should be noted that by cropping, the focus can be placed on the local area around the end mechanism of the picking and placing device, and the interference in the background can be eliminated. This makes it easier for the subsequent second preset model to focus on the actual area of ​​the object being held, avoiding the influence of background interference on the classification results, thereby improving the accuracy of classification.

[0155] Step 342: Input each cropped image into the subclassification network corresponding to each acquisition device, so as to output the probability distribution of the number of items in each image information at multiple preset quantity levels through each subclassification network.

[0156] For example, the second preset model may include multiple sub-classification networks. The number of sub-classification networks is related to the number of acquisition devices (i.e., multiple acquisition devices) performing the picking detection. Each sub-classification network can independently process the image information acquired by each acquisition device.

[0157] For example, taking a data acquisition device with four data acquisition devices as an example, the second preset model may include four parallel sub-classification networks. Each sub-classification network processes the image information acquired by one data acquisition device to output the probability distribution of the number of items in the image information at multiple preset quantity levels.

[0158] For example, multiple preset quantity levels can be pre-set as needed. These preset quantity levels include at least three levels: a first level, a second level, and a third level, with different preset quantity levels indicating different quantities. For instance, the first level indicates that the quantity of the second item is zero, the second level indicates that the quantity of the second item is one, and the third level indicates that the quantity of the second item is multiple.

[0159] In some examples, the probability distribution of item quantities across multiple preset quantity levels can indicate the probability or confidence level of item quantities satisfying each preset quantity level; that is, an estimate of the probability that item quantities belong to each preset quantity level. For example, the probability distribution could include the confidence level of item quantities belonging to a first level, the confidence level of item quantities belonging to a second level, and the confidence level of item quantities belonging to a third level.

[0160] It should be noted that the embodiments of this application use multiple preset quantity levels, including the first level, the second level and the third level, as examples for illustrative purposes. In actual applications, more or fewer preset quantity levels can be set according to needs, and the embodiments of this application do not limit this.

[0161] Step 343: Based on each probability distribution, determine the number of second items in each image information.

[0162] In some examples, after outputting the probability distribution of the number of items in each image information at multiple preset quantity levels through each subclassification network, the number of items in the image information can be determined based on the probability distribution, which is called the second item quantity.

[0163] In some embodiments, step 343 includes: determining the probability distribution of the number of items in the first image acquired by the first acquisition device across multiple preset quantity levels; determining the preset quantity level with the largest probability distribution as the target quantity level in the first image; and determining the number of second items in the first image based on the target quantity level.

[0164] In some examples, the first acquisition device is any one of multiple acquisition devices. Taking the first acquisition device as an example, the number of first items in the first image acquired by the first acquisition device can be determined.

[0165] For example, the classification subnetwork corresponding to the first acquisition device can input the probability distribution of the number of items in the first image across three preset quantity levels. For instance, the confidence level for the number of items in the first image belonging to the first level is the first confidence level, the confidence level for the number of items in the first image belonging to the second level is the second confidence level, and the confidence level for the number of items in the first image belonging to the third level is the third confidence level. The preset quantity level corresponding to the maximum value among the first, second, and third confidence levels is determined as the target quantity level. After determining the target quantity level, the quantity corresponding to the target quantity level can be determined as the second item quantity.

[0166] For example, when the second confidence level is the highest, the target quantity level is the second level, and the second item quantity is one corresponding to the second level; when the third confidence level is the highest, the target quantity level is the third level, and the second item quantity is multiple corresponding to the third level.

[0167] Step 344: Determine the second quantity based on the quantity of the second item in each image information.

[0168] In some examples, similar to the number of second items in the first image, the number of second items in each image information can be determined, and the number of items picked up by the pick-and-place device determined by the second branch, i.e., the second quantity, can be obtained based on the number of second items in each image information.

[0169] In some embodiments, the number of second items in each image information is compared, and the largest number of second items is determined as the second quantity.

[0170] For example, after determining the number of second items calculated by each of the acquisition devices, the control device can compare the number of second items and determine the largest number of second items as the second quantity, that is, determine the largest number of second items as the quantity determined by the second branch.

[0171] In some examples, similar to the first branch, besides determining the maximum quantity of the second item as the second quantity, embodiments of this application may also determine the second quantity in other ways. For example, other fusion strategies such as weighted voting, majority voting, and probability distribution weighted averaging may be used to fuse multiple quantities of the first item to obtain the second quantity, and embodiments of this application do not limit this.

[0172] In some examples, the sub-classification network described above can distinguish whether the number of items picked up by the pick-up and drop-off device is single or multiple. However, when it determines that the number of items belongs to the third level, i.e., multiple items, it cannot accurately distinguish the specific number of items. Therefore, when the sub-classification network determines that the number of items picked up in the image information (i.e., the second item quantity) is multiple, the quantity of items can be further verified based on depth information to further improve the accuracy of the quantity calculation.

[0173] Figure 8 This is a schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application. For example... Figure 8 As shown, the method further includes steps 810 to 840 as shown below.

[0174] Step 810: If the second acquisition device is a depth image acquisition device, acquire the second image acquired by the second acquisition device and determine the depth point cloud information corresponding to the second image.

[0175] For example, when a depth acquisition device (hereinafter referred to as the second acquisition device) with depth sensing capability is included among multiple acquisition devices, such as a depth camera, a structured light camera or a TOF camera, the control device determines the depth point cloud information in the second image when it acquires a depth image (such as the second image) acquired by the second acquisition device.

[0176] In some examples, the depth point cloud information includes the coordinates (x, y, z) of each pixel in the second image in three-dimensional space, where x and y represent the horizontal position and z represents the depth value, i.e., the distance of the pixel from the camera.

[0177] In some examples, taking the second acquisition device among multiple acquisition devices as a depth acquisition device, and the image information it acquires as the second image, the number of second items in the second image can be determined through steps 341 to 343 above. If there are multiple second items (i.e., belonging to the third level), the number of second items is further verified based on the depth point cloud information of the second image.

[0178] Step 820: Based on the gripping area of ​​the pick-and-place device in the second image, the spatial range of the depth point cloud information is extracted, and the extracted point cloud is determined as the point cloud within the gripping area of ​​the pick-and-place device.

[0179] For example, after determining the gripping area corresponding to the picking and placing device in the second image through the above steps 610 to 630, the gripping area can be used as a spatial template and superimposed on the depth point cloud information of the second image to perform a mask cropping operation, thereby extracting the point cloud in the gripping area.

[0180] In some examples, for each pixel (u, v) in the gripping area, the control device can find the corresponding 3D point (x, y, z) in the depth point cloud information, extract the 3D point, and add it to the point cloud set of the gripping area. This allows the point cloud belonging only to the gripping area to be extracted from the panoramic point cloud containing a large number of background point clouds, avoiding the miscalculation of the point clouds of other objects in the background into the clustering results.

[0181] Step 830: Perform three-dimensional spatial clustering on the point cloud within the holding area to determine the number of point cloud clusters within the holding area.

[0182] For example, the control device performs Euclidean distance clustering on the point cloud extracted within the gripping area. Euclidean distance clustering is a spatial distance-based clustering algorithm. In three-dimensional space, point clouds belonging to the same physical rigid body are relatively close to each other, while point clouds belonging to different physical rigid bodies have larger spatial gaps. Euclidean distance clustering can cluster point clouds with small distances (e.g., less than a preset distance threshold) into a single point cloud cluster, thereby obtaining multiple point cloud clusters within the gripping area and determining the number of point cloud clusters within the gripping area.

[0183] Step 840: Determine the number of second items in the second image based on the probability distribution of the number of items in the second image across multiple preset quantity levels and the number of point cloud clusters within the holding area.

[0184] In some examples, the number of items in the second image can be estimated based on the number of point cloud clusters in the holding area; that is, the number of point cloud clusters in the holding area can be used as the number of items in the second image determined by cluster analysis.

[0185] For example, the control device fuses the number of items determined by the subclassification network and the number of items determined by cluster analysis to determine the number of second items in the second image.

[0186] In some examples, when the probability distribution indicates that there are multiple items in the second image, i.e., the confidence level is the highest at the third level, if the number of point cloud clusters in the holding area is greater than 1, it indicates that the number of items determined by the subclassification network is consistent with the number of items determined by the cluster analysis, and the number of point cloud clusters can be determined as the number of second items.

[0187] For example, if the number of items in the second image belongs to the third level with the highest confidence, and the number of point cloud clusters in the holding area is 3, then the number of second items in the second image can be determined to be 3.

[0188] In other examples, when the probability distribution indicates that there are multiple items in the second image, i.e., the confidence level is the highest at the third level, if the number of point cloud clusters in the holding area is 1, it indicates that the number of items determined by the subclassification network is inconsistent with the number of items determined by the cluster analysis. In this case, it can be determined that there are multiple second items in the second image.

[0189] It should be noted that steps 810 to 840 above are supplementary calculations that can be performed when multiple acquisition devices include depth cameras. If there are no depth cameras among the multiple acquisition devices, steps 810 to 840 can be skipped, and the number of the second item can be determined directly based on the probability distribution output by the subclassification network.

[0190] Step 350: If the first quantity and the second quantity are the same, determine the picking result of the picking and placing device, and control the picking and placing device to perform the picking operation based on the picking result.

[0191] For example, after determining the first quantity through branch one of step 330 and the second quantity through branch two of step 340, the first quantity and the second quantity can be compared, and the picking result of the picking and placing device can be determined based on the comparison result.

[0192] In some examples, the first quantity can be the same as the second quantity when the comparison result is that the first quantity and the second quantity are the same. For example, the first quantity can be one and the second quantity can be multiple.

[0193] In some embodiments, when both the first quantity and the second quantity indicate that the number of items picked up by the pick-up and place device is one, the pick-up result of the pick-up and place device is determined to be a normal pick-up, and the pick-up and place device is controlled to transfer the picked-up item to the target location.

[0194] For example, when both the first quantity and the second quantity are 1, it indicates that the pick-and-place device has only picked up one item, and no double-pickup anomaly has occurred. Therefore, it can be determined that the picking result of the pick-and-place device is a normal pick. In the case of normal picking, the control device can send a first control command to the pick-and-place device to control the pick-and-place device to transfer the picked item to the target location, thereby completing the picking operation.

[0195] For example, the target location can be an order container, a designated workstation, or a conveyor line, etc., and this application embodiment does not limit this.

[0196] In some embodiments, when both the first quantity and the second quantity indicate that the number of items picked up by the pick-up and place device is multiple, the picking result of the pick-up and place device is determined to be a picking anomaly, and the pick-up and place device is controlled to transfer the picked-up items to the target container and pick up the target items again from the target container.

[0197] For example, when both the first and second quantities are greater than 1, it indicates that the pick-and-place device has picked up multiple items, meaning a double-picking anomaly has occurred. Therefore, the picking result of the pick-and-place device can be determined to be a picking anomaly. In the event of a picking anomaly, the control device can send a second control command to the pick-and-place device to control it to transfer all picked-up items back to the target container and release them. Afterward, the control device can then re-execute the picking operation for the target item from the target container. For example, steps 310 to 350 can be re-executed until it is determined that the pick-and-place device has picked up one item.

[0198] The picking anomaly detection method provided in this application firstly acquires picking images of the picking device from different angles simultaneously using multiple acquisition devices, overcoming the problem of some items being invisible due to occlusion under a single viewpoint, and providing a more comprehensive data foundation for subsequent item counting. Secondly, based on the mask information and pixel-by-pixel probability map of the output image from the first preset model (i.e., instance segmentation model), the picking device itself is accurately filtered, effectively avoiding false alarms caused by components on the picking device being misidentified as items. Then, the number of items picked by the picking device can be quickly determined through two parallel branches, branch one and branch two. When the number of items determined by the two branches is consistent, it is determined whether a double-picking anomaly has occurred based on the number of items. This application determines the number of picked items through a dual-branch method, which can avoid false detections and improve the accuracy and reliability of robot picking operations. Compared with the double-picking anomaly detection methods in related technologies, the detection method of this application has a faster response speed and is suitable for high-speed picking scenarios.

[0199] Figure 9 This is a schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application. For example... Figure 9 As shown, the picking anomaly detection method may further include steps 910 to 920 as shown below. It should be noted that step 910 can be performed after steps 330 and 340 above.

[0200] Step 910: If the first quantity and the second quantity are inconsistent, determine the third quantity of items picked up by the picking and placing device based on the gripping area corresponding to the picking and placing device, the preset prompt words, and the third preset model in each image information.

[0201] For example, when the first quantity and the second quantity are inconsistent, it indicates that there is a discrepancy in the calculation results of the two fast decision branches (i.e., branch one and branch two) in steps 330 and 340. In this case, to avoid false detection, the control device can trigger a slow channel to perform semantic verification.

[0202] In some examples, the inconsistency between the first quantity and the second quantity can include: the first quantity being one and the second quantity being multiple; or the first quantity being multiple and the second quantity being one.

[0203] In some embodiments, step 910 includes: determining a target image from multiple image information, cropping each target image based on the holding region in the target image to obtain a local image of the holding region corresponding to the target image. The local image of the holding region and a preset prompt word are input into a third preset model to obtain a third quantity output by the third preset model.

[0204] For example, multiple acquisition devices can acquire multiple image information, and the control device can select a target image from the multiple image information for verification through a slow channel. The target image can be an image with an optimal viewpoint among the multiple image information.

[0205] For example, the target image can be the image information that includes the largest number of first pixel regions, or it can be the image information that covers the most complete holding area, or it can be the image information with the highest clarity. This application embodiment does not limit this.

[0206] For example, after determining the target image, the target image can be cropped based on the holding region in the target image to obtain a partial image of the holding region. Then, the partial image of the holding region corresponding to the target image and the preset prompt word are input into a third preset model, and the third quantity in the target image is output by the third preset model.

[0207] It should be noted that determining the gripping area corresponding to the target image is similar to steps 610 to 630 above, and cropping the target image is similar to step 341 above. To avoid repetition, it will not be described again here.

[0208] In some examples, the third preset model is a large multimodal model with image understanding and natural language understanding capabilities, such as a Multimodal Vision-Language Model (VLM). The third preset model can process both image and text inputs simultaneously and understand the image content at a high-level semantic level, thereby outputting the number of items included in the local image of the holding region, i.e., the third quantity.

[0209] In some examples, the preset prompts are natural language query statements, the content of which can be dynamically configured according to the business scenario. For example, the preset prompts can be set to "Please identify how many items are held by the robot's end effector in this image?", or "How many individual items are gripped below the gripper in this image?", etc., and this application embodiment does not limit this.

[0210] Step 920: Based on the third quantity, determine the picking result of the pick-and-place device, and control the pick-and-place device to perform the picking operation based on the picking result.

[0211] In some examples, where the first quantity differs from the second quantity, a third quantity determined by a third preset model can be used as the final determined quantity of items. Therefore, the control device can determine the picking result of the picking and placing device based on the third quantity.

[0212] In some embodiments, if the third quantity is one, and the pickup result of the pickup device is determined to be a normal pickup, the pickup device is controlled to transfer the picked-up item to the target location. If the third quantity is greater than one, and the pickup result of the pickup device is determined to be a picking anomaly, the pickup device is controlled to transfer the picked-up item to the target container and pick up the target item again from the target container.

[0213] In some examples, when the third quantity is 1, it indicates that the slow channel semantic verification confirms that the pick-and-place device has picked up an item. The control device determines that the picking result is a normal picking and controls the pick-and-place device to transfer the picked item to the target location to complete the picking operation.

[0214] In some examples, if the third quantity is greater than 1, it indicates that the slow channel semantic review confirms that the pick-and-place device has picked up multiple items. The control device determines that the picking result is a picking anomaly, that is, the pick-and-place device has experienced a double picking anomaly. The control device triggers anomaly handling, controls the pick-and-place device to put all the picked items back into the target container, and re-executes the picking operation.

[0215] For example, the third preset model can output the confidence level of the third quantity along with the third quantity. When the confidence level of the third quantity is higher than a preset confidence threshold, it indicates that the third quantity output by the third preset model is reliable, and step 920 can continue. When the confidence level of the third quantity is lower than the preset confidence threshold, it indicates that the third quantity output by the third preset model is unreliable, and the control device may not accept the third quantity and control the output device in the warehousing system to output an abnormal prompt message to trigger manual review.

[0216] The picking anomaly detection method provided in this application can trigger a slow-channel review of the multimodal visual language model for accurate verification when the decision results of the two fast branches are inconsistent, thereby improving detection accuracy and correctness. This application embodiment can complete the decision with millisecond-level latency in most scenarios through two fast branches, and only triggers the slow-channel review of the multimodal visual language model as needed when the results are inconsistent. This concentrates the additional computational overhead on a few difficult scenarios, thus ensuring industrial-grade real-time performance while also considering detection accuracy in extreme scenarios, significantly improving the reliability and efficiency of robot picking operations in warehousing and logistics scenarios.

[0217] Figure 10 This is a schematic diagram illustrating another method for picking up anomalies provided in an embodiment of this application. For example... Figure 10 As shown, the anomaly detection method includes steps 1010 to 1060 as shown below. The following is in conjunction with... Figure 10 A specific implementation method provided by the embodiments of this application will be described.

[0218] In some examples, after the robot completes its grasping action, multiple vision cameras simultaneously capture images of the grasping scene from different angles. The system can create an independent inference namespace for each camera, allowing each camera to independently run instance segmentation model inference to obtain object masks and pixel-by-pixel robot probability maps, thereby filtering out the robot itself. Based on this, the system computes the results of two fast decision branches in parallel from each camera's perspective: Branch 1 calculates the number of grasped objects based on object mask inflation and adjacency matrix exponentiation; Branch 2 calculates the number of grasped objects based on object quantity probability distribution classification and combined with 3D spatial information provided by the depth camera. The two branches take the maximum value across all cameras to complete multi-camera fusion, resulting in the fusion result Count_A from Branch 1 and the fusion result Count_B from Branch 2. If the results of the two branches are consistent, the decision result can be directly output; if the results of the two branches are inconsistent, a slow channel based on a multimodal large model can be triggered for semantic verification, and the final decision is output based on the verification result. The specific process is as follows: Step 1010: Simultaneously capture scene images using multiple cameras.

[0219] Step 1020, Branch 1: Multi-camera segmentation reasoning, robot body filtering, calculate the number of objects held (Count_A).

[0220] In some examples, each camera triggers simultaneous shooting and then independently runs segmentation inference, outputting an object mask and a pixel-by-pixel robot probability map. For each candidate object, the average robot probability within its visible region is calculated; objects exceeding a preset threshold are identified as the robot itself and removed from the count. Further connected component analysis can be performed on the robot mask, retaining only the largest connected regions to suppress segmentation noise, thereby obtaining the robotic arm's gripping area.

[0221] In some examples, a branch pair performs morphological dilation on all non-robot object masks and the robot mask, causing adjacent but non-overlapping regions to overlap, and constructs an object adjacency matrix. The transitive closure of the adjacency matrix is ​​then calculated to obtain the set of objects directly or indirectly adjacent to the robot. The size of this set of objects is an estimate of the number of goods held by the robot from that camera's perspective.

[0222] Step 1030, Branch 2: Object Quantity Probability Distribution Classification. Combine the 3D bounding box / point cloud information from the depth camera to calculate the number of held objects, Count_B.

[0223] In some examples, branch two can output the probability of quantities such as parts and multiple items for the robotic arm's gripping area. When the probability of multiple items exceeds a preset probability threshold, the quantity of items can be determined to be two items; when the probability of parts exceeds a preset probability threshold, the quantity of items can be determined to be parts; otherwise, the quantity of items can be determined to be one item.

[0224] Step 1040: Are the branch results consistent?

[0225] Compare the result Count_A from branch one with the result Count_B from branch two to determine if Count_A and Count_B are the same. If Count_A and Count_B are the same, proceed to step 1006; if Count_A and Count_B are not the same, proceed to step 1005.

[0226] Step 1050: Perform semantic verification using a modal visual language model.

[0227] In some examples, when the decisions in branch one and branch two are inconsistent, a multimodal visual language model (VLM) can be triggered for semantic-level review. The current grasping scene image, along with a natural language query (such as "How many items are held by the robot's end effector in the image?"), is input into the VLM, which independently outputs a quantity estimate from a semantic understanding perspective. The quantity estimate output by the VLM serves as the basis for the final decision.

[0228] For example, manual review can be triggered when the confidence level of the number of VLM outputs is insufficient.

[0229] Step 1060: Output the decision result and execute the picking decision.

[0230] For example, the number of objects held by the robot can be used to determine whether double picking has occurred. If double picking occurs, the robot can put the objects back and pick them up again.

[0231] Figure 11 This is a schematic diagram of an anomaly detection device provided in an embodiment of this application. Figure 11 As shown, the abnormal pickup detection device 1100 includes an acquisition module 1110, a determination module 1120, and a control module 1130. Wherein: The acquisition module 1110 is configured to acquire multiple image information corresponding to the pick-up and place device through multiple acquisition devices after the pick-up and place device picks up the target item from the target container.

[0232] The determining module 1120 is configured to: determine multiple candidate pixel regions corresponding to each image information based on multiple image information and a first preset model; wherein, the multiple candidate pixel regions include a first pixel region corresponding to the item picked up by the picking and placing device and a second pixel region corresponding to the body of the picking and placing device; determine a first number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each image information; determine a second number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each image information and a second preset model; and determine the picking result of the picking and placing device if the first number and the second number are consistent.

[0233] The control module 1130 is configured to control the pick-and-place device to perform picking operations based on the picking results.

[0234] In some embodiments, the determining module 1120 is configured to: determine that the picking result of the picking device is normal picking when both the first quantity and the second quantity indicate that the number of items picked up by the picking device is one; the control module 1130 is configured to: control the picking device to transfer the picked item to the target location. Alternatively, the determining module 1120 is configured to: determine that the picking result of the picking device is abnormal when both the first quantity and the second quantity indicate that the number of items picked up by the picking device is multiple; the control module 1130 is configured to: control the picking device to transfer the picked item to the target container and pick up the target item again from the target container.

[0235] In some embodiments, the determining module 1120 is further configured to: perform dilation processing on the first pixel region and the second pixel region corresponding to each image information to obtain a plurality of effective pixel regions corresponding to each image information; wherein, the plurality of effective pixel regions include the dilated first pixel region and the dilated second pixel region; determine at least one target first pixel region that is directly or indirectly adjacent to the dilated second pixel region based on the pixel overlap relationship between each effective pixel region; and generate a gripping area corresponding to the picking and placing device in each image information based on the at least one target first pixel region.

[0236] In some embodiments, the determining module 1120 is further configured to: in the case that the first quantity and the second quantity are inconsistent, determine a third quantity of items picked up by the picking and placing device based on the gripping area corresponding to the picking and placing device in each image information, a preset prompt word, and a third preset model; and determine the picking result of the picking and placing device based on the third quantity. The control module 1130 is further configured to: control the picking and placing device to perform a picking operation based on the picking result.

[0237] In some embodiments, the determining module 1120 is configured to: determine a target image from multiple image information; perform cropping processing on each target image based on the holding region in the target image to obtain a local image of the holding region corresponding to the target image; input the local image of the holding region and a preset prompt word into a third preset model to obtain a third quantity output by the third preset model.

[0238] In some embodiments, the determining module 1120 is configured to: determine that the picking result of the picking and placing device is normal picking when the third quantity is one; the control module 1130 is configured to: control the picking and placing device to transfer the picked item to the target location. The determining module 1120 is configured to: determine that the picking result of the picking and placing device is picking abnormal when the third quantity is greater than one; the control module 1130 is configured to: control the picking and placing device to transfer the picked item to the target container and pick up the target item again from the target container.

[0239] In some embodiments, the determining module 1120 is configured to: determine the number of first items in each image information based on the number of at least one target first pixel region; and determine the first number of items picked up by the picking and placing device based on the number of first items in each image information.

[0240] In some embodiments, the determining module 1120 is configured to: compare the size of the first item quantity in each image information and determine the largest first item quantity as the first quantity.

[0241] In some embodiments, the determining module 1120 is configured to: crop each image based on the gripping area corresponding to the picking and placing device in each image to obtain a cropped image; input each cropped image into a sub-classification network corresponding to each acquisition device to output a probability distribution of the number of items in each image at multiple preset quantity levels through each sub-classification network; wherein, the second preset model includes multiple sub-classification networks; determine the number of second items in each image based on each probability distribution; and determine the second quantity based on the number of second items in each image.

[0242] In some embodiments, the determining module 1120 is configured to: determine the probability distribution of the number of items in a first image acquired by a first acquisition device in multiple preset quantity levels; wherein the first acquisition device is any one of multiple acquisition devices; determine the preset quantity level with the largest probability distribution as the target quantity level in the first image; and determine the number of second items in the first image based on the target quantity level; wherein the multiple preset quantity levels include a first level, a second level, and a third level, the first level indicating that the number of second items is zero, the second level indicating that the number of second items is one, and the second quantity level indicating that the number of second items is multiple.

[0243] In some embodiments, the determining module 1120 is configured to: compare the size of the second item quantity in each image information and determine the largest second item quantity as the second quantity.

[0244] In some embodiments, the acquisition module 1110 is configured to: acquire a second image acquired by the second acquisition device when the second acquisition device is a depth image acquisition device. The determination module 1120 is further configured to: determine the depth point cloud information corresponding to the second image; wherein the plurality of acquisition devices includes the second acquisition device. Based on the gripping area of ​​the pick-and-place device in the second image, spatial range extraction is performed on the depth point cloud information, and the extracted point cloud is determined as the point cloud within the gripping area of ​​the pick-and-place device; three-dimensional spatial clustering processing is performed on the point cloud within the gripping area to determine the number of point cloud clusters within the gripping area. Based on the probability distribution of the number of items in the second image at multiple preset quantity levels and the number of point cloud clusters within the gripping area, the number of second items in the second image is determined.

[0245] In some embodiments, the determining module 1120 is configured to: acquire a preset region of interest corresponding to each of the plurality of acquisition devices; input the image information corresponding to each of the plurality of acquisition devices and the preset region of interest corresponding to each acquisition device into a sub-segmentation network corresponding to each acquisition device, so as to output mask information and pixel-wise probability map corresponding to each image information through each sub-segmentation network; wherein, the first preset model includes multiple sub-segmentation networks, the mask information indicates the pixel region corresponding to the candidate object in the image information, and the pixel-wise probability map indicates the probability that each pixel in the image information belongs to the body of the picking and placing device; and determine multiple candidate pixel regions corresponding to each image information based on the mask information and pixel-wise probability map corresponding to each image information.

[0246] In some embodiments, the determining module 1120 is configured to: determine a candidate region corresponding to at least one candidate object in the first image based on the mask information corresponding to the first image acquired by the first acquisition device; wherein the first acquisition device is any one of a plurality of acquisition devices; determine the pixel probability of each pixel in each candidate region belonging to the body of the picking and placing device based on each candidate region and the pixel-by-pixel probability map corresponding to the first image; and filter out the first pixel region and the second pixel region corresponding to the first image from at least one candidate region based on the pixel probability of each pixel in each candidate region and a preset probability threshold.

[0247] In some embodiments, the determining module 1120 is configured to: determine the average pixel probability of the candidate regions based on the pixel probability of each pixel in each candidate region, delete the candidate regions whose average pixel probability exceeds a preset probability threshold, and obtain the first pixel region corresponding to the first image; in the first image, determine the pixels whose pixel probability exceeds the preset probability threshold as the main pixels of the pick-and-place device, and determine the second pixel region based on the main pixels of the pick-and-place device.

[0248] In some embodiments, the determining module 1120 is configured to: perform connected component processing on the body pixels of the picking and placing device to determine at least one connected component corresponding to the picking and placing device; determine the pixel area of ​​each connected component, and determine the connected component with the largest pixel area as the second pixel region in the first image.

[0249] In some embodiments, the acquisition module 1110 is configured to: acquire multiple image information corresponding to the end mechanism through multiple acquisition devices when the end mechanism of the pick-and-place device picks up the target item from the target container and transfers the target item to a preset acquisition area; wherein, the multiple acquisition devices include multiple acquisition devices set in the working area and / or at least one acquisition device on the pick-and-place device.

[0250] It should be noted that the picking anomaly detection device provided in this application embodiment is used to execute the corresponding picking anomaly detection method provided above. Therefore, the beneficial effects it can achieve can be referred to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0251] Figure 12 This is a schematic diagram of an electronic device provided in an embodiment of this application. In some embodiments, the electronic device includes one or more processors and a memory. The memory is configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the picking anomaly detection method in the above embodiments.

[0252] like Figure 12 As shown, the electronic device 1000 includes a processor 1001 and a memory 1002. Exemplarily, the electronic device 1000 may also include a communications interface 1003 and a communications bus 1004.

[0253] The processor 1001, memory 1002, and communication interface 1003 communicate with each other via communication bus 1004. Communication interface 1003 is used to communicate with other network elements such as clients or other servers.

[0254] In some embodiments, the processor 1001 is used to execute program 1005, specifically performing the relevant steps in the above-described embodiments of the anomaly detection method. Specifically, program 1005 may include program code, which includes computer-executable instructions.

[0255] For example, processor 1001 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. Electronic device 1000 may include one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0256] In some embodiments, memory 1002 is used to store program 1005. Memory 1002 may include high-speed RAM memory, and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0257] Specifically, program 1005 can be called by processor 1001 to cause electronic device 1000 to perform the operation of picking up the anomaly detection method.

[0258] This application provides a computer-readable storage medium storing at least one executable instruction. When the executable instruction is executed on an electronic device 1000, the electronic device 1000 performs the picking anomaly detection method described in the above embodiment.

[0259] The executable instructions can be used to cause the electronic device 1000 to perform the operation of picking up the anomaly detection method.

[0260] For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0261] In some embodiments, this application provides a computer program product including a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the picking anomaly detection method described in any of the above embodiments.

[0262] In some embodiments, this application also provides a computer program that, when executed by a processor, can implement the picking anomaly detection method described in any of the above embodiments.

[0263] The beneficial effects that the picking anomaly detection device, electronic device, computer-readable storage medium, computer program product, and computer program provided in this application embodiment can achieve are similar to the beneficial effects of the picking anomaly detection method provided above, and will not be repeated here.

[0264] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0265] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0266] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0267] For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0268] More specific examples of computer-readable media (a non-exhaustive list) include the following: electrical connections having one or more wires (electronic devices), portable computer disks (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM).

[0269] Furthermore, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory. It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof.

[0270] In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0271] The embodiments described above do not constitute a limitation on the scope of protection of this application.

Claims

1. A method for detecting anomalies, characterized in that, The method includes: After the pick-up and place device picks up the target item from the target container, multiple image information corresponding to the pick-up and place device is acquired by multiple acquisition devices; Based on the multiple image information and the first preset model, multiple candidate pixel regions corresponding to each of the image information are determined; wherein, the multiple candidate pixel regions include the first pixel region corresponding to the item picked up by the picking and placing device and the second pixel region corresponding to the body of the picking and placing device. Based on the plurality of candidate pixel regions corresponding to each of the image information, a first number of items picked up by the picking and placing device is determined; Based on the multiple candidate pixel regions corresponding to each of the image information and the second preset model, the second number of items picked up by the picking and placing device is determined; If the first quantity and the second quantity are the same, the picking result of the picking and placing device is determined, and the picking and placing device is controlled to perform a picking operation based on the picking result.

2. The method according to claim 1, characterized in that, When the first quantity and the second quantity are the same, determining the picking result of the pick-and-place device, and controlling the pick-and-place device to perform a picking operation based on the picking result, includes: If both the first quantity and the second quantity indicate that the pick-up and place device has picked up one item, the pick-up result of the pick-up and place device is determined to be a normal pick-up, and the pick-up and place device is controlled to transfer the picked-up item to the target location; or... If both the first quantity and the second quantity indicate that the picking device has picked up multiple items, the picking result of the picking device is determined to be a picking anomaly. The picking device is then controlled to transfer the picked items to the target container and pick up the target items again from the target container.

3. The method according to claim 1, characterized in that, After determining the multiple candidate pixel regions corresponding to each of the image information, the method further includes: The first pixel region and the second pixel region corresponding to each of the image information are dilated to obtain a plurality of effective pixel regions corresponding to each of the image information; wherein, the plurality of effective pixel regions include the dilated first pixel region and the dilated second pixel region; Based on the pixel overlap relationship between each effective pixel region, at least one target first pixel region is determined that is directly or indirectly adjacent to the dilated second pixel region. Based on the at least one target first pixel region, a gripping area corresponding to the picking and placing device in each image information is generated.

4. The method according to claim 3, characterized in that, The method further includes: If the first quantity and the second quantity are inconsistent, the third quantity of items picked up by the picking and placing device is determined based on the gripping area corresponding to the picking and placing device, the preset prompt words, and the third preset model in each image information. Based on the third quantity, the picking result of the picking and placing device is determined, and the picking and placing device is controlled to perform picking operation based on the picking result.

5. The method according to claim 4, characterized in that, The determination of the third quantity of items picked up by the picking and placing device based on the gripping area corresponding to the picking and placing device in each image information, the preset prompt words, and the third preset model includes: A target image is determined from the plurality of image information, and each target image is cropped based on the gripping area in the target image to obtain a local image of the gripping area corresponding to the target image; The local image of the gripping area and the preset prompt are input into the third preset model to obtain the third quantity output by the third preset model.

6. The method according to claim 4, characterized in that, The step of determining the picking result of the pick-and-place device based on the third quantity, and controlling the pick-and-place device to perform a picking operation based on the picking result, includes: If the third quantity is one, the pickup result of the pick-up and placement device is determined to be a normal pickup, and the pick-up and placement device is controlled to transfer the picked-up item to the target location; or, If the third quantity is greater than one, the picking result of the picking and placing device is determined to be a picking anomaly. The picking and placing device is then controlled to transfer the picked item to the target container and pick up the target item again from the target container.

7. The method according to claim 3, characterized in that, Determining the first number of items picked up by the picking and placing device based on the plurality of candidate pixel regions corresponding to each of the image information includes: Based on the number of the at least one target first pixel region, determine the number of first items for each of the image information; Based on the first number of items in each of the image information, the first number of items picked up by the picking and placing device is determined.

8. The method according to claim 7, characterized in that, Determining the first number of items picked up by the picking and placing device based on the first number of items in each of the image information includes: Compare the number of first items in each of the image information, and determine the largest number of first items as the first quantity.

9. The method according to claim 3, characterized in that, The determination of the second number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each of the image information and the second preset model includes: Based on the gripping area corresponding to the picking and placing device in each of the image information, each image information is cropped to obtain a cropped image; Each cropped image is input into the sub-classification network corresponding to each acquisition device, so as to output the probability distribution of the number of items in each image information at multiple preset quantity levels through each sub-classification network; wherein, the second preset model includes multiple sub-classification networks; Based on the probability distributions, determine the number of second items in each of the image information; The second quantity is determined based on the number of second items in each of the image information.

10. The method according to claim 9, characterized in that, The step of determining the quantity of the second item in each of the image information based on the probability distribution of the quantity of items in each of the image information at multiple preset quantity levels includes: Determine the probability distribution of the number of items in the first image acquired by the first acquisition device among the multiple preset quantity levels; wherein, the first acquisition device is any one of the multiple acquisition devices; The preset quantity level with the largest probability distribution is determined as the target quantity level in the first image; Based on the target quantity level, the quantity of the second item in the first image is determined; wherein, the plurality of preset quantity levels include a first level, a second level and a third level, the first level indicates that the quantity of the second item is zero, the second level indicates that the quantity of the second item is one, and the second quantity level indicates that the quantity of the second item is multiple.

11. The method according to claim 10, characterized in that, Determining the second quantity based on the quantity of the second items in each of the image information includes: Compare the number of second items in each of the image information, and determine the largest number of second items as the second quantity.

12. The method according to claim 9, characterized in that, The method further includes: When the second acquisition device is a depth image acquisition device, the second image acquired by the second acquisition device is obtained, and the depth point cloud information corresponding to the second image is determined; wherein, the plurality of acquisition devices includes the second acquisition device; Based on the gripping area of ​​the pick-and-place device in the second image, the spatial range of the depth point cloud information is extracted, and the extracted point cloud is determined as the point cloud within the gripping area of ​​the pick-and-place device. Perform three-dimensional spatial clustering on the point cloud within the gripping area to determine the number of point cloud clusters within the gripping area; Determining the number of second items in each of the image information based on each of the probability distributions includes: The number of second items in the second image is determined based on the probability distribution of the number of items in the second image across multiple preset quantity levels and the number of point cloud clusters within the gripping area.

13. The method according to any one of claims 1-12, characterized in that, The step of determining multiple candidate pixel regions corresponding to each of the multiple image information and the first preset model includes: Obtain the preset interest area corresponding to each of the plurality of acquisition devices; The image information corresponding to each of the plurality of acquisition devices, and the preset interest region corresponding to each acquisition device, are input into the sub-segmentation network corresponding to each acquisition device, so as to output the mask information and pixel-wise probability map corresponding to each image information through each sub-segmentation network; wherein, the first preset model includes a plurality of sub-segmentation networks, the mask information indicates the pixel region corresponding to the candidate object in the image information, and the pixel-wise probability map indicates the probability that each pixel in the image information belongs to the body of the picking and placing device; Based on the mask information corresponding to each of the image information and the pixel-by-pixel probability map, the plurality of candidate pixel regions corresponding to each of the image information are determined.

14. The method according to claim 13, characterized in that, The step of determining the plurality of candidate pixel regions corresponding to each of the image information based on the mask information corresponding to each of the image information and the pixel-by-pixel probability map includes: Based on the mask information corresponding to the first image acquired by the first acquisition device, a candidate region corresponding to at least one candidate object in the first image is determined; wherein, the first acquisition device is any one of the plurality of acquisition devices; Based on the pixel-by-pixel probability map corresponding to each candidate region and the first image, the pixel probability of each pixel in each candidate region belonging to the body of the picking and placing device is determined. Based on the pixel probability of each pixel in each of the candidate regions and a preset probability threshold, the first pixel region and the second pixel region corresponding to the first image are selected from the at least one candidate region.

15. The method according to claim 14, characterized in that, The step of selecting the first pixel region and the second pixel region corresponding to the first image from the at least one candidate region based on the pixel probability of each pixel in each of the candidate regions and a preset probability threshold includes: Based on the pixel probability of each pixel in each candidate region, the average pixel probability of the candidate region is determined, and candidate regions with an average pixel probability exceeding the preset probability threshold are deleted to obtain the first pixel region corresponding to the first image. In the first image, pixels with a probability exceeding the preset probability threshold are identified as the main pixels of the pick-and-place device, and the second pixel region is determined based on the main pixels of the pick-and-place device.

16. The method according to claim 15, characterized in that, Determining the second pixel region based on the body pixels of the pick-and-place device includes: Perform connected component processing on the main pixels of the pick-and-place device to determine at least one connected component corresponding to the pick-and-place device; Determine the pixel area of ​​each connected component, and identify the connected component with the largest pixel area as the second pixel region in the first image.

17. The method according to any one of claims 1-12, characterized in that, The process of acquiring multiple image information corresponding to the pick-and-place device through multiple acquisition devices includes: When the end mechanism of the picking and placing device picks up the target item from the target container and transfers the target item to the preset collection area, multiple image information corresponding to the end mechanism is acquired by the multiple collection devices. The plurality of acquisition devices include a plurality of acquisition devices set in the working area and / or at least one acquisition device on the pick-and-place device.

18. An anomaly detection device, characterized in that, include: The acquisition module is configured to acquire multiple image information corresponding to the pick-up and place device through multiple acquisition devices after the pick-up and place device picks up the target item from the target container; The determining module is configured to: determine multiple candidate pixel regions corresponding to each of the multiple image information and a first preset model; wherein the multiple candidate pixel regions include a first pixel region corresponding to the item picked up by the picking and placing device and a second pixel region corresponding to the body of the picking and placing device; determine a first number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each of the image information; determine a second number of items picked up by the picking and placing device based on the multiple candidate pixel regions corresponding to each of the image information and a second preset model; and determine the picking result of the picking and placing device when the first number and the second number are consistent. The control module is configured to control the pick-and-place device to perform a picking operation based on the picking result.

19. An electronic device, characterized in that, include: One or more processors and memory; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the picking anomaly detection method according to any one of claims 1-17.

20. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the picking anomaly detection method according to any one of claims 1-17.

21. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor, implements the picking anomaly detection method according to any one of claims 1-17.