Autonomous system and method performed by the autonomous system
By using depth cameras and clustering technology to determine whether the bin is empty, the problem of robots struggling to identify the state of their workspace in existing technologies is solved, improving grasping efficiency and protecting the reliability of robot components.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SIEMENS AG
- Filing Date
- 2023-03-24
- Publication Date
- 2026-04-17
AI Technical Summary
Existing robotic grasping methods struggle to effectively identify whether the workspace is empty in dynamic environments, leading to unnecessary grasping calculations and potential robot damage.
Depth images are generated using a depth camera, and k-means clustering is performed using a depth calculation module to determine the bottom position of the bin. The clustering map is evaluated using a workspace module to determine whether the bin is empty, thus avoiding unnecessary grabbing calculations.
It improves the efficiency and reliability of robot grasping, avoids unnecessary grasping calculations, saves time, and protects robot components.
Smart Images

Figure CN116803631B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of robotics, and more specifically to autonomous systems and methods performed by autonomous systems. Background Technology
[0002] Autonomous operation, such as robotic grasping and manipulation, faces various technical challenges in unknown or dynamic environments. Autonomous operation in dynamic environments can be applied to mass customization (e.g., high-mix, low-volume manufacturing), on-demand flexible manufacturing processes in smart factories, warehouse automation in smart stores, and automated delivery from distribution centers in smart logistics. To perform autonomous operations, such as grasping and manipulation, robots can, in some cases, use machine learning to learn skills, particularly deep neural networks or reinforcement learning.
[0003] In particular, for example, robots may interact with different objects in different situations. Some of these objects may be unknown to a particular robot. Bin picking is an example of an operation that robots can perform using artificial intelligence. Bin picking refers to a robot grasping objects from a container or bin that can be defined in random or arbitrary postures. The robot can move or transport these objects and place them in different locations for packaging or further processing. However, it should be recognized that current robotic grasping methods lack efficiency and capability. In particular, due to various ongoing technical challenges, current methods often fail to properly or effectively identify objects within certain workspaces of a particular robot. Summary of the Invention
[0004] Embodiments of the present invention address and overcome one or more of the described disadvantages or technical problems by providing methods, systems, and apparatus for determining, during robot runtime, whether a bin or container is empty or contains an object for robot grasping before performing any grasping calculations related to the bin.
[0005] In one example, the autonomous system includes a robot configured to operate during active industrial runtime to define the runtime. The robot includes an end effector configured to grasp multiple objects within a workspace. The autonomous system may include a depth camera configured to capture depth images of the workspace. The autonomous system also includes a processor and a memory storing instructions that, when executed by the processor, cause the autonomous system to perform various operations. In particular, the system can detect bins within the workspace. A bin can be defined as any container or pallet capable of holding one or more of multiple objects. Based on the depth image and without performing grasping calculations, in various examples, the system can determine whether the bin is empty or contains at least one object. Attached Figure Description
[0006] The above and other aspects of the invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings. For the purpose of illustrating the invention, presently preferred embodiments are shown in the drawings, but it should be understood that the invention is not limited to the specific tools disclosed. The drawings include the following figures:
[0007] Figure 1 An example autonomous system in an example physical environment is shown according to an example implementation, the autonomous system including a bin capable of accommodating various objects.
[0008] Figure 2 An example computing system is shown according to an example implementation of an autonomous system configured to determine whether a bin is empty.
[0009] Figure 3 A computing environment in which embodiments of the present disclosure can be implemented is shown. Detailed Implementation
[0010] As a starting point, robotic bin picking typically involves a robot equipped with sensors or cameras that can use a robotic end effector to grasp (pick up) objects in random poses from a container (bin). In the various examples described herein, the objects may be known or unknown to the robot, and the objects may be of the same type or a mixture. In some cases, the robot executes a bin picking algorithm before each pick to calculate and determine the next grasp to be performed. In particular, for example, a particular robotic system may use a deep neural network to calculate the grasp point. However, this paper recognizes that a technical problem involved in robotic grasping is assessing or determining whether a particular workspace of the robot is empty or contains an object. In one implementation, if the workspace is empty, a grasp calculation is not triggered. Alternatively, continuing to exemplify, a grasp calculation is triggered if the workspace, and particularly the bin within the workspace, contains an object. This paper recognizes that current grasping neural networks typically calculate and return grasp calculations regardless of whether the bin is empty. In such systems, when a particular workspace is empty, the quality score associated with the returned grasp may be low. However, this paper further recognizes that calculating a grasp score for an empty workspace can be unreliable. For example, a particular system might not be able to rely on its grasping neural network to correctly assess whether a particular bin is filled or empty. In particular, for tasks involving identifying empty workspaces, the error rate of the neural network might be too high, causing the system to generate false positives. For instance, the system might generate a grasping score (grasp score) associated with an empty bin, indicating that an object is present in the empty bin.
[0011] Furthermore, this paper recognizes that robots or autonomous systems may lose time calculating the grasping score for empty bins. Additionally, when attempting to grasp on empty bins, for example due to the associated grasping score calculation, the robot loses extra time unnecessarily spent on grasping attempts. In some cases, this use may wear down or damage the robot.
[0012] The implementation described herein can classify or determine whether a bin contains an object or is empty; for example, when a bin is empty, a grasping calculation is not performed. In some examples, the system classifies specific bins before each grasping calculation is performed during runtime. Therefore, the system described herein can avoid performing unnecessary grasping calculations, thereby saving processing time and overhead. Additionally, the system described herein can avoid attempting to grasp objects from empty bins based on fuzzy or incorrect grasping calculations, thereby saving operation time and protecting the robot and related components.
[0013] Now for reference Figure 1 This document illustrates an example of an industrial or physical environment or workspace 100. As used herein, a physical environment or workspace can refer to any unknown or dynamic industrial environment. Unless otherwise specified, physical environment and workspace are used interchangeably without limitation. Reconfiguration or modeling can define a virtual representation of the physical environment or workspace 100, or one or more objects 106 within the physical environment 100. For example, object 106 may be placed in a bin or container, such as bin 107, for positioning for gripping. Unless otherwise specified herein, bins, containers, pallets, boxes, or the like are used interchangeably without limitation. For example, object 106 may be picked up from bin 107 by one or more robots and transported or placed in another location, such as outside bin 107. Object 106 in the example is shown as a rectangular object, such as a box, but it is understood that object 106 may be an alternative shape, or an alternative structure may be defined as needed, and all such objects are considered to be within the scope of this disclosure.
[0014] Physical environment 100 may include a computerized autonomous system 102 configured to perform one or more manufacturing operations, such as assembly, transportation, or similar operations. Autonomous system 102 may include one or more robotic devices or autonomous machines, such as autonomous machine or robotic device 104, configured to perform one or more industrial tasks, such as bin picking, gripping, or similar tasks. System 102 may include one or more computing processors configured to process information and control the operation of system 102, particularly the operation of autonomous machine 104. Autonomous machine 104 may include one or more processors, such as processor 108, configured to process information and / or control various operations associated with autonomous machine 104. The autonomous system for operating the autonomous machine in the physical environment may further include memory for storage modules. The processor may be further configured to execute these modules to process information and generate models based on the information. It is understood that the illustrated environment 100 and system 102 are simplified for illustrative purposes. Environment 100 and system 102 may vary as needed, and all such systems and environments are considered to be within the scope of this disclosure.
[0015] Still refer to Figure 1 The autonomous machine 104 may further include a robotic arm or manipulator 110 and a base 112 configured to support the robotic manipulator 110. The base 112 may include wheels 114 or may be configured to otherwise move within the physical environment 100. The autonomous machine 104 may further include an end effector 116 connected to the robotic manipulator 110. The end effector 116 may include one or more tools configured to grasp and / or move object 106. Example end effectors 116 include finger grippers or vacuum-based grippers. The robotic manipulator 110 may be configured to move to change the position of the end effector 116, for example, to place or move object 106 within the physical environment 100. The system 102 may further include one or more cameras or sensors, such as a three-dimensional (3D) point cloud camera 118, configured to detect or record object 106 within the physical environment 100. Camera 118 may be mounted on robot manipulator 110 or configured to otherwise generate a 3D point cloud of a given scene (e.g., physical environment 100). Alternatively or additionally, one or more cameras of system 102 may include one or more standard 2D cameras that can record or capture images (e.g., RGB images or depth images) from different perspectives. These images can be used to construct a 3D image. For example, a 2D camera may be mounted on robot manipulator 110 to capture images from an angle along a given trajectory defined by manipulator 110.
[0016] Still referencing Figure 1Camera 118 can be configured to capture images of bin 107 along a first direction or transverse direction 120, and thus capture images of object 106. In some cases, a deep neural network is trained on a set of objects. Based on its training, the deep neural network can calculate a grasping score for a given area of an object (e.g., an object within bin 107). For example, robotic device 104 and / or system 102 can define one or more neural networks configured to learn various objects in order to identify the pose, grasping point (or location), and / or affordance of various objects that can be found in various industrial environments. According to various example embodiments, example systems or neural network models can be configured to learn objects and grasping positions (e.g., image-based). After the neural network is trained, for example, images of objects can be sent by robotic device 104 to the neural network for classification, particularly the classification of grasping positions or affordances.
[0017] This paper recognizes that in various state-of-the-art methods or algorithms, grasping computation scores from grasping neural networks are used to determine whether a particular bin is empty or contains objects for grasping. Specifically, for example, if a given grasping score associated with a particular bin is below a threshold, that bin is determined to be empty. However, this paper acknowledges that, among other technical drawbacks, this approach can lead to bins that are determined to be empty actually containing objects. For instance, one or more objects may be arranged or positioned in a way that makes them difficult to grasp, thus causing the grasping neural network to calculate a grasping score for the relevant bin below the threshold, thereby incorrectly determining that the bin is empty.
[0018] In another example method for evaluating whether a particular bin is empty (using a computer system), a robot can capture a color image of the bin, and if the bottom of the bin is defined as a uniform color, the bin is determined to be empty when the color image shows a uniform color. However, this paper recognizes that this method can lead to errors. For example, differences in shadow or lighting can cause a bin bottom that appears to contain an object to be actually empty based on observed color or shadow differences. Furthermore, this method requires the bottom of the bin to be defined as a uniform color, which is often not the case with existing bins.
[0019] In another example method for evaluating whether a particular bin is empty (by a computer system), a camera can be positioned at a predetermined distance from the bottom of the bin, and if the system measures a distance from the camera to the bottom of the bin that is less than the predetermined distance, the bin is determined to contain an object. That is, an object in the bin causes the depth measured from the depth camera to be less than the predetermined distance to the bin. Conversely, when there is no object in the bin, the depth image can indicate that the depth measurement from the camera is equal to the predetermined distance to the bin. However, this paper recognizes that evaluating bin emptyness in this way is not reliable enough for many industrial applications. For example, and in unrestricted situations, bins are often defined with slightly curved bottoms, making it impractical, or in some cases impossible, to define a predetermined or constant distance from the bottom of the bin to the camera.
[0020] Refer again Figure 1 Camera 118 may define a depth camera configured to capture depth images of workspace 100 from an angle along the lateral direction 120. For example, bin 107 may define a top end 109 and a bottom end 111 opposite to the top end 109 along the lateral direction 120. Bin 107 may further define a first side 113 and a second side 115 opposite to the first side 113 along a second direction or lateral direction 122 substantially perpendicular to the lateral direction 120. Bin 107 may further define a front end 117 and a rear end 119 opposite to the front end 117 along a third third direction or longitudinal direction 124 substantially perpendicular to the lateral direction 120 and the lateral direction 122, respectively. Although bin 107 is defined as square in the illustration, it is understood that the shape or size of bins or containers may be optional, and all such bins or containers are considered to be within the scope of this disclosure.
[0021] For further reference Figure 2Autonomous system 102 can define various computing systems. For example, example computing system 200 may include a depth camera, such as camera 118, which can be configured to generate depth images of a specific scene (e.g., a scene defined by the physical environment or workspace 100). Alternatively or additionally, example computing system 200 may include multiple cameras 118 configured to generate depth and color images of a specific scene. Autonomous system 102 can define various computing systems, such as computing system 200, which may include a depth computing module 204 communicatively coupled to workspace module 206. Depth computing module 204 may define one or more neural networks configured to predict or infer the position of the bottom end 111 of bin 107. In particular, depth computing module 204 may determine the distance from camera 118 to the bottom end 111 of bin 107 along the lateral direction 120. In various examples, depth computing module 204 determines the position of bottom end 111 at runtime without requiring any prior knowledge of the distance from camera 118 to bottom end 111. Runtime can refer to the period of time during which camera 118 observes the hopper. Therefore, system 200 performs calculations related to hopper 107 in real time because information about hopper 107 (such as information related to bottom 111) may not be available during the engineering or setup time prior to the runtime.
[0022] The depth calculation module 204 can perform unsupervised k-means clustering to determine the location of the bottom 111 of bin 107. Specifically, for example, the depth calculation module 204 can determine the correlation between pixels associated with the bottom 111 of bin 107. In some cases, the depth calculation module 204 determines pixels with similar depths in the depth image (e.g., distance from the camera 118 along the lateral direction 120). Therefore, the depth calculation module 204 can define a clustering problem. In particular, given a set of all pixels, the depth calculation module 204 can identify clusters of pixels. In one example, one pixel is clustered at the bottom 111 of bin 107, and another pixel is clustered not at the bottom 111 of bin 107. In various examples, the depth calculation module 204 can perform k-means clustering to solve the clustering problem.
[0023] Once the location of bottom 111, particularly the distance from bottom 111 to camera 118 along the lateral direction 120, is determined, information related to the location of depth calculation module 204 can be obtained by workspace module 206. Specifically, based on the depth image captured by camera 118, workspace module 206 can receive a clustering map associated with bin 107. Based on the depth image of workspace 100, depth calculation module 204 can generate a clustering map that defines sets (clusters) of data points grouped together based on certain similarities. Lateral direction 122 and longitudinal direction 124 can define the horizontal plane, while lateral direction 120 can define the depth; thus, the clustering map can indicate clusters in the space defined by the horizontal plane and depth. In one example, when bin 107 is empty, the clustering map can define clusters that are relatively close relative to lateral direction 120. Furthermore, these clusters can define patches, or sets, on the horizontal plane. Workspace module 206 can determine the distance of each cluster along lateral direction 120 from camera 118. Furthermore, the workspace module 206 can calculate a corresponding average value associated with each cluster. This average value can define the average distance of each data point in a particular cluster from the camera 118 along the lateral direction 120. In one example, the workspace module 206 can determine the position of the bottom 111 of the bin 107 by identifying the cluster with an average value that defines the maximum distance from the camera 118 along the lateral direction 120, compared with other average values of other clusters.
[0024] Therefore, in various examples, workspace module 206 associates clusters corresponding to the average value representing the maximum distance to camera 118 with the bottom 111 of bin 107. Workspace module 206 can assign (or classify) pixels to the corresponding clusters. For example, after clusters are identified or determined, workspace module 206 can assign each pixel in the image to the identified cluster. In particular, in the example where bin 107 is empty, workspace module 206 can assign the majority of pixels, such as more than 95% of the pixels, or pixels between 95% and 98%, to the cluster associated with the bottom 111 of bin 107. In some cases, workspace module 206 can identify other clusters with a corresponding average value indicating that these clusters are closer to camera 118 in the lateral direction 120 than the clusters associated with bottom 111. Such clusters can be caused by noise within workspace 100 or camera 118, or by uneven or imperfect surfaces defined by the bottom 111 of bin 107. Workspace module 206 can assign or associate pixels to clusters separate from the cluster associated with bottom 111. Further, workspace module 206 can compare the number of pixels associated with clusters not at bottom 111 to the number of pixels associated with clusters at bottom 111, thereby determining a depth pixel ratio not associated with bottom 111. Workspace module 206 can compare this ratio to a decision boundary to determine whether bin 107 is empty or contains an object for grabbing. In some examples, the decision boundary is less than about 5%, such as 2%. For example, when the ratio is less than 2%, such as at least 98% of the pixels are associated with the cluster at bottom 111, workspace module 206 determines that bin 107 is empty. Continuing with the example, when the ratio is greater than 2%, such as less than 98% of the pixels are associated with the cluster at bottom 111, workspace module 206 determines that bin 107 contains at least one object. In various examples, the ratio may represent a parameter adjusted by workspace module 206. For example, module 206 can be adjusted by collecting images of empty bins and bins containing small objects (e.g., the smallest object the system might encounter (e.g., 1cm × 1cm × 1cm)). Using these images, the decision boundaries can be adjusted until an empty bin is classified as empty and an image of at least one object is classified as non-empty. It is understood that the object size is presented as an example, and implementations can be adjusted using objects of various shapes and sizes.
[0025] In another example, after the depth calculation module 204 performs k-means clustering to generate a cluster map, the workspace module 206 can evaluate the clusters in the cluster map along a horizontal plane defined by the lateral direction 122 and the longitudinal direction 124. The workspace module 206 can also evaluate the clusters in the cluster map along a depth dimension defined by the lateral direction 120. Based on the evaluation of the cluster map, the workspace module 206 can determine whether the bin 107 is empty or contains at least one object. In one example, if the workspace module 206 determines that: few clusters are connected to each other along the horizontal plane to determine that the image defines a large plane (e.g., three or fewer spatial clusters are identified); the clusters have a distance from the camera 118 along the lateral direction 120 that defines the corresponding average values that are close together; and the percentage of pixels assigned to the cluster with the average value furthest from the camera 118 along the lateral direction 120 is greater than a predetermined threshold, such as 98%, then the workspace module 206 determines that the bin 107 is empty. For example, determining whether the averages are close together can rely on the minimum height of objects that might be in the bin. Additionally, in this example, if one of the three conditions mentioned above is not met, then workspace module 206 can determine that bin 107 contains at least one object. In various examples, workspace module 206 can evaluate the cluster map in less than 10 nanoseconds to determine whether a particular bin is empty. Furthermore, in various examples, workspace module 206 evaluates the cluster map to determine whether a particular bin is empty before calculating the grasping score associated with the bin (and therefore before running the grasping neural network). Thus, in various robotic bin picking applications, it is determined whether there is an object in the bin before executing any downstream algorithms that determine how to grasp it. This document recognizes that in some cases, bin picking algorithms can determine to grasp anything, including empty bins. Therefore, if a picking algorithm is performed on an empty bin, the system may attempt to lift the empty bin, which is undesirable for a variety of reasons, including, for example, wasting computational resources and time. In some cases, the robot may be harmed by lifting an object that it is not designed to grasp (e.g., an empty bin). According to various implementation methods, this undesirable result can be avoided by determining whether the bin is empty before performing the gripping calculation.
[0026] Refer again Figure 1 and Figure 2In one example, depth calculation module 204 obtains a depth image of a workspace (e.g., workspace 100 including bin 107 containing object 106) at 208. In different examples, system 200 may detect bin 107 within workspace 100 based on the depth image. In some cases, the bin defines a bottom 111 along a lateral direction 120 and a top 109 opposite the bottom 111, wherein the bottom 111 is farther from depth camera 118 along the lateral direction 120 than the top 109. Top 109 may define an opening 121, thereby further configuring depth camera 118 to capture a depth image of bin 107 from an angle along the lateral direction 120. Based on this depth image, depth calculation module 204 determines the distance from depth camera 118 to bottom 111 along the lateral direction 120. Specifically, the depth calculation module 204 can perform k-means clustering to determine a first cluster and a second cluster, wherein the first cluster is associated with a first distance, and the second cluster is associated with a second distance 120 from the depth camera 118 along the lateral direction, the second distance being less than the first distance. Therefore, at 210, based on the depth image, the workspace module 204 can generate an output indicating one or more clusters associated with the physical data points detected by the depth camera 118. In some cases, the output at 210 may include a cluster map, a heatmap, or other markers identifying the data points.
[0027] Furthermore, at 210, the output of the depth calculation module 204 is received by the workspace module 206. The workspace module 206 can process the clustering map to determine whether the bin 107 is empty or contains at least one object. Therefore, based on the depth image and without performing grasping calculations, in various examples, the system 200 can determine whether the bin 107 is empty or contains at least one object during the runtime of the autonomous system 102 (specifically, the robot 104). For example, the workspace module 206 can assign a first pixel to a first cluster and a second pixel to a second cluster during runtime. The workspace module 206 can compare the second pixel with the first pixel to determine a pixel ratio. Furthermore, the workspace module 206 can compare the pixel ratio with a decision boundary to determine whether the bin 107 is empty or contains at least one object.
[0028] In one example, when the pixel ratio is less than a decision boundary, workspace module 206 determines that the bin is empty. At 212, in response to determining that bin 107 is empty, workspace module 206 can send an instruction to the application controlling robot 104. This instruction can notify the application that bin 107 is empty so that no grasping calculation is performed on bin 107. In another example, when the pixel ratio is greater than a decision boundary, workspace module 206 determines that bin 107 contains at least one object. At 214, in response to determining that bin 107 contains at least one object, workspace module 206 can activate a neural network to calculate the grasping position for at least one object. For example, the neural network can calculate a grasping score for the object. When the grasping score is greater than a predetermined threshold, the area associated with the grasping score is classified as an area that the end effector 116 (e.g., a vacuum-based gripper) can grasp. Conversely, in one example, when the grasping score is below a predetermined threshold, the area associated with the grasping score is classified as an area that the end effector 116 (e.g., a vacuum-based gripper) cannot grasp (e.g., edges, negative space). The robot can then grasp the object 106 at the calculated grasping position.
[0029] Subsequently, according to one example, system 102 (specifically depth camera 118 or depth calculation 204) can detect different bins within the workspace 100 of robot 104. Furthermore, the computing system 200 can determine whether the different bin is empty or contains at least one object before activating the neural network, thus repeating the aforementioned steps each time autonomous system 102 detects a new bin.
[0030] Figure 3 An example of a computing environment in which embodiments of the present disclosure may be implemented is shown. The computing environment 600 includes a computer system 610, which may include communication mechanisms, such as a system bus 621 or other communication mechanisms, for information communication within the computer system 610. The computer system 610 further includes one or more processors 620 coupled to the system bus 621 for processing information. The autonomous system 102 (and therefore the computing system 200) may include or be coupled to one or more processors 620.
[0031] Processor 620 may include one or more central processing units (CPUs), graphics processing units (GPUs), or any other processor known in the art. More generally, a processor described herein is a device for performing tasks by executing machine-readable instructions stored on a computer-readable medium, and may include any or a combination of hardware and firmware. A processor may also include memory storing machine-readable instructions for performing the tasks. A processor acts on information by manipulating, analyzing, modifying, transforming, or transferring information for use by an executable program or information device, and / or transferring information to an output device. For example, a processor may use or include the capabilities of a computer, controller, or microprocessor, and be modulated using executable instructions to perform special-purpose functions that a general-purpose computer cannot perform. A processor may include any type of suitable processing unit, including but not limited to a central processing unit, microprocessor, reduced instruction set computer (RISC) microprocessor, complex instruction set computer (CISC) microprocessor, microcontroller, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), digital signal processor (DSP), and so on. Furthermore, the processor 620 can have any suitable microarchitecture design, including any number of components, such as registers, multiplexers, arithmetic logic units, cache controllers for controlling read / write operations on cache memory, branch predictors, or similar components. The processor's microarchitecture design can support any of a variety of instruction sets. The processor can be coupled (electrically and / or including executable components) to any other processor, enabling interaction and / or communication between them. The user interface processor or generator is a known element, including electronic circuitry or software, or a combination of both, for generating display images or portions thereof. The user interface includes one or more display images, enabling the user to interact with the processor or other devices.
[0032] System bus 621 may include at least one of a system bus, memory bus, address bus, and message bus, and may allow the exchange of information (e.g., data (including computer-executable code), signaling, etc.) between various components of computer system 610. System bus 621 may include, but is not limited to, a memory bus or memory controller, a peripheral bus, an accelerated graphics port, etc. System bus 621 may be associated with any suitable bus architecture, including but not limited to Industry Standard Architecture (ISA), Micro Channel Architecture (MCA), Enhanced ISA (EISA), Video Electronics Standards Association (VESA) architecture, Accelerated Graphics Port (AGP) architecture, Peripheral Component Interconnect (PCI) architecture, PCI-Express architecture, Personal Computer Memory Card International Association (PCMCIA) architecture, Universal Serial Bus (USB) architecture, etc.
[0033] Continue to refer to Figure 3 The computer system 610 may also include a system memory 630 coupled to a system bus 621, which stores information and instructions to be executed by the processor 620. The system memory 630 may include computer-readable storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) 631 and / or random access memory (RAM) 632. RAM 632 may include other dynamic storage devices (e.g., dynamic RAM, static RAM, and synchronous DRAM). ROM 631 may include other static storage devices (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). Furthermore, the system memory 630 may be used to store temporary variables or other intermediate information during the execution of instructions by the processor 620. A basic input / output system 633 (BIOS) containing, for example, basic routines that facilitate the transfer of information between elements within the computer system 610 during startup may be stored in ROM 631. RAM 632 may contain data and / or program modules that are immediately accessible and / or currently in operation by the processor 620. System memory 630 may additionally include, for example, an operating system 634, application programs 635, and other program modules 636. Application program 635 may also include a user portal for developing applications, allowing input of parameters and modification as necessary.
[0034] Operating system 634 can be loaded into memory 630 and can provide an interface between other application software executing on computer system 610 and the hardware resources of computer system 610. More specifically, operating system 634 may include a set of computer-executable instructions for managing the hardware resources of computer system 610 and providing common services to other applications (e.g., managing memory allocation among various applications). In some example implementations, operating system 634 may control the execution of one or more program modules described as being stored in data memory 640. Operating system 634 may include any operating system now known or that may be developed in the future, including but not limited to any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.
[0035] Computer system 610 may also include a disk / media controller 643 coupled to system bus 621 to control one or more storage devices for storing information and instructions, such as magnetic hard disk 641 and / or removable media drives 642 (e.g., floppy disk drives, optical disk drives, magnetic tape drives, flash memory drives, and / or solid-state drives). Storage device 640 may be added to computer system 610 using an appropriate device interface (e.g., Small Computer System Interface (SCSI), Integrated Device Electronics (IDE), Universal Serial Bus (USB), or FireWire). Storage devices 641, 642 may be external to computer system 610.
[0036] The computer system 610 may also include a field device interface coupled to the system bus 621 to control field devices, such as equipment used on a production line. The computer system 610 may include a user input interface 660 or a GUI 661, which may include one or more input devices, such as a keyboard, touchscreen, tablet, and / or pointing devices, for interacting with the computer user and providing information to the processor 620.
[0037] Computer system 610 can perform some or all of the processing steps of embodiments of the present invention in response to processor 620 executing one or more sequences of one or more instructions contained in memory (e.g., system memory 630). Such instructions can be read into system memory 630 from another computer-readable storage medium 640 (e.g., magnetic hard disk 641 or removable media drive 642). Magnetic hard disk 641 (or solid-state drive) and / or removable media drive 642 may contain one or more data storages and data files used in embodiments of the present disclosure. Data storage 640 may include, but is not limited to, databases (e.g., relational, object-oriented, etc.), file systems, flat files, distributed data storage (where data is stored on more than one node of a computer network), peer-to-peer network data storage, etc. Data storage can store various types of data, such as skill data, sensor data, or any other data generated according to embodiments of the present disclosure. Data storage contents and data files may be encrypted to improve security. Processor 620 may also employ a multiprocessing arrangement to execute one or more sequences of instructions contained in system memory 630. In alternative embodiments, hard-wired circuitry may be used instead of or combined with software instructions. Therefore, this embodiment is not limited to any specific combination of hardware circuitry and software.
[0038] As described above, computer system 610 may include at least one computer-readable medium or memory for storing instructions programmed according to embodiments of the present invention and data including data structures, tables, records, or other data as described herein. The term "computer-readable medium" as used herein refers to any medium that participates in providing instructions to processor 620 for execution. Computer-readable media may take many forms, including but not limited to non-transitory, non-volatile, volatile, and transmission media. Non-limiting examples of non-volatile media include optical discs, solid-state drives, magnetic disks, and magneto-optical disks, such as magnetic hard disk 641 or removable media drive 642. Non-limiting examples of volatile media include dynamic memory, such as system memory 630. Non-limiting examples of transmission media include coaxial cables, copper wires, and optical fibers, including conductors constituting system bus 621. Transmission media may also take the form of acoustic or optical waves, such as those generated during radio wave and infrared data communication.
[0039] Computer-readable medium instructions used to perform operations of this disclosure may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, or similar languages, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via an internet service provider using the internet). In some embodiments, electronic circuits including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuits and thereby perform various aspects of this disclosure.
[0040] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable medium instructions.
[0041] The computing environment 600 may further include a computer system 610 that operates in a network environment, using logical connections to one or more remote computers, such as remote computing devices 680. A network interface 670 enables communication, for example, with other remote devices 680 or systems and / or storage devices 641, 642 via a network 671. The remote computing device 680 may be a personal computer (laptop or desktop), mobile device, server, router, network PC, peer-to-peer device, or other common network node, and typically includes many or all of the elements described above relative to the computer system 610. When used in a network environment, the computer system 610 may include a modem 672 for establishing communication on the network 671 (e.g., the Internet). The modem 672 may be connected to the system bus 621 via the user network interface 670 or through other suitable mechanisms.
[0042] Network 671 can be any network or system commonly known in the art, including the Internet, intranet, local area network (LAN), wide area network (WAN), metropolitan area network (MAN), direct or serial connection, cellular telephone network, or any other network or medium capable of facilitating communication between computer system 610 and other computers (e.g., remote computing device 680). Network 671 can be wired, wireless, or a combination thereof. Wired connections can be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection commonly known in the art. Wireless connections can be implemented using Wi-Fi, WiMAX and Bluetooth, infrared, cellular networks, satellite, or any other wireless connection method commonly known in the art. Furthermore, several networks can operate independently or communicate with each other to facilitate communication within network 671.
[0043] It should be understood that Figure 3 The program modules, applications, computer-executable instructions, code, etc., stored in system memory 630 as described herein are illustrative and not exhaustive, and the processing described as being supported by any particular module may alternatively be distributed across multiple modules or executed by different modules. Furthermore, various program modules, scripts, plug-ins, application programming interfaces (APIs), or any other suitable computer-executable code may be provided locally hosted on computer system 610, on remote device 680, and / or hosted on other computer devices accessed via one or more networks 671 to support processing by… Figure 3 The functionality and / or additional or alternative functionality provided by the program modules, applications, or computer executable code described herein. Furthermore, functionality can be modularized differently, thus allowing it to be described as being provided by… Figure 3The processing collectively supported by the set of program modules described herein can be performed by fewer or more modules, or functionality described as being supported by any particular module can be at least partially supported by another module. Furthermore, the program modules supporting the functionality described herein can form part of one or more applications that can execute on any number of systems or devices according to any suitable computational model (e.g., client-server model, peer-to-peer model, etc.). Figure 3 The description states that any functionality supported by any program module can be at least partially implemented in hardware and / or firmware on any number of devices.
[0044] It should be further understood that computer system 610 may include alternative and / or additional hardware, software, or firmware components beyond those described or depicted without departing from the scope of this disclosure. More specifically, it should be understood that the software, firmware, or hardware components described as forming part of computer system 610 are merely illustrative, and in various embodiments, some components may be absent or additional components may be provided. While various illustrative program modules have been described as software modules stored in system memory 630, it should be understood that functionality described as supported by program modules can be enabled by any combination of hardware, software, and / or firmware. It should be further understood that each of the above modules may represent a logical partition of supported functionality in various embodiments. Such logical partitioning is described for the purpose of illustrating functionality and may not represent the structure of the software, hardware, and / or firmware used to implement that functionality. Therefore, it should be understood that in various embodiments, functionality described as provided by a particular module may be provided at least partially by one or more other modules. Furthermore, in some embodiments, one or more described modules may be absent, while in other embodiments, additional modules not described may be present and may support at least a portion of the described functionality and / or additional functionality. Furthermore, while some modules may be depicted and described as submodules of another module, in some implementations these modules may be provided as standalone modules or as submodules of other modules.
[0045] Although specific embodiments of this disclosure have been described, those skilled in the art will recognize that many other modifications and alternative embodiments are within the scope of this disclosure. For example, any functionality and / or processing capabilities described with respect to a particular device or component can be performed by any other device or component. Furthermore, while various illustrative embodiments and architectures have been described according to embodiments of this disclosure, those skilled in the art will understand that many other modifications to the illustrative embodiments and architectures described herein are also within the scope of this disclosure. Moreover, it should be understood that any operation, element, component, data, or like described herein that is based on another operation, element, component, data, or the like may additionally be based on one or more other operations, elements, components, data, or the like. Therefore, the phrase “based on” or variations thereof should be interpreted as “at least partially based on”.
[0046] Although the implementations have described structural features and / or methodological behaviors in specific language, it should be understood that this disclosure is not necessarily limited to the specific features or behaviors described. Rather, specific features and behaviors are disclosed as illustrative forms for practicing the invention. Conditional language, such as “may,” “can,” “may,” or “possibly,” among others, unless specifically stated or otherwise understood in the context of use, is generally intended to express that certain implementations may include certain features, elements, and / or steps, while other implementations do not. Therefore, such conditional language generally does not imply that features, elements, and / or steps are necessary in any way for one or more implementations, or that one or more implementations must include logic for determining whether to include or perform such features, elements, and / or steps in any particular implementation, with or without user input or prompting.
[0047] The flowcharts and block diagrams in the figures illustrate the structure, function, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the blocks may not appear in the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or, depending on the functionality involved, these blocks may sometimes be executed in reverse order. It will also be noted that each block in the block diagrams and / or flowchart descriptions, and combinations of blocks in the block diagrams and / or flowchart descriptions, may be implemented by a special-purpose hardware system that performs the specified function or behavior, or a combination of special-purpose hardware and computer instructions.
Claims
1. An autonomous system configured to operate in an active industrial environment to define runtime, said autonomous system comprising: A robot, defined as an end effector, is configured to grasp multiple objects within a workspace; A depth camera is configured to capture depth images of the workspace; One or more processors; as well as A memory for storing instructions, which, when executed by the one or more processors, enable the autonomous system to: A bin in a detection workspace, the bin being capable of accommodating one or more of a plurality of objects, wherein the bin defines a bottom end and a top end opposite the bottom end in a lateral direction, the bottom end being farther from the depth camera in the lateral direction than the top end, and the top end defining an opening, such that the depth camera is further configured to capture the depth image of the bin from an angle in the lateral direction through the opening; and Based on the depth image and without performing grasping calculations, determine whether the bin is empty or contains at least one object; Determining whether the bin is empty or contains at least one object includes: Determine a first distance along the lateral direction from the depth camera to the bottom end; Perform k-means clustering to determine a first cluster and a second cluster, the first cluster being associated with a first distance and the second cluster being associated with a second distance from the depth camera along the lateral direction, the second distance being less than the first distance; Assign the first pixel to the first cluster; Assign the second pixel to the second cluster; The second pixel is compared with the first pixel to determine the pixel ratio; and The pixel ratio is compared with a decision boundary to determine whether the bin is empty or contains at least one object.
2. The autonomous system according to claim 1, wherein the memory further stores instructions that, when executed by the one or more processors, cause the autonomous system to: When the pixel ratio is less than the decision boundary, it is determined that the bin is empty; and In response to determining that the bin is empty, an instruction is sent to the application controlling the robot, the instruction informing the application that the bin is empty.
3. The autonomous system according to claim 1, wherein the memory further stores instructions that, when executed by the one or more processors, cause the autonomous system to: When the pixel ratio is greater than the decision boundary, it is determined that the bin contains at least one object; and In response to determining that the bin contains at least one object, the neural network is activated to calculate the gripping position on the at least one object. wherein, The robot is further configured to grasp the at least one object at the grasping location.
4. A method executed by an autonomous system, said autonomous system comprising a robot operating in an active industrial environment, for defining runtime, said method comprising: A bin within the workspace of a robot with an autonomous detection system, the bin being capable of holding multiple objects, wherein the bin defines a bottom end and a top end opposite the bottom end in the lateral direction, and the top end defines an opening; A depth image of the workspace including the hopper is captured through the opening at an angle along the lateral direction by a depth camera, wherein the bottom end is farther from the depth camera along the lateral direction than the top end; and Based on the depth image and without performing grasping calculations, determine whether the bin is empty or contains at least one object. Determining whether the bin is empty or contains at least one object further includes: Determine a first distance along the lateral direction from the depth camera to the bottom end; Perform k-means clustering to determine a first cluster and a second cluster, the first cluster being associated with a first distance and the second cluster being associated with a second distance from the depth camera along the lateral direction, the second distance being less than the first distance; Assign the first pixel to the first cluster; Assign the second pixel to the second cluster; The second pixel is compared with the first pixel to determine the pixel ratio; and The pixel ratio is compared with a decision boundary to determine whether the bin is empty or contains at least one object.
5. The method according to claim 4, further comprising: When the pixel ratio is less than the decision boundary, it is determined that the hopper is empty; as well as In response to determining that the bin is empty, an instruction is sent to the application controlling the robot, the instruction informing the application that the bin is empty.
6. The method according to claim 4, further comprising: When the pixel ratio is greater than the decision boundary, it is determined that the bin contains at least one object; In response to determining that the bin contains at least one object, a neural network is activated to calculate a gripping position on the at least one object; and The robot grasps the at least one object at the grasping position.
7. The method according to claim 4, further comprising: Detect different bins within the robot's workspace; as well as Before activating the neural network, it is determined whether the different bins are empty or contain at least one object, such that the steps in claim 4 are repeated each time the autonomous system detects a new bin.
Citation Information
Patent Citations
Empty container detection
CN113710594A