Computer
system (110), comprising: a
communication interface (113) configured to communicate with at least one camera, comprising a first camera (270) with a first camera
field of view (272); a
control circuit (111) configured, when a stack (250, 750) with multiple objects is located in the first camera
field of view (272), to: receive camera data generated by the at least one camera, wherein the camera data describes a stack structure for the stack (250, 750), the stack structure being formed from at least one
object structure for a first object of the multiple objects;Identify, based on camera data generated by the at least one camera, a target feature (251B, 251C, 751B) of the
object structure or a target feature (251B, 251C, 751B) located on the
object structure, wherein the target feature (251B, 251C, 751B) is at least one of the following: a corner (251B) of the object structure, an edge (251C) of the object structure, a visual feature (751B) located on a surface (251A, 751A) of the object structure, or an outline of the surface (251A, 751A) of the object structure; Determine a two-dimensional, 2D, region (520, 620, 720) that is coplanar with the target feature (251B, 251C, 751B) and whose boundary is defined by the target feature (251B, 251C, 751B); Determining a three-dimensional, 3D, region (530, 630, 730) defined by connecting a location of the first camera (270) and the boundary of the 2D region (520, 620, 720), wherein the 3D region (530, 630, 730) is part of the first camera's
field of view (272);Determine, based on the camera data and the 3D region (530, 630, 730), a size of an
occlusion region (570, 670, 770), where the
occlusion region (570, 670, 770) is a region of the stack structure located between the target feature (251B, 251C, 751B) and the at least one camera (270) and within the 3D region (530, 630, 730); Determine a value of an
object detection confidence parameter based on the size of the
occlusion region (570, 670, 770); and performing an operation to control
robot interaction with the stack structure, wherein the operation is performed based on the value of the object recognition confidence parameter; wherein the first camera (270) with which the
communication interface (113) is configured to communicate is a 3D camera configured to generate, as part of the camera data, several 3D data points that specify corresponding depth values for locations on one or more faces of the stack structure.