System and method for compressing three-dimensional data
The 3D data compression mechanism addresses storage and bandwidth challenges by leveraging planar surfaces and preserving edges/corners, achieving efficient data reduction for robotic systems in environments like shipping and warehouse management.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MUJIN INC
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
Existing 3D image data compression methods do not provide adequate compression ratios for real-time processing and transmission to devices with limited computational resources, particularly in robotic systems, leading to substantial storage and bandwidth challenges.
A 3D data compression mechanism that leverages planar surfaces and known package dimensions to assign common height values, identifies regions of interest, and preserves edge and corner information through edge vector sets and data keys, combining structured and unstructured data formats for efficient down-sampling.
This approach achieves significant data size reduction while maintaining spatial accuracy, enabling efficient processing and transmission of 3D image data for robotic systems, particularly in environments like shipping and warehouse management.
Smart Images

Figure IB2025061411_15052026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR COMPRESSING THREE-DIMENSIONAL DATACROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to U.S. Provisional Application No. 63 / 718,661, titled "System and Method for Three-Dimensional Data Compression," filed November 10, 2024, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD
[0002] This disclosure relates generally to computer vision systems and, more specifically, to systems, processes, and techniques for compressing three-dimensional image data, such as in environments including packaging, boxing, shipping, warehouse management, and / or the like.BACKGROUND
[0003] Robotic systems have become increasingly prevalent across various industries, including manufacturing, warehouse management, shipping hubs, and distribution centers, where they perform tasks such as object manipulation, sorting, and transportation. These systems often rely on two-dimensional (2) and / or three-dimensional (3D) imaging technologies to perceive and understand their environment, enabling them to identify objects, determine spatial relationships, and plan appropriate actions. The 3D image data captured by these systems can be represented in various formats, including point clouds that contain spatial coordinate information for numerous data points within a scene.
[0004] The processing, storage, and transmission of 3D image data presents challenges due to the substantial file sizes associated with detailed spatial information. Point cloud data, whether structured or unstructured, can consume considerable storage space and bandwidth when transmitted between system components or to remote devices for analysis or visualization. While various compression algorithms exist for reducing data size, many approaches may not provide adequate compression ratios for certain applications, particularly those involving real-time processing or transmission to devices with limited computational resources or display capabilities. The balance between data compression and maintaining spatial accuracy remains a consideration in developing systems that can efficiently handle large volumes of 3D image data.14899-7240-8951\1BRIEF DESCRIPTION OF FIGURES
[0005] Many aspects of the present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale. Instead, emphasis is placed on illustrating clearly the principles of the present disclosure. The drawings should not be taken to limit the disclosure to the specific embodiments shown, but are provided for explanation and understanding.
[0006] FIG. 1 illustrates an example environment in which a system with a 3D may operate in accordance with one or more embodiments of the present technology.
[0007] FIG. 2 illustrates a block diagram of the robotic system in accordance with one or more embodiments of the present technology.
[0008] FIG. 3 illustrates an example robotic system in accordance with one or more embodiments of the present technology.
[0009] FIG. 4 illustrates a captured image in accordance with one or more embodiments of the present technology.
[0010] FIGS. 5A and 5B illustrate example 3D data compression mechanisms in accordance with one or more embodiments of the present technology.
[0011] FIG. 6 illustrates an example 3D image data within a reference segment in accordance with one or more embodiments of the present technology.
[0012] FIG. 7A illustrates an example capture data in accordance with one or more embodiments of the present technology.
[0013] FIG. 7B illustrates an example down-sampling result of the captured data of FIG. 7A in accordance with one or more embodiments of the present technology.
[0014] FIG. 8 A illustrates an example capture data from a perspective-view camera in accordance with one or more embodiments of the present technology.
[0015] FIG. 8B illustrates a down-sampled result of the captured data of FIG. 8 A in accordance with one or more embodiments of the present technology.24899-7240-8951\1
[0016] FIG. 8C illustrates a reconstruction result from the down-sampled result of FIG. 8B in accordance with one or more embodiments of the present technology.
[0017] FIGS. 9 A - 9C illustrate flow diagrams of example method for compressing 3D data in accordance with one or more embodiments of the present technology.DETAILED DESCRIPTION
[0018] Systems and methods for compressing 3D data are described herein. Embodiments of the present technology can include a mechanism for compressing 3D data (also referenced as "3D data compression mechanism"). The 3D data compression mechanism can be configured to leverage the targeted application in compressing the 3D data. For example, when utilized within a robotic system, the 3D data compression mechanism can be configured to leverage targets, such as package objects, operated on by the robotic system.
[0019] In some embodiments, the 3D data compression mechanism can be configured to leverage planar surfaces and / or known dimensions of packaged objects. The 3D data compression mechanism can assign a common height value for a given grouping or size of units (e.g., a lateral area) to reduce depth map data size. Moreover, the 3D data compression mechanism can determine a location and a size (e.g., dimensions) of a region of interest (ROI). The 3D data compression mechanism can provide depth values for the ROI, such as using a set or a variable grouping size, without providing the depth values outside of the ROI. Stated differently, the 3D data compression mechanism can ignore areas with not-a-number (NaN) values corresponding to detecting no objects within a scanning range of the sensor.
[0020] Additionally, the 3D data compression mechanism can preserve locations of edges and / or corners. For example, the 3D data compression mechanism can down-sample a captured image data and further generate an edge vector set that represents locations of the edges. Also, for example, the 3D data compression mechanism can generate a depth map that includes boundary values that represent presence / locations of detected corners and / or edges. Additionally or alternatively, the 3D data compression mechanism can generate the depth map as a variable depth map that has different sized and / or shaped areas that each have a common depth value. The 3D data compression mechanism can use predetermined down-sampling formats, such as data34899-7240-8951\1ordering pattern, data unit format, and / or the like to down-sample / compress that captured image and also interpolate / decompress the down-sampled data.
[0021] In the following description, specific details are set forth to provide a thorough understanding of aspects of the present technology. One skilled in the relevant art will recognize, however, that the systems, devices, and techniques described herein can be practiced without one or more of the specific details set forth herein, or with other methods, components, materials, etc.
[0022] Reference throughout this specification to an “example” or an “embodiment” means that a particular feature, structure, or characteristic described in connection with the example or embodiment is included in at least one example or embodiment of the present technology. Thus, use of the phrases “for example,” “as an example,” or “an embodiment” herein are not necessarily all referring to the same example or embodiment and are not necessarily limited to the specific example or embodiment discussed. Furthermore, features, structures, or characteristics of the present technology described herein may be combined in any suitable manner to provide further examples or embodiments of the present technology.
[0023] Spatially relative terms (e.g., “beneath,” “below,” “over,” “under,” “above,” “upper,” “top,” “bottom,” “left,” “right,” “center,” “middle,” and the like) may be used herein for ease of description to describe one element’s or feature’s relationship relative to one or more other elements or features as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of a device or system in use or operation, in addition to the orientation depicted in the figures. For example, if a device or system illustrated in the figures is rotated, turned, or flipped about a horizontal axis, elements or features described as “below” or “beneath” or “under” one or more other elements or features may then be oriented “above” the one or more other elements or features. Thus, the exemplary terms “below” and “under” are non-limiting and can encompass both an orientation of above and below. The device or system may additionally, or alternatively, be otherwise oriented (e.g., rotated ninety degrees about a vertical axis, or at other orientations) than illustrated in the figures, and the spatially relative descriptors used herein are interpreted accordingly. In addition, it will also be understood that when an element is referred to as being “between” two other elements, it can be the only element between the two other elements, or one or more intervening elements may also be present.Suitable Environments44899-7240-8951\1
[0024] FIG. 1 is an illustration of an example environment in which a robotic system 100 with a 3D image data compression mechanism may operate. The robotic system 100 can include and / or communicate with one or more units (e.g., robots) configured to execute one or more tasks. Aspects of the coordinated transfer mechanism can be practiced or implemented by the various units.
[0025] For the example illustrated in FIG. 1, the robotic system 100 can include an unloading unit 102, a transfer unit 104 (e.g., a palletizing robot and / or a piece-picker robot), a transport unit 106, a loading unit 108, or a combination thereof in a warehouse or a distribution / shipping hub. Each of the units in the robotic system 100 can be configured to execute one or more tasks. The tasks can be combined in sequence to perform an operation that achieves a goal, such as to unload objects from a truck or a van and store them in a warehouse or to unload objects from storage locations and prepare them for shipping. In some embodiments, the task can include placing the objects on a target location (e.g., on top of a pallet and / or inside a bin / cage / box / case).
[0026] In order to execute the tasks, the robotic system 100 can derive individual placement locations / orientations, calculate corresponding motion plans, or a combination thereof, such as for placing and / or stacking the objects. Each of the units can be configured to execute a sequence of actions (e.g., operating one or more components therein) according to the placement locations / orientations, the motion plans, and / or the like to execute a task. In some embodiments, the robotic system 100 can include (1) sensors 101 that provide various input data and (2) at least one controller 109 that processes the sensor outputs to compute the locations / orientations, motion plans, and / or the like. For example, the sensors 101 can include a camera configured to obtain 2D and / or 3D image data of various locations and / or objects. The controller 109 can process the image data to derive the locations / orientations, the motion plan, and / or the like.
[0027] One example of the task executed through the robotic system 100 can include manipulation (e.g., moving and / or reorienting) of a target object 112 (e.g., one of the packages, boxes, cases, cages, pallets, etc. corresponding to the executing task) from a start / source location 114 to a task / destination location 116. For example, the unloading unit 102 (e.g., a devanning robot) can be configured to transfer the target object 112 from a location in a carrier (e.g., a truck) to a location on a conveyor. Also, the transfer unit 104 can be configured to transfer the target object 112 from one location (e.g., the conveyor, a pallet, or a bin) to another location (e.g., a54899-7240-8951\1pallet, a bin, etc.). For another example, the transfer unit 104 (e.g., a palletizing robot) can be configured to transfer the target object 112 from a source location (e.g., a pallet, a pickup area, and / or a conveyor) to a destination pallet. In completing the operation, the transport unit 106 (e.g., a conveyor, an automated guided vehicle (AGV), a shelf-transport robot, etc.) can transfer the target object 112 from an area associated with the transfer unit 104 to an area associated with the loading unit 108, and the loading unit 108 can transfer the target object 112 (by, e.g., moving the pallet carrying the target object 112) from the transfer unit 104 to a storage location (e.g., a location on the shelves).
[0028] For illustrative purposes, the robotic system 100 is described in the context of a packaging and / or shipping center; however, it is understood that the robotic system 100 can be configured to execute tasks in other environments / for other purposes, such as for manufacturing, assembly, storage / stocking, healthcare, and / or other types of automation. It is also understood that the robotic system 100 can include other units, such as manipulators, service robots, modular robots, etc., not shown in FIG. 1. For example, in some embodiments, the robotic system 100 can include a depalletizing unit for transferring the objects from cage carts or pallets onto conveyors or other pallets, a container-switching unit for transferring the objects from one container to another, a packaging unit for wrapping / casing the objects, a sorting unit for grouping objects according to one or more characteristics thereof, a piece-picking unit for manipulating (e.g., for sorting, grouping, and / or transferring) the objects differently according to one or more characteristics thereof, or a combination thereof.Suitable System
[0029] FIG. 2 is a block diagram illustrating the robotic system 100 in accordance with one or more embodiments of the present technology. In some embodiments, for example, the robotic system 100 (e.g., at one or more of the units and / or robots described above) can include electronic / electrical devices, such as one or more processors 202, one or more storage devices 204, one or more communication devices 206, one or more input-output devices 208, one or more actuation devices 212, one or more transport motors 214, one or more sensors 216, or a combination thereof. The various devices can be coupled to each other via wire connections and / or wireless connections. For example, the robotic system 100 can include a bus, such as a system bus, a Peripheral Component Interconnect (PCI) bus or PCI-Express bus, a HyperTransport or industry64899-7240-8951\1standard architecture (ISA) bus, a small computer system interface (SCSI) bus, a universal serial bus (USB), an IIC (I2C) bus, or an Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus (also referred to as "Firewire"). Also, for example, the robotic system 100 can include bridges, adapters, processors, or other signal-related devices for providing the wire connections between the devices. The wireless connections can be based on, for example, cellular communication protocols (e.g., 3G, 4G, LTE, 5G, etc.), wireless local area network (LAN) protocols (e.g., wireless fidelity (Wi-Fi)), peer-to-peer or device-to-device communication protocols (e.g., Bluetooth, Near-Field communication (NFC), etc.), Internet of Things (loT) protocols (e.g., NB-IoT, LTE-M, etc.), and / or other wireless communication protocols.
[0030] The processors 202 can include data processors (e.g., central processing units (CPUs), special-purpose computers, and / or onboard servers) configured to execute instructions (e.g., software instructions) stored on the storage devices 204 (e.g., computer memory). In some embodiments, the processors 202 can be included in a separate / standalone controller (e.g., the controller 109 of FIG. 1) that is operably coupled to the other electronic / electrical devices illustrated in FIG. 2 and / or the robotic units illustrated in FIG. 1. The processors 202 can implement the program instructions to control / interface with other devices, thereby causing the robotic system 100 to execute actions, tasks, and / or operations.
[0031] The storage devices 204 can include non-transitory computer-readable mediums having stored thereon program instructions (e.g., software). Some examples of the storage devices 204 can include volatile memory (e.g., cache and / or random-access memory (RAM)) and / or nonvolatile memory (e.g., flash memory and / or magnetic disk drives). Other examples of the storage devices 204 can include portable memory and / or cloud storage devices.
[0032] In some embodiments, the storage devices 204 can be used to further store and provide access to processing results and / or predetermined data / thresholds. For example, the storage devices 204 can store master data 252 that includes descriptions of objects (e.g., boxes, cases, and / or products) that may be manipulated by the robotic system 100. In one or more embodiments, the master data 252 can include a dimension, a shape (e.g., templates for potential poses and / or computer-generated models for recognizing the object in different poses), a color scheme, an image, identification information (e.g., bar codes, quick response (QR) codes, logos, etc., and / or expected locations thereof), an expected weight, other physical / visual characteristics, or a74899-7240-8951\1combination thereof for the objects expected to be manipulated by the robotic system 100. In some embodiments, the master data 252 can include manipulation-related information regarding the objects, such as a center-of-mass (CoM) location on each of the objects, expected sensor measurements (e.g., for force, torque, pressure, and / or contact measurements) corresponding to one or more actions / maneuvers, or a combination thereof.
[0033] The communication devices 206 can include circuits configured to communicate with external or remote devices via a network. For example, the communication devices 206 can include receivers, transmitters, modulator s / demodulators (modems), signal detectors, signal encoders / decoders, connector ports, network cards, etc. The communication devices 206 can be configured to send, receive, and / or process electrical signals according to one or more communication protocols (e.g., the Internet Protocol (IP), wireless communication protocols, etc.). In some embodiments, the robotic system 100 can use the communication devices 206 to exchange information between units of the robotic system 100 and / or exchange information (e.g., for reporting, data gathering, analyzing, and / or troubleshooting purposes) with systems or devices external to the robotic system 100.
[0034] The input-output devices 208 can include user interface devices configured to communicate information to and / or receive information from human operators. For example, the input-output devices 208 can include a display 210 and / or other output devices (e.g., a speaker, a haptics circuit, or a tactile feedback device, etc.) for communicating information to the human operator. Also, the input-output devices 208 can include control or receiving devices, such as a keyboard, a mouse, a touchscreen, a microphone, a user interface (UI) sensor (e.g., a camera for receiving motion commands), a wearable input device, etc. In some embodiments, the robotic system 100 can use the input-output devices 208 to interact with the human operators in executing an action, a task, an operation, or a combination thereof.
[0035] The robotic system 100 can include physical or structural members (e.g., robotic manipulator arms) that are connected at joints for motion (e.g., rotational and / or translational displacements). The structural members and the joints can form a kinetic chain configured to manipulate an end-effector (e.g., the gripper) configured to execute one or more tasks (e.g., gripping, spinning, welding, etc.) depending on the use / operation of the robotic system 100. The robotic system 100 can include the actuation devices 212 (e.g., motors, actuators, wires, artificial84899-7240-8951\1muscles, electroactive polymers, etc.) configured to drive or manipulate (e.g., displace and / or reorient) the structural members about or at a corresponding joint. In some embodiments, the robotic system 100 can include the transport motors 214 configured to transport the corresponding units / chassis from place to place.
[0036] The robotic system 100 can include the sensors 216 (e.g., the sensors 101 of FIG. 1) configured to obtain information used to implement the tasks, such as for manipulating the structural members and / or for transporting the robotic units. The sensors 216 can include devices configured to detect or measure one or more physical properties of the robotic system 100 (e.g., a state, a condition, and / or a location of one or more structural members / joints thereof) and / or of a surrounding environment. Some examples of the sensors 216 can include accelerometers, gyroscopes, force sensors, strain gauges, tactile sensors, torque sensors, position encoders, etc.
[0037] In some embodiments, for example, the sensors 216 can include one or more imaging devices 222 (e.g., visual and / or infrared cameras, 2D and / or 3D imaging cameras, distance measuring devices such as LIDARs or radars, etc.) configured to detect the surrounding environment. The imaging devices 222 can generate representations of the detected environment, such as digital images and / or point clouds, which may be processed via machine / computer vision (e.g., for automatic inspection, robot guidance, or other robotic applications). As described in further detail below, the robotic system 100 (via, e.g., the processors 202) can process the digital image and / or the point cloud to identify the target object 112 of FIG. 1, the start location 114 of FIG. 1, the task location 116 of FIG. 1, a pose of the target object 112, a confidence measure regarding the start location 114 and / or the pose, or a combination thereof.
[0038] For manipulating the target object 112, the robotic system 100 (via, e.g., the various circuits / devices described above) can capture and analyze image data of a designated area (e.g., a pickup location, such as inside the truck or on the conveyor belt) to identify the target object 112 and the start location 114 thereof. Similarly, the robotic system 100 can capture and analyze image data of another designated area (e.g., a drop location for placing objects on the conveyor, a location for placing objects inside the container, or a location on the pallet for stacking purposes) to identify the task location 116. For example, the imaging devices 222 can include one or more cameras configured to generate image data of the pickup area and / or one or more cameras configured to generate image data of the task area (e.g., drop area). Based on the image data, as described below,94899-7240-8951\1the robotic system 100 can determine the start location 114, the task location 116, the associated poses, a packing / placement location, and / or other processing results.
[0039] In some embodiments, for example, the sensors 216 can include position sensors 224 (e.g., position encoders, potentiometers, etc.) configured to detect positions of structural members (e.g., the robotic arms and / or the end-effectors) and / or corresponding joints of the robotic system 100. The robotic system 100 can use the position sensors 224 to track locations and / or orientations of the structural members and / or the joints during execution of the task.Computer Vision System
[0040] Returning to image processing, the robotic system 100 can include a computer vision system implemented using the processor 202, the storage device 204, the imaging devices, portions thereof, or combinations thereof. The computer vision system can provide machine-level recognition, detection, identification, and / or related data processing for objects, things, persons, etc. depicted in the image data.
[0041] The computer vision system can process one or more of various types of data. For example, the computer vision system can process 2D and / or 3D image data. Even within one category, such as 3D image data, the computer vision system can process a variety of file types. For example, the computer vision system can receive and process a point cloud as the 3D image data. Moreover, the point cloud can have an unstructured format or a structured format.
[0042] Unstructured image data typically corresponds to 3D image data being stored with an unspecified predetermined structure. For instance, the unstructured image data can store valid data points along with their coordinates (e.g., x, y, and z coordinates) while excluding exclude data points and / or locations determined to be NaN, such as when no object / structure is detected within a sensing distance / range. Stated differently the unstructured image data can positively identify data associated with detected objects / structures and communicate through inference that locations that have not been named represent NaN or no detections.
[0043] On the other hand, structured image data typically corresponds to 3D image data being stored with a specified predetermined structure, such as a grid-like format. For instance, the structured image data communicates and / or store the data according to a predetermined pattern of the x and y coordinates. Hence, the structured image data can include a value for each104899-7240-8951\1predetermined location (e.g., a unique x-y pairing) within a field-of-view. Accordingly, the structured image data can infer the x-y coordinates of each data point / value according to the predetermined format. For locations without any detected objects / structures, the structured image data can include NaN values (e.g., a predetermined value, typically a maximum or a minimum bitvalue). In some usage cases, using an unstructured point cloud allows for more flexible applications and potentially greater accuracy, while a structured point cloud can result in reduced data usage and faster computation of that data. In other usage cases, such as for sparsely populated environment, the unstructured point cloud can reduce the overall data size, while the structure point cloud may provide additional details that may be missed or filtered out in the unstructured point cloud.
[0044] In communicating or processing the 3D image data, it may be desirable to reduce the size and / or complexity of the image data. Conventional systems use compression algorithms, such as standard (zstd). Zstd may be considered a lossless algorithm that provides relatively fast compression and decompression speeds. In exchange for the relatively fast compression / decompression, zstd trades away the overall compression rate, i.e., zstd provides relatively nominal decrease in the overall data size. Depending on the parameters selected, conventional systems can use zstd to reduce an unstructured point cloud data file of 26.09MB to 22.87MB.
[0045] In some situations, it may be possible to utilize structured point cloud data. Continuing with the example depiction above, the structured point cloud may depict the same scene with a data size of about 12.75MB. If a similar zstd compression algorithm is run on such data, the size may be further reduced to approximately 6.99MB. Despite the reduction in file size, zstd compression may still not be sufficient to serve certain purposes. Additionally, when the data compressed by zstd is to be used, it must first be decompressed before it can be used, which can take time and computational resources. Accordingly, more suitable compression methods may be desired.Example Image Processing
[0046] Embodiments of the technology described below includes a computer vision system configured to down-sample obtained image data based on known characteristics of an operating environment. For example, the computer vision system can be configured to leverage planar114899-7240-8951\1surfaces and known package dimensions for robotic systems operating in shipping / distribution hubs, warehouses, etc. The computer vision system can use a 3D data compression mechanism to down-sample the image data (e.g., the 3D image data, such as depth maps) by determining units or segments that correspond to an area covered by multiple data points and deriving a single value for each unit / segment. Additionally, the 3D data compression mechanism can be configured to determine ROIs (e.g., locations having depth values instead of NaN) within the captured image and provide the depth values for the ROIs while ignoring / deleting values outside of the ROIs. Within the ROIs, the 3D data compression mechanism can use one or more predetermined sizes, locations, reporting sequence, and / or the like to provide a reduced number of data points. Effectively, the 3D data compression mechanism can leverage and combine both the structured and the unstructured data formats for down sampling the image data. The 3D data compression mechanism (e.g., a hardware circuit, a set of software instructions, the processors 202 of FIG. 2, the instructions, or a combination thereof) can identify a location and size / shape of each ROI for mapping the provided depth values.
[0047] Additionally, the 3D data compression mechanism can preserve edge and corner information through data structures that maintain spatial accuracy while achieving compression. For example, an edge vector set can store location information about detected edges by recording endpoint pairs that define edges corresponding to the boundaries and transitions between different surfaces or regions within the compressed data. This approach may allow the system to retain precise edge locations without storing full resolution data for entire areas.
[0048] Additionally or alternatively, the 3D data compression mechanism can use data keys as a framework for including and locating edges and / or corners within compressed spatial information. The data keys can include edge identifiers that mark specific locations or patterns corresponding to detected edges within the 3D image data. Similarly, corner identifiers can designate positions where multiple edges intersect, preserving these geometrically significant features that may be important for object recognition and manipulation planning.
[0049] In some embodiments, the edge identifiers and corner identifiers can work in conjunction with the data keys to create a hierarchical representation where critical geometric features are explicitly preserved while less significant areas undergo more aggressive compression. This selective preservation approach may enable the robotic system to maintain sufficient spatial124899-7240-8951\1detail for accurate object manipulation while achieving substantial reductions in data size. The combination of the edge vector set and the identifier-based system can provide a comprehensive method for retaining essential geometric information during the compression process.
[0050] To describe the targeted environment for the 3D data compression mechanism, FIG. 3 illustrates an example robotic system 300 in accordance with one or more embodiments of the present technology. The robotic system 300 can correspond to a portion of the robotic system 100 of FIG. 1 (e.g., the transfer unit 104 of FIG. 1). The robotic system 300 can include a robotic arm 302 corresponding to a kinetic chain. The robotic arm 302 can have an end effector 304 at a distal end of the kinetic chain. The end effector 304 can be configured to interface with and / or manipulate one or more objects.
[0051] For the illustrated example, the end effector 304 can include a gripper (e.g., a vacuum gripper, a pincher, and / or the like) configured to grasp packaged objects 306. The robotic arm 302 can manipulate the kinetic chain to displace the packaged objects 306 between the start location 114 of FIG. 1 (e.g., a platform 308, such as a pallet, a cart, a bin, and / or the like, located at a reference location 310) and the destination location 116 of FIG. 1.
[0052] The robotic system 300 can include one or more of the imaging devices 222 of FIG. 2 configured to depict the packaged objects 306. The imaging devices 222 can provide 2D and / or 3D depictions of the reference locations 310 and / or the packaged objects 306. In some embodiments, one or more of the imaging devices 222 can have an overhead camera configuration 322 where the devices are placed directly over the reference location 310 and facing downward (e.g., -z direction). Accordingly, the overhead camera configuration 322 can provide a top-view of the packaged objects 306 and / or the reference location 310. Additionally or alternatively, one or more of the imaging devices 222 can have an angled camera configuration 324 where the devices are laterally displaced from the reference location 310 and are oriented at an angle away from the downward direction and toward the reference location 310. Accordingly, the angled camera configuration 324 can provide an angled view, such as a perspective view or a side view, of the reference locations 310 and / or the packaged objects 306. The imaging devices 222 can each have a camera location 326 and a camera orientation 328 corresponding to the configuration. The camera location 326 and the camera orientation 328 can be predetermined for the robotic system 300 or dynamically tracked / computed by the robotic system 300.134899-7240-8951\1
[0053] Based on the configuration of the imaging devices 222, the captured image data can have corresponding views, ranges, distances, and / or the like. For example, the space over the reference location 310 can be mapped to a predetermined location / grid map. Stated differently, the operating space for the robotic arm 302 can be pre mapped or identified using a coordinate system (e.g., unique x, y, z values). The 2D / 3D images captured by the imaging devices 222 can represent / identify the locations of the depicted features within the imaged space. For the 3D image data, the robotic system 300 can use the distances between the imaging devices 222 and the depicted portion of the object and / or a corresponding coordinate to locate and image the depicted features. Given the physical aspects of the packaged objects 306, the captured image data can depict planar surfaces 330, such as lateral surfaces 332 (e.g., horizontal surfaces) and vertical surfaces 334, respectively corresponding to top surfaces and peripheral services of the packaged objects 306.
[0054] FIG. 4 illustrates a captured image (e.g., a captured image data 402) in accordance with one or more embodiments of the present technology. The captured image data 402 can include a 3D image / depiction, such as a depth map or a point cloud, of the imaged scene. The captured image data 402 can correspond to a full set of data captured by the imaging devices 222 of FIG. 2, such as before any down-sampling / compression.
[0055] For the illustrated example, the captured image data 402 can correspond to an output of a 3D camera having the overhead camera configuration 322. The captured image data 402 can depict captured surfaces 404 of objects labeled a, b, c, d, and e in FIG. 4. The captured image data 402 can have data points / values corresponding to an imaging granularity / resolution of the corresponding camera. The depicted scene can include an object an adjacent / abutting an object b at a top portion of the image. In the depicted scene, a portion of an object c can be placed over and supported by an object d. Accordingly, the captured image data 402 is shown depicting a "top" surface ci and a "side surface" C2 of the object c. An object e is located at a lower right portion of the imaged region and can have an orientation that is rotated relative to other objects and / or the edges of the platform 308 of FIG. 3.
[0056] The robotic system 100 of FIG. 1 (e.g., using the controller 109 of FIG. 1, a circuit within the imaging device, or a separate computer vision system) can process the captured image data 402 to detect edges. The detected edges 406 may correspond to boundaries or transitions144899-7240-8951\1between different surfaces, objects, or regions within the captured image data 402. In some embodiments, the detected edges 406 can represent locations where there are significant changes (e.g., exceeding a predetermined threshold) in depth values, surface orientations, or material properties between adjacent areas. For example, the detected edges 406 may occur at the boundaries between the captured surfaces 404 of different packaged objects 306, and / or at transitions between horizontal and vertical surfaces of individual objects. The detection of edges 406 can be accomplished through various image processing techniques, such as gradient analysis or using a Sobel filter / operator, where rapid changes in depth or intensity values indicate the presence of an edge.
[0057] The detected corners 408 may represent locations where multiple detected edges 406 intersect or converge. In some aspects, the detected corners 408 can correspond to geometric features such as the corners of rectangular or cubic packaged objects 306, where two or more edges and / or three or more surfaces meet at a common point. The detected corners 408 may be identified through corner detection algorithms that analyze the intersection patterns of the detected edges 406. These corner locations can be particularly valuable for object recognition, pose estimation, and manipulation planning, as they provide distinctive geometric landmarks that can be used to determine the orientation and position of objects within the captured scene.
[0058] As described in further detail below, the 3D data compression mechanism can segment the captured image data 402 using a segmenting grid 410 that includes a set of segment units 412. Each segment unit 412 can correspond to an area (e.g., a predetermined area according to the camera configuration / orientation) in the depicted region. The segmenting grid 410 can have a predetermined number of segment units 412 arranged across the field of view. Accordingly, each of the segment units 412 can have a unique identifier that corresponds to a relative location within the segmenting grid 410. In some embodiments, the grid can correspond to a x-y grid for the overhead camera configuration 322 of FIG. 3 with 0 corresponding to a center of the imaged area. As shown, a maximum x value is represented as +n, a minimum x value is represented as -n, a maximum y value is represented as +m, and minimum y value is represented as -m.
[0059] In some embodiments, the captured image data 402 can be stored as a set of depth values or height values (e.g., distances from the camera or corresponding z coordinates). The lateral positions (e.g., x and y coordinates) of the depth values can be communicated and / or stored154899-7240-8951\1according to a readout format 420a (e.g., a predetermined format or sequence for the data). For example, the captured image data 402 can include a string / stream of values where a predetermined number / set of sequential values represent predetermined locations across a row. Also, for example, the captured image data 402 can include a 2D array that correspond to the x and y coordinates, with the value in each array corresponding to the depth / height value for the corresponding x-y location.
[0060] For illustrated purposes, the examples FIG. 4 shows the segments aligned with the captured image (e.g., the platform and / or one or more packages thereon). However, it is understood that the segments can be oriented differently. For example, the 3D data compression can align the segments to a detected feature (e.g., an edge), an ROI, and / or the like. Using FIG. 4 as an illustrative example, the 3D data compression can tilt or reshape the segments similar to (e.g., parallel to) a slope of the top surface CL Also, using FIG. 4 as an illustrative example, the 3D data compression can rotate the segments 90 degrees in a clockwise direction for alignment with object e.Example Down Sampling Mechanisms
[0061] FIGS. 5A and 5B illustrate example 3D data compression mechanisms (e.g., processing results of the 3D data compression mechanism) in accordance with one or more embodiments of the present technology. FIG. 5 A illustrates a first example output (e.g., a first down-sample format 500) corresponding to the captured image data 402.
[0062] The 3D data compression mechanism can process and down-sample the captured image data 402 using the segmenting grid 410 of FIG. 4. For example, the 3D data compression mechanism can divide or segment the captured image data 402 according to the segmenting grid 410 and combine the data points within each of the segment units 412. Thus, the 3D data compression mechanism can compute a single height / depth value representative of the corresponding data points in the captured image data 402. The 3D data compression mechanism can combine the initial data points according to a predetermined method or equation, such as by computing a mean, a median, a weighted average, an adjusted average, an outlier value, or a combination thereof. Details regarding the corresponding computation are described below.
[0063] As a result of segmenting the captured image data 402 and combining the data points in each of the segment units 412, the 3D data compression mechanism can generate a down-164899-7240-8951\1sampled depth map 502 having a combined depth value 506 for each down-sampled unit 504. Stated differently, in down sampling the image data, the combined depth value 506 can replace and / or represent the set of data points in the corresponding portion of the initial captured image data 402. The down-sampled depth map 502 can be communicated and / or stored according to the readout format 420 of FIG. 4, such as a string or an array of values corresponding to a sequence / pattern of the combined depth values 506.
[0064] The 3D data compression mechanism additionally can include edge information that may be lost in down sampling the captured image data 402. In some embodiments, the 3D data compression mechanism can generate an edge vector set 510 along with the down-sampled depth map 502. The edge vector set 510 may include endpoint pairs 512 that define the locations of the detected edges 406. Each endpoint pair 512 can specify the coordinates of two points that together define a line segment corresponding to one of the detected edges 406. In some embodiments, the endpoint pairs 512 may be expressed in terms of the segmenting grid 410 coordinates, allowing the edge information to be referenced relative to the down-sampled units 504. In other embodiments, the endpoint pairs 512 can be expressed using the coordinate system, such as x, y, and z values.
[0065] The edge vector set 510 may preserve spatial accuracy of edge locations even when the surrounding area has been compressed through the down-sampling process. For example, while multiple data points within a segment unit 412 may be represented by a single combined depth value 506, the edge vector set 510 can maintain the precise locations where edges were detected within that segment. This approach may allow the robotic system 100 to retain geometric detail that could be important for object recognition and manipulation planning.
[0066] For the example illustrated in FIGS. 4 and 5 A, the top surfaces of objects 'a' and 'b' can be two units / measures away from the camera. The surfaces ci and C2 can be slanted, with the corner between the two surfaces being at 3 measures away from the camera and the opposite edge being 8 measures away from the camera. Accordingly, the slanted surface ci can have values ranging from 7 measures to 4 measures. Moreover, the surface C2 can overlap the top surface of object 'd'. The units corresponding to the overlap can have height values that are combined (e.g., averaged). Other partially filled locations, such as the left-most edge of object 'a', left and right edges of object 'd', and the edges of object 'd' can result in adjusted (e.g., averaged) depth values given the location174899-7240-8951\1of the edge within the corresponding down-sampled units 504. Hence, the edge information and locations can be hidden by the adjusted depth values. Similarly, the edge data can be lost when two objects present coplanar surfaces, such as for the edge between object 'a' and 'b'. The edge vector set 510 can use the endpoint pairs 512 to retain the presence and the location for each of the edges.
[0067] Additionally, the 3D data compression mechanism can selectively specify regions of interest (ROIs) 522 within the captured image data 402. The ROIs 522 may correspond to areas within the captured image data 402 that contain depictions of objects or features. In some embodiments, the ROIs 522 can be identified by analyzing the captured image data 402 to locate regions that contain depth values corresponding to detected objects, as opposed to areas that contain NaN values (shown as 'oo' in FIG. 5A) or background surfaces. The 3D data compression mechanism can determine the ROIs 522 using the detected edges 406 or depth values outside of NaN.
[0068] The 3D data compression mechanism may determine the location, size, and shape of each ROI 522 based on the spatial distribution of valid depth measurements within the captured scene. The ROIs 522 can each be defined by a reference location (e.g., top left corner) and a size / dimension according to a predetermined shape (e.g., a rectangle). The 3D data compression mechanism can generate a ROI set 520 including a region identifier 524 for each ROI 522. The region identifier 524 can include a region location 526 representative of the reference location for the corresponding ROI and a region size 528 for representing the dimensions (e.g., number of units along x and y directions) the corresponding ROI.
[0069] In some aspects, the 3D data compression mechanism may provide detailed depth information only for areas within the ROIs 522, while excluding or providing reduced information for areas outside these regions. For example, the corresponding depth map 502 can have a string of values corresponding to the combined depth values in the ROIs 522. The ROI set 520 can identify how the string of values is to be arranged for the corresponding ROIs and where the ROIs are located relative to the segmenting grid 410. Effectively, the 3D data compression mechanism can combine the structured and unstructured aspects of the image data format for additional data savings while preserving relevant details. This selective approach may allow the system to focus computational resources and data storage on areas that are relevant for object manipulation tasks.184899-7240-8951\1
[0070] FIG. 5B illustrates a second example output (e.g., a second down-sample format 550) corresponding to the captured image data 402 of FIG. 4. The second down-sample format 550 can use a set of data keys 540 to generate, store, and use a variable depth map 552 that include multiple types / sizes of units.
[0071] The data keys 540 may include information that defines the structure and interpretation of the variable depth map 552. In some embodiments, the data keys 540 can specify unit sizes 542 that indicate the dimensions of different data units 554 within the variable depth map 552. The unit sizes 542 may allow for variable-sized segments rather than the uniform grid structure illustrated in FIG. 5A, enabling more efficient compression by adapting segment sizes to the content of different regions within the captured image data 402. In some embodiments, the unit sizes 542 can correspond to surface dimensions of known or expected packages and smaller finer tuning dimensions to adjust for real-world variations and edges.
[0072] The data keys 540 may also include edge identifiers 544 that mark existence and / or patterns of detected edges within the corresponding location / segment. The edge identifiers 544 can provide a mechanism for preserving edge information directly within the compressed data structure, rather than storing edge information separately as in the edge vector set 510. In some aspects, the edge identifiers 544 may be embedded within the variable depth map 552 as special values or flags that indicate the presence of edges at particular locations.
[0073] Additionally, the data keys 540 can include corner identifiers 546 that designate existence and / or orientations of corners (e.g., intersection of separate edges). The corner identifiers 546 may work in conjunction with the edge identifiers 544 to create a comprehensive representation of important geometric features within the compressed data. In some embodiments, the corner identifiers 546 may be implemented as specific bit patterns or numerical codes that can be distinguished from regular depth values within the variable depth map 552.
[0074] The data keys 540 may provide a framework for creating a hierarchical compression scheme where different types of information are encoded using different identifiers or patterns. This approach may allow the 3D data compression mechanism to maintain spatial accuracy for geometrically significant features while achieving substantial data reduction in areas with less critical information. The data keys 540 can enable the system to reconstruct the original spatial194899-7240-8951\1relationships and geometric features when the compressed data is processed for object manipulation or analysis tasks.
[0075] The variable depth map 552 can include a set of data units 554 that each correspond to one location and area / portion within the captured image data 402. Each data unit 554 can have a size identifier 556 and a depth value 558.
[0076] The size identifier 556 may specify the spatial dimensions of the corresponding data unit 554 within the variable depth map 552, and the corresponding depth value 558 can be attributed to the area covered by the size identifier 556. In some embodiments, the size identifier 556 can indicate the area coverage of the data unit 554 in terms of the number of original data points or pixels from the captured image data 402 that are represented by the single data unit 554. For example, the size identifier 556 may specify that a particular data unit 554 represents a 2x2, 4x4, or 8x8 area of the segment units 412, allowing for adaptive compression based on the content characteristics of different regions.
[0077] The size identifier 556 may be determined based on the depth values / height of the region being compressed. In areas with relatively uniform depth values or minimal geometric features, the size identifier 556 may be larger to achieve greater compression ratios. Conversely, in areas containing edges, corners, or significant depth variations, the size identifier 556 may be smaller to preserve spatial detail and accuracy.
[0078] In some aspects, the size identifier 556 can be encoded using a value corresponding to one of the predetermined size categories or patterns that are referenced in the data keys 540. The variable sizing approach may enable the 3D data compression mechanism to balance compression efficiency with preservation of geometrically significant features within the captured image data 402.
[0079] The data unit 554 can also include a boundary value 560 representative of one of the edge identifiers 544 or one of the corner identifiers 546 corresponding to the pattern of depth values found in the corresponding portion of the captured image data 402. The boundary values 560 can be included in the data stream (e.g., the string or the array) corresponding to the variable depth map 552.204899-7240-8951\1
[0080] The 3D data compression mechanism can use the readout format 420 of FIG. 4 or a corresponding pattern to organize / arrange the data units 554 in generating the variable depth map 552. Accordingly, the 3D data compression mechanism can use the readout format 420 in communicating, storing, and / or further processing (e.g., reconstructing / decompressing from) the variable depth map 552. Further, given the variable sizes of the represented segments, the 3D data compression mechanism can dynamically track a full-row end 562 and a partial-row end 564 in organizing or spatially arranging the data units 554. In generating the variable depth map 552 or in using it to reconstruct the captured image, the 3D data compression mechanism can first satisfy a row width of the scene (e.g., from -n to +n and fill empty spots or indentations provided by the different column heights of the data units 554 in the same row format starting from the most- indented location and until all the indentations are filled. The full-row end 562 can correspond to the greatest column height for given row, and the partial-row end 564 can correspond to an end of the shorter rows used to fill the indentations.
[0081] Using the example illustrated in FIG. 5B, the 3D data compression mechanism can follow predetermine rules to fill the first row at +m. In doing so, the 3D data compression mechanism can set the full-row end 562 to four rows down (+m - 4) based on the tallest segment in the first row corresponding to the '8' size identifier 556. The indentation begins at row +m - 1 and the left most position of -n. Accordingly, the 3D data compression mechanism can assign the next data unit to the indentation location and fill the remaining width of the partial row until the size '8' cell. When the size '8' cell is reached, the 3D data compression mechanism can identify the partial-row end 564 to mark an end for the corresponding partial row. Filling the first partial row (e.g., second line entry in the variable depth map 552) can result in two size T indentations under the two edge identifiers 554 with 'a' values. Thus, the following partial row (e.g., the third line entry) can be used to fill the two sized ' indentations. The process can repeat to fill the next row (e.g., the fourth row) and so forth.
[0082] In some instances, filling the indentation can result in the filler region in the indentation extending past the full-row end 562 along the y-direction. For example, in filling the fourth row, the size '4' region can extend past the initially set full-row end 562 shown as T. Accordingly, the 3D data compression mechanism can adjust the full-row end 562 to the '2' position associated with the newly added size '4' regions. The same adjustment can occur due to another size '4' region used in the next partial row, thus causing the 3D data compression mechanism can adjust the full-row214899-7240-8951\1end 562 to the '3' position. For each row, the data unit 554 controlling the full-row end 562 is shown in bold and underlined.
[0083] Alternatively, the variable depth map 552 can be combined with the ROI set 520 of FIG. 5A. For example, instead of storing or communicating the NaN values, the 3D data compression mechanism can first identify the ROIs 522 of FIG. 5 A and retain the depth measures within the ROIs 522. The 3D data compression mechanism can subsequently down-sample the depth values within the ROIs 522. In down sampling, the 3D data compression mechanism can utilize variable unit sizes 542 that can best capture the necessary details while minimizing the amount of data. In some embodiments, the 3D data compression mechanism can use a predetermined process / routine, a predetermined model, and / or the like in computing the sizes of the data unit that minimizes the overall data size while maintaining a minimum granularity / accuracy for the edges for the corresponding input captured image.
[0084] Also, as an alternative combination, the variable depth map 552 can use the various unit sizes 542 and use the edge vector set 510 instead of the edge identifiers 544 and the corner identifiers 546. In yet another embodiment, the variable depth map 552 can use the various unit sizes 542 to provide depth measures up to the detected edges 406 and rely on the transition in the depth values across data units to depict the edges. Stated differently, the variable depth map 552 can use finely coordinated unit sizes 542 and corresponding depth values to depict the detected edges 406 and the detected corners 408 without using the edge identifiers 544 and the corner identifiers 546. Similarly, various features of the first down sample format 500 of FIG. 5 A and the second down sample format 550 can be combined.
[0085] FIG. 6 illustrates an example 3D image data within a reference segment in accordance with one or more embodiments of the present technology. FIG. 6 depicts a representation of 3D image data that demonstrates the segmentation and data organization approach used by the 3D data compression mechanism. The reference segment can consist of multiple cells or data units arranged in a grid pattern, where each cell can be identified by a unique cell reference number that corresponds to its position within the segment. The cell reference numbers can use x-y coordinate system that uses row numbers and column numbers. Also, each cell or region can be given a unique number, such as 1, 2, 7, 118, etc. shown within the cells in FIG. 6. These cell numbers can be used to identify the location. The locations with the shaded circle can have a depth value. Stated224899-7240-8951\1differently, the shaded circles can represent a depth value, ranging from -128 to +127, as shown using the bar on the right side of FIG. 6.
[0086] The cells may be arranged in a predetermined pattern where each cell reference number can be conceptualized as representing an integer coordinate on an x-y plane. The coordinate system may use a specific number of bits to represent the x-coordinate and y-coordinate positions, allowing each cell reference number to be encoded as a multi-bit identifier. The filled circles within certain cells represent depth data points corresponding to 3D image data at those particular locations. Each of these data points has an associated depth value that indicates the distance or height measurement for that specific cell location.
[0087] Cells that do not contain data points may be designated as having NaN (not-a-number) values, which can correspond to locations where no depth data was detected or where depth values fall outside a region of interest or exceed predetermined thresholds. The 3D data compression mechanism can process this reference segment by analyzing the distribution of valid data points versus NaN locations to determine appropriate compression strategies. The segment structure allows the system to evaluate whether sufficient data points exist within the segment to warrant compression using representative values, or whether the segment should be designated as a nonsegment or processed using alternative compression approaches.
[0088] The reference segment illustration demonstrates how the 3D data compression mechanism can organize and analyze spatial data within defined boundaries, enabling efficient processing and compression of 3D image data while maintaining spatial relationships and coordinate information for subsequent reconstruction or analysis tasks.
[0089] For brevity, the reconstruction / decompression process of the 3D data compression mechanism is not illustrated. However, it is understood that the 3D data compression mechanism can generate an estimate of the captured image by organizing / placing the data units / cells and populating additional features / pixel values with the same depth value for the cell. Effectively, the 3D data compression mechanism can extrapolate and / or interpolate the removed data points (e.g., according to predetermined granularity / point density) within each of the segments, thereby generating the contour from the down sampled data.Example Down Sampling Results234899-7240-8951\1
[0090] FIGS. 7A, 7B and 8A - 8C illustrate various stages of down sampling using the 3D data compression mechanism as described above. For example, FIG. 7A illustrates an example capture data (e.g., 2D and 3D combination image) in accordance with one or more embodiments of the present technology. FIG. 7B illustrates an example down-sampling result of the captured data of FIG. 7A in accordance with one or more embodiments of the present technology. Each dot within the down-sampling result represents a height value for a corresponding segment / data unit.
[0091] FIGS. 8A - 8C illustrate a sequence of image processing results that compress the data and then decompress the compressed result to recreate or estimate the initial image data. FIG. 8A illustrates an example capture data 802 from a perspective-view camera in accordance with one or more embodiments of the present technology. The captured data 802 can include a 3D point cloud from a 3D camera. In some embodiments, the 2D image (e.g., visual picture) can be superimposed on the 3D point cloud to provide a combined image.
[0092] FIG. 8B illustrates a down-sampled result 804 of the captured data 802 of FIG. 8 A in accordance with one or more embodiments of the present technology. The 3D data compression mechanism can process the captured data 802 to generate the down-sampled result 804 as described herein. For example, the 3D data compression mechanism can segment the captured data 802 and then compute a representative value (shown as a point or a dot in FIG. 8B) for each segment. Accordingly, the overall size of the image file, such as from the captured data 802 to the down-sampled result 804, can be reduced by a double-digit factor or by a triple-digit factor (e.g., reduced to 1 / 100 or 1 / 200 of the original capture data). In some embodiments, the 2D visual image can be compressed and / or stored separately.
[0093] FIG. 8C illustrates a reconstruction result 806 from the down-sampled result 804 of FIG. 8B in accordance with one or more embodiments of the present technology. The 3D data compression mechanism can effectively reverse the down-sampling process to decompress the down-sampled result 804. In doing so, the 3D data compression mechanism can generate a reconstruction result 806 that corresponds to an estimate or a recreation of the captured data 802 of FIG. 8A. For example, the 3D data compression mechanism can determine a number of data points within each segment and then assign the representative data for the corresponding segment to the data points therein in generating the reconstruction result 806. In other embodiments, the 3D data compression mechanism can be configured to compute slopes / gradients between data244899-7240-8951\1points and interpolate the depth values according to the computed slopes / gradients. As described above, the 3D data compression mechanism can overlay or separately combine the 2D image data or a reconstruction thereof over the repopulated depth values.Control Flow
[0094] FIGS. 9 A - 9C illustrate flow diagrams of example method for compressing 3D data in accordance with one or more embodiments of the present technology. The method can be implemented using the 3D data compression mechanism, such as implemented through the controller 109 of FIG. 1, the processor 202 of FIG. 2, the imaging device 222 of FIG. 2, a different portion of the robotic system 100 of FIG. 1, a different computer vision system, or a combination thereof.
[0095] FIG. 9A illustrates a method 900 for compressing the 3D data. At block 902, the robotic system 100 may obtain image data (e.g., the captured image data 402 of FIG. 4) from one or more of the imaging devices 222. The image data can include point cloud data, depth maps, or other 3D representations of a scene captured by the imaging devices 222. The obtained 3D image data may represent captured surfaces 404 of FIG. 4 of objects within the field of view of the imaging devices 222, such as the packaged objects 306 of FIG. 3 positioned on the platform 308 of FIG. 3.
[0096] At block 904, the 3D data compression mechanism may segment the obtained image data into a plurality of segments. The segmentation process can divide the captured image data 402 into discrete units or regions according to a predetermined segmentation scheme. In some embodiments, the 3D data compression mechanism can segment the captured image data using the segmenting grid 410 of FIG. 4, which can include the segment units 412 of FIG. 4 that each correspond to a specific area within the depicted region.
[0097] At block 905, the 3D data compression mechanism may obtain edges and / or corners detected from the captured image data. The edge detection process can analyze the captured image data 402 to identify boundaries or transitions between different surfaces, objects, or regions within the scene. In some embodiments, the edge detection may be performed by evaluating changes in depth values, surface orientations, or material properties between adjacent areas within the captured image data 402.254899-7240-8951\1
[0098] The edge detection process may utilize various image processing techniques to identify locations where significant changes occur. For example, the 3D data compression mechanism can apply gradient analysis methods that calculate the rate of change in depth values across neighboring data points. In some aspects, the system may employ edge detection operators such as Sobel filters, Canny edge detectors, or other gradient-based algorithms that can identify rapid transitions in the 3D image data.
[0099] The 3D data compression mechanism can implement the edge detection process or obtain the results thereof. In some embodiments, the 3D data compression mechanism can store the edge detection results as part of the edge vector set 510 of FIG. 5 A, such as using the endpoint pairs 512 of FIG. 5 A.
[0100] At block 906, the 3D data compression mechanism may determine representative data for each segment. The representative data determination process can analyze the depth values within each segment to compute a single value that represents the spatial characteristics of that segment. In some embodiments, the 3D data compression mechanism can calculate the representative data using statistical methods such as computing a mean, median, mode, or weighted average of the depth values within each segment. Additionally, the 3D data compression mechanism may derive groupings of adjacent segments. The 3D data compression mechanism may calculate one representative data for each of the derived groupings.
[0101] FIG. 9B illustrates an example detailed flow for the block 906 of FIG. 9A. In some embodiments, the 3D data compression mechanism can determine the representative data by iteratively selecting the segments (e.g., the segment units 412 of FIG. 4) as shown in block 912. In some embodiments, the segment selection may be performed iteratively and sequentially according to a predetermined order, such as following the readout format 420 of FIG. 4. The segment selection may correspond to moving the pointer or loading the next grouping of data values into an evaluation buffer.
[0102] At decision block 914, the 3D data compression mechanism can determine whether the loaded values include an end of file (EOF) indicator. When the selected segment does not correspond to the EOF indicator, the 3D data compression mechanism can determine whether the selected segment unit includes at least one data point (e.g., at least one valid depth value), as shown at decision block 916. The selected segment unit may include no data points when the264899-7240-8951\1corresponding portion of the scene does not include or depict an object. Stated differently, the selected segment unit may include no data points when the corresponding portion of the captured image data 402 includes NaN without other values / depth measures. When the selected segment unit includes no data points, as shown in block 918, the 3D data compression mechanism can set a default value, such as NaN or a corresponding value (e.g., +128, all 1 bits, and / or the like), for the selected segment unit.
[0103] Otherwise, when the selected segment unit 412 includes at least one depth value, the 3D data compression mechanism can compute a representative value for the selected segment unit as shown at block 920. Details regarding the computation are described below.
[0104] After computing the representative value or setting the default value, the 3D data compression mechanism can select the next segment or grouping of data as shown by the feedback loop to block 912. Accordingly, the 3D data compression mechanism can iteratively process each of the segment units 412 for the captured image data 402 until the EOF is reached. Effectively, the 3D data compression mechanism can implement a first level of down-sampling by computing / assigning one value for each of the segment units 412.
[0105] After computing the representative / default values for the segment units 412, the 3D data compression mechanism can determine groupings. For example, the 3D data compression mechanism can group adjacent segment units having the same representative / default values or values within a threshold range of each other as a single and larger unit. The 3D data compression mechanism can use the unit sizes 542 as a template in determining the groupings. Upon grouping the adjacent segment units, the 3D data compression mechanism can assign a representative value for the grouping. Effectively, the 3D data compression mechanism can utilize the same processes described for blocks 916, 918, and 920 to assign the representative value.
[0106] Additionally or alternatively, the 3D data compression mechanism can determine the groupings that correspond to the ROIs 522 of FIG. 5. The 3D data compression mechanism can determine the ROIs 522 by identifying the segment units and / or groupings that have valid depth values instead of NaN. As described above, the 3D data compression mechanism can generate the down-sampled output (e.g., a depth map) using the determined groupings, the ROIs, and / or the segment units, with or without specifying the locations having the NaN or default values, as described above.274899-7240-8951\1
[0107] Referring back to block 920, FIG. 9C illustrates an example detailed flow for the block 920 of FIG. 9B. In computing a representative value, the 3D data compression mechanism can determine whether the selected segment unit includes or overlaps a portion of a detected edge or a detected corner. The 3D data compression mechanism can use the edge vector set 510 to determine whether the selected segment unit includes or overlaps the edge / corner. In other embodiments, the 3D data compression mechanism can implement the edge detection within the selected segment unit. When one or more edges are detected for the selected segment unit, as shown at block 934, the 3D data compression mechanism can determine the edge pattern. For example, the 3D data compression mechanism can determine an orientation of the corresponding edge / corner, such as by comparing the edge to one or more template patterns. At block 936, the 3D data compression mechanism can select an edge value from the edge identifiers 544 of FIG. 5B or the corner identifiers 546 of FIG. 5B according to the determined orientation.
[0108] Otherwise, when the selected segment unit does not correspond to an edge, the 3D data compression mechanism can determine a minimum value and a maximum value therein as shown at block 942. For example, the 3D data compression mechanism can iteratively evaluate / compare the values within the selected segment unit to determine the minimum and maximum values. At block 944, the 3D data compression mechanism can compute a difference between the minimum and maximum values for the selected segment unit.
[0109] At decision block 946, the 3D data compression mechanism can determine whether computed difference exceeds a predetermined threshold. When the difference exceeds the threshold, the 3D data compression mechanism can iteratively parse the values and adjust the difference value. At decision block 948, the 3D data compression mechanism can determine whether a threshold for the iterative parse and adjustment has been reached. When the maximum iteration is reached, the 3D data compression mechanism can set the representative value as an outlier value (e.g., a predetermined value representative of an error, a likely estimate, NaN, or the like). In some embodiments, the outlier value can be stored as is; the 3D data compression mechanism can store the full set of data points within the corresponding segment without compression.
[0110] Otherwise, if the number of iterations has not reached the maximum iteration count, the 3D data compression mechanism can adjust the difference value as shown in block 950.284899-7240-8951\1In some embodiments, the 3D data compression mechanism can first parse the values within the selected segment unit into a predetermined number of (e.g., two) groupings. For example, the 3D data compression mechanism can calculate a statistical median or an average value among the depth measures and then group the values as above and below the calculated median / average. The 3D data compression mechanism can be configured to adjust the difference up or down based on the number of values in each group. As an illustrative example, the 3D data compression mechanism can compute a separation between the number of values in the above / below groupings and then decrease the difference value by an amount proportional to the separation. Accordingly, the 3D data compression mechanism can decrease the difference value by a larger amount when a larger number of depth measures are clustered closer together and by a smaller amount when the depth measures are more evenly spread across the dividing line / value. After the adjustment, the 3D data compression mechanism can evaluate the adjusted difference value using the threshold, as shown by the feedback loop to the decision block 946.[OHl] When the difference value is less than the threshold (e.g., the range of depth values are within an expected range of each other), the 3D data compression mechanism can determine whether the number / count of the data values are above a maximum count as shown at decision block 954. If the number / count is not greater than the maximum count, the 3D data compression mechanism can similarly determine whether the number / count is less than a minimum count as shown at block 958. When the number / count is greater than the maximum or not less than the minimum, the 3D data compression mechanism can compute a representative data value that effectively combines the depth values for the selected segment unit as shown at block 956. The 3D data compression mechanism can compute the representative data value using one or more of various methods. For instance, if the datapoints correspond to depth values, the 3D data compression mechanism may determine that the representative value is equal to the datapoint with the highest value. In other embodiments, the 3D data compression mechanism can compute the representative value as the mean or the median value of the datapoints within the segment or the larger set of parsed datapoints. As another example, the representative value may be determined based on both the depth values and the positional values of the 3D image data.
[0112] Otherwise, when the number / count of the values is less than the minimum threshold, the 3D data compression mechanism can set the representative value as an outlier as described294899-7240-8951\1above for block 952. The subroutine can return after selecting the edge value, setting the outlier, or computing the representative data value for the selected segment unit.Examples
[0113] The present technology is illustrated, for example, according to various aspects described below. Various examples of aspects of the present technology are described as numbered examples (1, 2, 3, etc.) for convenience. These are provided as examples and do not limit the present technology. It is noted that any of the dependent examples can be combined in any suitable manner, and placed into a respective independent example. The other examples can be presented in a similar manner.1. A method for compressing image data, comprising: obtaining the image data including a plurality of data points each representing a depiction of a spatial location corresponding to a set of coordinates; segmenting / dividing / grouping the image data into a plurality of segments with each segment corresponding to a portion of the image data; and determining representative data for each segment by computing a reduced number of one or more values that represent a greater number of data points within the segment.2. A method for compressing three-dimensional (3D) image data, comprising: obtaining the 3D image data from an imaging device, the 3D image data including a plurality of data points each representing a set of spatial coordinates; segmenting the 3D image data into a plurality of segments across a two-dimensional (2D) arrangement, each segment corresponding to a predetermined area within the 3D image data; and determining representative data for each segment by computing a single value that represents multiple data points within the segment to down sample the 3D image data.304899-7240-8951\13. The method of one or more examples herein, one or more portions thereof, or a combination thereof, wherein the representative data includes a statistical measure that combines the multiple data points within each segment.4. The method of one or more examples herein, one or more portions thereof, or a combination thereof, wherein the statistical measure comprises one of a mean, a median, and a weighted average of depth values within the segment.5. The method one or more examples herein, one or more portions thereof, or a combination thereof, further comprising: generating a depth map including data units that each correspond to one of the segments, wherein each data unit includes a corresponding one of the representative data; detecting edges within the 3D image data; and generating an edge vector set based on the detected edges, wherein the edge vector set is for locating the detected edges in addition to the depth map.6. The method of one or more examples herein, one or more portions thereof, or a combination thereof, further comprising: generating a depth map including data units that each correspond to one of the segments, wherein each data unit includes a corresponding one of the representative data, wherein the representative data for one or more of the data units include an edge identifier or a corner identifier representing that an edge or a corner, respectively, is depicted at a corresponding location within the 3D image data.7. The method of one or more examples herein, one or more portions thereof, or a combination thereof, further comprising: detecting edges within the 3D image data; and determining the representative data includes identifying that the corresponding one of the data units overlaps at least a portion of one of the detected edges.314899-7240-8951\18. The method of one or more examples herein, one or more portions thereof, or a combination thereof, wherein determining the edge identifier or the corner identifier for the representative data includes performing edge detection using the data points within the corresponding one of the data units.9. The method of one or more examples herein, one or more portions thereof, or a combination thereof, further comprising: identifying regions of interest (ROIs) within the 3D image data, wherein each of the ROIs include a unique and non- overlapping grouping of adjacent segments that have valid values instead of a not-a-number (NaN) designation; for each of the ROIs, determining a region location and a region size for locating and describing the corresponding one of the ROIs within the 3D image data; and generating a depth map using the region location, the region size, and the representative data for each of the ROIs and without including the NaN outside of the ROIs.10. The method of one or more examples herein, one or more portions thereof, or a combination thereof, wherein determining representative data for each segment includes: grouping a set of adjacent segments as a single unit using one of predetermined unit sizes; and computing one representative value for the single unit that replaces the data points for the set of adjacent segments.11. The method of one or more examples herein, one or more portions thereof, or a combination thereof, further comprising: generating a variable depth map including units having two or more unique sizes, wherein each unit includes the one representative value for a corresponding set of adjacent segments, wherein the units are arranged according to (1) a full-row end that corresponds to a greatest column height for a given row within the variable depth map and (2) a partial-row end that corresponds to an end of one or more shorter rows used to fill indentations within the given row.324899-7240-8951\112. The method of one or more examples herein, one or more portions thereof, or a combination thereof, wherein determining representative data for each segment includes: determining a maximum value and a minimum value among data points within a corresponding segment; computing a difference between the maximum value and the minimum value; and assigning an outlier value for the representative data when the difference is greater than a predetermined threshold.13. The method of one or more examples herein, one or more portions thereof, or a combination thereof, wherein determining representative data for each segment includes: determining a quantity of data points within a corresponding segment; and assigning an outlier value for the representative data when the quantity is less than a predetermined minimum.14. A robotic system comprising: at least one processor; at least one memory including processor instructions that, when executed, causes the at least one processor to perform the method of one or more of examples 1-13, one or more portions thereof, or a combination thereof.15. A non-transitory computer readable medium including processor instructions that, when executed by one or more processors, causes the one or more processors to perform the method of one or more of examples 1-13, one or more portions thereof, or a combination thereof.Remarks
[0114] The above detailed descriptions of embodiments of the technology are not intended to be exhaustive or to limit the technology to the precise form disclosed above. Although specific embodiments of, and examples for, the technology are described above for illustrative purposes, various equivalent modifications are possible within the scope of the technology as those skilled in the relevant art will recognize. For example, although steps are presented in a given order above,334899-7240-8951\1alternative embodiments may perform steps in a different order. Furthermore, the various embodiments described herein may also be combined to provide further embodiments.
[0115] From the foregoing, it will be appreciated that specific embodiments of the technology have been described herein for purposes of illustration, but well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments of the technology.
[0116] Where the context permits, singular or plural terms may also include the plural or singular term, respectively. In addition, unless the word “or” is expressly limited to mean only a single item exclusive from the other items in reference to a list of two or more items, then the use of “or” in such a list is to be interpreted as including (a) any single item in the list, (b) all of the items in the list, or (c) any combination of the items in the list. Furthermore, as used herein, the phrase “and / or” as in “A and / or B” refers to A alone, B alone, and both A and B. Additionally, the terms “comprising,” “including,” “having,” and “with” are used throughout to mean including at least the recited feature(s) such that any greater number of the same features and / or additional types of other features are not precluded. Moreover, as used herein, the phrases “based on,” “depends on,” “as a result of,” and “in response to” shall not be construed as a reference to a closed set of conditions. For example, an exemplary step that is described as “based on condition A” may be based on both condition A and condition B without departing from the scope of the present disclosure. In other words, as used herein, the phrase “based on” shall be construed in the same manner as the phrase “based at least in part on” or the phrase “based at least partially on.”
[0117] From the foregoing, it will also be appreciated that various modifications may be made without deviating from the disclosure or the technology. For example, one of ordinary skill in the art will understand that various components of the technology can be further divided into subcomponents, or that various components and functions of the technology may be combined and integrated. In addition, certain aspects of the technology described in the context of particular embodiments may also be combined or eliminated in other embodiments. Furthermore, although advantages associated with certain embodiments of the technology have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments need necessarily exhibit such advantages to fall within the scope of the technology.344899-7240-8951\1Accordingly, the disclosure and associated technology can encompass other embodiments not expressly shown or described herein.354899-7240-8951\1
Claims
CLAIMSWhat is claimed is:
1. A method for compressing three-dimensional (3D) image data, comprising: obtaining the 3D image data from an imaging device, the 3D image data including a plurality of data points each representing a set of spatial coordinates; segmenting the 3D image data into a plurality of segments across a two-dimensional (2D) arrangement, each segment corresponding to a predetermined area within the 3D image data; and determining representative data for each segment by computing a single value that represents multiple data points within the segment to down sample the 3D image data.
2. The method of claim 1, wherein the representative data includes a statistical measure that combines the multiple data points within each segment.
3. The method of claim 2, wherein the statistical measure comprises one of a mean, a median, and a weighted average of depth values within the segment.
4. The method of claim 1, further comprising: generating a depth map including data units that each correspond to one of the segments, wherein each data unit includes a corresponding one of the representative data; detecting edges within the 3D image data; and generating an edge vector set based on the detected edges, wherein the edge vector set is for locating the detected edges in addition to the depth map.
5. The method of claim 1, further comprising: generating a depth map including data units that each correspond to one of the segments, wherein each data unit includes a corresponding one of the representative data,364899-7240-8951\1wherein the representative data for one or more of the data units include an edge identifier or a corner identifier representing that an edge or a corner, respectively, is depicted at a corresponding location within the 3D image data.
6. The method of claim 5, further comprising: detecting edges within the 3D image data; and determining the representative data includes identifying that the corresponding one of the data units overlaps at least a portion of one of the detected edges.
7. The method of claim 5, wherein determining the edge identifier or the corner identifier for the representative data includes performing edge detection using the data points within the corresponding one of the data units.
8. The method of claim 1, further comprising: identifying regions of interest (ROIs) within the 3D image data, wherein each of the ROIs include a unique and non- overlapping grouping of adjacent segments that have valid values instead of a not-a-number (NaN) designation; for each of the ROIs, determining a region location and a region size for locating and describing the corresponding one of the ROIs within the 3D image data; and generating a depth map using the region location, the region size, and the representative data for each of the ROIs and without including the NaN outside of the ROIs.
9. The method of claim 1, wherein determining representative data for each segment includes: grouping a set of adjacent segments as a single unit using one of predetermined unit sizes; and computing one representative value for the single unit that replaces the data points for the set of adjacent segments.
10. The method of claim 9, further comprising:374899-7240-8951\1generating a variable depth map including units having two or more unique sizes, wherein each unit includes the one representative value for a corresponding set of adjacent segments, wherein the units are arranged according to (1) a full-row end that corresponds to a greatest column height for a given row within the variable depth map and (2) a partial-row end that corresponds to an end of one or more shorter rows used to fill indentations within the given row.
11. The method of claim 1, wherein determining representative data for each segment includes: determining a maximum value and a minimum value among data points within a corresponding segment; computing a difference between the maximum value and the minimum value; and assigning an outlier value for the representative data when the difference is greater than a predetermined threshold.
12. The method of claim 1, wherein determining representative data for each segment includes: determining a quantity of data points within a corresponding segment; and assigning an outlier value for the representative data when the quantity is less than a predetermined minimum.
13. A computer vision system configured to provide down sampled data used for controlling a robotic system to transfer packages, the computer vision system comprising: an interface configured to obtain three-dimensional (3D) image data; a processor coupled to the interface and configured to generate a down sampled depth map from the 3D image data; and a memory coupled to the processor and including instructions that, when executed by the processor, causes the processor to: segment the 3D image data into a plurality of non-overlapping two-dimensional segments according to a predetermined segmentation scheme, wherein the384899-7240-8951\1two-dimensional segments correspond to depictions of planar surfaces of the packages in the 3D image data; determine a representative value for each of the segments by computing a combined depth value from multiple depth values within the corresponding one of the segments; and generate the down sampled depth map based on arranging combined depth values for the segments.
14. The robotic system of claim 13, wherein the representative value includes wherein a mean, a median, or a weighted average of the multiple depth values within the segment, an edge identifier, a corner identifier, or a not-a-number designation.
15. The robotic system of claim 13, wherein the instructions are further configured to cause the processor to: generate an edge vector set that identifies locations of edges depicted in the 3D image data, wherein the edge vector set is for preserving the edges in down sampling the 3D image data.
16. The robotic system of claim 13, wherein the instructions are further configured to cause the processor to: identify regions of interest (ROIs) within the 3D image data, wherein each of the ROIs include a unique and non- overlapping grouping of adjacent segments that have valid values instead of a not-a-number (NaN) designation; for each of the ROIs, determining a region location and a region size for locating and describing the corresponding one of the ROIs within the 3D image data; and generate the down sampled depth map based on including the region location, the region size, and the representative value for each of the ROIs without including data points outside of the ROIs.
17. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method comprising:394899-7240-8951\1obtaining a three-dimensional (3D) image data depicting one or more packages, the 3D image data including a plurality of data points each representing a set of spatial coordinates; segmenting the 3D image data into a plurality of segments across a two-dimensional (2D) arrangement, each segment corresponding to a unique area within the 3D image data; and determining representative data for each segment by computing a single value that represents multiple data points within the segment to down sample the 3D image data.
18. The non-transitory computer-readable medium of claim 17, wherein the method further comprises: grouping a set of adjacent segments as a single unit using one of two or more unit sizes; computing one representative value for the single unit that replaces the data points for the set of adjacent segments; and generating a variable depth map including units having two or more unique sizes, wherein each unit includes the one representative value for a corresponding set of adjacent segments, wherein the units are arranged according to (1) a full-row end that corresponds to a greatest column height for a given row within the variable depth map and (2) a partial-row end that corresponds to an end of one or more shorter rows used to fill indentations within the given row.
19. The non-transitory computer-readable medium of claim 17, wherein the method further comprises: determining a maximum value and a minimum value among data points within a corresponding segment; computing a difference between the maximum value and the minimum value; and assigning an outlier value for the representative data when the difference is greater than a predetermined threshold.404899-7240-8951\120. The non-transitory computer-readable medium of claim 17, wherein the method further comprises: determining a quantity of data points within a corresponding segment; and assigning an outlier value for the representative data when the quantity is less than a predetermined minimum.414899-7240-8951\1