Method and apparatus for picking up object having high transmittance or reflectance

The method and device address the challenge of accurately recognizing and picking transparent or highly reflective objects by using detection and restoration models to generate accurate point cloud information, enabling improved robotic picking accuracy and efficiency.

WO2025116521A1PCT designated stage expired Publication Date: 2025-06-05CMES INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018983
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-26
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing depth sensor and camera technologies struggle to accurately obtain point cloud information for transparent, translucent, or highly reflective objects, often resulting in noise or lack of information, which hinders robots' ability to accurately recognize and pick such objects.

Method used

A method and device that acquire an image and depth information, detect an object region using a detection model, generate a point cloud or restore depth information using a restoration model to remove noise, and estimate a picking position and direction based on the cleaned data.

Benefits of technology

The solution enables accurate recognition of the position and direction of transparent, translucent, or highly reflective objects, allowing robots to automatically determine the picking position and direction, thereby improving work speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018983_05062025_PF_FP_ABST
    Figure KR2024018983_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and an apparatus for picking up an object having high transmittance or reflectance are disclosed. According to an aspect of the present disclosure, a method by which a picking apparatus picks up an object having high transmittance or reflectance is provided, the method comprising the steps of: acquiring an image and depth information; using a detection model to detect, from the image, an object region in which an object appears; generating a point cloud generated from the depth information, a point cloud from which noise is removed by inputting the depth information into a restoration model, or depth information from which noise is removed; and estimating the picking position and direction from the point cloud from which the noise is removed or the depth information from which the noise is removed.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for picking objects with high transmittance or reflectance

[0001] The present disclosure relates to a method and device for picking an object with high transmittance or reflectivity.

[0002]

[0003] The content described below merely provides background information related to the present embodiment and does not constitute prior art.

[0004] The use of robots is increasing in smart factories and logistics automation systems, and cameras and depth sensors working alongside robots play a crucial role in determining the location and orientation of objects. When the location and orientation of an object are fixed or known in advance, picking is possible without depth information. However, in environments where objects are randomly placed, a point cloud containing depth information and object orientation is essential.

[0005] Depth information, or point clouds, play a crucial role in enabling robots to accurately determine the location and orientation of objects. However, for transparent, translucent, or highly reflective objects, conventional depth sensor and camera technologies struggle to obtain accurate point cloud information, often resulting in noise or incomplete information. Therefore, a technology is needed to restore depth information or point clouds to enable accurate recognition and picking of transparent or highly reflective objects.

[0006]

[0007] The present disclosure aims to restore an accurate point cloud of a transparent, translucent or reflective object, and enable a robot to automatically estimate a picking position and direction based on the point cloud.

[0008] The present disclosure aims to restore accurate depth information of a transparent, translucent or reflective object and enable a robot to automatically estimate a picking position and direction based on the depth information.

[0009] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0010]

[0011] According to one aspect of the present disclosure, a method is provided, performed by a picking device, for picking an object having a high transmittance or reflectance, comprising: a step of acquiring an image and depth information; a step of detecting an object region in which an object appears from the image using a detection model; a step of generating a point cloud generated from the depth information or inputting the depth information into a restoration model to generate a point cloud from which noise has been removed or depth information from which noise has been removed; and a step of estimating a picking position and direction from the point cloud from which noise has been removed or the depth information from which noise has been removed.

[0012] According to another aspect of the present disclosure, there is provided a device comprising at least one memory; and at least one processor, wherein the at least one processor executes instructions to acquire an image and depth information, detect an object region in which an object appears from the image using a detection model, generate a point cloud generated from the depth information or input the depth information into a restoration model to generate a point cloud from which noise has been removed or depth information from which noise has been removed, and estimate a picking position and direction from the point cloud from which noise has been removed or the depth information from which noise has been removed.

[0013]

[0014] According to one embodiment of the present disclosure, a point cloud of a transparent, translucent, or highly reflective object that is difficult to detect with existing sensors can be restored, thereby accurately recognizing the position and direction of the object.

[0015] According to one embodiment of the present disclosure, by using restored point cloud information, the robot automatically determines the picking position and direction of an object, thereby improving work speed and accuracy.

[0016] According to one embodiment of the present disclosure, a robot can perform picking by considering the actual direction of an object through a restored point cloud, and can perform efficient work by accurately placing the picked object in a desired location and direction.

[0017] According to one embodiment of the present disclosure, depth information of transparent, translucent, or highly reflective objects that are difficult to detect with existing sensors can be restored, thereby accurately recognizing the spatial characteristics and direction of the objects.

[0018] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0019]

[0020] FIG. 1 is a schematic block diagram of a picking device according to one embodiment of the present disclosure.

[0021] FIG. 2 is a block diagram illustrating an execution device according to one embodiment of the present disclosure.

[0022] FIG. 3a is a diagram schematically illustrating a process in which a learning device according to one embodiment of the present disclosure learns a detection model.

[0023] FIG. 3b is a diagram schematically illustrating a process in which a learning device according to one embodiment of the present disclosure learns a restoration model.

[0024] FIG. 4a is a diagram schematically illustrating a process in which a learning device according to one embodiment of the present disclosure learns a detection model.

[0025] FIG. 4b is a diagram schematically illustrating a process in which a learning device according to one embodiment of the present disclosure learns a restoration model.

[0026] FIG. 5a is a schematic diagram illustrating a process in which a second detection unit detects a picking position and direction according to one embodiment of the present disclosure.

[0027] FIG. 5b is a diagram schematically illustrating a process in which a second detection unit detects a picking position and direction according to one embodiment of the present disclosure.

[0028] FIG. 6 is a flowchart showing the operation process of a picking system according to one embodiment of the present disclosure.

[0029] FIG. 7 is a flowchart showing the operation process of a picking system according to another embodiment of the present disclosure.

[0030] FIG. 8 is a flowchart showing the operation process of a picking system according to another embodiment of the present disclosure.

[0031] FIG. 9 is a block diagram schematically illustrating an exemplary computing device that can be used to implement a method or device according to the present disclosure.

[0032]

[0033] Hereinafter, some embodiments of the present disclosure will be described in detail using exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, when describing the present disclosure, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present disclosure.

[0034] In describing components of embodiments according to the present disclosure, symbols such as first, second, i), ii), a), b) may be used. These symbols are only for distinguishing the components from other components, and the nature, order, or sequence of the components are not limited by the symbols. When a part in the specification is said to "include" or "have" a component, this does not mean that other components are excluded, but rather that other components may be included, unless explicitly stated otherwise.

[0035] The detailed description set forth below, together with the accompanying drawings, is intended to explain exemplary embodiments of the present disclosure and is not intended to represent the only embodiments in which the present disclosure may be practiced.

[0036]

[0037] FIG. 1 is a schematic block diagram of a picking device (10) according to one embodiment of the present disclosure.

[0038] A picking device (10) according to one embodiment of the present disclosure may include all or part of an execution device (100) and a learning device (120). The components illustrated in FIG. 1 represent functionally distinct elements, and at least one of the components may be implemented in an integrated form in an actual physical environment.

[0039] The execution device (100) recognizes an object and direction from a two-dimensional image, and detects object area and / or direction information. The object area refers to a specific portion of the image that must be detected to recognize the object. Specifically, it does not simply refer to the overall outline or boundary of the object, but may also include important sub-areas of the object so that the object's direction and posture can be identified. The direction information may include various directional elements such as the object's rotation, inclination, front / back, left / right, etc. The execution device (100) can acquire depth information from a depth sensor and generate a point cloud. The execution device (100) can input the point cloud within the detected object area into the reconstruction model (122b) to generate a restored point cloud. As another example, the execution device (100) can input the two-dimensional image and the entire depth information into the reconstruction model (122b) to generate restored depth information. Furthermore, the point cloud can be generated from the restored depth information. Accordingly, the execution device (100) can estimate the picking location and direction based on the restored point cloud, the restored depth information, or the point cloud generated from the restored depth information. This can be utilized for bin picking, which requires identifying the exact shape and orientation of an object. Alternatively, it can be utilized for order picking, which requires performing picking on any arbitrary object.

[0040] The learning device (120) is a device that performs data processing, model training, performance evaluation, etc. to train artificial intelligence models (122). The learning device (120) may include a high-performance computing device such as a GPU (Graphics Processing Unit) or a TPU (Tensor Processing Unit), software for executing a learning algorithm, and a system for data processing and storage. The artificial intelligence model (122) may include a detection model (122a), a restoration model (122b), etc. The process by which the learning device (120) trains the artificial intelligence models (122) is described below with reference to FIGS. 3a, 3b, 4a, and 4b.

[0041] FIG. 2 is a block diagram illustrating an execution device (100) according to one embodiment of the present disclosure. To explain FIG. 2, FIG. 1 may also be referred to.

[0042] In one embodiment of the present disclosure, the execution device (100) may include all or part of the first detection unit (102), the restoration unit (104), and the second detection unit (106).

[0043] The first detection unit (102) uses a detection model (122a) to detect object region and / or direction information in which the object appears from an image of the target object. The detection model (122a) used by the first detection unit (102) is an artificial intelligence model and may be an object detection neural network including at least one structural feature. That is, the detection model (122a) may be any one of a convolutional neural network (CNN), a recurrent neural network (RNN), a transformer, YOLO (You Only Look Once), Faster R-CNN (Region-based Convolutional Neural Networks), SSD (Single Shot MultiBox Detector), and a reinforcement learning model.

[0044] The detection model (122a) used by the first detection unit (102) is an artificial intelligence model that includes at least one of a convolution operation, a recursive information processing, and an attention mechanism. A convolutional neural network using a convolution operation is a type of artificial neural network specialized in image processing and pattern recognition. It is effective in tasks such as image recognition, object detection, and segmentation, and is designed to understand and process the spatial structure of images. A recurrent neural network, which processes information recursively, is a type of artificial neural network specialized in processing sequential data. Sequential data refers to a series of data arranged in order. A transformer using an attention mechanism is one of the deep learning model architectures used in the fields of natural language processing and computer vision. This model is used in various natural language understanding tasks such as machine translation, sentence generation, summarization, and question answering, as well as in fields such as object detection in images. YOLO is a single-step object detection algorithm. The YOLO algorithm divides the original image into grids of equal size. For each grid, the number of bounding boxes with a predefined shape centered around the grid center is predicted, and the confidence level is calculated based on this. This includes whether the image contains an object or is just a background, and the location with high object confidence is selected to determine the object category. Faster R-CNN is a two-stage object detection algorithm. Faster R-CNN uses a backbone network based on a convolutional neural network to extract image features. The backbone network generates feature maps at various levels from the image. SSD is a neural network model for object detection, and it uses a base network based on a convolutional neural network to extract image features. Typically, networks such as VGG16 and ResNet are used as base networks.SSD uses multiple feature maps generated from the base network to detect objects of different sizes and proportions. A reinforcement learning model is a machine learning model in which an agent learns optimal actions to maximize rewards while interacting with the environment.

[0045] The first detection unit (102) transmits the detected object area and / or direction information to the restoration unit (104).

[0046] The restoration unit (104) receives depth information from a sensor. A sensor is a device that collects data from the physical environment and converts it into a form that can be recognized by a robotic system or automated device. The sensor may be at least one of a depth sensor, an ultrasonic sensor, or an infrared sensor.

[0047] The restoration unit (104) outputs restored depth information or restored point cloud using the restoration model (122b). At this time, the restored depth information or restored point cloud is in a form in which noise has been removed.

[0048] The restoration model (122b) is an artificial intelligence model, and may be any one of a convolutional neural network (CNN), a recurrent neural network (RNN), a transformer, an autoencoder, a generative adversarial network (GAN), and a neural network model or generative model having an encoder-decoder structure based on a convolution operation.

[0049] Convolutional neural networks (CNNs) are a type of artificial neural network specialized in image processing and pattern recognition. They are effective for tasks such as image recognition, object detection, and segmentation, and are designed to understand and process the spatial structure of images. Recurrent neural networks (RCNs) are a type of AI specialized in processing sequential data. Sequential data refers to a series of data arranged in order. Transformers are one of the deep learning model architectures used in natural language processing. Autoencoders use a neural network structure that encodes and then decodes input data to learn important features of the data. Generative adversarial networks (GANs) are models in which two neural networks, a generator and a discriminator, compete to learn.

[0050] The restoration unit (104) preprocesses the object region and / or direction information and the point cloud. The point cloud within the object region preprocessed by the restoration unit (104) becomes the input of the restoration model (122b). That is, the restoration model (122b) can receive the point cloud within the object region. In addition, the restoration model (122b) can receive a two-dimensional image and depth information. Accordingly, the restoration model (122b) outputs a restored point cloud using the point cloud within the object region. Alternatively, the restoration model (122b) outputs restored depth information using the two-dimensional image and depth information.

[0051] The second detection unit (106) can detect the picking position and direction considering the object's posture based on the restored point cloud. In addition, the second detection unit (106) can detect the picking position and direction based on the restored depth information.

[0052] FIG. 3a is a diagram schematically illustrating a process in which a learning device (120) according to one embodiment of the present disclosure learns a detection model (122a).

[0053] The learning device (120) collects a two-dimensional image of the target object.

[0054] The learning device (120) labels the correct object area in the collected two-dimensional image. The learning device (200) can label the entire area and / or a partial area. For example, the learning device (200) can label the body area and the lid area as separate partial areas. In this case, the correct object area can be manually performed by a person or automatically acquired through a virtual environment simulator.

[0055] When a labeled image is input to a detection model (122a), the detection model (122a) outputs a predicted object region. At this time, the learning device (120) trains the detection model (122a) so that the loss between the correct object region and the predicted object region is reduced. In other words, the parameters of the detection model (122a) are updated so that the detection model (122a) can predict a region identical to and / or similar to the correct object region from a given image.

[0056] FIG. 3b is a schematic diagram illustrating a process in which a learning device (120) according to one embodiment of the present disclosure learns a restoration model (122b). To explain FIG. 3b, FIG. 3a may also be referred to.

[0057] The learning device (120) collects a noisy point cloud of the target object from the sensor. The point cloud can be limited to point clouds located within the object area obtained using the learned detection model (122a). For this purpose, the sensor coordinate system and the two-dimensional image coordinate system may be pre-aligned.

[0058] The learning device (120) labels the collected point cloud with a point cloud with less noise. For example, the point cloud with less noise can be acquired from a 3D model of the target object. The 3D model can be automatically generated based on point clouds collected at various viewpoints or can be pre-made. Additionally, during the point cloud labeling process, the 3D model can be aligned to a noisy point cloud area, and then a point cloud with less noise can be acquired.

[0059] When a noisy point cloud is input to a restoration model (122b), the restoration model (122b) outputs a restored point cloud. At this time, the restoration model (122b) is trained so that the loss of the restored point cloud and the point cloud with less noise is reduced. In other words, the parameters of the restoration model (122b) are updated so that the restoration model (122b) can remove noise from the given point cloud in the same and / or similar form as the point cloud with less noise.

[0060] FIG. 4a is a diagram schematically illustrating a process in which a learning device (120) according to one embodiment of the present disclosure learns a detection model (122a).

[0061] The learning device (120) collects a two-dimensional image of the target object.

[0062] The learning device (120) labels the correct object area in the collected two-dimensional image. At this time, the correct object area may be manually labeled by a person, but is not limited thereto.

[0063] When a labeled image is input to a detection model (122a), the detection model (122a) outputs a predicted object region. At this time, the learning device (120) trains the detection model (122a) so that the loss between the correct object region and the predicted object region is reduced. In other words, the parameters of the detection model (122a) are updated so that the detection model (122a) can predict a region identical to and / or similar to the correct object region from a given image.

[0064] FIG. 4b is a diagram schematically illustrating a process in which a learning device (120) according to one embodiment of the present disclosure learns a restoration model (122b).

[0065] The learning device (120) collects noisy depth information of a target object from a sensor. In some examples, the sensor coordinate system and the two-dimensional image coordinate system may be pre-aligned.

[0066] The learning device (120) labels the collected depth information with less noise. For example, to obtain less noise depth information, objects that cause noise are made opaque or have low reflectivity, and then a sensor is used to acquire less noise depth information. At this time, special sprays or general paint sprays that assist in 3D scanning of objects that cause noise may be used.

[0067] When noisy depth information is input to a restoration model (122b), the restoration model (122b) outputs restored depth information. At this time, the restoration model (122b) is trained so that the loss of the restored depth information and the depth information with little noise is reduced. In other words, the parameters of the restoration model (122b) are updated so that the restoration model (122b) can remove noise from the given depth information in a form identical to and / or similar to the depth information with little noise.

[0068] FIG. 5a is a schematic diagram illustrating a process in which a second detection unit (106) detects a picking position and direction according to one embodiment of the present disclosure.

[0069] The second detection unit (106) can detect the picking position and direction considering the object's posture based on the restored point cloud. The restored point cloud may be a point cloud generated from restored depth information.

[0070] The second detection unit (106) can acquire a direction based on the object's area information and acquire a picking location and direction using the restored point cloud and a normal vector, which is a geometric feature. This can be utilized for bin picking, which requires identifying and performing the exact shape and orientation of an object.

[0071] FIG. 5b is a schematic diagram illustrating a process in which a second detection unit (106) according to one embodiment of the present disclosure detects a picking position and direction.

[0072] The second detection unit (106) can detect the picking position and direction based on the restored depth information.

[0073] The second detection unit (106) can acquire the picking position and direction using the object's area information, restored depth information, and the normal vector, which is a geometric feature. Therefore, it can be utilized for order picking, which must be performed on any object.

[0074] FIG. 6 is a flowchart showing the operation process of a picking system (10) according to one embodiment of the present disclosure.

[0075] The execution device (100) can acquire depth information from a two-dimensional image and a depth sensor and generate a point cloud (S602).

[0076] The execution device (100) recognizes an object and direction from a two-dimensional image, and uses a detection model (122a) to detect an object region and / or direction information where the object appears from the image (S604). The object region refers to a specific portion of the image that must be detected to recognize the object. Specifically, it does not simply refer to the overall outline or boundary of the object, but may also include important subregions of the object so that the object's direction and posture can be identified. The direction information may include various directional elements such as the object's rotation, inclination, front / back, or left / right.

[0077] The learning device (120) is a device that performs data processing, model training, performance evaluation, etc. to train artificial intelligence models (122).

[0078] The execution device (100) can input a point cloud generated from depth information into a restoration model (122b) to generate a restored point cloud (S606).

[0079] The execution device (100) can estimate the picking position and direction considering the pose of the object based on the restored point cloud (S608).

[0080] FIG. 7 is a flowchart showing the operation process of a picking system (10) according to another embodiment of the present disclosure.

[0081] The execution device (100) can acquire depth information from a two-dimensional image and depth sensor (S702).

[0082] The execution device (100) recognizes an object and direction from a two-dimensional image and detects an object area using a detection model (122a) (S704).

[0083] The learning device (120) is a device that performs data processing, model training, performance evaluation, etc. to train artificial intelligence models (122).

[0084] The execution device (100) can input a two-dimensional image and full depth information into the restoration model (122b) to generate restored depth information (S706).

[0085] The execution device (100) can estimate the picking position and direction based on the restored depth information (S708).

[0086] Figure 8 is a flowchart showing the operation process of a picking system (10) according to another embodiment of the present disclosure.

[0087] The execution device (100) can acquire depth information from a two-dimensional image and depth sensor (S802).

[0088] The execution device (100) recognizes an object and direction from a two-dimensional image and detects an object area using a detection model (122a) (S804).

[0089] The learning device (120) is a device that performs data processing, model training, performance evaluation, etc. to train artificial intelligence models (122).

[0090] The execution device (100) can input a two-dimensional image and full depth information into a restoration model (122b) to generate restored depth information (S806). Furthermore, a point cloud can be generated from the restored depth information (S808).

[0091] The execution device (100) can estimate the picking position and direction based on the point cloud generated from the restored depth information (S810).

[0092] FIG. 9 is a block diagram schematically illustrating an exemplary computing device (90) that can be used to implement a method or device according to the present disclosure.

[0093] The computing device (90) may include some or all of a memory (900), a processor (920), storage (940), an input / output interface (960), and a communication interface (980). The computing device (90) may be a stationary computing device such as a desktop computer, a server, or the like, as well as a mobile computing device such as a laptop computer, a smart phone, or the like. The computing device (90) may also include any specialized hardware accelerator capable of efficiently processing operations for an artificial intelligence model. For example, the computing device (90) may include a graphic processing unit (GPU), a Tensor Processing Unit (TPU), or a neural processing unit (NPU).

[0094] The memory (900) may store a program that causes the processor (920) to perform a method or operation according to various embodiments of the present disclosure. For example, the program may include a plurality of instructions executable by the processor (920), and the above-described method or operation may be performed by executing the plurality of instructions by the processor (920). The memory (900) may be a single memory or a plurality of memories. In this case, information required to perform the method or operation according to various embodiments of the present disclosure may be stored in a single memory or may be divided and stored in a plurality of memories. When the memory (900) is composed of a plurality of memories, the plurality of memories may be physically separated. The memory (900) may include at least one of a volatile memory and a non-volatile memory. The volatile memory includes a static random access memory (SRAM) or a dynamic random access memory (DRAM), and the non-volatile memory includes a flash memory.

[0095] The processor (920) may include at least one core capable of executing at least one instruction. The processor (920) may execute instructions stored in the memory (900). The processor (920) may be a single processor or multiple processors.

[0096] Storage (940) maintains stored data even when power supplied to the computing device (90) is cut off. For example, storage (940) may include non-volatile memory, or may include storage media such as magnetic tape, optical disk, or magnetic disk. A program stored in storage (940) may be loaded into memory (900) before being executed by processor (920). Storage (940) may store a file written in a programming language, and a program generated from the file by a compiler or the like may be loaded into memory (900). Storage (940) may store data to be processed by processor (920) and / or data processed by processor (920).

[0097] The input / output interface (960) may provide an interface with an input device such as a keyboard, mouse, etc. and / or an output device such as a display device, printer, etc. A user may trigger the execution of a program by the processor (920) through an input device and / or check the processing result of the processor (920) through an output device.

[0098] The communication interface (980) may provide access to an external network. The computing device (90) may communicate with other devices via the communication interface (980).

[0099] At least some of the components described in the exemplary embodiments of the present disclosure may be implemented as hardware elements including at least one or a combination of a Digital Signal Processor (DSP), a processor, a controller, an Application-Specific Integrated Circuit (ASIC), a programmable logic device (FPGA, etc.), and other electronic devices. In addition, at least some of the functions or processes described in the exemplary embodiments may be implemented as software, and the software may be stored on a recording medium. At least some of the components, functions, and processes described in the exemplary embodiments of the present disclosure may be implemented as a combination of hardware and software.

[0100] The method according to exemplary embodiments of the present disclosure can be written as a program that can be executed on a computer, and can also be implemented in various recording media such as a magnetic storage medium, an optical readable medium, a digital storage medium, etc.

[0101] Implementations of the various technologies described herein may be implemented as digital electronic circuitry, or as computer hardware, firmware, software, or combinations thereof. Implementations may be implemented as a computer program product, i.e., a computer program tangibly embodied in an information carrier, e.g., a machine-readable storage medium (computer-readable medium) or a radio signal, for processing by the operation of a data processing device, e.g., a programmable processor, a computer, or multiple computers, or for controlling the operation thereof. A computer program, such as the computer program(s) described above, may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be processed on one computer or multiple computers at a single site, or to be distributed across multiple sites and interconnected by a communications network.

[0102] Processors suitable for processing a computer program include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random-access memory, or both. Components of a computer may include at least one processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer may include, or be coupled to receive data from, transmit data to, or both, one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data. Information carriers suitable for embodying computer program instructions and data include, for example, semiconductor memory devices, magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as compact disk read only memory (CD-ROM), digital video disks (DVD), magneto-optical media such as floptical disks, read only memory (ROM), random access memory (RAM), flash memory, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), etc. The processor and memory may be supplemented by, or included in, special purpose logic circuitry.

[0103] A processor can execute an operating system and software applications running on the operating system. Furthermore, the processor device can access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processor device is sometimes described as being used singly; however, those skilled in the art will appreciate that the processor device can include multiple processing elements and / or multiple types of processing elements. For example, the processor device can include multiple processors, or one processor and one controller. Other processing configurations, such as parallel processors, are also possible.

[0104] Additionally, non-transitory computer-readable media can be any available media that can be accessed by a computer, and can include both computer storage media and transmission media.

[0105] While this specification contains details of a number of specific implementations, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be unique to particular embodiments of particular inventions. Certain features described herein in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, either individually or in any suitable subcombination. Furthermore, although features may operate in a particular combination and may initially be described as being claimed as such, one or more features from a claimed combination may in some cases be excluded from that combination, and the claimed combination may be modified into a subcombination or variation of a subcombination.

[0106] Likewise, while operations are depicted in the drawings in a particular order, this should not be construed as requiring that those operations be performed in the particular or sequential order depicted to achieve desired results, or that all depicted operations be performed. In certain instances, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various device components of the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the program components and devices described may generally be integrated together in a single software product or packaged into multiple software products.

[0107] Meanwhile, the embodiments of the present invention disclosed in this specification and drawings are merely specific examples presented to aid understanding and are not intended to limit the scope of the present invention. It will be apparent to those skilled in the art that other modifications based on the technical concepts of the present invention are possible in addition to the embodiments disclosed herein.

[0108] The scope of protection of this embodiment should be interpreted by the claims below, and all technical ideas within the equivalent scope should be interpreted as being included in the scope of rights of this embodiment.

[0109]

[0110] CROSS-REFERENCE TO RELATED APPLICATION

[0111] This patent application claims priority to Korean Patent Application No. 10-2023-0166755, filed in Korea on November 27, 2023, and Korean Patent Application No. 10-2024-0170394, filed in Korea on November 26, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. A method performed by a picking device to pick an object with high transmittance or reflectivity, The process of acquiring image and depth information; A process of detecting an object region in which an object appears from the image using a detection model; A process of generating a point cloud generated from the above depth information or inputting the depth information into a restoration model to generate a point cloud or depth information from which noise has been removed; and A method comprising a process of estimating a picking location and direction from a point cloud from which the noise has been removed or from depth information from which the noise has been removed.

2. In paragraph 1, The above detection process is, A method for further detecting information indicating the direction of the object from the image.

3. In paragraph 1, The above restoration model is, A method learned in the form of removing noise from the point cloud or the depth information.

4. In paragraph 1, The above generating process is, A method for generating depth information with noise removed by inputting the depth information and image information into the restoration model.

5. In paragraph 1, The above generating process is, A method further comprising a process of generating a point cloud from the depth information from which the noise has been removed.

6. At least one memory; and Containing at least one processor, At least one of said processors executes instructions, Acquire image and depth information, Using the detection model, the object area where the object appears is detected from the image, A point cloud generated from the above depth information or a point cloud with noise removed or depth information with noise removed is generated by inputting the depth information into a restoration model, A device for estimating a picking position and direction from a point cloud from which the noise has been removed or from depth information from which the noise has been removed.

7. In paragraph 6, The above detection process is, A device for further detecting information indicating the direction of the object from the image.

8. In paragraph 6, The above restoration model is, A device learned in the form of removing noise from the point cloud or the depth information.

9. In paragraph 6, The above generating process is, A device that inputs the depth information and image information into the restoration model to generate depth information with the noise removed.

10. In paragraph 6, The above generating process is, A device further comprising a process of generating a point cloud from the depth information from which the noise has been removed.

Citation Information

Patent Citations

  • A Vehicle Target Detection Method and System Based on Adaptive Fusion of Ravages Semantic Segmentation

    CN114724120B

  • A semiconductor device, and a method of fabricating of the same

    KR1020220111758A

  • System for providing child care

    KR102138144B1

  • Apparatus and method for estimating the attitude of a picking object

    KR102261498B1

  • KR20210082281A