Training Method of Deep Network
Patent Information
- Application Number
- JP2022503981
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-13
- Filing Date
- 2020-06-05
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2040-06-05
AI Technical Summary
Conventional systems for training robotic devices to detect objects in real-world environments fail to account for variations in deformations, object articulations, viewing angles, and lighting, leading to ineffective object detection due to environmental diversity.
A method involving a 3D camera to capture images from different viewpoints, adjust lighting conditions, and generate paired 3D images to create reference images with embedded descriptors, which are used to train deep neural networks for robust object detection.
Enables robotic devices to accurately identify objects in unknown environments by correlating real-time images with trained reference images, overcoming variations in deformation, articulation, and lighting.
Smart Images

Figure 00000016_0000 
Figure 00000017_0000 
Figure 00000018_0000
Abstract
Description
Technical Field
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 877,792, filed on July 23, 2019, entitled "Keyframe Matcher", U.S. Provisional Patent Application No. 62 / 877,791, filed on July 23, 2019, entitled "Visual Instruction and Repetition for Operations - Instruction VR", and U.S. Provisional Patent Application No. 62 / 877,793, filed on July 23, 2019, entitled "Visualization", and claims the benefit of U.S. Patent Application No. 16 / 570,813, filed on September 13, 2019, entitled "Method for Training a Deep Network", the entire contents of which are incorporated herein by reference.
[0002] Certain aspects of the present disclosure generally relate to object detection training, and more specifically, to systems and methods for training deep networks.
Background Art
[0003] A robotic device may use one or more sensors (e.g., as a camera) to identify objects in an environment based on training of the robotic device using real-world images. In real-life situations, however, the encountered images may be different from the real images used to train the robotic device. That is, the deformation, object articulation, viewing angle, and lighting diversity in the image data used for training may prevent object detection in real-world operations.
[0004] Conventional systems collect training images in the real world in the actual situations where observations are expected. For example, the training of a robotic device is limited to the actual situations used when collecting training images, including the actual lighting level and a specific viewing angle at which the training images were collected. These conventional systems do not consider the diversity of the environment. These differences between the training data and real-world objects are particularly problematic when training a deep neural network for a robotic device to perform object detection.
Summary of the Invention
[0005] A method for training a deep neural network in a robotic device is described. The method involves constructing a 3D model using images captured via a 3D camera in the training environment. The method also involves generating pairs of 3D images from the 3D model by artificially adjusting the parameters of the training environment to form a manipulated image using the deep neural network. The method further involves processing the pairs of 3D images to form a reference image containing embedded descriptors of objects common to the pairs of 3D images. The method further involves using the reference image from the neural network training to identify detected objects in future images and determine correlations.
[0006] A method for controlling a robotic device based on the identification of detected objects in an unknown environment is described. The method includes detecting objects in an unknown environment. The method also includes selecting a corresponding reference image containing an embedded descriptor corresponding to a trained object manipulated by artificially adjusting the parameters of the image acquisition environment. The method further includes identifying the detected object based on the embedded descriptor of the corresponding reference image.
[0007] A system for controlling a robotic device based on the identification of detected objects in an unknown environment is described. The system includes a pre-trained object identification module. The object identification module selects a corresponding reference image to identify a detected object in the captured image. The corresponding reference image includes an embedded descriptor based on the trained object, which is manipulated by artificially adjusting the parameters of the image acquisition environment. The system also includes a controller that selects autonomous actions for the robotic device based on the identity of the detected object.
[0008] The above provides a broad overview of the features and technical advantages of this disclosure in order to better understand the detailed explanation that follows. Additional features and advantages of this disclosure are described below. It will be understood by those skilled in the art that this disclosure can be readily used as a basis for modifying or designing other structures to accomplish the same purposes as this disclosure. It will also be recognized by those skilled in the art that such equivalent configurations do not deviate from the teachings of this disclosure as set forth in the attached claims. The new features that are considered to be features of this disclosure, along with their configuration and operation, will be better understood from the following description when considered in conjunction with the attached drawings, along with further purposes and advantages. However, it should be clearly understood that each drawing is provided for illustrative and explanatory purposes only and is not intended to define the limits of this disclosure. [Brief explanation of the drawing]
[0009] The functions, nature, and benefits of this disclosure will become more apparent from the detailed description below, when considered in conjunction with the drawings, which use similar reference letters to correspond throughout.
[0010] [Figure 1] Original images of the environment used for training robots in the manner of this disclosure are shown. [Figure 2] Examples of manipulated images created using 3D models, used for training a robot in a training environment, according to aspects of this disclosure, are shown. [Figure 3A] The following are pairs of images of a training environment generated for training a robot according to an aspect of this disclosure. [Figure 3B] The following are pairs of images of a training environment generated for training a robot according to an aspect of this disclosure. [Figure 4A] This disclosure shows images of a real-world environment captured by a robot in the manner described herein. [Figure 4B] This disclosure shows images of a real-world environment captured by a robot in the manner described herein. [Figure 5]This figure shows an example of a hardware implementation of an object recognition system according to the embodiments of this disclosure. [Figure 6] This is a flowchart showing a method for training a deep neural network of a robotic device according to the embodiments of this disclosure. [Modes for carrying out the invention]
[0011] The detailed descriptions relating to the accompanying drawings provided below are intended to describe various configurations and not to present a single configuration that implements the concepts described herein. The detailed descriptions include certain details for the purpose of providing a complete understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. In some cases, well-known structures and components are shown in block diagrams to avoid obscuring such concepts.
[0012] A robotic device may use one or more sensors to identify objects in its environment. These sensors may include red-green-blue (RGB) cameras, radar detection and ranging (RADAR) sensors, light detection and ranging (LiDAR) sensors, or other types of sensors. In the images captured by the sensors, one or more objects are identified by the robotic device based on training of its deep neural network to perform object identification. In real-world situations, however, the images encountered may differ from the real-world images used to train the robotic device. That is, deformations, object articulation, viewing angles, and changes in lighting in the image data used for training may interfere with object detection in real-world operation.
[0013] Conventional systems collect training images in the real world under actual conditions where observation is expected. For example, the real-world conditions under which training images are collected include the actual lighting levels and specific viewing angles at which the training images were taken. These conventional systems do not take into account changes in the environment. These changes in training data and real-world objects become particularly problematic when training a deep neural network for a robotic device to perform object detection.
[0014] This disclosure relates to providing data for training deep networks by taking into account environmental changes. Changes include object deformation, object articulation, changes in viewing angle, and / or changes in lighting.
[0015] For the sake of simplification, robotic devices may be referred to as robots in this disclosure. In addition, objects may include static and dynamic objects in the environment. Objects may include artificial objects (e.g., chairs, desks, cars, books, etc.), natural objects (e.g., rocks, trees, animals, etc.), and humans.
[0016] Figure 1 shows an original image 101 of a training environment 102 used to train a robot 100 according to an aspect of this disclosure. In the example in Figure 1, the robot 100 is a humanoid robot, and the training environment 102 is a kitchen. The aspects of this disclosure are not limited to humanoid robots. The robot 100 may be any type of autonomous or semi-autonomous device, such as a drone or a vehicle. In addition, the robot 100 may be in any environment.
[0017] In one configuration, robot 100 acquires an original image 101 of the training environment 102 using one or more of its sensors. Robot 100 may detect and locate one or more objects in the original image 101. Locatement is the determination of the position (e.g., coordinates) of the detected object in the original image 101. Conventional object detection systems may use bounding boxes to indicate the position of the detected object in the original image 101. The detected object may be one or more objects of a specific class, such as a table 104, a pushed-in chair 106, a closed window 108, a bottle 110, utensils 120 and 122, a counter 140, a sink 142, a cabinet 130 with a handle 132, or all objects in the original image 101. The objects may be detected and identified using an object detection system, such as a pre-trained object detection neural network.
[0018] In one configuration, a 3D camera on robot 100 captures images of the training environment 102 from different viewpoints / angles. A 3D model of the training environment 102 is generated from the captured images. The 3D model is used to create images from a viewpoint different from that of the original image 101 captured by the 3D camera. The 3D model is also used to change lighting conditions in the created images (e.g., adjust the lighting level). In addition, the 3D model may create images that include the object being manipulated. For example, the 3D model may create a scene with a drawer / window that is open or closed. Furthermore, the system associates images with common features. The captured images and images created by the 3D model (e.g., training data) are used to train a deep network object detection system.
[0019] Figure 2 shows an example of a manipulated image 200 created from a 3D model used to train the robot 100 in a training environment 202 according to an aspect of this disclosure. In the example in Figure 2, the training environment 202 is the kitchen in Figure 1, with the elements horizontally inverted to provide different viewpoints. The robot 100 may detect and identify objects in each image via an object detection system such as a pre-trained object recognition neural network using the original image 101 of Figure 1 and the manipulated image 200.
[0020] In this configuration, the system generates a manipulated image 200 and pairs it with the original image 101 of the training environment 102 shown in Figure 1. According to aspects of this disclosure, linked elements are identified between the original image 101 and the manipulated image 200. That is, such elements in the training environment 202 may be given pixel coordinates. Overlapping pixel coordinates indicate overlapping portions (e.g., linked elements). For example, the pulled-out chair 206 is a linked element between the manipulated image 200 and the pushed-in chair 106 of the original image 101. The link indicates that the same element is drawn with different articulations. The linked portion may be defined by the correspondence of points between the original image 101 and the manipulated image 200 (e.g., the same viewpoint as the original image 101).
[0021] In this example, the closed window 108 in the original image 101 is paired with the open window 208 in the operation image 200. For example, the glass of the open window 208 is linked between the original image 101 and the operation image 200. In addition, the table 204 in the operation image 200 is also linked to the table 104 in the original image 101. Similarly, the bottle 210 in the operation image 200 is also linked to the bottle 110 in the original image 101. The bottle 110 is placed on counter 240, which is linked to counter 140 in the original image 101. The sink 242 is also linked between the operation image 200 and the original image 101. In addition, the cabinet 230 and handle 232 in the operation image 200 are also linked to the cabinet 130 and handle 132 in the original image 101.
[0022] The robot 100 is trained to detect the pulled-out chair 206, which is horizontally inverted, from the pushed-in chair 106 shown in FIG. 1. Similarly, the robot 100 is trained to follow the bottle 210 placed on the table 204 after being moved from the counter 240. Additionally, the robot 100 is trained to follow the area of the appliances 220 and 222 moved from the table 204 to the counter 240. Although the original image 101 and the manipulated image 200 are shown, it should be recognized that aspects of the present disclosure including the generation of additional manipulated images are possible under various lighting conditions, viewing angles, deformations, etc.
[0023] According to aspects of the present disclosure, paired images in a 3D environment are processed by an image-to-image neural network. The network receives an RGB image as input and outputs an embedding or descriptor image that includes values assigned to each pixel. The embedding / descriptor image may provide a digital "fingerprint" represented by numbers to distinguish one feature from another by encoding information into a series of numbers. Ideally, this information should be invariant even when image transformation is performed. Unfortunately, traditional systems are generally trained without considering environmental changes, so traditional feature descriptors are not invariant even when image transformation is performed.
[0024] In this embodiment of the Disclosure, the embedded / descriptor image determines its correlation to future images (e.g., images taken in real time as the robot 100 operates) that define objects and points in the environment. That is, after training, when placed in a new environment, the robot identifies the location of objects in the new environment that can be manipulated, such as chairs, windows, bottles, utensils (e.g., spoons), cabinets, etc. The robot 100 may also identify various elements regardless of deformation, object articulation, viewing angle, and lighting. For example, an object detected in a different orientation than the original image 101 can be easily identified based on elements linked to a descriptor image created from paired images (e.g., the original image 101 and the manipulated image 200).
[0025] Figures 3A and 3B show pairs of images of a training environment 302 generated for training a robot 100 according to an aspect of this disclosure. As shown in Figures 3A and 3B, the training system automatically generates pairs of images in which similar elements in different images are linked. For example, Figure 3A shows the original image 300 of the training environment 302. The original image 300 further shows a cabinet 330 including a counter 340, a sink 342, and a handle 332. In this example, the cabinet 330 is closed.
[0026] Figure 3B shows an operational image 350 of a training environment 302 according to an aspect of the present disclosure. In this example, the handle 332 of the cabinet 330 in a scene where the cabinet 330 is closed (e.g., Figure 3A) is linked to a scene where the cabinet 330 is open. In addition, instruments 320 and 322 are paired between the original image 300 (e.g., inside the cabinet 330) and the operational image 350 (e.g., showing the cabinet 330 in an open state). The pairing of the original image 300 and the operational image 350 links objects that are similar elements but depicted with different articulations. The linked portion is defined by the point correspondence between the operational image 350 and the original image 300. Corresponding elements between paired images may be determined by identifying overlapping portions of the training environment 302 captured in each image (i.e., scene).
[0027] The pair of images is then processed by an image-to-image neural network, which receives an RGB image as input and outputs an embedded or descriptor image generated by assigning a value to each pixel of the image. According to aspects of this disclosure, the embedded is used to determine correlation with future images (e.g., images taken in real time during robot operation). For example, the embedded may define objects and points in the environment to identify correlated objects. In other words, the system can quickly determine the location in the environment and identify objects through the correlation of the embedded with real-time images, for example, as shown in Figures 4A and 4B.
[0028] Figures 4A and 4B show images of an unknown environment 402 captured by the robot 100 according to an aspect of the present disclosure. In the example of Figures 4A and 4B, the unknown environment 402 is a restaurant including a table 404, a pulled-out chair 406, an open window 408, bottles 410, utensils 420 and 422, and a cabinet 430. In one configuration, the robot 100 uses reference images based on pairs of original and manipulated images of a training environment, such as the kitchen training environment shown in Figures 1, 2, 3A, and 3B. Using the reference images, the robot 100 detects the pulled-out chair 406 using a pre-trained object detection neural network. In addition, the reference images enable the robot 100 to detect the open window 408.
[0029] Figure 4A shows an image 400 of an unknown environment 402 captured by a 3D camera of robot 100 according to an aspect of the present disclosure. In the example of Figure 4A, the unknown environment 402 is a restaurant including a table 404, a retracted chair 406, and an open window 408. In one configuration, robot 100 uses reference images based on pairings of images of a training environment, such as the kitchen training environment shown in Figures 1, 2, 3A, and 3B. Using the reference images, robot 100 uses a pre-trained object detection neural network to locate the retracted chair 406. In addition, the reference images enable robot 100 to identify the open window 408.
[0030] As further shown in Figure 4A, the reference image enables the robot 100 to detect utensils 420 and 422 on table 404. In addition, the reference image enables the robot 100 to detect jar 410 on cabinet 430. Detection is not limited to the position and / or orientation of an object in the environment. According to aspects of this disclosure, the robot 100 is trained to track the movement of an object over time. For simplicity, kitchen items are used as examples of objects to be detected. Nevertheless, aspects of this disclosure are not limited to the detection of kitchen items, and other objects are also considered.
[0031] Figure 4B shows an image 450 of an unknown environment 402 captured by a 3D camera of robot 100 according to an aspect of the present disclosure. In the example of Figure 4B, the unknown environment 402 is also a restaurant, including a table 404, a pulled-out chair 406, an open window 408, and a cabinet 430. In one configuration, robot 100 uses a reference image to track utensils 420 and 422 in addition to a bottle 410. Using a pre-trained object detection neural network, robot 100 is able to track the movement of utensils 420 and 422 as well as the bottle 410. That is, between Figure 4A and Figure 4B, the bottle 410 moves from the cabinet 430 to the table 404. Similarly, between Figure 4A and Figure 4B, the utensils 420 and 422 move from the table 430 to the cabinet 404.
[0032] According to aspects of this disclosure, a pre-trained object detection neural network uses embeddings (e.g., object descriptors) to determine the correlation of future images (e.g., images taken in real time as robot 100 operates) that define objects and points in the environment. In other words, the system can quickly determine the location in an unknown environment through the correlation of embeddings to real-time images. This disclosure provides a method for generating and training a deep network by collecting training images using a 3D camera, artificially adjusting the lighting levels, and automatically creating pairs of images linked by common features. Consequently, object detection in an unknown environment is not limited to the pose or location of objects in the unknown environment.
[0033] Figure 5 shows an example of a hardware implementation of the object identification system 500 according to an embodiment of the present disclosure. The object identification system 500 may be a component of a vehicle, a robotic device, or other device. For example, as shown in Figure 5, the object identification system 500 is a component of a robot 100 (e.g., a robotic device).
[0034] The embodiments of this disclosure are not limited to the object recognition system 500 which is a component of the robot 100. Other devices such as buses, boats, drones, or vehicles may also be considered as using the object recognition system 500. The robot 100 may operate in at least an autonomous mode and a manual mode.
[0035] The object recognition system 500 may be implemented by a bus architecture generally represented as bus 550. Bus 550 may include any number of interconnection buses and bridges depending on the specific application and overall design constraints of the object recognition system 500. Bus 550 connects various circuits such as one or more processors and / or hardware modules represented as processor 520, communication module 522, position module 524, sensor module 502, motion module 526, navigation module 528, and computer-readable media 530. Bus 550 may also connect various other circuits known to those skilled in the art, such as timing sources, peripherals, voltage controllers, and power management circuits, which will not be described further.
[0036] The object identification system 500 includes a transceiver 540 connected to a processor 520, a sensor module 502, an object identification module 510, a communication module 522, a position module 524, a movement module 526, a navigation module 528, and a computer-readable medium 530. The transceiver 540 is connected to an antenna 542. The transceiver 540 communicates with various devices via a transmission medium. For example, the transceiver 540 may receive commands from a user or remote device via communication. As another example, the transceiver 540 may transmit statistics and other information from the object identification module 510 to a server (not shown).
[0037] The object identification system 500 includes a processor 520 connected to a computer-readable medium 530. The processor 520 performs processing, including the execution of software stored in the computer-readable medium 530 that provides the functions described herein. When executed by the processor 520, the software causes the object identification system 500 to perform various functions described for a specific device such as the robot 100 or modules 502, 510, 512, 514, 516, 522, 524, 526, and 528. The computer-readable medium 530 may also be used to store data manipulated by the processor 520 when the software is executed.
[0038] The sensor module 502 may be used to obtain measurements via different sensors, such as a first sensor 504 and a second sensor 506. The first sensor 504 may be a visual sensor, such as a stereo camera or a red-green-blue (RGB) camera for capturing 3D images. The second sensor 506 may be a distance sensor, such as a light detection and ranging (LiDAR) sensor or a radio wave detection ranging (RADAR) sensor. Of course, embodiments of this disclosure are not limited to the sensors described above, and other types of sensors, such as temperature, sound waves, and / or lasers, can also be considered as either the first sensor 504 or the second sensor 506.
[0039] The measurements from the first sensor 504 and the second sensor 506 may be processed in conjunction with a computer-readable medium 530 by one or more of the following: the processor 520, the sensor module 502, the object identification module 510, the communication module 522, the position module 524, the movement module 526, and the navigation module 528, in order to implement the functions described herein. In some configurations, the data captured by the first sensor 504 and the second sensor 506 may be transmitted to an external device via a transceiver 540. The first sensor 504 and the second sensor 506 may be connected to the robot 100, or may be in communication with the robot 100.
[0040] The position module 524 may be used to determine the position of the robot 100. For example, the position module 524 may use the Global Positioning System (GPS) to determine the position of the robot 100. The communication module 522 may be used to facilitate communication via the transceiver 540. For example, the communication module 522 may provide communication capabilities via different radio protocols such as Wi-Fi, Long Term Evolution (LTE), 5G, etc. The communication module 522 may also be used to communicate with other components of the robot 100 that are not modules of the object identification system 500.
[0041] The mobility module 526 may be used to facilitate the movement of the robot 100. As an alternative example, the mobility module 526 may be in communication with one or more power sources of the robot 100, such as motors and / or batteries. Mobility may be demonstrated by wheels, movable limbs, propellers, treads, fins, jet engines, and / or other sources of mobility.
[0042] The object identification system 500 includes a navigation module 528 for planning a path or controlling the movement of the robot 100, via a movement module 526. The path may be planned based on data provided via the object identification module 510. The modules may be software modules running within the processor 520, modules resident / stored on a computer-readable medium 530, one or more hardware modules connected to the processor 520, or a combination thereof.
[0043] The object identification module 510 may communicate with the sensor module 502, the transceiver 540, the processor 520, the communication module 522, the position module 524, the movement module 526, the navigation module 528, and the computer-readable medium 530. In one configuration, the object identification module 510 receives sensor data from the sensor module 502. The sensor module 502 may receive sensor data from a first sensor 504 and a second sensor 506. According to aspects of this disclosure, the sensor module 502 may filter the data to remove noise, encode the data, decode the data, merge the data, extract frames, or perform other functions. In an alternative configuration, the object identification module 510 may receive sensor data directly from the first sensor 504 and the second sensor 506.
[0044] In one configuration, the object identification module 510 identifies a detected object based on information from the processor 520, the position module 524, the computer-readable medium 530, the first sensor 504, and / or the second sensor 506. Identification of detected objects from the object detection module 512 may be performed using the embedding correlation module 514. Based on the identified object, the object identification module 510 may control one or more actions of the robot 100 through the action module 516.
[0045] For example, the action may be a security action such as tracking a moving object among various images of a landscape captured by the robot 100 and contacting a security service. The object identification module 510 may perform the action via the processor 520, position module 524, communication module 522, computer-readable medium 530, movement module 526, and / or navigation module 528.
[0046] In this embodiment of the Disclosure, an embedding correlation module 514 is used to determine the correlation between an embedding / descriptor image and a future image that defines an object and a point in an unknown environment, from training to future images. That is, after training, when placed in a new environment, the robot 100 identifies the location in the new environment where it can be manipulated, such as a chair, window, bottle, utensil (e.g., a spoon), cabinet, etc. The robot 100 may identify various elements regardless of deformation, object articulation, viewing angle, and lighting.
[0047] Figure 6 is a flowchart illustrating a method for training a deep neural network in a robotic device according to an aspect of this disclosure. For simplicity, the robotic device is referred to as a robot.
[0048] As shown in Figure 6, Method 600 begins with block 602, and a 3D model is constructed using images captured in the training environment via a 3D camera on the robotic device. For example, as shown in Figure 1, the robot 100 captures an original image 101 of the training environment 102. The object may be captured by one or more sensors of the robot 100, such as a LiDAR, RADAR, and / or RGB camera. The object may be observed over a period of time, such as several hours or several days.
[0049] In block 604, a neural network is used to artificially adjust the parameters of the training environment to create a manipulated image, forming a pair of 3D images from a 3D model. For example, Figure 3B shows a manipulated image 350 in which a cabinet 330 has a handle 332 in the landscape, and the cabinet 330 is open, linked to a landscape (e.g., Figure 3A) in which the cabinet 330 is closed. The pairing of the original image 300 and the manipulated image 350 links objects that have similar elements but are depicted with different (e.g., artificial) articulations.
[0050] In block 606, pairs of 3D images are processed to generate a reference image containing embedded descriptors for objects common to the pair of 3D images. For example, Figures 4A and 4B show images taken of an unknown environment 402. In one configuration, robot 100 uses a reference image based on a pair of original and manipulated images of a training environment, such as the kitchen training environment shown in Figures 1, 2, 3A, and 3B. In block 608, the reference image obtained from training the neural network is used to determine its correlation to future images. For example, as shown in Figures 4A and 4B, using the reference image, robot 100 uses a pre-trained object detection neural network to detect a pulled-out chair 406. In addition, the reference image allows robot 100 to detect an open window 408.
[0051] Aspects of this disclosure describe a method for controlling a robotic device based on the identification of detected objects in an unknown environment. The method includes detecting objects in an unknown environment. For example, as shown in Figure 4A, the robot 100 detects instruments 420 and 422 on a table 404. This detection may be performed by selecting a corresponding reference image containing an embedded descriptor corresponding to a trained object, which is manipulated by artificially adjusting the parameters of the image acquisition environment.
[0052] The method further includes identifying the detected object based on the embedded descriptor of the corresponding reference image. For example, using a pre-trained object detection neural network, the robot 100 can track instruments 420 and 422 as well as bottle 410. That is, between Figure 4A and Figure 4B, bottle 410 moves from cabinet 430 to table 404. Similarly, between Figure 4A and Figure 4B, instruments 420 and 422 move from table 404 to cabinet 430.
[0053] Based on the teachings, it should be understood by those skilled in the art that the scope of this disclosure is intended to include any aspect of the disclosure, whether implemented independently or in combination with other aspects of the disclosure. For example, an apparatus may be implemented using any number of the aspects disclosed, or a method may be carried out. In addition, the scope of this disclosure is intended to include, in addition to or other structures and functions, or such apparatus or method carried out using structures and functions, in addition to the various aspects disclosed in this disclosure. It should be understood that any aspect of the disclosure may be embodied by one or more elements of the claims.
[0054] In this specification, the term “exemplary” is used to mean “serving as an example, illustration, or demonstrative.” Any aspect of this specification described as “exemplary” should not necessarily be understood as being preferable or advantageous to other aspects.
[0055] While specific embodiments are described herein, the scope of this disclosure includes numerous variations and substitutions to these embodiments. Although some advantages and benefits of preferred embodiments are described, the scope of this disclosure is not intended to be limited to any particular advantage, use, or purpose. Rather, embodiments of this disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated in the figures and descriptions of preferred embodiments. Detailed descriptions and drawings are for illustrative purposes only, rather than limiting purposes, and the scope of this disclosure is defined by the appended claims and equivalents.
[0056] As used herein, “decision” encompasses a wide range of actions. For example, “decision” may include calculation, computation, processing, derivation, investigation, retrieval (e.g., searching within tables, databases, or other structures), investigation, etc. In addition, “decision” may include reception (e.g., receiving information), access (e.g., accessing data in memory), etc. Furthermore, “decision” may include resolution, selection, choice, establishment, etc.
[0057] As used herein, the phrase “at least one of” refers to any combination of items from the list of items, including a single item. For example, “at least one of a, b, or c” is intended to include a, b, c, ab, ac, bc, and abc.
[0058] Various exemplary logic blocks, modules, and circuits described in connection with this disclosure may be implemented or run by a processor specially configured to perform the functions discussed herein. The processor may be a neural network processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination of the above designed to perform the functions described herein. Alternatively, the processing system may comprise one or more neuromorphic processors to implement the neuron models and neural system models described herein. The processor may be a microprocessor, controller, microcontroller, or state machine configured as described herein. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or other special configurations described herein.
[0059] Steps or algorithms of methods described in connection with this disclosure may be embodied directly in hardware, software modules executed by a processor, or a combination of the two. Software modules may reside in storage devices or machine-readable machine media, including random access memory (RAM), read-only memory (ROM), flash memory, Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), registers, hard disks, removable disks, CD-ROMs or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium accessible by a computer and usable to carry or store desired program code in the form of instructions or data structures. Software modules may comprise a single instruction or a number of instructions, and may be distributed across multiple different code segments, different programs, and multiple storage media. Storage media may be connected to a processor so that the processor can write information to and read information from the storage media. Alternatively, storage media may be integrated with the processor.
[0060] The methods disclosed herein include one or more steps or actions for carrying out the disclosed method. The steps and / or actions of the method may be replaced with each other without departing from the claims. In other words, unless a particular order of steps or actions is specified, the order and / or use of a particular step and / or action may be changed without departing from the claims.
[0061] The described functions may be implemented by hardware, software, firmware, or any combination thereof. When implemented in hardware, an example of a hardware configuration may include a processing system within the device. The processing system may be implemented using a bus architecture. The bus may include any number of interconnection buses and bridges, depending on the specific application and overall design constraints of the processing system. The bus may connect various circuits, including processors, machine-readable media, and bus interfaces. The bus interface may also be used to connect network adapters to the processing system via the bus, among other things. Network adapters may be used to implement signal processing functions. In certain embodiments, user interfaces (e.g., keypads, displays, mice, joysticks, etc.) may also be connected to the bus. The bus may also connect various other circuits known to those skilled in the art, such as timing sources, peripherals, voltage control, and power management circuits, which will not be described further.
[0062] The processor may be responsible for bus management and processing including the execution of software stored in machine-readable media. Software, whether called software, firmware, middleware, microcode, hardware description language, or otherwise, is construed to mean instructions, data, or any combination thereof.
[0063] In hardware implementations, machine-readable media may be part of a processing system separate from the processor. However, as will be readily apparent to those skilled in the art, machine-readable media, or any part thereof, may be outside the processing system. For example, machine-readable media may include communication lines, data-modulated carriers, and / or computer products isolated from the device, all of which may be accessed by the processor via a bus interface. Alternatively, or in addition, machine-readable media, or any part thereof, may be integrated into the processor, such as in the case of caches and / or special register files. Although the various components discussed have been described as having a special location, like local components, they may be configured in various ways, like certain components configured as part of a distributed computing system.
[0064] A machine-readable medium may contain several software modules. These software modules may include transmit modules and receive modules. Each software module may reside in a single storage device or may be distributed across multiple storage devices. For example, a software module may be loaded from a hard drive into RAM when a triggering event occurs. While a software module is executing, the processor may load several instructions into a cache to increase access speed. One or more cache lines may then be loaded into a special-purpose register file for execution by the processor. The following functions of a software module will be understood to be performed by the processor when instructions are executed by the software module. Furthermore, it should be understood that the functions of a processor, computer, machine, or other system implementing such a function are improved by the embodiments of this disclosure.
[0065] When implemented in software, functions may be stored or transferred on a computer-readable medium as one or more instructions or codes. Computer-readable mediums include both computer storage devices and communication media, which include any storage devices that facilitate the transfer of computer programs from one location to another.
[0066] Furthermore, it should be understood that modules and / or other suitable means of performing the methods and techniques described herein may be obtained as needed by download and / or by user terminals and / or base stations. For example, such devices may be connected to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein may be provided via storage means, such that user terminals and / or base stations can obtain the various methods by connecting storage means to the device or by providing storage means to the device. Furthermore, any other techniques for providing the methods and techniques described herein to the device may be used.
[0067] It should be understood that the claims are not limited to the exact configuration and components shown above. Various modifications, changes, and variations may be made to the arrangement, operation, and details of the methods and apparatus described above without departing from the claims.
Claims
1. constructing a 3D model using images captured using a 3D camera of the robotic device in the training environment; forming a manipulation image using a deep neural network by artificially adjusting parameters of the training environment and generating a pair of 3D images from the 3D model; processing the pair of 3D images to form a reference image that includes embedded descriptors of objects common between the pair of 3D images; using the reference images from training the neural network to identify and correlate detected objects in future images; A method for training a deep neural network of a robotic device, comprising:
2. The generation of the pair of 3D images comprises: Pairing 3D images with linked elements; manipulating linked elements between the pair of 3D images to create a scene with different object articulations; The method of claim 1 , comprising:
3. Artificial adjustment of parameters modifying object articulations between the original 3D image and the manipulated 3D image; The method of claim 1.
4. The change in object articulation may include: modifying lighting between the original 3D image and the manipulated 3D image; The method of claim 3.
5. The change in object articulation may include: changing the viewing angle between the original 3D image and the manipulated 3D image; The method of claim 3.
6. Identifying an object in an unknown environment that can be manipulated regardless of deformation, object articulation, viewing angle, and lighting in the unknown environment; Manipulating an identified object; The method of claim 1 further comprising:
7. Detecting objects in an unknown environment; selecting a corresponding reference image containing an embedding descriptor corresponding to the trained object manipulated by artificially adjusting parameters of the image capture environment; identifying a detected object based on the embedded descriptor of the corresponding reference image; A method for controlling a robotic device based on identifying detected objects in an unknown environment, comprising:
8. The method of claim 7 further comprising following the identified object for a period of time.
9. determining that the identified object can be manipulated; manipulating the identification object; The method of claim 7 further comprising:
10. overlaying the corresponding reference image onto a captured image of the scene; determining an identity of the detected object based on point correspondences between the corresponding reference image and the captured image; The method of claim 7 further comprising:
11. 1. A non-transitory computer-readable medium having recorded thereon program code for training a deep neural network of a robotic device, the medium comprising: The program code is executed by a processor, program code for generating a pair of 3D images from a 3D model by artificially adjusting parameters of a training environment to form manipulated images using the deep neural network; program code for processing the pair of 3D images to form a reference image that includes embedded descriptors of objects common between the pair of 3D images; program code that uses the reference images from training a neural network to identify and correlate detected objects in future images; 1. A non-transitory computer-readable medium, comprising:
12. The program code for generating the pair of 3D images comprises: program code for pairing 3D images at linked elements; program code for manipulating linked elements between the pair of 3D images to create a scene with different object articulations; 12. The non-transitory computer-readable medium of claim 11, comprising:
13. The program code for generating the pair of 3D images comprises: program code for modifying object articulations between the original 3D image and the manipulated 3D image; The non-transitory computer-readable medium of claim 11.
14. The program code for modifying an object articulation comprises: program code for modifying lighting between the original 3D image and the manipulated 3D image; 14. The non-transitory computer-readable medium of claim 13.
15. The program code for modifying the object articulation further comprises: program code for modifying a viewing angle between the original 3D image and the manipulated 3D image; 14. The non-transitory computer-readable medium of claim 13.
16. a pre-trained object identification module configured to select a corresponding reference image for identifying a detected object in the captured image, the corresponding reference image including a trained object-based embedded descriptor manipulated by artificially adjusted parameters of the image capture environment; a controller configured to select an autonomous operation of the robotic device based on the identity of the detected object; A system for controlling a robotic device based on identifying detected objects in an unknown environment.
17. 17. The system of claim 16, wherein the pre-trained object identification module is configured to track an identified object over a period of time.
18. The system of claim 16 , wherein the controller is further configured to manipulate an identifying object.
19. 17. The system of claim 16, wherein the pre-trained object identification module is configured to overlay the corresponding reference image onto a captured image of a scene and determine an identification of the detected object based on point correspondences between the corresponding reference image and the captured image.
20. 17. The system of claim 16, wherein the pre-trained object identification module is configured to detect common objects between the corresponding reference and captured images based on correlations that identify the detected objects in future images.