Deep learning system
By using sparse volumetric data structures and hardware accelerators, the latency and storage limitations of computer systems when processing large datasets in AR, VR, and MR applications are solved, enabling low-latency real-time processing and efficient 3D rendering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MOVIDIUS LTD
- Filing Date
- 2019-05-21
- Publication Date
- 2026-05-29
Smart Images

Figure CN122114035A_ABST
Abstract
Description
[0001] Related applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 675,601, filed May 23, 2018, which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure generally relates to the field of computer systems, and more specifically to machine learning systems. Background Technology
[0003] The world of computer vision and graphics is rapidly converging with the emergence of augmented reality (AR), virtual reality (VR), and mixed reality (MR) products, such as those from Magic Leap. TM Microsoft TM HoloLens TM , Those products, and such as those from Valve TM and HTC TM Other VR systems, such as those mentioned above, utilize separate graphics processing units (GPUs) and computer vision subsystems that operate in parallel. These parallel systems can be assembled from existing GPUs in parallel with a computer vision pipeline implemented in software running on an array of processors and / or programmable hardware accelerators. Attached Figure Description
[0004] The various objectives, features, and advantages of the disclosed subject matter can be more fully understood when considered in conjunction with the following detailed description of the subject matter, wherein like reference numerals identify like elements. The drawings are schematic and not intended to be drawn to scale. For clarity, not every component is labeled in every figure. Nor is every component of every embodiment of the disclosed subject matter shown where illustration is unnecessary to allow those skilled in the art to understand the disclosed subject matter.
[0005] Figure 1 This demonstrates a conventional augmented reality or mixed reality rendering system; Figure 2 A voxel-based augmented reality or mixed reality rendering system according to some embodiments is shown; Figure 3 The differences between dense and sparse volume representations according to some embodiments are illustrated; Figure 4 A composite view of a scene according to some embodiments is shown; Figure 5The levels of detail in an example element tree structure according to some embodiments are shown; Figure 6 This illustrates applications of the data structures and voxel data of this application according to some embodiments; Figure 7 An example network for recognizing 3D digits according to some embodiments is shown; Figure 8 This illustrates multiple classifications performed on the same data structure using implicit levels of detail, according to some embodiments; Figure 9 Operation elimination performed by a 2D convolutional neural network according to some embodiments is illustrated; Figure 10 Experimental results from analysis of example test images according to some embodiments are shown; Figure 11 Hardware for a culling operation according to some embodiments is shown; Figure 12 Improvements to the hardware for the culling operation are shown according to some embodiments; Figure 13 Hardware according to some embodiments is shown; Figure 14 An example system employing an example training set generator according to at least some embodiments is shown; Figure 15 An example of synthetic training data generation according to at least some embodiments is shown; Figure 16 An example Siamese Network is shown according to at least some example embodiments; Figure 17 An example use of SiamNet performing autonomous comparisons according to at least some example embodiments is shown; Figure 18 An example voxelization of a point cloud according to at least some embodiments is shown; Figure 19 It is a simplified block diagram of an example machine learning model based on at least some embodiments; Figure 20 A simplified block diagram illustrating various aspects of example training of a model, based on at least some embodiments; Figure 21 An example robot is shown that uses a neural network to generate a 3D map for navigation, according to at least some embodiments; Figure 22 A block diagram illustrating an example machine learning model for use with inertial measurement data, according to at least some embodiments; Figure 23A block diagram illustrating an example machine learning model for use with image data is shown according to at least some embodiments; Figure 24 It shows the combination Figure 22 and Figure 23 The example model is shown in the block diagram of an example machine learning model. Figures 25A-25B It shows the relationship with Figure 24 A graph showing the results of similar machine learning models to the example machine learning model; Figure 26 An example system including an example neural network optimizer is shown according to at least some embodiments; Figure 27 This is a block diagram illustrating example optimizations of a neural network model according to at least some embodiments; Figure 28 This is a table showing example results generated and used during the optimization of the example neural network model; Figure 29 This is a table showing example results generated and used during the optimization of the example neural network model; Figure 30A This is a simplified block diagram of an example of hybrid neural network pruning according to at least some embodiments; Figure 30B It is a simplified flowchart of example trimming of a neural network based on at least some embodiments; Figure 31 It is a simplified block diagram of example weight quantization performed in conjunction with neural network pruning according to at least some embodiments; Figure 32 This is a table comparing the results of example neural network pruning techniques; Figures 33A-33F It is a simplified flowchart of an example computer implementation of a technology associated with machine learning according to at least some embodiments; Figure 34 An example multi-slot vector processor according to some embodiments is described; Figure 35 Example volume acceleration hardware according to some embodiments is shown; Figure 36 The organization of voxel cubes according to some embodiments is shown; Figure 37 A two-level sparse voxel tree according to some embodiments is shown; Figure 38 A two-level sparse voxel tree according to some embodiments is shown; Figure 39 The storage of example voxel data according to some embodiments is illustrated; Figure 40The insertion of voxels into an example volume data structure according to some embodiments is illustrated; Figure 41 The projection of an example 3D volumetric object according to some embodiments is shown; Figure 42 Example operations involving example volume data structures are shown; Figure 43 The use of projection to generate simplified maps is illustrated according to some embodiments; Figure 44 Example aggregations of example volumetric 3D measurement results and / or simple 2D measurement results from embedded devices according to some embodiments are shown; Figure 45 An example acceleration of 2D path finding on a 2D 2×2 bitmap is shown according to some embodiments; Figure 46 An example acceleration of collision detection using an example volume data structure is shown according to some embodiments; Figure 47 This is a simplified block diagram of an exemplary network having devices according to at least some embodiments; Figure 48 This is a simplified block diagram of an exemplary fog or cloud computing network according to at least some embodiments; Figure 49 This is a simplified block diagram of a system including an example device according to at least some embodiments; Figure 50 This is a simplified block diagram of an example processing device according to at least some embodiments; Figure 51 This is a block diagram of an exemplary processor according to at least some embodiments; and Figure 52 This is a block diagram of an exemplary computing system according to at least some embodiments. Detailed Implementation
[0006] In the following description, numerous specific details are set forth in relation to the systems and methods of the disclosed subject matter and the environments in which such systems and methods may operate, in order to provide a thorough understanding of the disclosed subject matter. However, it will be apparent to those skilled in the art that the disclosed subject matter can be practiced without these specific details, and that certain features well-known in the art are not described in detail in order to avoid complicating the disclosed subject matter. Furthermore, it should be understood that the embodiments provided below are exemplary, and other systems and methods are contemplated within the scope of the disclosed subject matter.
[0007] Various technologies based on and incorporating augmented reality, virtual reality, mixed reality, autonomous devices, and robotics have emerged, utilizing volumetric data models representing three-dimensional space and geometry. Descriptions of various real and virtual environments using such 3D or volumetric data have traditionally involved large datasets, which some computer systems have struggled to process in the desired manner. Furthermore, as devices such as drones, wearables, and virtual reality systems become smaller, the memory and processing resources of such devices may also be limited. As an example, AR / VR / MR applications may require high frame rates for graphical representations generated using hardware-enabled systems. However, in some applications, the GPUs and computer vision subsystems of such hardware may need to process data (e.g., 3D data) at high rates (e.g., up to 130 fps (7 milliseconds)) to produce desired results (e.g., generating believable graphical scenes with frame rates that produce believable results, preventing motion sickness in users due to excessively long wait times, and other example objectives). Similarly, additional applications may face the challenge of satisfactorily processing large volumes of data while meeting the limitations of the corresponding system's processing, memory, power, application requirements, and other example problems.
[0008] In some implementations, logic can be provided to the computing system to generate and / or use sparse volumetric data defined according to a specific format. For example, a defined volumetric data structure can be provided to unify computer vision and 3D rendering across various systems and applications. A volumetric representation of an object can be captured using optical sensors, such as stereo cameras or depth cameras. The volumetric representation of an object can include multiple voxels. An improved volumetric data structure can be defined that allows the corresponding volumetric representation to be recursively subdivided to obtain the target resolution of the object. During subdivision, blank spaces in the volumetric representation (and supporting operations) can be culled from the volumetric representation (and may be included in one or more voxels). Blank spaces can be regions in the volumetric representation that do not include the geometric properties of the object.
[0009] Therefore, in the improved volume data structure, each voxel within a corresponding volume can be marked as either "occupied" or "blank" (indicating that the corresponding volume consists of blank space) (by some geometry existing within the corresponding volume space). Such a label can be additionally interpreted as specifying that one or more sub-volumes within its corresponding sub-volume are also occupied (e.g., if the parent voxel or a higher-level voxel is marked as occupied), or that all of its sub-volumes are blank space (i.e., in the case where the parent voxel or a higher-level voxel is marked as blank). In some implementations, marking a voxel as blank allows that voxel and / or its corresponding sub-volume voxel to be efficiently removed from the operations used to generate the corresponding volume representation. The volume data structure can be a sparse tree structure, such as a sparse sixty-fourth-quarter tree (SST) format. Furthermore, this sparse volume data structure approach can utilize relatively less storage space than traditional methods used to store volume representations of objects. Additionally, compression of the volume data can increase the transport viability of such representations and enable faster processing of such representations, among other examples of benefits.
[0010] Volumetric data structures can be hardware-accelerated to rapidly allow updates to 3D renderers, eliminating latency that can occur in separate computer vision and graphics systems. Such latency can lead to latency that, when used in AR, VR, MR, and other applications, can cause motion sickness and other additional drawbacks in users. The ability to quickly test the occupancy of voxels' geometric properties in accelerated data structures allows for low-latency construction of AR, VR, MR, or other systems that can be updated in real time.
[0011] In some embodiments, the capabilities of volumetric data structures can also provide intra-frame warnings. For example, in AR, VR, MR, and other applications, when a user may collide with a real or synthetic object in the imaged scene, or in computer vision applications for drones or robots, when such devices may collide with a real or synthetic object in the imaged scene, the processing speed provided by volumetric data structures allows for warnings of impending collisions.
[0012] Embodiments of this disclosure can be associated with the storage and processing of volumetric data in applications such as robotics, head-mounted displays for augmented and mixed reality headsets, and telephones and tablets. Embodiments of this disclosure represent each volume element (e.g., a voxel) within a set of voxels, along with optional physical properties related to the voxel's geometry, as a single bit. Additional parameters associated with a set of 64 voxels can be associated with the voxels, such as corresponding red-green-blue (RGB) or other color coding, transparency, truncated signed distance function (TSDF) information, etc., and stored in an associated, optional 64-bit data structure (e.g., such that two or more bits are used to represent each voxel). Such a representation scheme can achieve minimal memory requirements. Furthermore, representing voxels with a single bit allows for the execution of numerous simplified computations to logically or mathematically combine elements from the volume representation. Combining elements from the volume representation can include, for example, performing an OR-ing operation on planes in the volume to create a 2D projection of 3D volume data, and calculating surface area by counting the number of voxels occupying a 2.5D manifold, etc. For comparison, XOR logic can be used to compare 64-bit sub-volumes (e.g., 4^3 sub-volumes), and volumes can be inverted, where objects can be merged to create hybrid objects by performing an OR operation on them together, and other examples.
[0013] Figure 1A conventional augmented reality or mixed reality system is illustrated, comprising parallel graphics rendering and computer vision subsystems. The computer vision subsystem has post-rendering connectivity to account for changes caused by rapid head movements and variations in the environment that may result in occlusion and shadows in the rendered graphics. In one example implementation, the system may include a host processor 100 supported by host memory 124 to control the execution of the graphics pipeline, computer vision pipeline, and post-rendering correction devices via bus 101, on-chip network-on-chip interconnects, or other interconnects. This interconnect allows the host processor 100 to run appropriate software to control the execution of the graphics processing unit (GPU) 106, associated graphics memory 111, computer vision pipeline 116, and associated computer vision memory 124. In one example, rendering of graphics using the GPU 106 via OpenGL graphics shaders 107 (e.g., operations on triangle list 105) may occur at a slower rate compared to the computer vision pipeline. As a result, post-render corrections can be performed via the bending engine 108 and the display / occlusion processor 109 to account for changes in head pose and occlusion in the scene geometry that may have occurred since the graphics GPU 106 rendered. The output of the GPU 106 is timestamped, allowing it to combine correction control signals 121 from the head pose pipeline 120 and correction control signals 123 from the occlusion pipeline 123 to produce correct graphics output that accounts for any changes in head pose 119 and occlusion geometry 113, among other examples.
[0014] Multiple sensors and cameras (e.g., including active and passive stereo cameras for depth and vision processing 117) running in parallel with GPU 106 can be connected to computer vision pipeline 116. Computer vision pipeline 116 can include one or more of at least three stages, each of which can contain lower-level processing from multiple stages. In one example, the stages in computer vision pipeline 116 could be an image signal processing (ISP) pipeline 118, a head pose pipeline 120, and an occlusion pipeline 122. ISP pipeline 118 can take the outputs of input camera sensors 117 and modulate them so that they are usable for subsequent head pose and occlusion processing. Head pose pipeline 120 can take the output of ISP pipeline 118 and use it in conjunction with the output of inertial measurement unit (IMU) in head-mounted device 110 to compute changes in head pose since the corresponding output graphics frame rendered by the free GPU 106. The output 121 of the head pose pipeline (HPP) 120 can be applied to the bending engine 108 along with a user-specified mesh to warp the GPU output 102 so that it matches the updated head pose position 119. The occlusion pipeline 122 can take the output of the head pose pipeline 121 and search for new objects in the field of view, such as a hand 113 (or other example object) entering the field of view, which should produce a corresponding shadow 114 on the scene geometry. The output 123 of the occlusion pipeline 122 can be used by the display and occlusion processor 109 to correctly overlay the viewport on top of the output 103 of the bending engine 108. The display and occlusion processor 109 uses the calculated head pose 119 to generate a shadow mask for compositing shadow 114, and the display and occlusion processor 109 can compose the occlusion geometry of the hand 113 on top of the shadow mask to generate a graphic shadow 114 on top of the output 103 of the bending engine 108, and generate multiple final output frames 104 for display on the augmented / mixed reality head-mounted device 110, as well as other example use cases and features.
[0015] Figure 2 A voxel-based augmented reality or mixed reality rendering system according to some embodiments of the present disclosure is illustrated. Figure 2The apparatus depicted may include a host system comprised of a host CPU 200 and an associated host memory 201. Such a system may communicate via a unified computer vision and graphics pipeline 223 and an associated unified computer vision and graphics memory 213 via a bus 204, an on-chip network, or other communication mechanisms. The computer vision and graphics pipeline 223 and the associated unified computer vision and graphics memory 213 contain real and synthetic voxels to be rendered in the final scene for display on a head-mounted augmented reality or mixed reality display 211. The AR / MR display 211 may also include multiple active and passive image sensors 214, and an inertial measurement unit (IMU) 212 for measuring changes in head pose 222 orientation.
[0016] In the combined rendering pipeline, composite geometry can be generated starting from a list of triangles 204, which is processed by an OpenGL JiT (Just-in-Time) translator 205 to produce composite voxel geometry 202. For example, composite voxel geometry can be generated by selecting the principal plane of triangles from the triangle list. 2D rasterization (e.g., in the X and Z directions) can then be performed on each triangle in the selected plane. A third coordinate (e.g., Y) can be created as an attribute to be interpolated across the triangles. Each pixel of the rasterized triangles can lead to the definition of a corresponding voxel. This processing can be performed by the CPU or GPU. When performed by the GPU, each rasterized triangle can be read back from the GPU to create a voxel where the GPU draws the pixel, among other example implementations. For example, a 2D buffer of a list can be used to generate composite voxels, where each entry in the list stores depth information of the polygon rendered at that pixel. For example, an orthogonal viewpoint (e.g., top-down) can be used to render the model. For example, each (x, y) provided in the example buffer can represent a column at (x, y) in the corresponding voxel volume (e.g., from (x, y, 0) to (x, y, 4095)). The information in each list can then be used to render each column from the information into a 3D scanline.
[0017] continue Figure 2For example, in some implementations, the synthetic voxel geometry 202 may be combined with a measured geometry voxel 227 constructed using a Simultaneous Localization and Mapping (SLAM) pipeline 217. The SLAM pipeline may use active and / or passive image sensors 214 (e.g., 214.1 and 214.2), which are first processed using an image signal processing (ISP) pipeline 215 to produce an output 225, which may be converted into a depth image 226 by a depth pipeline 216. The active or passive image sensors 214 (214.1 and 214.2) may include active or passive stereo sensors, structured light sensors, time-of-flight sensors, and other examples. For example, the depth pipeline 216 may process depth data from a structured light or time-of-flight sensor 214.1 or alternatively, a passive stereo sensor 214.2. In one example implementation, the stereo sensor 214.2 may include a pair of passive stereo sensors, and other example implementations exist.
[0018] The depth image generated by the depth pipeline 215 can be processed by the dense SLAM pipeline 217 using a SLAM algorithm (e.g., Kinect Fusion) to produce a voxelized model of the measured geometry voxels 227. A ray tracing accelerator 206 can be provided, which can combine the measured geometry voxels 227 (e.g., real voxel geometry) with the synthesized voxel geometry 202 to produce a 2D rendering of the scene for output to a display device (e.g., a head-mounted display 211 in a VR or AR application) via the display processor 210. In such an implementation, a complete scene model can be constructed from the real voxels of the measured geometry voxels 227 and the synthesized geometry 202. Therefore, bending of the 2D rendered geometry is not required (e.g., as in...). Figure 1 (in the middle). Such implementations can be combined with head pose tracking sensors and corresponding logic to correctly align the actual geometry with the measured geometry. For example, example head pose pipeline 221 can process head pose measurement results 232 from IMU 212 mounted in head-mounted display 212, and the output 231 of the head pose measurement results pipeline can be taken into account during rendering via display processor 210.
[0019] In some examples, the unified rendering pipeline can also use measured geometry voxels 227 (e.g., real voxel models) and synthetic geometry 202 (e.g., synthetic voxel models) to render audio reverberation models and model the physics of real-world, virtual reality, or mixed reality scenes. As an example, the physics pipeline 218 can take measured geometry voxels 227 and synthetic geometry 202 voxel geometries and use a raycasting accelerator 206 to compute output audio samples for the left and right earpieces in a head-mounted display (HMD) 211, calculating output sample 230 using acoustic reflection coefficients built into the voxel data structure. Similarly, the unified voxel model composed of 202 and 227 can also be used to determine physical updates for synthetic objects in a composite AR / MR scene. The physics pipeline 218 takes composite scene geometry as input and uses the raycasting accelerator 206 to compute conflicts before computing updates 228 for rendering on the synthetic geometry 202 and as the basis for future iterations of the physics model.
[0020] In some implementations, the system (such as Figure 2 The system shown may additionally be provided with one or more hardware accelerators to implement and / or utilize a convolutional neural network (CNN) that can process RGB video / image input from the output of ISP pipeline 215, volumetric scene data from the output of SLAM pipeline 217, and other examples. The neural network classifier may exclusively run using hardware (HW) convolutional neural network (CNN) accelerator 207 or in a combination of a processor and HW CNN accelerator 207 to produce output classification 237. The availability of the HW CNN accelerator 207 for inference on volumetric representations may allow groupings of voxels in measured geometry 227 to be labeled as belonging to a specific object category, and other examples are used.
[0021] Labeling voxels (e.g., using CNNs and supporting hardware acceleration) allows the objects to which those voxels belong to to be recognized by the system as corresponding to known objects, and source voxels can be removed from the measured geometric voxels 227 and replaced with bounding boxes corresponding to the objects and / or with information related to the object's origin, pose, object descriptor, and other example information. This can lead to a semantically much more meaningful description of the scene, for example, as input to interacting with objects in the scene via robots, drones, or other computing systems, or via audio systems to find the sound absorption coefficients of objects in the scene and reflect them in the scene's acoustic model, and other examples.
[0022] One or more processor devices and hardware accelerators can be provided to implement in Figure 2 The pipeline of the example system shown and described herein. In some implementations, all hardware and software elements of the combined rendering pipeline may share access to DRAM controller 209, which in turn allows data to be stored in a shared DDR memory device 208, and other example implementations.
[0023] Presentation Figure 3 The differences between dense volume representation and sparse volume representation will be illustrated according to some embodiments. Figure 3 As shown in the examples, real-world objects or synthetic objects 300 (e.g., a rabbit statue) can be described in the form of voxels in a dense manner as shown in 302 or a sparse manner as shown in 304. The advantage of a dense representation such as 302 is the uniform rate of access to all voxels in the volume, but the disadvantage is the potentially large amount of storage required. For example, for a dense representation of a volume such as 512^3 elements (e.g., corresponding to 5m in a volume scanned at 1cm resolution using a Kinect sensor), 512 megabytes would be needed to store the relatively small volume using a 4-byte truncated signed distance function (TSDF) for each voxel. Octree representation 304 embodies a sparse representation, which, on the other hand, allows storage only of those voxels with realistic geometry in the real-world scene, thereby reducing the amount of data required to store the same volume.
[0024] Go to Figure 4 This illustrates a composite view of an example scenario according to some embodiments. Specifically, Figure 4 The composite view of scene 404 is shown to be preserved, displayed, or further processed using parallel data structures to represent, respectively, a synthetic voxel 401 within an equivalent bounding box 400 for synthetic voxel data and a real-world measured voxel 403 within an equivalent bounding box 402 for real-world voxel data. Figure 5 The diagram illustrates the level of detail in a uniform 4^3 element tree structure according to some embodiments. In some implementations, as little as one bit can be used to describe each voxel in a volume represented using an octree, such as in Figure 5 The example illustrates this. However, a drawback of octree-based techniques can be the number of indirect memory accesses required to access a specific voxel within the octree. In the case of sparse voxel octrees, the same geometry can be implicitly represented at multiple levels of detail, advantageously allowing for the operation of techniques such as ray casting, game physics, CNNs, and others to allow for the removal of blank areas of the scene from further computation. This results not only in a reduction in overall memory requirements but also in overall power consumption and computational load, among other example advantages.
[0025] In one implementation, an improved voxel descriptor (also referred to herein as a “volume data structure”) can be provided to organize volume information into 4^3 (or 64-bit) unsigned integers, as shown in 501 with a memory requirement of 1 bit per voxel. In this example, 1 bit per voxel is insufficient (compared to the TSDF in SLAMbench / KFusion using 64 bits) to store truncated symbolic distance function values. In this example, an additional (e.g., 64-bit) field 500 can be included in the voxel descriptor. This example can be further enhanced such that when the TSDF in the 64-bit field 500 is 16 bits, an additional 2 decimal places of resolution in x, y, and z can be implicitly provided in the voxel descriptor 501, so that the combination of the voxel TSDF in the 64-bit field 500 and the voxel position 501 is equivalent to a much higher-resolution TSDF, as used in SLAMbench / KFusion or other examples. For example, additional data in the 64-bit field 500 (voxel descriptor) can be used to store 1 byte of RGB color information for each item (e.g., from a scene via a passive RGB sensor) and an 8-bit transparency value alpha, as well as two application-specific reserved fields R1 and R2 that can be used to store, for example, acoustic reflectivity for audio applications, stiffness for physical applications, object material type, and other examples.
[0026] like Figure 5 As shown, voxel descriptors 501 can be logically grouped into four 2D planes, each of which includes 16 voxels 502. (As...) Figure 5 These 2D planes (or voxel planes) can be used to describe each level of an octree-type structure based on a successive decomposition of powers of 4. In this example implementation, a 64-bit voxel descriptor is chosen because it matches well with the 64-bit bus infrastructure used in the corresponding system implementation (although other voxel descriptor sizes and formats can be provided in other system implementations, and the size should be designed according to the bus or other infrastructure of the system). In some implementations, the size of the voxel descriptor can be designed to reduce the number of memory accesses required to access voxels. For example, a 64-bit voxel descriptor can be used to reduce the number of memory accesses required to access voxels at any level in an octree by half compared to a traditional octree that operates on 2^3 elements, among other example considerations and implementations.
[0027] In one example, an octree can be described starting with a 4^3 root volume 503, and each non-zero entry in the example 256^3 volume is depicted as an encoding of the geometry used to represent the layers 504, 505, and 506 below. In this particular example, four memory accesses are needed to access the lowest level of the octree. In cases where such overhead is prohibitive, an alternative approach is to encode the highest level of the octree as a larger volume, such as 64^3, as shown in 507. In this case, each non-zero entry in 507 can indicate the presence of the lower 4^3 octree in the 256^3 volume 508 below. Compared to the alternative formulations shown in 503, 504, and 505, this alternative organization results in only two memory accesses required to access any voxel in the 256^3 volume 508. The latter approach is advantageous when the device hosting the octree structure has a large amount of embedded memory, allowing only lower- and less frequently accessed portions of the voxel octree 508 to reside in external memory. This method is more expensive in terms of storage, for example, in the case of storing a full, large (e.g., 64^3) volume in on-chip memory, but this trade-off allows for faster memory access (e.g., 2x faster) and much lower power consumption, among other advantages.
[0028] Go to Figure 6 The diagram illustrates a sample application, according to some embodiments, demonstrating the data structures and voxel data that can be utilized by this application. In one example, such as Figure 5 As shown, additional information can be provided through example voxel descriptor 500. Voxel descriptors enable a wide range of applications that can utilize voxel data, such as... (The sentence is incomplete and requires further context to translate accurately.) Figure 6 The volumetric data represented in the figure. For example, a shared volumetric representation 602 (such as that generated using a dense SLAM system 601 (e.g., SLAMbench)) can be used to render a scene using graphics ray casting or ray tracing 603, for audio ray casting 604, and other implementations. In yet another example, the volumetric representation 602 can also be used in convolutional neural network (CNN) inference 605 and can be backed up via cloud infrastructure 607. In some instances, cloud infrastructure 607 can contain detailed volumetric descriptors of objects (such as a tree, a piece of furniture, or other objects (e.g., 606)) that can be accessed via inference. Based on inference or other means of identifying the object, the corresponding detailed description can be returned to the device, allowing the voxels of volumetric representation 602 to be replaced by bounding box representations with pose information and descriptors containing object attributes and other example features.
[0029] In yet another embodiment, in some systems, the voxel model discussed above may be additionally or alternatively utilized to construct a 2D map 608 of the example environment from volumetric representation 602 using 3D-to-2D projection. These 2D maps can be shared again via communication machines through cloud infrastructure and / or other network-based resources 607 and aggregated (e.g., using the same cloud infrastructure) to build higher-quality maps using crowd-sourced techniques. These maps can be shared by cloud infrastructure 607 to connected machines and devices. In yet another further example, the projection can be used to optimize the 2D map for ultra-low bandwidth applications by subsequently simplifying it segment by segment 609 (e.g., assuming the width and height of the vehicle or robot are fixed). The simplified path can then have only a single X, Y coordinate pair in each linear segment of the path, thereby reducing the amount of bandwidth required to transmit the path of vehicle 609 to cloud infrastructure 607 and aggregate it in the same cloud infrastructure 607 to build higher-quality maps using crowdsourcing techniques. These maps can be shared by cloud infrastructure 607 to connected machines and devices.
[0030] To achieve these different applications, common functionalities can be provided in some implementations, such as through a shared software library, which in some embodiments can be accelerated using hardware accelerators or processor instruction set architecture (ISA) extensions, and other examples. For example, such functionalities may include inserting voxels into descriptors, deleting voxels, or finding voxels 610. In some implementations, collision detection functionality 620 and the deletion of points / voxels from volume 630 may also be supported, and other examples. As described above, a system can be provided that has the capability to rapidly generate 2D projections 640 in the X, Y, and Z directions from a corresponding volume representation 602 (3D volume) (e.g., which can be used as the basis for path or collision determination). In some cases, it may also be advantageous to be able to generate a list of triangles from volume representation 602 using a histogram pyramid 650. Furthermore, a system can be provided that has the capability to rapidly determine free paths 660 in the 2D and 3D representations of volume space 602. Such functionalities may be useful in a range of applications. Further features can be provided, such as illustrating the number of voxels in a volume, using a population counter to count the number of bits with a value of 1 in the mask region of volume representation 602 to determine the surface of an object, and other examples.
[0031] Go to Figure 7 A simplified block diagram illustrates an example network according to at least some embodiments, which includes a system equipped with the ability to recognize 3D digits. For example, Figure 6 One of the applications shown is a volumetric CNN application 605, which in Figure 7The example network is described in more detail below, where it is used to recognize 3D digits generated from datasets such as the hybrid National Institute of Standards and Technology (MNIST) dataset 700. Digits from such datasets can be used to train a CNN-based convolutional network classifier 710 by applying appropriate rotations and translations to the digits in the X, Y, and Z directions prior to training. When used for inference in embedded devices, the trained network 710 can be used to classify 3D digits in a scene with high accuracy, even when the digits undergo rotations and translations in the X, Y, and Z directions 720, as well as other examples. In some implementations, this can be achieved by... Figure 2 The HW CNN accelerator 207 shown accelerates the operation of the CNN classifier. As the first layer of the neural network performs multiplication operations using voxels in the volume representation 602, these arithmetic operations can be skipped because multiplying by zero always equals zero and the data value A multiplied by 1 (voxel) equals A.
[0032] Figure 8 This illustrates multiple classifications performed on the same data structure using implicit level of detail. Further improvements to the CNN classification using a volume representation of 602 could be made, since the octree representation is as follows... Figure 5 The octree structure shown implicitly contains multiple levels of detail, so such as Figure 8 As shown, multiple classifications can be performed on the same data structure using a single classifier 830 or multiple classifiers in parallel, employing implicit levels of detail 800, 810, and 820. In conventional systems, comparable parallel classification can be slow due to the image scaling required between classifications. Because the same octree can contain the same information at multiple levels of detail, such scaling can be predetermined in implementations applying the voxel structures discussed in this paper. In fact, a single training dataset based on a volumetric model can cover all levels of detail, rather than the scaled training datasets required in conventional CNN networks.
[0033] Go to Figure 9 For example, according to some embodiments, example operation elimination is illustrated by a 2D CNN. Such as Figure 9 As shown, operation elimination can be used for both 3D volumetric CNNs and 2D CNNs. For example, in Figure 9 In the first layer, a bitmap mask 900 can be used to describe the expected "shape" of the input 910 and can be applied to the incoming video stream 920. In one example, operation cancellation can be used not only for 3D volumetric CNNs but also for 2D volumetric CNNs. For example, in... Figure 9In the example 2D CNN, a bitmap mask 900 can be applied to the first layer of the CNN to describe the expected "shape" of the input 910, and can also be applied to the CNN's input data, such as the incoming video stream 820. As an example, Figure 9 The diagram illustrates the effect of applying a bitmap mask to an image of a pedestrian for training or inference in a CNN network, where 901 represents the original image of pedestrian 901, and 903 represents the corresponding version with the bitmap mask applied. Similarly, 902 shows an image without pedestrians, and 904 shows the corresponding bitmap mask version. The same approach can be applied to any type of 2D or 3D object to reduce the number of operations required for CNN training or inference by leveraging knowledge of the expected 2D or 3D geometry anticipated by the detector. An example of a 3D volumetric bitmap is shown in 911. 920 illustrates the use of a 2D bitmap for inference in a real-world scene.
[0034] exist Figure 9 In the example implementation, a conceptual bitmap is shown (at 900), while the actual bitmap is generated by averaging a series of training images for a specific category of object 910. The example shown is two-dimensional; however, similar bitmap masks can also be generated for 3D objects in the proposed volumetric data format with 1 bit per voxel. In fact, the method can potentially be extended to use additional bits per voxel / pixel to specify the desired color gamut or other characteristics of the 2D or 3D object, among other example implementations.
[0035] Figure 10 This is a table illustrating the results of example experiments involving the analysis of 10,000 CIFAR-10 test images, according to some embodiments. In some implementations, operation cancellation can be used to eliminate the effects caused by changes in CNN networks (such as...). Figure 10 The intermediate computations in 1D, 2D, and 3D CNNs caused by frequent Rectified Linear Unit (ReLU) operations in LeNet 1000 (as shown). Figure 10 As shown in the experiment using 10,000 CIFAR-10 test images, the percentage of zero values generated by the ReLU unit in relation to the data can be as high as 85%. This means that in the case of zero values, a system can be provided that recognizes the zero value and, in response, does not acquire the corresponding data and perform the corresponding multiplication operation. In this example, the 85% represents the percentage of ReLU dynamic zero values generated from a modified National Institute of Standards and Technology (MNIST) database test dataset. The elimination of corresponding operations for these zero values can be used to reduce power consumption and memory bandwidth requirements, among other examples of benefits.
[0036] Unimportant operations can be eliminated based on bitmaps. For example, the use of such bitmaps can be based on the principles and embodiments discussed and illustrated in U.S. Patent No. 8,713,080, entitled "Circuit for compressing data and a processor employing the same," the contents of which are incorporated herein by reference in their entirety. Some implementations can provide hardware capable of using such bitmaps, such as the systems, circuits, and other implementations discussed and illustrated in U.S. Patent No. 9,104,633, entitled "Hardware for performing arithmetic operations," the contents of which are also incorporated herein by reference in their entirety.
[0037] Figure 11 Hardware, according to some embodiments, can be incorporated into a system to provide functionality for bitmap-based culling of unimportant operations. In this example, a multi-layer neural network is provided, comprising repeating convolutional layers. The hardware may include one or more processors, one or more microprocessors, one or more circuits, one or more computers, etc. In this particular example, the neural network includes an initial convolutional processing layer 1100, followed by pooling processing 1110, and finally activation function processing, such as a rectified linear unit (ReLU) function 1120. The output of the ReLU unit 1120, which provides a ReLU output vector 1131, can be concatenated to a subsequent convolutional processing layer 1180 (e.g., possibly via delay 1132), which receives the ReLU output vector 1131. In one example implementation, a ReLU bitmap 1130, representing which elements in the ReLU output vector 1131 are zero and which are non-zero, can also be generated in parallel with the connection from the ReLU unit 1120 to the subsequent convolutional unit 1180.
[0038] In one implementation, a bitmap (e.g., 1130) can be generated or otherwise provided to inform enabled hardware to eliminate the opportunity for operations involving computations of the neural network. For example, bits in the ReLU bitmap 1130 can be interpreted by a bitmap scheduler 1160, which, given a corresponding binary zero in the ReLU bitmap 1130, will always produce zero as output, instructing the multipliers in subsequent convolution units 1180 to skip zero-value entries in the ReLU output vector 1131. In parallel, memory fetches from address generator 1140 for data / weights corresponding to zero values in the ReLU bitmap 1130 can also be skipped, as the fetched weights contain few values and will be skipped by subsequent convolution units 1180. If weights are fetched from the attached DDR DRAM memory device 1170 via DDR controller 1150, the latency may be too high to save only some on-chip bandwidth and associated power consumption. On the other hand, if the weights are retrieved from the on-chip RAM 1180, it may be possible to bypass / skip the entire weight retrieval operation, especially if a delay corresponding to the RAM / DDR retrieval delay 1132 is added to the input of the subsequent convolutional unit 1180.
[0039] Go to Figure 12 Simplified block diagrams according to some embodiments are presented to illustrate improvements to example hardware equipped with circuitry and other logic for eliminating unimportant operations (or performing operation elimination). Figure 12 As shown in the example, additional hardware logic can be provided to predict the sign of the input to ReLU unit 1220 from the leading max-pooling unit 1210 or convolution unit 1200. Adding sign prediction and ReLU bitmap generation to max-pooling unit 1210 allows for earlier prediction of ReLU bitmap information from a timing perspective, overriding potential delays that may occur via address generator 1240, external DDR controller 1250 and DDR memory 1270, or internal RAM memory 1271. If the delay is sufficiently low, the ReLU bitmap can be interpreted in address generator 1240, and memory fetches associated with the zero values of the ReLU bitmap can be completely skipped, as it can be determined that the result fetched from memory will never be used. Figure 11 This modification of the scheme can save additional power, and if the latency via the DDR access path (e.g., 1240 to 1250 to 1270) or RAM access path (e.g., 1240 to 1271) is low enough that the latency stage 1232 is not guaranteed, it can also allow the removal of the latency stages (e.g., 1132, 1232) at the input of the subsequent convolutional unit 1280, as well as other example features and functions.
[0040] Figure 13 This is another simplified block diagram of illustrative example hardware according to some embodiments. For example, a CNN ReLU layer can produce a high number of zero output values corresponding to negative inputs. In fact, negative ReLU inputs can be obtained by looking at previous layers (e.g., ...). Figure 13 The pooling layer in the example uses (multiple) signed inputs to predictively determine the sign. Floating-point and integer arithmetic can explicitly mark the sign in the manner of the most significant bit (MSB), allowing a simple bitwise XOR operation on the input vectors to be multiplied across the convolutional layer to predict which multiplications will produce zero output values, such as... Figure 13 As shown in the diagram. The resulting ReLU bitmap vector with predicted symbols can be used as a basis for determining the subset of multiplications to be eliminated and the associated coefficients read from memory, as described in the other examples above.
[0041] Providing a way to revert the generation of the ReLU bitmap to a previous pooling or convolution stage (i.e., the stage preceding the corresponding ReLU stage) can result in additional power consumption. For example, sign prediction logic can be provided to disable the multiplier when it will produce a negative output that will eventually be set to zero by the ReLU activation logic. This is illustrated, for example, in the case where the two sign bits 1310 and 1315 of the inputs 1301 and 1302 of multiplier 1314 are logically combined via an XOR gate to form the pre-ReLU bitmap bit 1303. This same signal can be used to disable the operation of multiplier 1314, which would otherwise unnecessarily consume energy to generate a negative output that will be set to zero by the ReLU logic before being input for multiplication in the next convolution stage 1390, and other examples.
[0042] Note that the representations of 1300, 1301, 1302, and 1303 (symbol A) indicate that... Figure 13 The B-marked representation in the diagram represents a higher-level view of the view shown. In this example, the input to block 1302 may include two floating-point operands. Input 1301 may include an explicit sign bit 1310, a mantissa of multiple bits 1311, and an exponent of multiple bits 1312. Similarly, input 1302 may also include a sign bit 1315, a mantissa 1317, and an exponent 1316. In some implementations, the mantissa and exponent may have different precisions because the sign of the result 1303 depends only on the signs of 1301 and 1302 or 1310 and 1315, respectively. In practice, neither 1301 nor 1302 needs to be floating-point numbers, but they can be any integer or fixed-point format as long as they are signed numbers and their most significant bit (MSB) is explicitly or implicitly (e.g., if the number is in one's complement or two's complement, etc.) effectively a sign bit.
[0043] continue Figure 13 For example, an XOR (sometimes alternatively referred to as ExOR or EXOR in this document) gate can be used to combine two signed inputs 1310 and 1315 to generate a bitmap bit 1303, which can then be processed by hardware to identify downstream multiplications that can be omitted in the next convolutional block (e.g., 1390). In the case where two input numbers 1313 (e.g., corresponding to 1301) and 1318 (e.g., corresponding to 1302) have opposite signs and produce a negative output 1304, which will be set to zero by the ReLU block 1319, resulting in a zero value in the ReLU output vector 13191 to be input to the subsequent convolutional stage 1390, the same XOR output 1303 can also be used to disable the multiplier 1314. Therefore, in some implementations, the pre-ReLU bitmap 1320 can be passed in parallel to a bitmap scheduler 1360, which can schedule (and / or omit) multiplications to be performed on the convolution unit 1390. For example, for each zero value in the bitmap 1320, the corresponding convolution operation can be skipped in the convolution unit 1390. In parallel, the bitmap 1320 can be consumed by an example address generator 1330, which controls the acquisition of weights used in the convolution unit 1390. The list of addresses corresponding to the 1 values in the bitmap 1320 can be compiled in the address generator 1330 and controlled via a DDR controller 1350 to either DDR memory 1370 or on-chip RAM 1380. In any case, the weights corresponding to the 1 values in the previous ReLU bitmap 1320 (e.g., after some wait time in clock cycles to the weight input 1371) can be obtained and presented to the convolution block 1390, while the acquisition of the weights corresponding to the zero values can be omitted, and so on, in other examples.
[0044] As described above, in some implementations, a delay (e.g., 1361) can be positioned between the bitmap scheduler 1360 and the convolution unit 1390 to balance the delays of the path through the address generator 1330, the DDR controller 1350, and DDR 1350, or through the address generator 1330 and the internal RAM 1380. This delay allows the convolution driven by the bitmap scheduler to be correctly aligned in time with the corresponding weights used for convolution computation in the convolution unit 1390. In effect, from a timing perspective, generating the ReLU bitmap earlier than the output of the ReLU block 1319 allows for additional time that can be used to intercept reads from the memory (e.g., RAM 1380 or DDR 1370) before the address generator 1330 generates them, allowing some reads (e.g., corresponding to zero values) to be predetermined. Because memory reads can be much more expensive than on-chip logic operations, excluding such memory accesses can result in very significant energy savings, among other example advantages.
[0045] In some implementations, if the clock cycle savings are still insufficient to cover DRAM access time, block-oriented techniques can be used to read groups of sign bits (e.g., 1301) from DDR in advance. These groups of sign bits can be used with a block of sign bits from the input image or intermediate convolutional layer 1302 to generate a pre-ReLU bitmap block using a set(s) of XOR gates 1300 (e.g., to compute the difference between sign bits in a 2D or 3D convolution between 2D or 3D arrays / matrices, and other examples). In such implementations, an additional 1-bit storage in DDR or on-chip RAM can be provided to store the sign for each weight, but this allows for covering the latency of multiple cycles by avoiding reading weights that will be multiplied by zero from the ReLU stage. In some implementations, the additional 1-bit storage per weight in DDR or on-chip RAM can be avoided because the signs can be stored in a way that is independent of exponent and mantissa addressable, and other examples of considerations and implementations exist.
[0046] In some implementations, accessing readily available training sets to train machine learning models (including those discussed above) can be particularly difficult. In fact, in some cases, the training set may not exist for a specific machine learning application or correspond to the type of sensor that will generate inputs for the model to be trained, among other example problems. In some implementations, synthetic training sets can be developed and utilized to train neural networks or other deep reinforcement learning models. For example, instead of acquiring or capturing a training dataset consisting of hundreds or thousands of images of specific people, animals, objects, products, etc., a synthetic 3D representation of the subject can be generated manually (e.g., using graphic design or 3D photo editing tools) or automatically (e.g., using a 3D scanner), and the resulting 3D model can serve as the basis for automatically generating training data associated with the 3D model subject. This training data can be combined with other training data to form a training dataset that is at least partially composed of synthetic training data, and this training dataset can be used to train one or more machine learning models.
[0047] As an example, deep reinforcement learning models or other machine learning models (such as those described herein) can be used to allow autonomous machines to scan the shelves of stores, warehouses, or other businesses to assess the availability of certain products within the store. Thus, machine learning models can be trained to allow autonomous machines to detect individual products. In some cases, the machine learning model can identify not only the products on the shelf but also the quantity of products on the shelf (e.g., using a deep model). Instead of training the machine learning model using a series of real-world images of every single product or item carried in the store (e.g., from the same or different stores), every single product or every configuration of a product (e.g., every pose or view (full or partial) of a product under various lighting conditions on various displays, views of product packaging in various orientations, etc.), synthetic 3D models of each product (or at least some of the products) can be generated (e.g., from the product provider, the machine learning model provider, or other sources). The 3D models can be photorealistic or near-photorealistic in their detail and resolution. 3D models can be provided together with other 3D models for consumption to generate various different views of a given subject (e.g., a product) or even to generate collections of different subjects (e.g., a collection of products on a store shelf, various combinations of products placed next to each other in different orientations under different lighting, etc.) to generate synthetic training data image sets, and other example applications.
[0048] Go to Figure 14A simplified block diagram 1400 of an example computing system (e.g., 1415) is shown, which implements a training set generator 1420 to generate synthetic training data for a machine learning system 1430 to train one or more machine learning models (e.g., 1435) (such as deep reinforcement learning models, Siamese neural networks, convolutional neural networks, and other artificial neural networks). For example, a 3D scanner 1405 or other tools can be used to generate a set of 3D models 1410 (e.g., people, indoor and / or outdoor buildings, landscape elements, products, furniture, transportation elements (e.g., road signs, cars, traffic hazards, etc.), and these 3D models 1410 can be consumed by the training set generator 1420 or can be provided as input to the training set generator 1420. In some implementations, the training data generator 1420 can automatically render training image sets, point clouds, depth maps, or other training data 1425 from the 3D models 1410. For example, the training set generator 1420 can be programmed to automatically tilt, rotate, and scale the 3D model, and capture a set of images of the 3D model in different orientations and poses and different (e.g., computer-simulated) lighting, wherein all or part of the entire 3D model subject is captured in the images, to capture a large number and variety of images to satisfy a “complete” and diverse set of images to capture the subject.
[0049] In some implementations, synthetic training images generated from 3D models can have photographic-realistic resolution comparable to the real subjects(s) on which they are based. In some cases, the training set generator 1420 can be configured to automatically render or generate images or other training data from the 3D model in a manner that intentionally reduces the resolution and quality of the resulting images (compared to a high-resolution 3D model). For example, image quality can be reduced by adding noise, applying filters (e.g., Gaussian filters), introducing noise by adjusting one or more rendering parameters, reducing contrast, reducing resolution, changing brightness levels, and bringing the image to a quality level comparable to that produced by sensors (e.g., 3D scanners, cameras, etc.) intended to provide input to the machine learning model to be trained.
[0050] When constructing datasets specifically for training deep neural networks, training set generator systems can define and consider many different conditions or rules. For example, CNNs traditionally require large amounts of data for training to produce accurate results. Synthetic data can avoid situations where the available training dataset is too small. Therefore, the target number of training data samples can be identified for a specific machine learning model, and the training set generator can satisfy the desired training sample size based on the number and type of training samples generated. Furthermore, the training set generator can be designed and considered to generate sets with variations exceeding a threshold in the samples. This is to minimize overfitting of the machine learning model and provide the necessary generalization to perform well in scenarios with a large amount of high variation. Such variations can be achieved through adjustable parameters applied by the training set generator, such as camera angle, camera height, field of view, lighting conditions, etc., used to generate individual samples from the 3D model.
[0051] In some embodiments, a sensor model (e.g., 1440) may be provided that defines aspects of a particular type or model of sensor (e.g., a specific 2D or 3D camera, LiDAR sensor, etc.), wherein model 1440 defines filtering and other modifications to be made to raw images, point clouds, or other training data (e.g., generated from a 3D model) to simulate data generated by the modeled sensor (e.g., resolution, glare sensitivity, light / dark sensitivity, noise sensitivity, etc.). In such instances, the training set generator may artificially degrade samples generated from the 3D model to simulate equivalent images or samples generated by the modeled sensor. In this way, samples in synthetic training data can be generated that are comparable in quality to the data to be input into a trained machine learning model (e.g., generated from real-world versions of (multiple) sensors).
[0052] Go to Figure 15 The diagram illustrates an example block diagram 1500 generated from synthetic training data. For example, a 3D model 1410 of a specific subject can be generated. The 3D model can be a realistic or photorealistic representation of the subject. Figure 15In the example, model 1410 represents a cardboard package containing a set of glass bottles. An image set (e.g., 1505) can be generated based on 3D model 1410, capturing various views of the 3D model 1410, including views of the 3D model under various lighting, environments, and conditions (e.g., in use / idle, open / closed, damaged, etc.). Image set 1505 can be processed using sensor filters (e.g., defined in the sensor model) from the example training data generator. Processing of image 1505 can result in image 1505 being modified to degrade it, generating a "realistic" image 1425 that simulates the quality and features of an image captured using a real sensor.
[0053] In some implementations, to help generate degraded versions of synthetic training data samples, the model (e.g., 1410) can include metadata indicating the material and other properties of the model's subject. In such implementations, the training data generator (e.g., incorporating a sensor model) can consider the properties of the subject defined in the model to determine how a particular sensor might generate a realistic image (or point cloud) for a given illumination, the sensor's position relative to the modeled subject, the subject's properties (e.g., (multiple) materials), and other considerations. For example, Figure 15 In a specific example, model 1410 (modeling the packaging of a bottle) can include metadata to define which parts of the 3D model (e.g., which pixels or polygons) correspond to the glass material (bottle) and which parts correspond to the cardboard (packaging). Therefore, when the training data generator generates a synthetic image and applies sensor filtering 1510 to image 1505 (e.g., modeling how light reflects from various surfaces of the 3D model), the modeled sensor response to these characteristics can be more realistically applied to generate believable training data that more closely resembles the data created by an actual sensor when used to generate training data. For example, the material modeled in the 3D model can allow the generation of training data images that simulate the sensor's sensitivity to generate images with glare, noise, or other defects, such as... Figure 15 The image in the example corresponds to the reflection from the glass surface of the bottle, but with less noise or glare compared to the less reflective cardboard surface. Similarly, the material type, temperature, and other properties modeled in the 3D representation of the subject may have different effects on different sensors (e.g., camera sensors versus LiDAR sensors). Therefore, the example test data generator system can take into account the metadata of the 3D model and the specific sensor model to automatically determine which filters or processes to apply to image 1505 in order to generate a downgraded version (e.g., 1425) of the image that is very likely generated by a real-world sensor.
[0054] Additionally, in some implementations, further post-processing of image 1505 may include depth-of-field adjustment. In some 3D rendering programs, the virtual camera used in the software is perfect, capturing objects at both close and distant points with perfect focus. However, this may not be true for real-world cameras or sensors (and may be defined in the properties of the corresponding sensor model used by the training set generator). Therefore, in some implementations, depth-of-field effects can be applied to the image during post-processing (e.g., using a training set generator to automatically identify and select points on the background where the camera will focus, causing features of the modeled subject to be out of focus, thus creating instances of flawed but more realistic photographs (e.g., 1425)). Additional post-processing can include adding noise to the image to simulate noise artifacts that occur in photography. For example, the training set generator can include adding noise by limiting the number of light bounces calculated on the object by the ray tracing algorithm, as well as other example techniques. Additionally, slight pixelation can be applied on top of the rendered model to eliminate any overly or unrealistically smooth edges or surfaces resulting from the compositing process. For example, a light blur layer can be added to average out the “blocks” of pixels, which, combined with other post-processing operations (e.g., based on the corresponding sensor model), can produce more realistic synthetic training samples.
[0055] like Figure 15 As shown, when generating training data samples (e.g., 1425) from example 3D model 1410, samples (e.g., images, point clouds, etc.) can be added to or included in other real or synthetically generated training samples to construct a training dataset for a deep learning model (e.g., 1435). Also as... Figure 14 As illustrated in the examples, (e.g., the model trainer 1455 of example machine learning system 1430) can be used to train one or more machine learning models 1435. In some cases, synthetic depth images can be generated from the 3D models(s). The trained machine learning model 1435 can then be used by an autonomous machine to perform various tasks, such as object recognition, automated inventory management, navigation, etc.
[0056] exist Figure 14 In the example, the computing system 1415 implementing the example training set generator 1420 may include one or more data processing devices 1445, one or more computer-readable storage elements 1450, and logic implemented in hardware and / or software to implement the training set generator 1420. It should be understood that, although Figure 14The examples show that computing system 1415 (and its components) is separate from machine learning system 1430 for training and / or executing machine learning models (e.g., 1435), but in some implementations, a single computing system may be used to implement a combination of two or more of the functions of model generator 1405, training set generator 1420 and machine learning system 1430, as well as other alternative implementations and example architectures.
[0057] In some implementations, a computational system can be provided that is capable of one-time learning using synthetic training data. Such systems can classify objects without requiring training on hundreds or thousands of images. One-time learning allows classification from a small number of training images, or even, in some cases, a single training image. This saves time and resources in developing training sets for specific machine learning models. In some implementations, such as in Figure 16 and Figure 17 As shown, a machine learning model can be a neural network that learns to distinguish between two inputs, rather than a neural network that learns to classify its inputs. The output of such a machine learning model can identify a measure of similarity between two inputs provided to the model.
[0058] In some implementations, machine learning systems can be provided, such as in Figure 14 In the example, this machine learning system can simulate the ability to classify object categories from a small number of training examples. Such systems can also eliminate the need to create large datasets of multiple classes to efficiently train the corresponding machine learning models. Similarly, machine learning models that do not require training on multiple classes can be chosen. The machine learning model can be used to identify an object by feeding a single image of the object (e.g., a product, person, animal, or other object) along with a comparison image. If the system cannot identify the comparison images, a machine learning model (e.g., Siam Network) is used to determine that the objects do not match.
[0059] In some implementations, the Siamese network can be used as a machine learning model trained with synthetic training data, as illustrated in the examples above. For instance, Figure 16A simplified block diagram 1600 is shown, illustrating an example Siamese network consisting of two identical neural networks 1605a and 1605b, each with the same weights after training. A comparison block (e.g., 1620) can be provided to evaluate the similarity of the outputs of the two identical networks and compare the determined similarity to a threshold. If the similarity is within the threshold range (e.g., below or above a given threshold), the output of the Siamese network (consisting of the two neural networks 1605a and 1605b and comparison block 1620) can indicate (e.g., at 1625) whether the two inputs reference a common object. For example, two samples 1610 and 1615 (e.g., images, point clouds, depth images, etc.) can be provided as inputs to each of the two identical neural networks 1605a and 1605b, respectively. In one example, neural networks 1605a and 1605b can be implemented as ResNet-based networks (e.g., ResNet50 or another variant), and the output of each network can be a feature vector input to comparison block 1620. In some implementations, the comparison block can generate similarity vectors from two feature vector inputs to indicate the similarity between the two inputs 1610 and 1615. In some implementations, the inputs (e.g., 1610, 1615) of such Siamese network implementations can constitute two images representing the agent's current observation (e.g., an image or depth map generated by the autonomous machine's sensors) and the target. Deep Siamese networks are a type of two-stream neural network model for discriminative embedded learning, which can learn in one go using synthetic training data. For example, at least one of the two inputs (e.g., 1610, 1615) can be a synthetically generated training image or a reference image, as discussed above.
[0060] In some implementations, the execution of the Siam network or other machine learning models trained using synthetic data can leverage specialized machine learning hardware (such as machine learning accelerators, e.g., Intel Motor Neural Compute Stick (NCS)) that can interface with a general-purpose microcomputer, among other example implementations. The system can be used in a variety of applications. For example, the network can be used in security or authentication applications, such as applications that identify a person, animal, or vehicle before allowing the triggering of an actuator that allows access to that person, animal, or vehicle. As a specific example, a smart door can be equipped with image sensors to identify a person or animal approaching the door and can grant access only to a user matching a set of authorized users (using a machine learning model). Such machine learning models (e.g., trained using synthetic data) can also be used in industrial or commercial applications, such as product verification, inventory and other applications that utilize product identification in a store to determine the presence (or quantity) of a product or its location within a specific area (e.g., on the appropriate shelf), among other examples. For example, as in Figure 17 As shown in the example of the simplified block diagram 1700, two sample images 1705 and 1710 related to consumer products can be provided as input to a Siamese network with threshold determination logic 1720. The Siamese network model 1720 can determine whether the two sample images 1705 and 1710 are likely images of the same product (at 1715). In practice, in some implementations, such a Siamese network model 1720 can identify products at various rotations and occlusions. In some instances, the 3D model can be used to generate multiple reference images for more complex products and objects for additional levels of verification, as well as other exemplary considerations and features.
[0061] In some implementations, the computing system may be equipped with logic and hardware suitable for performing machine learning tasks to register or merge two or more independent point clouds. To perform point cloud merging, a transformation needs to be found that aligns the contents of the point clouds. This type of problem is common in applications involving autonomous machines, such as in robotic perception applications, creating maps for unknown environments, and other use cases.
[0062] In some implementations, convolutional networks can be used as a solution to find relative poses between 2D images, thus providing results comparable to traditional feature-based methods. Advances in 3D scanning technology allow for the creation of multiple datasets where 3D data aids in training neural networks. In some implementations, a machine learning model can be provided that accepts two or more streams of different inputs, each containing its own three-dimensional (3D) point cloud. The two 3D point clouds can be representations of the same physical (or a virtualized version of physical) space or objects measured from two different corresponding poses. The machine learning model can accept these two 3D point cloud inputs and generate, as output, an indication of the relative or absolute pose between the sources of the two 3D point clouds. The relative pose information can then be used to generate a global 3D point cloud representation of the environment from multiple snapshots (3D point clouds) of the environment (from one or more different sensors and devices, e.g., multiple drones or moving the same drone to scan the environment)). Relative pose can also be used to compare a 3D point cloud input measured by a particular machine with a previously generated global 3D point cloud representation of the environment to determine the relative position of the particular machine within the environment, among other example uses.
[0063] In one example, a voxelized point cloud processing technique is used to create a 3D mesh to sort the points, where convolutional layers, such as..., can be applied. Figure 18 As shown in the example. In some implementations, 3D meshes or point clouds can be implemented or represented as voxel-based data structures, such as those discussed in this paper. For example, in Figure 18In the example, a voxel-based data structure (represented by 1810) can be generated from a point cloud 1805 generated by scanning an example 3D environment with an RGB-D camera or LiDAR. In some implementations, two point cloud inputs can be provided to a comparative machine learning model; for example, the point cloud input using a Siamese network can be a pair of 3D voxel meshes (e.g., 1810). In some instances, the two inputs can first be voxelized at any of a plurality of potential voxel resolutions (and converted into a voxel-based data structure (e.g., 1810)), as discussed above.
[0064] In the Figure 19 In one example represented by the simplified block diagram 1900, the machine learning model can consist of a representation part 1920 and a regression part 1925. As described above, the Siamese network-based machine learning model can be configured to directly estimate the relative camera pose from a pair of 3D voxel grid inputs (e.g., 1905, 1910) or other point cloud data. In some implementations, given the organization of the voxel grid data, the 3D voxel grid can be advantageously used in a neural network with classic convolutional layers. The relative camera pose determined using the network can be used to merge the corresponding point clouds of the voxel grid inputs.
[0065] In some implementations, the representation part 1920 of the example network may include a Siamese network with shared weights and biases. Each branch (or channel of the Siamese network) is formed by successive convolutional layers to extract feature vectors from their respective inputs 1905 and 1910. Furthermore, in some implementations, a rectified linear unit (ReLU) can be provided as the activation function after each convolutional layer. In some cases, pooling layers can be omitted to ensure that spatial information of the data is preserved. The feature vectors output from the representation part 1920 of the network can be combined to enter the regression part 1925. The regression part 1925 includes a set of fully connected layers capable of producing an output 1930 representing the relative pose between the two input point clouds 1905 and 1910. In some implementations, the regression part 1925 may consist of two sets of fully connected layers, one set responsible for generating rotation values for pose estimation and the second set responsible for generating translation values for pose. In some implementations, the fully connected layer of the regression part 1925 can be followed by a ReLU activation function (except for the final layer, since the output may have negative values), as well as other exemplary features and implementations.
[0066] In some implementations, self-supervised learning can be performed on the machine learning model during the training phase, such as the one mentioned above. Figure 19 In the example. For example, due to the network (in Figure 19The goal (as shown in the example) is to solve a regression problem, so a loss function can be provided to guide the network to achieve its solution. A training phase can be provided to derive the loss function, such as a label-based loss function or a loss function quantized by aligning two point clouds. In one example, an example training phase 2020 can be implemented, such as... Figure 20 A simplified block diagram 2000 is shown, where input 2005 is provided to a method 2015 based on Iterative Nearest Point (ICP) (e.g., related to the corresponding CNN 2010), which obtains a ground truth based on data for the pose at which y is compared with the network prediction by the prediction loss function 2025. In such examples, a dataset with labeled ground truth is not required.
[0067] A training model based on Siamese networks (such as...) Figures 19-20 The methods discussed herein can be used in applications such as 3D map generation, navigation, and positioning. For example, Figure 21 As illustrated in the examples, such networks can be used by mobile robots (e.g., 2105) or other autonomous machines (in conjunction with machine learning hardware (e.g., NCS device 2110)) to assist in navigation within an environment. The network can also be used to generate 3D maps of the environment 2115 (e.g., which can later be used by the robot or autonomous machine), as well as other simultaneous localization and mapping (SLAM) applications.
[0068] In some implementations, edge-to-edge machine learning can be used to perform sensor fusion within an application. Such solutions can regress a robot's movement over time by fusing data from different sensors. While this is a well-studied problem, current solutions suffer from time-varying drift or are computationally expensive. In some examples, machine learning methods can be used for computer vision tasks while being less sensitive to noise, lighting variations, and motion blur in the data, among other advantages. For instance, convolutional neural networks (CNNs) can be used for object recognition and to compute optical flow. The system hardware implementing CNN-based models can employ hardware components and subsystems, such as blocks of long short-term memory (LSTM), to identify additional efficiencies, such as good results in signal regression, among others.
[0069] In one example, a system could be provided that utilizes a machine learning model capable of accepting input from multiple sources of different types of data (e.g., RGB and IMU data) to independently overcome the weaknesses of each source (e.g., monocular RGB: lack of scaling, IMU: drift over time, etc.). The machine learning module could include appropriate neural networks (or other machine learning models) tailored to the analysis of each type of data source, which could be connected and fed into stages of fully connected layers to generate results (e.g., poses) from multiple data streams. Such systems could be used, for example, in computing systems for enabling autonomous navigation of machines (such as robots, drones, or vehicles), among other example applications.
[0070] For example, such as Figure 22 As shown in the example, IMU data can be provided as input to a network tailored to the IMU data. IMU data can provide a method for tracking the movement of a subject by measuring acceleration and orientation. However, in some cases, when used alone in machine learning applications, IMU data may drift over time. In some implementations, LSTM can be used to track the relationship of this data over time, thereby helping to reduce drift. Figure 22 In one example shown in the simplified block diagram 2200, a subsequence of n raw accelerometer (acc) and gyroscope (gyr) data elements 2205 is used as the input to the example LSTM 2210 (e.g., where each data element consists of 6 values (from the IMU's accelerometer and gyroscope's 3 axes, respectively)). In other instances, the input 2205 may include a subsequence of n (e.g., 10) relative poses between image frames (e.g., frame f). i and frame f i+1 n relative poses between them, Fully connected layers (FC) 2215 can be provided in the network to extract the rotation and translation components of the transformation. For example, after the fully connected layer 2215, the resulting output can be fed into each of the fully connected layer 2220 for extracting the rotation value 2230 and the fully connected layer 2225 for extracting the translation value 2225. In some implementations, the model can be constructed to include fewer LSTM layers (e.g., 1 LSTM layer) and a greater number of LSTM units (e.g., 512 or 1024 units, etc.). In some implementations, a set of three fully connected layers 2215 is used, followed by fully connected layers for rotation 2220 and translation 2225.
[0071] Go to Figure 23The example illustrates a simplified block diagram 2300 of a network capable of processing data streams of image data, such as monocular RGB data. Therefore, an RGB CNN portion can be provided, trained to compute optical flow and featuring dimensionality-reduced characteristics for pose estimation. For example, several fully connected layers can be provided to reduce dimensionality and / or the feature vectors can be shaped into matrices, and a set of four LSTMs can be used to find correspondences between features and reduce dimensionality, etc.
[0072] exist Figure 23 In the example, a pre-trained optical flow CNN (e.g., 2310) can be provided to accept a pair of consecutive RGB images as input 2305. This pre-trained optical flow CNN can be such as FlowNetSimple, FlowNetCorr, guided optical flow, VINet, or other optical networks. The model can be further constructed to extract feature vectors from the image pair 2305 via the optical flow CNN 2310, and then reduce these vectors to obtain a pose vector corresponding to the input 2305. For example, the output of the optical network portion 2310 can be fed to a set of one or more additional convolutional layers 2315 (e.g., for reducing the dimensionality from the output of the optical network portion 2310 and / or removing information used for estimating the flow vector but not needed for pose estimation), whose output can be a matrix that can be flattened into a corresponding vector (at 2320). This vector can be fed to a fully connected layer 2325 to perform dimensionality reduction of the flattened vector (e.g., reduction from 1536 to 512). The reduced vector can be fed back to the shaping block 2330 to transform or shape the vector into a matrix. Then, a set of four LSTM 2335s (one for each direction of the shaped matrix (e.g., left-to-right, top-to-bottom; right-to-left, bottom-to-top; top-to-bottom, left-to-right; and bottom-to-top, right-to-left)) can be used to track feature correspondences over time and reduce dimensionality. The output of the LSTM set 2335 can then be fed to a rotation fully connected layer 2340 to generate rotation values 2350 based on the image pair 2305, and to a translation fully connected layer 2345 to generate rotation values 2355 based on the image pair 2305, among other example implementations.
[0073] Go to Figure 24 A simplified block diagram 2400 of a sensor fusion network 2405 is shown, which integrates the IMU neural network portion 2410 (e.g., as shown in the diagram). Figure 22 (as shown in the example) and RGB neural network part 2415 (e.g., as shown in the example) Figure 23The results are connected as shown in the example. Such machine learning models 2405 can further enable sensor fusion by taking the optimal from each sensor type. For example, the machine learning model can combine CNNs and LSTMs to obtain more reliable results (e.g., CNNs are able to extract features from a pair of consecutive images, and LSTMs are able to acquire information about the progressive movement of the sensors). In this respect, the outputs of CNNs and LSTMs are complementary, enabling the machine to accurately estimate the differences (and their relative transformations) between two consecutive frames and their representations in real-world units.
[0074] exist Figure 24 In the example, the results from each sensor-specific part (e.g., 2405, 2410) can be concatenated and fed into a fully connected layer of the combined machine learning model 2405 (or sensor fusion network) to generate pose results including rotational and translational poses. While IMU data and monocular RGB alone may not seem to provide sufficient information for a reliable solution to a regression problem, combining these data inputs (as shown and discussed in this paper) can provide more robust and reliable results (e.g., such as...). Figures 25A-25B (Example results shown in 2500a and 2500b). This type of network 2405 utilizes useful information from two sensor types (e.g., RGB and IMU). For example, in this particular example, the RGB CNN part 2415 of network 2405 can extract information about the relative transformations between consecutive images, while the IMU LSTM-based part 2410 provides the scale of the transformation. The corresponding feature vectors output by each part 2410, 2415 can be fed into the concatenator block 2420 to concatenate the vectors and the result is fed into the core fully connected layer 2425, then into the rotation fully connected layer 2430 to generate rotation values 2440 and into the translation fully connected layer 2435 to generate translation values 2445, the generation of which is based on a combination of RGB image 2305 and IMU data 2205. It should be understood that, although Figures 22-24 The example illustrates the fusion of RGB and IMU data. Based on the principles discussed in this paper, other data types can be substituted (e.g., using GPS data to replace and supplement IMU data) and combined into machine learning models. In fact, in other example implementations, more than two data streams (and corresponding neural network parts fed into the connector) can be provided to allow for more robust solutions in other example implementations, as well as other example modifications and alternatives.
[0075] In some implementations, a neural network optimizer can be provided that identifies one or more recommended neural networks for a specific application and the hardware platform for performing machine learning tasks using these neural networks, to the user or system. For example, ... Figure 26 As shown, a computing system 2605 may be provided, which includes a microprocessor 2610 and computer memory 2615. The computing system 2605 may implement a neural network optimizer 2620. The neural network optimizer 2620 may include an execution engine to enable the execution of a set of machine learning tasks (e.g., 2635), such as those described herein, and other example hardware, using machine learning hardware (e.g., 2625). The neural network optimizer 2620 may additionally include one or more probes (e.g., 2630) to monitor the execution of the machine learning tasks, as these tasks are performed using one of a set of neural networks selected by the neural network optimizer 2620 for execution on and monitored by the machine learning hardware 2625. Probe 2630 can measure attributes such as the power consumed by machine learning hardware 2625 during task execution, the temperature of machine learning hardware during execution, the speed or time taken to complete a task using a specific neural network, the accuracy of the task results using a specific neural network, the amount of memory used (e.g., storing the neural network in use), and other example parameters.
[0076] In some implementations, the computational system 2605 may interface with a neural network generation system (e.g., 2640). In some implementations, the computational system for evaluating neural networks (e.g., 2605) and the neural network generation system 2640 may be implemented on the same computational system. The neural network generation system 2640 allows users to manually design neural network models (e.g., CNNs) for various tasks and solutions. In some implementations, the neural network generation system 2640 may additionally include a repository 2645 of previously generated neural networks. In one example, the neural network generation system 2640 (e.g., a system such as CAFFE, TensorFlow, etc.) may generate a set of neural networks 2650. This set may be generated randomly, generating new neural networks from scratch (e.g., based on some generalized parameters suitable for a given application, or according to a general neural network type or genus) and / or randomly selecting neural networks from the repository 2645.
[0077] In some implementations, a set of neural networks 2650 may be generated by a neural network generation system 2640 and provided to a neural network optimizer 2620. The neural network optimizer 2620 enables a specific set of machine learning hardware (e.g., 2625) to use each of the neural networks in the set 2650 to perform a normalized set of one or more machine learning tasks. The neural network optimizer 2620 may monitor the execution of tasks related to the hardware 2625 using each of the neural networks in the set 2650. The neural network optimizer 2620 may additionally accept data as input to identify which parameters or features, measured by the probes (e.g., 2630) of the neural network optimizer, will be given the highest weight or priority when the neural network optimizer determines which one in the set is "best". Based on these criteria and observations made by the neural network optimizer during a period of use (e.g., randomly generated) of the neural network set, the neural network optimizer 2620 may identify and provide the best-performing neural network for a specific machine learning hardware (e.g., 2625) based on the provided criteria. In some implementations, the neural network optimizer may automatically provide this best-performing neural network to the hardware for additional use and training, etc.
[0078] In some implementations, the neural network optimizer may employ evolutionary exploration to iteratively improve the results identified from an initial (e.g., randomly generated) set of neural networks evaluated by the neural network optimizer (e.g., 2620). For example, the neural network optimizer may identify the characteristics of one or more neural networks that perform best from the initial set evaluated by the neural network optimizer. The neural network optimizer may then send a request to a neural network generator (e.g., 2640) to generate another set of different neural networks with performance similar to those identified in the best-performing neural networks for a particular hardware (e.g., 2625). The neural network optimizer 2620 may then repeat its evaluation using the next set or next generation of neural networks generated by the neural network generator based on the best-performing neural networks from the initial batch evaluated by the neural network optimizer. Again, the neural network optimizer 2620 may identify which of the second-generation neural networks performs best according to provided criteria, and again determine the characteristics of the best-performing neural network in the second generation as the basis for sending a request to the neural network generator to generate a third-generation neural network for evaluation, thus utilizing the neural network optimizer 2620 to iteratively evaluate neural networks that evolve (theoretically improve) from one generation to the next. As in the previous examples, the neural network optimizer 2620 can provide instructions or copies of the latest generation of best-performing neural networks for use by machine learning hardware (e.g., 2625), as well as other example implementations.
[0079] As a specific example, such as Figure 27 As shown in block diagram 2700, machine learning hardware 2625 (such as the Movidius NCS) can be used with a neural network optimizer as a design space exploration tool, which can use a neural network generator 2640 or a provider (such as Caffe) to find networks with the highest accuracy under hardware constraints. Such design space exploration (DSX) tools can be provided to utilize a complete API, including bandwidth measurements and network graphs. Furthermore, extensions can be made to the machine learning hardware API or provided within the machine learning hardware API to extract additional parameters useful for design space exploration, such as temperature measurements, inference time measurements, and other examples.
[0080] To demonstrate the power of the DSX concept, this paper provides examples exploring the neural network design space for small, always-on face detectors, such as face detection wake-up implemented in recent mobile phones. Various neural networks can be provided to machine learning hardware, and performance can be monitored for each network's usage, such as the power consumption of networks trained during the inference phase. DSX tools (or neural network optimizers) can generate different neural networks for a given classification task. Data can be transferred to hardware (e.g., in the case of NCS, via USB to and from the NCS using the NCS API). Through the implementation methods explained above, the optimal model for different purposes can be the result of design space exploration, rather than being found by manually editing, copying, and pasting any files. As an illustrative example, Figure 28 Table 2800 is shown, which illustrates example results of the DSX tool's evaluation of several different, randomly generated neural networks, including performance characteristics of the machine learning hardware (e.g., a general-purpose microprocessor connected to the NCS) during the execution of machine learning tasks using each neural network (e.g., showing example parameters such as accuracy, execution time, temperature, size of the neural network in memory, and measured power, as well as other example parameters that can be measured by the DSX tool). Figure 29 The comparison results of validation accuracy with execution time (at 2900) and size (at 2905) are shown. These relationships and ratios can be taken into account by NCS when determining which of the evaluated neural networks is "optimal" for a particular machine learning platform.
[0081] Deep neural networks (DNNs) have delivered state-of-the-art accuracy in a variety of computer vision tasks, such as image classification and object detection. However, the success of DNNs is often achieved through a significant increase in computational and memory requirements, making them difficult to deploy on resource-constrained edge inference devices. In some implementations, network compression techniques such as pruning and quantization can reduce computational and memory demands. This also helps prevent overfitting, especially for transfer learning on small, custom datasets, with minimal loss of accuracy.
[0082] In some implementations, a neural network optimizer (e.g., 2620) or other tools may be provided to dynamically and automatically reduce the size of the neural network for use by a specific machine learning hardware. For example, the neural network optimizer may perform fine-grained pruning (e.g., connection or weight pruning) and coarse-grained pruning (e.g., convolutional kernel, neuron, or channel pruning) to reduce the size of the neural network stored and operated on by the given machine learning hardware. In some implementations, the machine learning hardware (e.g., 2625) may be equipped with arithmetic circuitry capable of performing sparse matrix multiplication, enabling the hardware to efficiently process weight-pruned neural networks.
[0083] In one implementation, the neural network optimizer 2620 or other tools can perform hybrid pruning of the neural network (e.g., ... Figure 30A (As shown in block diagram 3000a), pruning can be performed at both the convolutional kernel level and the weight level. For example, the neural network optimizer 2620 or other tools may consider one or more algorithms, rules, or parameters to automatically identify the set of convolutional kernels or channels that can be pruned from a given neural network 3010 (at 3005). After the first channel pruning step is completed (at 3015), weight pruning 3020 can be performed on the remaining channels 3015, as shown in block diagram 3000a. Figure 30A The example illustration is shown below. For example, a rule can control weight pruning by setting a threshold, thereby pruning weights below the threshold (e.g., reassigning weights "0") (at 3025). The pruned network can then be run or iteratively mixed to allow the accuracy of the network to be recovered from the pruning, resulting in a compact version of the network without harmfully reducing the model's accuracy. Additionally, in some implementations, the remaining weights after pruning can be quantized to further reduce the amount of memory required to store the weights of the pruned model. For example, as... Figure 31As shown in block diagram 3100, logarithmic scaling 3110 can be performed such that the floating-point weight value (at 3105) is replaced by its nearest base-2 counterpart (at 3115). In this way, a 32-bit floating-point value can be replaced with a 4-bit base-2 value, significantly reducing the amount of memory required to store network weights while sacrificing only minimally the accuracy of the compact neural network, as well as other examples of quantization and feature (e.g., as...) Figure 32 (As shown in Table 3200, the example results are shown). In fact, in Figure 32 In specific examples, research on the application of hybrid pruning is shown as being applied to example neural networks such as ResNet50. Additionally, the application of weight quantization on pruned sparse thin ResNet50 is shown to further reduce the model size for greater model friendliness.
[0084] like Figure 30B As shown in the simplified block diagram 3000b, in one example, a hybrid pruning of the example neural network can be performed by accessing an initial or reference neural network model 3035 and (optionally) training that model using regularization (L1, L2, or L0) 3040. The importance of individual neurons (or connections) within the network can be evaluated 3045, where neurons determined to be of lower importance (at 3050) are pruned from the network. The pruned network can be fine-tuned 3055, resulting in a final compact (or sparse) network 3060 from this pruning. Figure 30B As shown in the examples, in some cases, the network can be iteratively pruned, where additional training and pruning (e.g., 3040-3050) are performed after fine-tuning the pruned network, and other example implementations exist. In some implementations, hybrid pruning techniques such as those described above can be used to determine the importance of neurons. For example, fine-grained weight pruning / sparsening can be performed (e.g., using (mean + standard deviation)). Global progressive pruning (of factors). Coarse-grained channel pruning can be performed layer-by-layer based on sensitivity testing and / or the target number of MACs (e.g., weight sum pruning). Coarse pruning can be performed before sparse pruning. For example, weight quantization can also be performed to set constraints on non-zero weights to powers of 2 or 0 and / or to use 1 bit to represent 0 and 4 bits to represent weights. In some cases, low-precision (e.g., weight and activation) quantization can be performed, as well as other example techniques. Pruning techniques (such as those discussed above) can yield a wide variety of example benefits. For example, compact matrices can reduce the size of stored network parameters, while runtime weight decompression can reduce DDR bandwidth. Accelerated computation can also be provided, as well as other example benefits.
[0085] Figure 33AThis is a simplified flowchart 3300a of an example technique for generating a training dataset that includes synthetic training data samples (e.g., synthetically generated images or synthetically generated point clouds). For example, a digital 3D model 3302 can be accessed from computer memory, and multiple training samples 3304 can be generated from various views of the digital 3D model. The training samples can be modified 3306 to add defects to the training samples to simulate training samples generated from one or more real-world samples. A training dataset 3308 is generated to include the modified synthetically generated training samples. The generated training dataset can be used to train one or more neural networks 3310.
[0086] Figure 33B This is a simplified flowchart 3300b for an example technique to perform one-off classification using an example Siamese neural network model. The subject input 3312 can be provided as input to the first part of the Siamese neural network model, and the reference input 3314 can be provided as input to the second part of the Siamese neural network model. The first and second parts of the model can be identical and have the same weights. 3316 The output of the Siamese network can be generated from the outputs of the first and second parts based on the subject input and the reference input (such as a difference vector). For example, 3318 whether the output indicates that the subject input is sufficiently similar to the reference input (e.g., indicating that the subject of the subject input is the same as the subject of the reference input) can be determined based on a similarity threshold of the output of the Siamese neural network model.
[0087] Figure 33C This is a simplified flowchart 3300c of an example technique for determining relative pose using an example Siamese neural network model. For example, a first input 3320 can be received as input to a first part of the Siamese neural network model, representing a view of 3D space (e.g., point cloud data, depth map data, etc.) from a first pose (e.g., an autonomous machine). A second input 3322 can be received as input to a second part of the Siamese neural network model, representing a view of 3D space from a second pose. The output 3324 of the Siamese network can be generated based on the first and second inputs, representing the relative pose between the first and second poses. The machine's position associated with the first and second poses and / or the 3D map where the machine resides can be determined 3326 based on the determined relative pose.
[0088] Figure 33DThis is a simplified flowchart 3300d illustrating an example technique involving a sensor fusion machine learning model that combines at least portions of two or more machine learning models, these portions being adapted for use with a corresponding type of two or more different data types. First sensor data of a first type can be received 3330 as input to the first of two or more machine learning models in the sensor fusion machine learning model. Second sensor data of a second type (e.g., generated simultaneously with the first sensor data (e.g., generated by sensors on the same or different machines)) can be received 3332 as input to the second of the two or more machine learning models. The outputs of the first and second machine learning models can be concatenated 3334, and the concatenated output is provided 3336 to a set of fully connected layers of the sensor fusion machine learning model. Based on the first and second sensor data, the sensor fusion machine learning model can generate 3338 an output defining the pose of a device (e.g., a machine on which the sensors that generated the first and second sensor data are located).
[0089] Figure 33E This is a simplified flowchart 3300e of an example technique for generating improved or optimized neural networks tailored to specific machine learning hardware based on an evolutionary algorithm. For example, 3340 can be accessed or a set of neural networks can be generated (e.g., automatically generated based on randomly selected attributes). The machine learning task can be performed by specific hardware using the set of neural networks 3342, and 3344 the performance attributes of these hardware-specific tasks can be monitored. Based on the results of this monitoring, 3346 one or more of the best-performing neural networks in the set can be identified. The characteristics of the best-performing neural networks (for specific hardware) can be determined, and 3350 another set of neural networks including such characteristics can be generated. In some instances, this new set of neural networks can also be tested (e.g., via steps 3342-3348) to iteratively improve the set of neural networks considered in conjunction with the hardware until one or more neural networks with sufficiently good or adequately optimized performance are identified for the specific hardware.
[0090] Figure 33F This is a simplified flowchart 3300f of an example technique for pruning neural networks. For example, a neural network can be identified as 3352, and a subset 3354 of the convolutional kernels of the neural network can be identified as less important or otherwise good candidates for pruning. This subset of convolutional kernels can be pruned 3356 to generate a pruned version of the neural network. The remaining convolutional kernels 3358 can then be further pruned to trim a subset of the weights from these remaining kernels, further pruning the neural network at both coarse-grained and fine-grained levels.
[0091] Figure 34This is a simplified block diagram of a multi-slot vector processor (e.g., a Very Long Instruction Word (VLIW) vector processor) according to some embodiments. In this example, the vector processor may include multiple (e.g., nine) functional units (e.g., 3403-3411), which may be fed by a multi-port memory system 3400 and backed up by a vector register file (VRF) 3401 and a general-purpose register file (GRF) 3402. The processor includes an instruction decoder (IDEC) 3412 that decodes instructions and generates control signals for the control functional units 3403-3411. Functional units 3403-3411 are the asserted Execution Unit (PEU) 3403, Branch and Repeat Unit (BRU) 3404, Load Storage Port Unit (e.g., LSU0 3405 and LSU1 3406), Vector Operation Unit (VAU) 3407, Scalar Operation Unit (SAU) 3410, Compare and Move Unit (CMU) 3408, Integer Operation Unit (IAU) 3411, and Volume Acceleration Unit (VXU) 3409. In this particular implementation, VXU 3409 can accelerate operations on volumetric data, including both store / retrieval operations, logical operations, and arithmetic operations. Although in Figure 34 In the example, the VXU circuit 3409 is shown as a single component, but it should be understood that the functionality of the VXU (and any of the other functional units 3403-3411) can be distributed among multiple circuits. Furthermore, in some implementations, the functionality of the VXU 3409 can be distributed within one or more other functional units of the processor (e.g., 3403-3408, 3410, 3411), as in other example implementations.
[0092] Figure 35 This is a simplified block diagram illustrating an example implementation of the VXU 3500 according to some embodiments. For example, the VXU 3500 may provide at least one 64-bit input port 3501 to accept input from a vector register file 3401 or a general-purpose register file 3402. This input may be connected to a plurality of functional units to operate on a 64-bit unsigned integer volume bitmap, including a register file 3503, an address generator 3504, point addressing logic 3505, point insertion logic 3506, point deletion logic 3507, 3D-to-2D projection logic in the X dimension 3508, 3D-to-2D projection logic in the Y dimension 3509, 3D-to-2D projection logic in the X dimension 3510, a 2D histogram pyramid generator 3511, a 3D histogram pyramid generator 3512, a group counter 3513, 2D pathfinding logic 3514, 3D pathfinding logic 3515, and possibly additional functional units. The output from block 3502 can be written back to either the vector register file VRF 3401 or the general-purpose register file GRF 3402.
[0093] Go to Figure 36 The example shows a representation of the organization of a 4^3 voxel cube 3600. A second voxel cube 3601 is also shown. In this example, voxel cubes can be defined in data such as a 64-bit integer 3602, where each individual voxel within the cube is represented by a single corresponding bit in the 64-bit integer. For example, voxel 3512 at address {x,y,z}={3,0,3} can be set to "1" to indicate the presence of geometry at that coordinate within the volume space represented by voxel cube 3601. Furthermore, in this example, all other voxels (except voxel 3602) can correspond to "empty" space and can be set to "0" to indicate the absence of physical geometry at those coordinates, and so on. Go to Figure 37 An example two-level sparse voxel tree 3700 is shown according to some embodiments. In this example, only a single “occupied” voxel is included in the volume (e.g., at position {15,0,15}). In this case, the upper level 0 of tree 3701 contains a single voxel entry {3,0,3}. This voxel then points to the next level of tree 3702 that contains a single voxel in element {3,0,3}. The entry in the data structure corresponding to level 0 of the sparse voxel tree is a 64-bit integer 3703 with one voxel set to occupied. This set voxel means that an array of 64-bit integers is then assigned to level 1 of the tree corresponding to the set of voxel volumes in 3703. In the level 1 subarray 3704, only one voxel is set to occupied, while all other voxels are set to unoccupied. In this example, because the tree is two-level, level 1 represents the bottom of the tree, causing the hierarchical structure to terminate there.
[0094] Figure 38A two-layer sparse voxel tree 3800 according to some embodiments is shown, containing occupied voxels at positions {15,0,3} and {15,0,15} in a specific volume. In this case (the specific volume is subdivided into 64 upper-level 0 voxels), the upper-level 0 of tree 3801 contains two voxel entries {3,0,0} and {3,0,3} with corresponding data 3804 indicating that the two voxels are set (or occupied). The next layer of the sparse voxel tree (SVT) is provided as an array of 64-bit integers containing two sub-cubes 3802 and 3803, each sub-cube for a set of voxels in a level 0. In the level 1 subarray 3805, two voxels are set to occupied: v15 and v63, and all other voxels are set to unoccupied and tree. This format is flexible because the 64 entries in the next level of the tree are always assigned a corresponding voxel to each one that is set in the upper level of the tree. This flexibility allows dynamically changing scene geometries to be inserted into the existing volumetric data structure in a flexible manner (i.e., rather than in a fixed order, such as randomly), provided the corresponding voxels in the upper level are set. Without setting corresponding voxels in the upper level, either a table of pointers would be maintained, leading to higher memory requirements, or at least a partial reconstruction of the tree would be required to insert unforeseen geometries.
[0095] Figure 39 An illustration shows a method for storing data from [various sources] according to some embodiments. Figure 38 A voxel replacement technology. In this example, the overall volume of 3900 contains, for example, voxels. Figure 23 The two voxels stored in the global coordinates {15,0,3} and {15,0,15} are used. In this method, instead of allocating an array of 64 entries to represent all sub-cubes in level 1 below level 0, only those elements in level 1 that actually contain geometry (e.g., as indicated by whether the corresponding level 0 voxel is occupied or not) are allocated as corresponding 64-bit level 1 records. In this example, this makes level 1 have only two 64-bit entries instead of 64 entries (i.e., for each of the 64 level 1 voxels, whether occupied or empty). Therefore, in this example, level 0 3904 is equivalent to... Figure 38 The 3804 is one of the higher-level 3804, while the next level, the 3905, has higher memory requirements. Figure 38 The corresponding 3905 is 1 / 62. In some implementations, if no space has already been allocated in level 1 for the new geometry to be inserted into level 0, the tree must be copied and rearranged.
[0096] exist Figure 39In the example, a subvolume can be derived by counting the occupied voxels in the layers above the current layer. In this way, the system can determine where a higher layer in the voxel data terminates and where the next lower layer begins. For example, if three level 0 voxels are occupied, the system can expect three corresponding level 1 entries to follow in the voxel data, and (after these three) the next entry corresponds to the first entry in level 2, and so on. Such optimal compaction can be very useful when parts of the scene do not change over time, or in applications where volume data needs to be transmitted remotely (e.g., from a space probe scanning the surface of Pluto, where the transmission of each bit is expensive and time-consuming).
[0097] Figure 40 This illustrates how voxels can be inserted into a 4^3 cube, represented as entries in a 64-bit integer volume data structure, according to some embodiments, to reflect variations in geometry within the corresponding volume. In one example, as shown in 4000, each voxel block can be organized as four logical 16-bit planes within a 64-bit integer. Each plane corresponds to a Z value from 0 to 3, and each y value within each plane is encoded for a 4-bit shift of 0 to 3 for the four logical values, and finally, each bit within each 4-bit y-plane encodes four possible values (0 to 3) for x, and other example organization. Thus, in this example, to insert a voxel into a 4^3 volume, the x value from 0 to 3 can first be shifted by one bit, then that value can be shifted by 0 / 4 / 8 / 12 bits to encode the y value, and finally, the z value can be represented by shifts of 0 / 16 / 32 / 48 bits, as shown in the C code expression in 4001. Finally, since each 64-bit integer can be a combination of up to 64 voxels, and each voxel is written separately, as shown in 4002, the new bitmap must be logically combined with the old 64-bit value read from the sparse voxel tree by performing an OR operation between the old bitmap value and the new bitmap value.
[0098] Go to Figure 41 The illustration shows a representation according to some embodiments to explain how a 3D volume object stored in a 64-bit integer 4100 can be projected by a logical OR operation in the X direction to produce a 2D pattern 4101, by a logical OR operation in the Y direction to produce a 2D output 4102, and finally by a logical OR operation in the Z direction to produce the pattern shown in 4103. Figure 42The diagram illustrates how bits from an input 64-bit integer are logically ORed according to some embodiments to produce output projections in X, Y, and Z. In this example, Table 4201 shows, column by column, which element indices from input vector 4200 are ORed to produce the x-projection output vector 4202. Table 4203 shows, column by column, which element indices from input vector 4200 are ORed to produce the y-projection output vector 4204. Finally, Table 4205 shows, column by column, which element indices from input vector 4200 are ORed to produce the z-projection output vector 4206.
[0099] The X-projection performs a logical OR operation on the 0th, 1st, 2nd, and 3rd bits of the input data 4200 to produce the 0th bit of the X-projection 4201. For example, the 1st bit of 4201 can be produced by XORing the 4th, 5th, 6th, and 7th bits of 4200, and so on. Similarly, the 0th bit of the Y-projection 4204 can be produced by ORing the 0th, 4th, 8th, and 12th bits of 4200 together. And the 1st bit of 4204 can be produced by ORing the 1st, 5th, 9th, and 13th bits of 4200 together. Finally, the 0th bit of the Z-projection 4206 can be produced by ORing the 0th, 16th, 32nd, and 48th bits of 4200 together. And the 1st bit of 4206 can be produced by ORing the 1st, 17th, 33rd, and 49th bits of 4200 together, and so on.
[0100] Figure 43An example of how projection can be used to generate a simplified map, according to some embodiments, is shown. In this scenario, the objective may be to generate a compact 2D map of a path from voxel volume 4302, along which a vehicle 4300 with height h 4310 and width w 4301 travels. Here, Y-projection logic can be used to generate an initial coarse 2D map 4303 from voxel volume 4302. In some implementations, this map can be processed to check whether a specific vehicle (e.g., a car (or autonomous vehicle), drone, etc.) in a specific dimension can traverse the path's width constraint 4301 and height constraint 4310. This can be performed by performing a projection in Z to check the width constraint 4301, and the projection in Y can be masked to limit the calculation of the vehicle's height 4310, thus ensuring that the path is traversable. As can be seen through (e.g., in software) additional post-processing, for a path that is passable and only satisfies the width and height constraints at X and Z, the coordinates of points A 4304, B 4305, C 4306, D 4307, E 4308, and F 4309 along the path can be stored or transmitted over the network to fully reconstruct the legal path along which a vehicle can travel. Given that the path can be parsed into such segmented pieces, it is possible to fully describe the path using only one or two bytes for each segment of the linear region. This can facilitate the rapid transmission and processing of such path data (e.g., via autonomous vehicles), and other examples.
[0101] Figure 44This paper illustrates how, according to some embodiments, volumetric 3D measurements or simple 2D measurements from embedded devices can be mathematically aggregated to generate high-quality crowdsourced maps as an alternative to precision measurements using LiDAR or other expensive equipment. In the proposed system, multiple embedded devices 4400, 4401, etc., can be equipped with various sensors capable of acquiring measurement results, which can be sent to a central server 4410. Software running on the server performs aggregation 4402 of all measurement results, and a nonlinear solver 4403 performs numerical computation on the resulting matrix to produce a highly accurate map, which can then be redistributed back to the embedded devices. In practice, data aggregation can also include high-accuracy survey data from satellites 4420, aerial LiDAR surveys 4421, and ground-based LiDAR measurements 4422 (where such high-fidelity datasets are available) to improve the accuracy of the resulting map. In some implementations, the measurement results of maps and / or records can be generated using a sparse voxel data structure with a format such as that described herein, the measurement results of maps and / or records can be converted into a sparse voxel data structure with a format such as that described herein, or the measurement results of maps and / or records can be otherwise expressed using a sparse voxel data structure with a format such as that described herein, and other example implementations.
[0102] Figure 45 This is an illustration of how 2D path finding on a 2D 2×2 bitmap can be accelerated according to some embodiments. The principle of operation is that, in order for connectivity to exist between points on the map in the same grid cell, the values of cells consecutively arranged in x or y, or x and y, must all be set to one value. Therefore, a logical AND operation of the bits drawn from those cells can be instantiated to test whether a valid path exists in the bitmap in the grid, and a different AND gate can be instantiated for each valid path through the N×N grid. In some instances, this method can introduce combinatorial complexity because even an 8×8 2D grid can contain 2 64-1 valid path. Therefore, in some improved implementations, the grid can be reduced to 2×2 or 4×4 tiles, which can be used to test connectivity hierarchically. The 2×2 bitmap 4500 contains four bits labeled b0, b1, b2, and b3. These four bits can take values from 0000 to 1111, with corresponding labels 4501 to 4517. Each of these bit patterns represents a different level of connectivity between faces of the 2×2 grid labeled 4521 to 4530. For example, when a 2×2 grid 4500 contains bitmaps 1010 (7112), 1011 (7113), 1110 (7116), or 1111 (7117), there exists 4521 or v0 representing the vertical connectivity between x0 and y0 in 4500. A 2-input logical AND-OR operation, as shown in row 1 of Table 4518, generates v0 in the connectivity graph, which can be used in higher-level hardware or software to determine global connectivity through a global grid that has been subdivided into 2×2 subgrids. If the global graph contains an odd number of grid points on the x-axis or y-axis, the top-level grid will need to be padded to the next highest even number of grid points (e.g., such that an additional row of zero values is added to the x-axis and / or y-axis of the global grid). Figure 45 An exemplary 7×7 grid 4550 is further illustrated, showing how it can be filled to 8×8 by adding additional rows 4532 and columns 4534 filled with zero values. To accelerate pathfinding compared to other techniques (e.g., depth-first search, breadth-first search, Dijkstra's algorithm, or other graph-based methods), this example progressively subsamples the N×N graph 4550 to a 2×2 graph. For example, in this example, cell W in 4540 is filled by performing an OR operation on the contents of cells A, B, C, and D in 4550, and so on. Furthermore, the bits in the 2×2 cells of 4540 are ORed to fill the cells in 4542. For pathfinding, the algorithm starts with the smallest 2×2 representation of grid 4542 and tests each bit. Since we know that a zero-valued bit means there is no corresponding 2×2 grid cell in 4540, the portion of the 4×4 grid (composed of four 2×2 grids) in 4540 that corresponds only to one bit in the 2×2 grid 4542 needs to be tested for connectivity. This method can also be used to search an 8×8 grid in 4520; for example, if cell W in 4540 contains a zero value, then we subsequently know there is no path ABCD in 4520, etc. This method prunes branches from the graph search algorithm used, regardless of whether the graph search algorithm used is A... Dijkstra's algorithm, DFS, BFS, and their variants are also mentioned. Furthermore, the use of a hardware-based pathfinder 4518 with a 2×2 organization can further limit the associated computations. In fact, a 4×4-based hardware element can be composed of five 2×2 hardware blocks with the same arrangement as 4540 and 4542 to further limit the amount of graph search that needs to be performed. Additionally, an 8×8 hardware-based search engine can be constructed using 21 2×2 HW blocks (7118) with the same arrangement as 4542, 4540, 4500, etc., for potentially any N×N topology.
[0103] Figure 46 This is a simplified block diagram based on some embodiments, illustrating how the proposed volumetric data structure can be used to accelerate collision detection. A 3D N×N×N map of the geometry can be subsampled into a pyramid consisting of a lowest level of detail (LoD) 2×2×2 volume 4602, the next highest 4×4×4 volume 4601, an 8×8×8 volume 4600, and so on up to N×N×N. If the location of a drone, vehicle, or robot 4605 is known in 3D space via a positioning method such as GPS, or via repositioning from a 3D map, then that location can be quickly used to test whether the geometry exists or not in the quadrant of the relevant 2×2×2 sub-volume by appropriately scaling the drone / robot's x, y, and z positions (dividing them by the number of relevances), and querying 4602 for the presence of the geometry (e.g., checking if the corresponding bitmap bit is a 1 indicating a possible collision). If a potential conflict exists (e.g., a "1" is found), further checks can then be performed in volumes 4601, 4600, etc., to confirm whether the drone / robot can move. However, if the voxel in 4602 is free (e.g., a "0"), the robot / drone can then interpret the same voxel as free space and manipulate directional controls to move freely across most of the map.
[0104] While some of the systems and solutions described and illustrated herein have been described as comprising or associated with multiple elements, not all elements explicitly shown or described are used in every alternative implementation of this disclosure. Additionally, one or more elements may be located outside the system, and in other instances, certain elements may be included within, or as part of, one or more of the other described elements and other elements not described in the illustrated implementations. Furthermore, certain elements may be combined with other components and components for alternative or additional purposes beyond those described herein.
[0105] Furthermore, it should be understood that the examples presented above are non-limiting examples provided solely for the purpose of illustrating certain principles and features, and do not necessarily limit or constrain possible embodiments of the concepts described herein. For example, various different embodiments can be implemented using various combinations of the features and components described herein (including combinations implemented through various implementations of the components described herein). Other implementations, features, and details should be understood from the content of this specification.
[0106] Figures 47-52 This is a block diagram of an exemplary computer architecture that can be used according to the embodiments disclosed herein. In fact, the computing devices, processors, and other logic and circuitry of the systems described herein may include all or part of the functionality and supporting software and / or hardware circuitry for implementing such functionality. Furthermore, other computer architecture designs known in the art for processors and computing systems, beyond the examples shown herein, may also be used. Generally, the computer architectures suitable for the embodiments disclosed herein may include, but are not limited to, those shown herein. Figures 47-52 The configuration shown in the image.
[0107] Figure 47 An example domain topology of a corresponding Internet of Things (IoT) network, coupled to a corresponding gateway via a link, is shown. The Internet of Things (IoT) is a concept in which a large number of computing devices are interconnected and connected to the Internet to provide very low-level functionality and data acquisition. Therefore, as used herein, IoT devices can include semi-autonomous devices that perform functions such as sensing or control to communicate with other IoT devices and the wider network (such as the Internet). Such IoT devices can be equipped with logic and memory to implement and use hash tables (such as those described above).
[0108] Typically, IoT devices are limited in memory, size, or functionality, allowing for the deployment of a large number of IoT devices at a cost similar to a smaller number of larger devices. However, IoT devices can be smartphones, laptops, tablets, PCs, or other larger devices. Furthermore, IoT devices can be virtual devices, such as applications on smartphones or other computing devices. IoT devices may include IoT gateways for coupling IoT devices to other IoT devices and to cloud applications for data storage, process control, and more.
[0109] IoT device networks can include commercial and home automation equipment such as water supply systems, power distribution systems, pipeline control systems, factory control systems, light switches, thermostats, locks, cameras, alarms, motion sensors, etc. IoT devices can be accessed via remote computers, servers, or other systems, for example, to control systems or access data.
[0110] The future growth of the Internet and similar networks is likely to involve an enormous number of IoT devices. Therefore, in the context of the technologies discussed in this paper, many innovations for this future networking will address the need for unhindered growth of all these layers, the discovery and accessibility of connected resources, and the ability to support the hiding and partitioning of connected resources. Any number of network protocols and communication standards can be used, each designed to address a specific objective. Furthermore, protocols are part of the infrastructure that supports human-accessible services that operate regardless of location, time, or space. These innovations include service delivery and associated infrastructure, such as hardware and software, security enhancements, and the provision of services based on Quality of Service (QoS) specified in service levels and service delivery protocols. As will be understood, IoT devices and networks (such as…) Figure 47 and Figure 48 The use of those described in the document presents many new challenges in heterogeneous connectivity networks that combine wired and wireless technologies.
[0111] Figure 47 A simplified diagram of the domain topology is provided, which can be used for multiple Internet of Things (IoT) networks including IoT devices 4704, where IoT networks 4756, 4758, 4760, and 4762 are coupled to corresponding gateways 4754 via a backbone link 4702. For example, multiple IoT devices 4704 can communicate with gateway 4754 and communicate with each other through gateway 4754. For simplicity, not every IoT device 4704 or communication link (e.g., links 4716, 4722, 4728, or 4732) is labeled. The backbone link 4702 can include any number of wired or wireless technologies (including optical networks) and can be part of a local area network (LAN), wide area network (WAN), or the Internet. Furthermore, this communication link facilitates optical signal paths between both IoT devices 4704 and gateway 4754, including the use of multiplexing / demultiplexing (MUXing / deMUXing) components to facilitate interconnection of various devices.
[0112] The network topology can include any number of IoT networks of various types, such as a mesh network provided using Bluetooth Low Energy (BLE) link 4722 along with network 4756. Other possible types of IoT networks include a wireless local area network (WLAN) 4758 for use via IEEE 802.11. Link 4728 communicates with IoT device 4704; cellular network 4760 is used to communicate with IoT device 4704 via LTE / LTE-A (4G) or 5G cellular network; and low-power wide-area (LPWA) network 4762, such as an LPWA network compatible with the LoRaWan specification issued by the LoRa Alliance, or IPv6 on a low-power wide-area (LPWAN) network compatible with the specifications issued by the Internet Engineering Task Force (IETF). Furthermore, the corresponding IoT network can use any number of communication links to communicate with external network providers (e.g., Layer 2 or Layer 3 providers), such as LTE cellular links, LPWA links, or links based on IEEE 802.15.4 standards (such as...). The corresponding IoT network can also operate using various network and Internet application protocols, such as CoAP (Co-Restricted Application Protocol). The corresponding IoT network can also be integrated with a coordinator device that provides the link chains that form a cluster tree of linked devices and networks.
[0113] Each of these IoT networks offers opportunities for new technological features, such as those described in this paper. Improved technologies and networks can enable exponential growth in devices and networks, including the use of IoT networks as fog devices or systems. As the use of these improved technologies increases, IoT networks can be developed for self-management, functional evolution, and collaboration without direct human intervention. Improved technologies can even enable IoT networks to function without a central control system. Therefore, the improved technologies described in this paper can be used to automate and enhance network management and operational capabilities far beyond current implementations.
[0114] In the example, communication between IoT devices 4704, such as communication via backbone link 4702, can be protected by a decentralized system for authentication, authorization, and accounting (AAA). In a decentralized AAA system, distributed payment, lending, auditing, authorization, and authentication systems can be implemented across interconnected heterogeneous network infrastructures. This allows systems and networks to move towards autonomous operation. In these types of autonomous operation, systems can even contract with human resources and negotiate partnerships with other machine networks. This can allow for the achievement of common goals and balanced service delivery relative to outlined, planned service level agreements, as well as solutions that provide metering, measurement, traceability, and trackability. Creating new supply chain structures and methods can create, extract value from, and break down large numbers of services without human involvement.
[0115] This IoT network can be further enhanced by integrating sensing technologies such as sound, light, electronic traffic, facial and pattern recognition, odor, and vibration into the autonomous organization between IoT devices. The integration of sensing systems allows for autonomous communication and coordination of service delivery relative to contractual service objectives, based on service orchestration and quality of service (QoS) resource aggregation and fusion. Some examples of network-based resource processing include the following.
[0116] For example, a mesh network 4756 can be enhanced by a system that performs inline data-to-information transformation. For instance, a self-forming chain of processing resources, including a multi-link network, can efficiently allocate the transformation of raw data into information, and has the ability to differentiate between assets and resources, and the associated management of each asset and resource. Furthermore, appropriate components of the infrastructure and resource-based trust and service indexes can be inserted to improve data integrity, quality, assurance, and provide data confidence metrics.
[0117] For example, WLAN network 4758 can use a system that performs standard conversion to provide multi-standard connectivity, enabling IoT devices 4704 to communicate using different protocols. Other systems can provide seamless interconnection across multi-standard infrastructures, including both visible and hidden Internet resources.
[0118] For example, communication in cellular network 4760 can be enhanced by migrating data, extending communication to more remote devices, or migrating data and extending communication to more remote devices. LPWA network 4762 may include systems performing non-IP to IP interconnection, addressing, and routing. Furthermore, each IoT device 4704 may include a suitable transceiver for wide-area communication. Further, each IoT device 4704 may include additional transceivers using additional protocols and frequencies for communication. Regarding... Figure 49 and Figure 50 The communication environment and hardware of the IoT processing device described in the paper further discuss this point.
[0119] Finally, IoT device clusters can be configured to communicate with other IoT devices and cloud networks. This allows IoT devices to form ad-hoc networks, enabling them to function as individual devices, which can be referred to as fog devices. The following is about... Figure 48 The configuration was discussed further.
[0120] Figure 48A cloud computing network is shown communicating with a mesh network of IoT devices (device 4802) operating as fog devices at the edge of a cloud computing network. This mesh network of IoT devices, which may be referred to as fog 4820, operates at the edge of cloud 4800. For simplicity, not every IoT device 4802 is labeled.
[0121] Fog 4820 can be viewed as a large-scale interconnected network where several IoT devices 4802 communicate with each other, for example, via radio lines 4822. As an example, the Open Connectivity Foundation (OPF) can be used. TM The OCF (Optimized Link-State Routing) has released an interconnection specification to facilitate this interconnection network. This standard allows devices to discover each other and establish communication for interconnection. Other interconnection protocols can also be used, including, for example, Optimized Link-State Routing (OLSR), Best Mobile Ad Hoc Networks (BATMAN) routing protocol, or OMA Lightweight M2M (LWM2M) protocol.
[0122] Although three types of IoT devices 4802 are shown in this example—gateway 4804, data aggregator 4826, and sensor 4828—any combination of IoT devices 4802 and their functionalities can be used. Gateway 4804 can be an edge device providing communication between cloud 4800 and fog 4820, and can also provide back-end processing capabilities for data acquired from sensor 4828, such as motion data, flow data, temperature data, etc. Data aggregator 4826 can collect data from any number of sensors 4828 and perform back-end processing functions for analysis. Resulting data, raw data, or both can be transmitted to cloud 4800 via gateway 4804. Sensor 4828 can be a complete IoT device 4802, for example, capable of both collecting and processing data. In some cases, sensor 4828 can be more functionally limited, for example, collecting data and allowing data aggregator 4826 or gateway 4804 to process it.
[0123] Communication from any IoT device 4802 can be transmitted to the gateway 4804 via a convenient path (e.g., the most convenient path) between any IoT devices 4802. The number of interconnections in these networks provides significant redundancy, allowing communication to be maintained even if several IoT devices 4802 are lost. Furthermore, using a mesh network allows for the use of very low-power IoT devices 4802 or those located far from the infrastructure, as the distance to connect to another IoT device 4802 can be much smaller than the distance to connect to the gateway 4804.
[0124] The fog 4820 provided by these IoT devices 4802 can be presented to devices in the cloud 4800 (such as server 4806) as individual devices located at the edge of the cloud 4800, such as fog devices. In this example, an alarm from the fog device can be sent without being identified as originating from a specific IoT device 4802 within the fog 4820. In this way, the fog 4820 can be considered a distributed platform that provides computing and storage resources to perform processing or data-intensive tasks, such as data analysis, data aggregation, and machine learning.
[0125] In some examples, an imperative programming style can be used to configure IoT devices 4802, where each IoT device 4802 has specific functionalities and communication partners. However, the IoT devices 4802 forming the fog device can be configured in a declarative programming style, which allows the IoT devices 4802 to reconfigure their operation and communication, such as determining the required resources in response to conditions, queries, and device failures. As an example, a query from a user located on server 4806 related to the operation of a subset of devices monitored by IoT devices 4802 may cause the fog device 4820 to select the IoT devices 4802 (such as specific sensors 4828) required to answer the query. Data from these sensors 4828 can then be aggregated and analyzed by any combination of sensors 4828, data aggregator 4826, or gateway 4804 before being sent by the fog device 4828 to server 4806 to answer the query. In this example, the IoT devices 4802 in fog 4820 can select the sensors 4828 to use based on a query (such as adding data from a flow sensor or temperature sensor). Furthermore, if some of the IoT devices in IoT device 4802 are inoperable, other IoT devices 4802 in fog device 4820 can provide similar data (if available).
[0126] In other examples, the operations and functions described above can be embodied by an IoT device machine, exemplified as an electronic processing system, which, according to an example embodiment, can execute a set or sequence of instructions to cause the electronic processing system to perform any of the methods discussed herein. The machine can be an IoT device or IoT gateway, including aspects embodied by a personal computer (PC), tablet PC, personal digital assistant (PDA), mobile phone, or smartphone, or any machine capable of (sequentially or otherwise) executing instructions specifying the actions to be taken by that machine. Furthermore, while only a single machine is described and referenced in the examples above, such machines should also be considered as any collection of machines that individually or jointly execute a set (or more) of instructions to perform any one or more of the methods discussed herein. Moreover, these examples and examples similar to processor-based systems should be considered as any collection of machines controlled or operated by a processor (e.g., a computer) to individually or jointly execute instructions to perform any one or more of the methodologies discussed herein. In some implementations, one or more devices can operate collaboratively to achieve functionality and perform the tasks described herein. In some cases, one or more host devices can supply data, provide instructions, aggregate results, or otherwise facilitate joint operations and functions provided by multiple devices. When a function is implemented by a single device, it can be considered a device-local function; in implementations where multiple devices operate as a single machine, the function can be considered device-commonly local, and the set of devices can provide or consume results provided by other remote machines (implemented as a single device or a set of devices), among other example implementations.
[0127] For example, Figure 49A diagram illustrates a cloud computing network or cloud 4900 communicating with multiple Internet of Things (IoT) devices. Cloud 4900 can represent the Internet, or it can be a local area network (LAN) or a wide area network (WAN), such as a company's private network. IoT devices can include any number of different types of devices, grouped in various combinations. For example, traffic control group 4906 can include IoT devices along streets in a city. These IoT devices can include traffic lights, traffic flow monitors, cameras, weather sensors, etc. Traffic control group 4906 or other subgroups can communicate with cloud 4908 via wired or wireless links 4900 (such as LPWA links, optical links, etc.). Furthermore, wired or wireless subnetwork 4912 can allow IoT devices (such as via LAN, WLAN, etc.) to communicate with each other. IoT devices can use another device (such as gateway 4910 or gateway 4928) to communicate with remote locations (such as cloud 4900); IoT devices can also use one or more servers 4930 to facilitate communication with cloud 4900 or gateway 4910. For example, one or more servers 4930 can operate as intermediate network nodes to support local edge cloud or fog implementations between local area networks. Furthermore, the depicted gateway 4928 can operate in a cloud-to-gateway-to-many edge device configuration (such as various IoT devices 4914, 4920, 4924, where the allocation and use of resources in the cloud 4900 is restricted or dynamic).
[0128] Other example groups of IoT devices may include a remote weather station 4914, a local information terminal 4916, an alarm system 4918, an ATM 4920, an alarm panel 4922, or mobile vehicles (such as an emergency vehicle 4924 or other vehicles 4926, etc.). Each of these IoT devices can interact with other IoT devices, with a server 4904, with another IoT fog device or system (not shown, but...). Figure 48 (as depicted in the image), or a combination thereof, communicating. This group of IoT devices can be deployed in a variety of residential, commercial, and industrial settings (including both private and public environments).
[0129] As from Figure 49As can be seen, a large number of IoT devices can communicate through the cloud 4900. This allows different IoT devices to autonomously request or provide information to other devices. For example, a group of IoT devices (e.g., traffic control group 4906) can request the current weather forecast from a group of remote weather stations 4914, which can provide forecasts without human intervention. Furthermore, an emergency vehicle 4924 can be alerted by an ATM 4920 that theft is in progress. As the emergency vehicle 4924 moves toward the ATM 4920, it can access the traffic control group 4906 to request permission to proceed to that location, for example, by turning a traffic light red to block cross traffic at an intersection for a sufficient time to allow the emergency vehicle 4924 to enter the intersection unimpeded.
[0130] IoT device clusters (such as remote weather station 4914 or traffic control group 4906) can be configured to communicate with other IoT devices and the cloud 4900. This allows IoT devices to form self-organizing networks among themselves, allowing them to function as individual devices, which can be referred to as fog devices or systems (e.g., as referenced above). Figure 48 The above).
[0131] Figure 50 This is a block diagram illustrating examples of components that may exist in an IoT device 5050 for implementing the techniques described herein. The IoT device 5050 may include any combination of the components shown in the examples or those referenced in the foregoing disclosure. These components may be implemented as ICs, portions of ICs, discrete electronic devices or other modules, logic, hardware, software, firmware, or combinations thereof suitable for use in the IoT device 5050, or implemented as components further incorporated into the chassis of a larger system. Additionally, Figure 50 The block diagram is intended to depict a high-level view of the components of an IoT device 5050. However, some of the components shown may be omitted, additional components may be present, and different arrangements of the components shown may occur in other implementations.
[0132] The IoT device 5050 may include a processor 5052, which may be a microprocessor, multi-core processor, multi-threaded processor, ultra-low voltage processor, embedded processor, or other known processing element. The processor 5052 may be part of a system-on-a-chip (SoC), where the processor 5052 and other components are formed as a single integrated circuit or a single package (such as Intel's Edison). TM ) or Galileo TM (SoC board). As an example, the processor 5052 may include an Intel architecture-based core ( Architecture Core TM processors such as Quark TM Atom TM i3, i5, i7, or MCU-level processors, or those from Intel Corporation in Santa Clara, California. Another such processor may be obtained from Apple Inc. However, any number of other processors may be used, such as processors from Advanced Micro Devices, Inc. (AMD) of Sunnyvale, California; processors based on MIPS designs from MIPS Technologies, Inc. (MIPS) of Sunnyvale, California; processors based on ARM designs licensed from ARM Holdings, Ltd. or its customers, or its licensees or adopters. Processors may include those from Apple Inc. The A5-A10 processors from Qualcomm Incorporated are from Qualcomm Technologies, Inc. Snapdragon by Technologies, Inc. TM The processor is either an OMAP from Texas Instruments, Inc. TM The processor unit.
[0133] Processor 5052 can communicate with system memory 5054 via interconnect 5056 (e.g., a bus). Any number of memory devices can be used to provide a fixed amount of system memory. As an example, the memory can be random access memory (RAM) designed according to the Joint Electron Devices Engineering Council (JEDEC) standards (such as DDR or mobile DDR standards (e.g., LPDDR, LPDDR2, LPDDR3, or LPDDR4)). In various implementations, the individual memory devices can be any number of different package types, such as single-die package (SDP), dual-die package (DDP), or quad-die package (Q17P). In some examples, these devices can be directly soldered to the motherboard to provide a thin solution, while in other examples, the device is configured as one or more memory modules that are coupled to the motherboard via a given connector. Any number of other memory implementations can be used, such as other types of memory modules, for example, different kinds of dual in-line memory modules (DIMMs), including but not limited to microDIMMs or MiniDIMMs.
[0134] To provide persistent storage for information such as data, applications, and operating systems, storage 5058 can also be coupled to processor 5052 via interconnect 5056. In the example, storage 5058 can be implemented via a solid-state drive (SSDD). Other devices that can be used for storage 5058 include flash memory cards (such as SD cards, microSD cards, xD graphics cards, etc.) and USB flash drives. In a low-power implementation, storage 5058 can be on-die memory or on-die registers associated with processor 5052. However, in some examples, storage 5058 can be implemented using a micro hard disk drive (HDD). Furthermore, in addition to, or in lieu of, the described technologies, any number of new technologies can be used for storage 5058, such as resistance-changing memory, phase-change memory, holographic memory, or chemical memory.
[0135] Components can communicate via interconnect 5056. Interconnect 5056 can include any number of technologies, including Industry Standard Architecture (ISA), Extended ISA (EISA), Peripheral Component Interconnect (PCI), Extended Peripheral Component Interconnect (PCIx), Fast PCI (PCIe), or any other technologies. Interconnect 5056 can be a dedicated bus, such as a dedicated bus used in a SoC-based system. Other bus systems can be included, such as I2C interfaces, SPI interfaces, point-to-point interfaces, and power buses.
[0136] Interconnect 5056 couples processor 5052 to mesh transceiver 5062 for communication with other mesh devices 5064. Mesh transceiver 5062 can use any number of frequencies and protocols (such as 2.4 GHz under the IEEE 802.15.4 standard) for transmission, using the Bluetooth Special Interest Group (SIG) protocol. The Bluetooth Low Energy (BLE) standard defined by the Special Interest Group, or (Standards, etc.). Any number of radios configured for a specific wireless communication protocol can be used for connectivity to the mesh device 5064. For example, a WLAN unit can be used to implement Wi-Fi according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard. TM Communication. Additionally, (for example, according to cellular or other wireless wide area protocols) wireless wide area communication can occur via a WWAN unit.
[0137] The mesh transceiver 5062 can communicate using multiple standards or radios for communication at different ranges. For example, the IoT device 5050 can use a BLE-based local transceiver or another low-power radio to communicate with nearby devices (e.g., within approximately 10 meters) to save power. More distant mesh devices 5064 (e.g., within approximately 50 meters) can be reached via ZigBee or other intermediate-power radios. The two communication technologies can occur at different power levels on a single radio, or they can occur on separate transceivers, such as a local transceiver using BLE and a separate mesh transceiver using ZigBee.
[0138] The device may include a wireless transceiver 5066 to communicate with devices or services in the cloud 5000 via local area network (LAN) or wide area network (WAN) protocols. The wireless transceiver 5066 may be an LPWA transceiver conforming to standards such as IEEE 802.15.4 or IEEE 802.15.4g. The IoT device 5050 may use LoRaWAN, developed by Semtech and the LoRa Alliance. TM (Long-Range Wide Area Network) enables communication over a wide area. The techniques described herein are not limited to these and can be used with any number of other cloud transceivers (such as Sigfox and other technologies) to achieve long-distance, low-bandwidth communication. Furthermore, other communication techniques described in the IEEE 802.15.4e specification, such as time-slotted channel hopping, can be used.
[0139] In addition to the systems mentioned for mesh transceiver 5062 and wireless network transceiver 5066, any number of other radio communications and protocols, as described herein, can be used. For example, radio transceivers 5062 and 5066 may include LTE or other cellular transceivers using spread spectrum (SPA / SAS) communication to achieve high-speed communication. Furthermore, any number of other protocols (such as...) can be used. (Network) is used for medium-speed communication and to provide network communication.
[0140] Radio transceivers 5062 and 5066 may include radios compatible with any number of 3GPP (Third Generation Partnership Project) specifications, particularly Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), and Long Term Evolution-Advanced Pro (LTE-A Pro). It can be noted that radios compatible with any number of other fixed communication technologies and standards, mobile communication technologies and standards, or satellite communication technologies and standards can be selected. These can include, for example, any cellular wide-area radio communication technology, such as fifth-generation (5G) communication systems, Global System for Mobile Communications (GSM) wireless communication technology, General Packet Radio Service (GPRS) wireless communication technology, or GSM Evolution Enhanced Data Rate (EDGE) wireless communication technology, Universal Mobile Telecommunications System (UMTS) communication technology. In addition to the standards mentioned above, any number of satellite uplink technologies can be used with the 5066 wireless network transceiver, including radios conforming to standards published by, for example, the International Telecommunication Union (ITU) or the European Telecommunications Standards Institute (ETSI). The examples provided herein are therefore to be understood as applicable to a variety of other existing and undeveloped communication technologies.
[0141] A network interface controller (NIC) 5068 may be included to provide wired communication to the cloud 5000 or other devices, such as the mesh device 5064. Wired communication may provide an Ethernet connection or may be based on other types of networks, such as a controller area network (CAN), local interconnect network (LIN), device network, control network, data highway+, PROFIBUS, or PROFINET. Additional NICs 5068 may be included to allow connection to a second network; for example, one NIC 5068 provides communication to the cloud via Ethernet, and a second NIC 5068 provides communication to other devices via another type of network.
[0142] Interconnect 5056 couples processor 5052 to external interface 5070 for connecting external devices or subsystems. External devices may include sensors 5072, such as accelerometers, level sensors, flow sensors, optical light sensors, camera sensors, temperature sensors, GPS sensors, pressure sensors, barometric pressure sensors, etc. External interface 5070 can further be used to connect IoT devices 5050 to actuators 5074, such as power switches, valve actuators, audible sound generators, visual warning devices, etc.
[0143] In some alternative examples, various input / output (I / O) devices may be present within or connected to the IoT device 5050. For example, a display or other output device 5084 may be included to display information such as sensor readings or actuator positions. Input devices 5086, such as touchscreens or keypads, may be included to accept input. Output devices 5084 may include any number of various forms of audio or visual displays, including simple visual outputs (such as binary status indicators (e.g., LEDs) and multi-character visual outputs), or more complex outputs (such as displays (e.g., LCD screens)) where the output of characters, graphics, multimedia objects, etc., is generated or produced by the operation of the IoT device 5050.
[0144] Battery 5076 can power IoT device 5050, although in an example where IoT device 5050 is installed in a fixed location, it can have a power source coupled to the grid. Battery 5076 can be a lithium-ion battery or a metal-air battery (such as zinc-air batteries, aluminum-air batteries, lithium-air batteries, etc.).
[0145] The IoT device 5050 may include a battery monitor / charger 5078 to track the state of charge (SoCh) of the battery 5076. The battery monitor / charger 5078 can be used to monitor other parameters of the battery 5076, such as the state of health (SoH) and state of function (SoF) of the battery 5076, to provide failure prediction. The battery monitor / charger 5078 may include a battery monitoring integrated circuit, such as the LTC4020 or LTC2990 from Linear Technologies, the ADT7488A from ON Semiconductor in Phoenix, Arizona, or the UCD90xxx series IC from Texas Instruments in Dallas, Texas. The battery monitor / charger 5078 can transmit information about the battery 5076 to the processor 5052 via interconnect 5076. The battery monitor / charger 5078 may also include an analog-to-digital converter (ADC) that allows the processor 5052 to directly monitor the voltage of the battery 5076 or the current from the battery 5076. Battery parameters can be used to determine actions that the IoT device 5050 can perform, such as transmission frequency, mesh network operation, sensing frequency, etc.
[0146] Power block 5080 or another power source coupled to the grid can be coupled to battery monitor / charger 5078 to charge battery 5076. In some examples, power block 5080 can be replaced with a wireless power receiver to wirelessly obtain power (e.g., via a loop antenna in IoT device 5050). Battery monitor / charger 5078 may include wireless battery charging circuitry (such as the LTC4020 chip from Linear Technologies in Milpitas, California). The specific charging circuitry chosen depends on the size of battery 5076 and, therefore, on the required current. Charging can be performed using the Airfuel Alliance standard, the Qi wireless charging standard from the Wireless Power Consortium, or the Rezence charging standard from the Alliance for Wireless Power, etc.
[0147] Storage 5058 may include instructions 3582 in the form of software, firmware, or hardware commands to implement the techniques described herein. Although these instructions 5082 are shown as blocks of code included in memory 5054 and storage 5058, it is understood that any block of code may be replaced by hard-wired circuitry, for example, embedded in an application-specific integrated circuit (ASIC).
[0148] In the example, instructions 5082 provided via memory 5084, storage 5058, or processor 5052 can be embodied in a non-transient, machine-readable medium 5060, which includes code for directing processor 5052 to perform electronic operations within IoT device 5050. Processor 5052 can access the non-transient, machine-readable medium 5060 via interconnect 5056. For example, the non-transient, machine-readable medium 5060 can be provided by… Figure 50 The storage 5058 described herein may embody a device or may include a specific storage unit (such as an optical disc, a flash drive, or any number of other hardware devices). The non-transient, machine-readable medium 5060 may include instructions that direct the processor 5052 to perform a specific sequence of actions or processes, for example, as described in the flowcharts and block diagrams(s) above regarding the operation and function.
[0149] Figure 51 This is an example illustration of a processor according to an embodiment. Processor 5100 is an example of a hardware device that can be used in conjunction with the above implementation. Processor 5100 can be any type of processor, such as a microprocessor, embedded processor, digital signal processor (DSP), network processor, multi-core processor, single-core processor, or other device for executing code. Although Figure 51 The demonstration showed only one processor 5100, but the processing element can alternatively include more than one. Figure 51 The processor 5100 shown in the figure. The processor 5100 may be a single-threaded core, or for at least one embodiment, the processor 5100 may be multi-threaded, as the processor may include multiple hardware thread contexts (or "logical processors") per core.
[0150] Figure 51 Also shown is a memory 5102 coupled to processor 5100 according to an embodiment. Memory 5102 can be any of a wide variety of memories (including different layers of memory hierarchy) as known to those skilled in the art or otherwise available. Such memory elements can include, but are not limited to, random access memory (RAM), read-only memory (ROM), logic blocks of field-programmable gate arrays (FPGAs), erasable programmable read-only memory (EPROM), and electrically erasable programmable ROM (EEPROM).
[0151] Processor 5100 can execute any type of instructions associated with the algorithms, processes, or operations described in detail herein. Generally, processor 5100 is capable of transforming elements or items (e.g., data) from one state or thing to another.
[0152] Code 5104 (which may be one or more instructions to be executed by processor 5100) may be stored in memory 5102, or it may be stored in software, hardware, firmware, or any suitable combination thereof, or in any other internal or external component, device, element, or suitable, demand-specific object. In one example, processor 5100 may follow a program sequence of instructions indicated by code 5104. Each instruction enters front-end logic 5106 and is processed by one or more decoders 5108. The decoders may generate micro-operations, such as fixed-width micro-operations of a predetermined format, as their output, or may generate other instructions, micro-instructions, or control signals that reflect the original code instructions. Front-end logic 5106 also includes register renaming logic 5110 and scheduling logic 5112, which typically allocate resources and queue operations corresponding to the instructions to be executed.
[0153] The processor 5100 may also include execution logic 5114 having a set of execution units 5116a, 5116b, 5116n, etc. Some embodiments may include several execution units dedicated to a specific function or group of functions. Other embodiments may include only one execution unit or a single execution unit capable of performing a specific function. Execution logic 5114 performs the operations specified by code instructions.
[0154] After the operation specified by the code instructions has been executed, the back-end logic 5118 may decommission the instructions of code 5104. In one embodiment, processor 5100 allows out-of-order execution but requires ordered instruction decommissioning. Decommissioning logic 5120 may take various known forms (e.g., reordering buffers, etc.). In this way, processor 5114 can be switched during the execution of code 5104, at least in terms of the output generated by the decoder, the hardware registers and tables utilized by register renaming logic 5100, and any registers (not shown) modified by execution logic 5110.
[0155] although Figure 51 Not shown, but the processing element may include other elements on the chip having processor 5100. For example, the processing element may include memory control logic along with processor 5100. The processing element may include I / O control logic, and / or may include I / O control logic integrated with the memory control logic. The processing element may also include one or more caches. In some embodiments, non-volatile memory (such as flash memory or fuses) may also be included on the chip having processor 5100.
[0156] Figure 52 A computing system 5200 configured as a point-to-point (PtP) according to an embodiment is shown. Specifically, Figure 52A system is illustrated in which a processor, memory, and input / output devices are interconnected via a number of point-to-point interfaces. Generally, one or more computing systems described herein can be configured in the same or similar manner as computing system 5200.
[0157] Processors 5270 and 5280 may also each include integrated memory controller logic (MC) 5272 and 5282 for communicating with memory elements 5232 and 5234. In alternative embodiments, the memory controller logic 5272 and 5282 may be discrete logic separate from processors 5270 and 5280. Memory elements 5232 and / or 5234 may store various data to be used by processors 5270 and 5280 for acquiring operations and functions outlined herein.
[0158] Processors 5270 and 5280 can be any type of processor, such as those discussed in connection with the other accompanying figures. Processors 5270 and 5280 can exchange data via point-to-point (PtP) interface 5250 using point-to-point (PtP) interface circuits 5278 and 5288, respectively. Processors 5270 and 5280 can each exchange data with chipset 5290 via separate point-to-point interfaces 5252 and 5254 using point-to-point interface circuits 5276, 5286, 5294, and 5298, respectively. Chipset 5290 can also exchange data with high-performance graphics circuit 5238 via high-performance graphics interface 5239 using interface circuit 5292 (which may be a PtP interface circuit). In an alternative embodiment, [the following can be omitted]. Figure 52 Any or all PtP links shown are implemented as multi-station buses rather than having PtP links.
[0159] Chipset 5290 can communicate with bus 5220 via interface circuitry 5296. Bus 5220 may have one or more devices communicating through it, such as bus bridge 5218 and I / O device 5216. Bus bridge 5218 can communicate with other devices via bus 5210, such as user interface 5212 (e.g., keyboard, mouse, touchscreen, or other input device), communication device 5226 (e.g., modem, network interface device, or other type of communication device that can communicate via computer network 5260), audio I / O device 5214, and / or data storage device 5228. Data storage device 5228 can store code 5230, which can be executed by processor 5270 and / or 5280. In alternative embodiments, any part of the bus architecture can be implemented using one or more PtP links.
[0160] Figure 52The computer system depicted herein is a schematic illustration of an embodiment of a computing system that can be used to implement the various embodiments discussed herein. It will be understood that... Figure 52 The various components described herein can be combined in a system-on-a-chip (SoC) architecture capable of achieving the functionality and features of the examples and implementations provided herein, or in any other suitable configuration.
[0161] In a further example, a machine-readable medium also includes any tangible medium capable of storing, encoding, or carrying instructions executable by the machine and causing the machine to perform any or more of the methods of this disclosure, or capable of storing, encoding, or carrying data structures utilized by or associated with such instructions. Thus, "machine-readable medium" can include, but is not limited to, solid-state memory, as well as optical and magnetic media. Specific examples of machine-readable media include non-volatile memory, including, but not limited to, semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; disks (such as internal hard disks and removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Instructions embodied in the machine-readable medium can be further transmitted or received over a communication network via a network interface device using any of several transport protocols (e.g., HTTP).
[0162] It should be understood that the functional units or capabilities described in this specification may be referred to or labeled as “components” or “modules” to more specifically emphasize their implementation independence. These components may be embodied in any number of software or hardware forms. For example, a component or module may be implemented as hardware circuitry including custom-designed very large-scale integration (VLSI) circuitry or gate arrays, or off-the-shelf semiconductors (e.g., logic chips, transistors, or other discrete components). Components or modules may also be implemented in programmable hardware devices such as field-programmable gate arrays, programmable array logic, or programmable logic devices. Components or modules may also be implemented in software executed by various types of processors. For example, the identified components or modules of executable code may include one or more physical or logical blocks of computer instructions, such as those organized as objects, processes, or functions. However, the executable files of the identified components or modules do not need to be physically located together, but may include different instructions stored in different locations that, when logically linked together, include the components or modules and achieve the specified purpose of the components or modules.
[0163] In practice, executable code components or modules can be single instructions or many instructions, and can even be distributed across several different code segments, different programs, and across several memory devices or processing systems. In particular, some aspects of the described processes (such as code rewriting and code analysis) can occur on a different processing system (e.g., a computer in a data center) than the processing system on which the code is deployed (e.g., in a computer embedded in a sensor or robot). Similarly, operational data described herein can be identified and described within components or modules, and can be embodied in any suitable form and organized within any suitable type of data structure. Components or modules can be passive or active, including agents operable to perform desired functions.
[0164] Additional examples of embodiments of the methods, systems, and devices described herein include the following non-limiting configurations. Each of the following non-limiting examples may exist independently, or may be combined in any permutation or combination with any one or more other examples provided below or throughout this disclosure.
[0165] Although this disclosure has been described in terms of certain implementations and generally associated methods, modifications and substitutions to these implementations and methods will be apparent to those skilled in the art. For example, the actions described herein may be performed in a different order than those described, and still achieve the desired result. As an example, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous. Additionally, other user interface layouts and functionalities may be supported. Other variations are within the scope of the following claims.
[0166] Although this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of the particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Moreover, while features may be described above as functioning in certain combinations and even so initially claimed, one or more features from a claimed combination may, in some cases, be separate from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0167] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring that such operations be performed in the specific order shown or in an ordered sequence, or that all the operations shown can be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Moreover, the separation of different system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0168] The following examples relate to embodiments according to this specification. Example 1 is a method comprising: accessing a synthetic three-dimensional (3D) graphical model of an object from memory, wherein the 3D graphical model has a photorealistic resolution; generating a plurality of different training samples from views of the 3D graphical model, wherein the plurality of training samples are generated to add defects to the plurality of training samples to simulate the characteristics of real-world samples generated by a real-world sensor device; and generating a training set including the plurality of training samples, wherein the training data is for training an artificial neural network.
[0169] Example 2 includes the subject of Example 1, wherein the plurality of training samples include digital images and the sensor device includes a camera sensor.
[0170] Example 3 includes a topic from any of Examples 1-2, wherein the plurality of training samples include a point cloud representation of the object.
[0171] Example 4 includes the subject of Example 3, wherein the sensor includes a LIDAR sensor.
[0172] Example 5 includes the subject matter of any one of Examples 1-4, and further includes: accessing data to indicate parameters of the sensor device; and determining defects to be added to the plurality of training samples based on the parameters.
[0173] Example 6 includes the subject of Example 5, wherein the data includes a model of the sensor device.
[0174] Example 7 includes the subject of any one of Examples 1-6, and further includes: accessing data to indicate the characteristics of one or more surfaces of the object modeled by the 3D graphics model; and determining, based on the characteristics, defects to be added to the plurality of training samples.
[0175] Example 8 includes the subject of Example 7, wherein the 3D graphical model includes the data.
[0176] Example 9 includes the subject of any one of Examples 1-8, wherein the defect includes one or more of noise or glare.
[0177] Example 10 includes the subject of any one of Examples 1-9, wherein generating the plurality of different training samples includes: applying different lighting settings to the 3D graphics model to simulate lighting in an environment; determining the defects of a subset of the plurality of training samples generated during a particular period of applying the different lighting settings, wherein the defects of the subset of the plurality of training samples are based on the specific lighting settings.
[0178] Example 11 includes the subject of any one of Examples 1-10, wherein generating the plurality of different training samples includes: placing the 3D graphics model in different graphics environments, wherein the graphics environments model their respective real-world environments; and generating a subset of the plurality of training samples when the 3D graphics model is placed in the different graphics environments.
[0179] Example 12 is a system comprising means for performing the method as described in any one of claims 1-11.
[0180] Example 13 includes the subject matter of Example 12, wherein the system includes an apparatus comprising hardware circuitry for performing at least a portion of the method of any one of claims 1-11.
[0181] Example 14 is a computer-readable storage medium for storing processor-executable instructions for performing the method as described in any one of claims 1-11.
[0182] Example 15 is a method comprising: receiving a subject input and a reference input at a Siamese neural network, wherein the Siamese neural network includes a first network portion and a second network portion, the first network portion including a first plurality of layers and the second network portion including a second plurality of layers, the weights of the first network portion being the same as the weights of the second network portion, the subject input being provided as input to the first network portion, and the reference input being provided as input to the second network portion; and generating an output of the Siamese neural network based on the subject input and the reference input, wherein the output of the Siamese neural network is used to indicate the similarity between the reference input and the subject input.
[0183] Example 16 includes the subject of Example 15, wherein generating the output includes: determining a difference amount between the reference input and the subject input; and determining whether the difference amount satisfies a threshold, wherein the output identifies whether the difference amount satisfies the threshold.
[0184] Example 17 includes the subject of Example 16, wherein determining the amount of difference between the reference input and the subject input includes: receiving a first feature vector output by the first network portion and a second feature vector output by the second network portion; and determining a difference vector based on the first feature vector and the second feature vector.
[0185] Example 18 includes the topic of any one of Examples 15-17, wherein generating the output includes a one-time classification.
[0186] Example 19 includes the subject of any one of Examples 15-18, and further includes training the Siamese neural network using one or more synthetic training samples.
[0187] Example 20 includes the subject of Example 19, wherein the method of any one of claims 1-11 generates the one or more synthetic training samples.
[0188] Example 21 includes the subject of any one of Examples 15-20, wherein the reference input includes synthetically generated samples.
[0189] Example 22 includes the subject of Example 21, wherein the synthesized sample is generated by the method according to any one of claims 1-11.
[0190] Example 23 includes the subject of any one of Examples 15-22, wherein the subject input includes a first digital image and the reference input includes a second digital image.
[0191] Example 24 includes the subject of any one of Examples 15-22, wherein the subject input includes a first point cloud representation and the reference input includes a second point cloud representation.
[0192] Example 25 is a system comprising means for performing the method as described in any one of claims 15-24.
[0193] Example 26 includes the subject matter of Example 25, wherein the system includes an apparatus comprising hardware circuitry for performing at least a portion of the method of any one of claims 15-24.
[0194] Example 27 includes the subject of Example 25, wherein the system includes one of a robot, a drone, or an autonomous vehicle.
[0195] Example 28 is a computer-readable storage medium storing instructions that can be executed by a processor to perform the method as described in any one of claims 15-24.
[0196] Example 29 is a method comprising: providing first input data to a Siamese neural network, wherein the first input data includes a first representation in 3D space from a first pose; providing second input data to the Siamese neural network, wherein the second input data includes a second representation in 3D space from a second pose, the Siamese neural network including a first network portion and a second network portion, the first network portion including a first plurality of layers and the second network portion including a second plurality of layers, the weights of the first network portion being the same as the weights of the second network portion, the first input data being provided as input to the first network portion, the second input data being provided as input to the second network portion; and generating an output of the Siamese neural network, wherein the output includes a relative pose between the first pose and the second pose.
[0197] Example 30 includes the subject of Example 29, wherein the first representation of the 3D space includes a first 3D point cloud, and the second representation of the 3D space includes a second 3D point cloud.
[0198] Example 31 includes the subject of any one of Examples 29-30, wherein the first representation of the 3D space includes a first point cloud, and the second representation of the 3D space includes a second point cloud.
[0199] Example 32 includes the subject of Example 31, wherein the first point cloud and the second point cloud each include a corresponding voxelized point cloud representation.
[0200] Example 33 includes the subject of any one of Examples 29-32, and further includes generating a 3D map of the 3D space from at least the first input data and the second input data based on the relative pose.
[0201] Example 34 includes the subject of any one of Examples 29-32, and further includes determining the position of an observer in the 3D space based on the relative pose.
[0202] Example 35 includes the subject of Example 34, wherein the observer includes an autonomous machine.
[0203] Example 36 includes the subject matter of Example 35, wherein the system includes one of a drone, a robot, or an autonomous vehicle.
[0204] Example 37 is a system comprising means for performing the method as described in any one of claims 29-36.
[0205] Example 38 includes the subject matter of Example 37, wherein the system includes an apparatus comprising hardware circuitry for performing at least a portion of the method of any one of claims 29-36.
[0206] Example 39 is a computer-readable storage medium storing instructions that can be executed by a processor to perform the method as described in any one of claims 29-36.
[0207] Example 40 is a method comprising: providing first sensor data as input to a first part of a machine learning model; providing second sensor data as input to a second part of the machine learning model, wherein the machine learning model includes a connector and a set of fully connected layers, the first sensor data being data of a first type generated by a device, and the second sensor data being data of a different second type generated by the device, wherein the connector takes the output of the first part of the machine learning model as a first input and takes the output of the second part of the machine learning model as a second input, and the output of the connector is provided to the set of fully connected layers; and generating an output of the machine learning model from the first data and the second data, including the pose of the device in an environment.
[0208] Example 41 includes the subject of Example 40, wherein the first sensor data includes image data, and the second sensor data identifies the movement of the device.
[0209] Example 42 includes the subject of Example 41, wherein the image data includes red, green and blue (RGB) data.
[0210] Example 43 includes the subject of Example 41, wherein the image data includes inertial measurement unit (IMU) data.
[0211] Example 44 includes the subject of Example 41, wherein the second sensor data includes inertial measurement unit (IMU) data.
[0212] Example 45 includes the subject of Example 41, wherein the second sensor data includes global positioning data.
[0213] Example 46 includes the subject of any one of Examples 40-45, wherein the first part of the machine learning model is adjusted for the first type of sensor data, and the second part of the machine learning model is adjusted for the second type of sensor data.
[0214] Example 47 includes the subject matter of any one of Examples 40-46, and further includes providing third-type third sensor data as input to a third part of the machine learning model, and further generating the output based on the third data.
[0215] Example 47 includes the subject of any one of Examples 40-46, wherein the output of the pose includes rotational and translational components.
[0216] Example 49 includes the subject of Example 48, wherein one fully connected layer in the set of fully connected layers includes a fully connected layer for determining the rotation component, and another fully connected layer in the set of fully connected layers includes a fully connected layer for determining the translation component.
[0217] Example 50 includes the subject of any one of Examples 40-49, wherein one or both of the first and second parts of the machine learning model include corresponding convolutional layers.
[0218] Example 51 includes the subject of any one of Examples 40-50, wherein one or both of the first and second portions of the machine learning model include one or more corresponding Long Short-Term Memory (LSTM) blocks.
[0219] Example 52 includes the subject of any one of Examples 40-51, wherein the device includes an autonomous machine and navigates the autonomous machine within the environment based on the gesture.
[0220] Example 53 includes the subject of Example 52, wherein the system includes one of a drone, a robot, or an autonomous vehicle.
[0221] Example 54 is a system comprising means for performing the method as described in any one of claims 40-52.
[0222] Example 55 includes the subject matter of Example 54, wherein the system includes an apparatus comprising hardware circuitry for performing at least a portion of the method of any one of claims 40-52.
[0223] Example 56 is a computer-readable storage medium storing instructions that can be executed by a processor to perform the method as described in any one of claims 40-52.
[0224] Example 57 is a method comprising: requesting the random generation of a set of neural networks; performing a machine learning task using each of the neural networks in the set, wherein the machine learning task is performed using specific processing hardware; monitoring attributes for each neural network in the set regarding the performance of the machine learning task, wherein the attributes include the accuracy of the result of the machine learning task; and identifying the best-performing neural network in the set based on the attributes of the best-performing neural network when the machine learning task is performed using the specific processing hardware.
[0225] Example 58 includes the subject of Example 57, and further includes providing the said best-performing neural network for use by the machine when performing machine learning applications.
[0226] Example 59 includes the subject matter of any one of Examples 57-58, and further includes: determining the performance of the best-performing neural network; generating a second set of neural networks based on the feature request, wherein the second set of neural networks includes a plurality of different neural networks, each neural network including one or more features; performing the machine learning task using each of the second set of neural networks, wherein the machine learning task is performed using the specific processing hardware; monitoring the attributes of performing the machine learning task for each neural network in the second set of neural networks; and identifying the best-performing neural network in the second set of neural networks based on the attributes.
[0227] Example 60 includes the subject of any one of Examples 57-59, and further includes a parameter-based acceptance criterion, wherein the best-performing neural network is based on the criterion.
[0228] Example 60 includes the subject of any one of Examples 57-60, wherein the attribute includes the attribute of the particular processing hardware.
[0229] Example 62 includes the subject matter of Example 61, wherein the attributes of the particular processing hardware include the power consumed by the particular processing hardware during the execution of the machine learning task, the temperature of the particular processing hardware during the execution of the machine learning task, and one or more memories for storing the neural network on the particular processing hardware.
[0230] Example 63 includes the subject of any one of Examples 57-62, wherein the attribute includes the time taken to complete the machine learning task using a corresponding one of the neural networks in the set.
[0231] Example 64 is a system comprising means for performing the method as described in any one of claims 57-63.
[0232] Example 65 includes the subject matter of Example 64, wherein the system includes an apparatus comprising hardware circuitry for performing at least a portion of the method of any one of claims 57-63.
[0233] Example 66 is a computer-readable storage medium storing instructions that can be executed by a processor to perform the method as described in any one of claims 57-63.
[0234] Example 67 is a method comprising: identifying a neural network comprising a plurality of convolutional kernels, wherein each of the convolutional kernels comprises a corresponding set of weights; pruning a subset of the plurality of convolutional kernels according to one or more parameters to reduce the plurality of convolutional kernels to a particular set of convolutional kernels; pruning a subset of weights in the particular set of convolutional kernels to form a pruned version of the neural network, wherein the pruning of the subset of weights assigns one or more non-zero weights in the subset of weights to zero, wherein the subset of weights is selected based on the original values of the weights.
[0235] Example 68 includes the subject of Example 67, wherein the subset of weights is pruned based on the values of the subset of weights that are below a threshold.
[0236] Example 69 includes the subject of any of Examples 67-68, and further includes performing one or more iterations of a machine learning task using a pruned version of the neural network to recover at least a portion of the accuracy lost due to the pruning of the convolutional kernels and weights.
[0237] Example 70 includes the subject of any one of Examples 67-69, and further includes quantifying the values of the unpruned weights in the pruned version of the neural network to generate a compact version of the neural network.
[0238] Example 71 includes the subject of Example 70, wherein the quantization includes log-base quantization.
[0239] Example 72 includes the subject of Example 71, wherein the weights are quantized from floating-point values to base-2 values.
[0240] Example 72 includes the subject of any one of Examples 67-72, and further includes providing the pruned version of the neural network to perform the machine learning task using hardware suitable for sparse matrix algorithms.
[0241] Example 74 is a system comprising means for performing the method as described in any one of claims 67-73.
[0242] Example 75 includes the subject matter of Example 64, wherein the system includes an apparatus comprising hardware circuitry for performing at least a portion of the method of any one of claims 67-73.
[0243] Example 76 is a computer-readable storage medium storing instructions that can be executed by a processor to perform the method as described in any one of claims 67-73.
[0244] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions listed in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result.
Claims
1. A calculation method, comprising: Generate one or more synthetic training samples, the synthetic training samples including synthetic images; One or more photorealistic training samples are generated by reducing the quality of the one or more synthetic training samples. A training set is formed using one or more photorealistic training samples. as well as The training set is used to train a neural network, wherein the neural network includes a first network and a second network, and after the training, the first network and the second network are used to process two different input images in order to evaluate a similarity measure between the two different input images.
2. The method of claim 1, wherein reducing the quality of the one or more synthetic training samples comprises: Apply a filter to the one or more synthetic training samples.
3. The method of claim 1, wherein the photorealistic training samples have a photorealistic resolution lower than the resolution of the synthetic training samples used to generate the photorealistic training samples.
4. The method of claim 1, wherein the quality of the one or more synthetic training samples includes brightness, noise level, or contrast.
5. The method of claim 1, wherein the first network and the second network have the same weights.
6. The method of claim 1, wherein the synthesized image comprises a three-dimensional model of an object, and reducing the quality of the one or more synthesized training samples comprises: The quality of the synthesized image is reduced based on one or more materials of the object.
7. The method of claim 1, wherein the first network is used to generate a first output from one of the two different images, the second network is used to generate a second output from the other of the two different images, and the neural network is used to determine the similarity measure based on the first output and the second output.
8. A non-transitory computer-readable medium storing one or more instructions, said instructions being executable to perform operations, said operations including: Generate one or more synthetic training samples, the synthetic training samples including synthetic images; One or more photorealistic training samples are generated by reducing the quality of the one or more synthetic training samples. A training set is formed using one or more photorealistic training samples. as well as The training set is used to train a neural network, wherein the neural network includes a first network and a second network, and after the training, the first network and the second network are used to process two different input images in order to evaluate a similarity measure between the two different input images.
9. One or more non-transient computer-readable media according to claim 8, wherein reducing the quality of the one or more synthetic training samples comprises: Apply a filter to the one or more synthetic training samples.
10. One or more non-transient computer-readable media according to claim 8, wherein the photorealistic training samples have a photorealistic resolution lower than the resolution of the synthetic training samples used to generate the photorealistic training samples.
11. One or more non-transient computer-readable media according to claim 8, wherein the quality of the one or more synthetic training samples includes brightness, noise level, or contrast.
12. One or more non-transient computer-readable media according to claim 8, wherein the first network and the second network have the same weights.
13. The one or more non-transient computer-readable media of claim 8, wherein the synthesized image comprises a three-dimensional model of an object, and reducing the quality of the one or more synthesized training samples comprises: The quality of the synthesized image is reduced based on one or more materials of the object.
14. The one or more non-transient computer-readable media of claim 8, wherein the first network is configured to generate a first output from one of the two different images, the second network is configured to generate a second output from the other of the two different images, and the neural network is configured to determine the similarity measure based on the first output and the second output.
15. A computing system, comprising: The memory stores instructions; as well as One or more processors, The instruction is responsive to execution by the one or more processors to cause the one or more processors to perform the method according to any one of claims 1 to 7.
16. A computing device comprising means for performing the method according to any one of claims 1-7.
17. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-7.