A COMPUTER-IMPLEMENTED METHOD AND NETWORK FOR DENSE PREDICTION

The iterative architecture for deep prediction, which cascades sub-backbones with bottom-up and top-down paths and utilizes lateral skip connections, addresses the inefficiencies in existing frameworks by improving multi-scale feature extraction and fusion, thereby enhancing performance without additional computational cost.

DE112022007710T5Pending Publication Date: 2025-06-12ROBERT BOSCH GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112022007710
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing deep prediction frameworks, such as Feature Pyramid Networks (FPNs), decouple feature learning into separate extraction and fusion stages, which may not be optimal for deep prediction tasks and can be insufficient under limited computational budgets.

Method used

The proposed solution involves an iterative architecture that splits a single backbone into cascaded sub-backbones with bottom-up sub-backbones for multi-scale feature extraction and top-down feedback paths for multi-scale feature fusion, utilizing lateral skip connections within and across sub-networks.

Benefits of technology

This approach enhances the performance of multi-scale processing for deep prediction tasks without increasing computational effort, maintaining similar model size and complexity as FPNs while improving performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to a computer-implemented network for dense prediction, the network comprising: a plurality of subnetworks, wherein the plurality of subnetworks are cascaded, and wherein each of the plurality of subnetworks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion; lateral skip connections between feature maps at corresponding feature levels in a bottom-up sub-backbone and in a top-down feedback path of a subnetwork; and lateral skip connections across adjacent subnetworks.
Need to check novelty before this filing date? Find Prior Art

Description

REGIONAspects of the present disclosure generally relate to artificial intelligence, and more particularly to a method and network for deep prediction.BACKGROUNDMulti-scale feature maps are essential for complex deep prediction tasks on images, including object recognition, instance segmentation, semantic segmentation, and the like. For example, the purpose of instance segmentation is to categorize and locate individual objects with pixel-by-pixel instance masks. Instance segmentation is widely used including autonomous driving, robotics, video surveillance, and the like. Learning of multi-scale feature representations is of great importance, since in high-performance instance segmentation, different numbers of instances must be recognized in a broad spectrum of scales and locations.A Feature Pyramid Network (FPN) is widely used in existing frameworks for multi-scale processing. An FPN generally utilizes an inherent feature hierarchy generated by a classification network to construct a feature pyramid that has strong semantics on all scales by fusing adjacent features from the inherent feature hierarchy through lateral links and a top-down path.However, an FPN and many related methods decouple the feature learning process into two consecutive parts, i.e., first, multi-scale feature maps are extracted through the classification network, and then these feature maps are fused by a lightweight module consisting of one or more paths. For deep prediction tasks, such a design may not be optimal. Moreover, the feature fusion process may not be sufficient for a limited computational budget, as most parameters and computational budget have been allocated to the classification network as backbone.Accordingly, it may be desirable to provide an improved architecture or technique to improve the performance of multi-scale processing for deep prediction.SUMMARYThe following provides a simplified summary of one or more aspects in accordance with the present disclosure to provide a basic understanding of such aspects. This summary is not a comprehensive overview of all aspects considered and is not intended to identify key or critical elements of all aspects, nor to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in simplified form as a prelude to the more detailed description presented below.In one aspect of the disclosure, a computer-implemented network for deep prediction is provided, the network comprising: a plurality of sub-networks, wherein the plurality of sub-networks are cascaded, and wherein each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion; lateral skip connections between feature maps at respective feature levels in a bottom-up sub-backbone and in a top-down feedback path of a sub-network; and lateral skip connections across adjacent sub-networks.In another aspect of the disclosure, there is provided a computer-implemented method for deep prediction using a network, the method comprising: providing an input to the network consisting of a plurality of sub-networks, wherein the plurality of sub-networks are cascaded and each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion, and wherein the network further comprises lateral skip connections between feature maps at corresponding feature levels in a bottom up sub-backbone and in a top down feedback path of a sub-network, and lateral skip connections across adjacent sub-networks; and providing output feature maps of the network in headers for a task; and wherein the input comprises a digital image of a scale and the output feature maps have different scales, and wherein the task comprises image classification, instance segmentation, object detection, or semantic segmentation.In another aspect of the disclosure, there is provided an apparatus for deep prediction, the apparatus comprising a memory and at least one processor coupled to the memory and configured to perform a computer-implemented method for deep prediction using a network, the method comprising: providing an input to the network consisting of a plurality of sub-networks, wherein the plurality of sub-networks are cascaded and each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion, and wherein the network further comprises lateral skip connections between feature maps at corresponding feature levels in a bottom up sub-backbone and in a top down feedback path of a sub-network, and lateral skip connections across adjacent sub-networks; and providing output feature maps of the network in headers for a task; and wherein the input comprises a digital image of a scale and the output feature maps have different scales, and wherein the task comprises image classification, instance segmentation, object detection, or semantic segmentation.In another aspect of the disclosure, there is provided a computer program product for deep prediction, the computer program product comprising processor executable computer code for performing a computer implemented method for deep prediction using a network, the method comprising: providing an input to the network comprised of a plurality of sub-networks, wherein the plurality of sub-networks are cascaded and each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion, and wherein the network further comprises lateral skip connections between feature maps at corresponding feature levels in a bottom up sub-backbone and in a top down feedback path of a sub-network, and lateral skip connections across adjacent sub-networks; and providing output feature maps of the network in headers for a task; and wherein the input comprises a digital image of a scale and the output feature maps have different scales, and wherein the task comprises image classification, instance segmentation, object detection, or semantic segmentation.In another aspect of the disclosure, there is provided a computer readable medium storing computer code for deep prediction, the computer code, when executed by a processor, causing the processor to perform a computer implemented method for deep prediction using a network, the method comprising: providing an input to the network comprised of a plurality of sub-networks, wherein the plurality of sub-networks are cascaded, and each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion, and wherein the network further comprises lateral skip connections between feature maps at corresponding feature levels in a bottom up sub-backbone and in a top down feedback path of a sub-network, and lateral skip connections across adjacent sub-networks; and providing output feature maps of the network in headers for a task; and wherein the input comprises a digital image of a scale and the output feature maps have different scales, and wherein the task comprises image classification, instance segmentation, object detection, or semantic segmentation.By dividing a single backbone into a plurality of sub-backbone and connecting them in cascade with a top-down feedback path following each of the sub-backbone, extraction of multi-scale feature maps and merging of multi-scale feature maps may be performed in an iterative manner, which may increase performance of the multi-scale process without causing additional computational effort.Other aspects or variations of the disclosure, as well as other advantages, will become apparent upon consideration of the following detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGSThe disclosed aspects are described below in conjunction with the accompanying drawings, which are provided to illustrate and not limit the disclosed aspects. FIGS. 1A, 1B, 1C, and 1D illustrate several current architectures for prior art density prediction tasks. FIG. 2 illustrates a schematic diagram of a feature pyramid network (FPN) for obtaining multi-scale feature maps. FIG. 3 illustrates a schematic diagram of an iterative architecture for obtaining multi-scale feature maps, in accordance with one or more aspects of the present disclosure. FIGS. 4A, 4B, and 4C illustrate schematic diagrams of three alternatives for a sub-network architecture that may be implemented in the proposed iterative architecture according to one or more aspects of the present disclosure. FIG. 5 illustrates an example schematic diagram of a common structure in the bottom-up path, in accordance with one or more aspects of the present disclosure. FIG. 6 illustrates a schematic diagram of a design for a network block that may be adopted in the sub-backbone, in accordance with one or more aspects of the present disclosure. FIG. 7 illustrates a schematic diagram of another design for a central network block that may be adopted in the sub-backbone, in accordance with one or more aspects of the present disclosure. FIG. 8 illustrates an example workflow of a method for deep prediction using a network of the proposed iterative architecture according to one or more aspects of the present disclosure. FIG. 9 illustrates an example of a hardware implementation for a device according to one or more aspects of the present disclosure.DETAILED DESCRIPTIONThe present disclosure will now be discussed with reference to several example implementations. It should be appreciated that these implementations are discussed only to enable those skilled in the art to better understand and thus implement the embodiments of the present disclosure, and not to suggest limitations on the scope of the present disclosure.Convolutional neural networks (CNN) and transformer-based networks have achieved promising results in many computer vision tasks, including image classification, object recognition, semantic segmentation, etc. For these tasks, the most common CNN (e.g., AlexNet, VGGNet, ResNet, ConvNeXt) and recently developed transformer-based networks (e.g., Swin Transformer, Focal Transformer, PoolFormer) follow a sequential approach in design of the architecture, i.e., gradually reduce the spatial size of feature maps and make predictions based on the coarsest feature scale (with the lowest resolution). However, for many deep prediction tasks such as object recognition and instance segmentation, multi-scale features are needed to be able to process objects with different scales.FIGS. 1A, 1B, 1C, and 1D illustrate several current architectures for deep prediction tasks such as object detection and the like, in which images are characterized by shadows and feature maps by outlines, and thicker outlines indicate semantically stronger feature maps. An architecture 110 of FIG. 1A illustrates feature pyramids built on image pyramids, where feature maps on each of the image scales are calculated independently of one another, which is slow. An architecture 120 of FIG. 1B illustrates the aforementioned architecture that uses only a single scale of feature maps for faster detection (e.g., via a prediction 105). An architecture 130 of FIG. 1C illustrates an alternative that reuses a pyramidal feature hierarchy computed by a ConvNet (Convolutional Network), as being feature pyramids built on image pyramids. This intra-network feature hierarchy (e.g., as shown in FIGS. 1B and 1C ) generates feature maps with different spatial resolutions, but introduces large semantic gaps therebetween (e.g., as shown by different thicknesses of the outlines in FIGS. 1B and 1C ) caused by different depths in the network. In general, high resolution feature maps have low-level features that can impair their display capacity for object recognition. However, to avoid the use of low-level features, certain methods do not use the high-resolution feature maps, thereby failing to use the high-resolution feature maps to facilitate detecting small objects.An architecture 140 of FIG. 1D illustrates the FPN combining low resolution semantically strong features with high resolution semantically weak features from the pyramid shape of the feature hierarchy of a ConvNet via a top-down path and lateral connections. Based on an FPN, numerous methods have been proposed for extracting multi-scale feature maps and fusing the multi-scale feature maps, for example by adding further bottom-up and / or top-down paths.However, these methods decouple the feature learning process into two consecutive parts, i.e., first, multi-scale feature maps are extracted through a classification network (e.g., ConvNet), and then these feature maps are fused through a lightweight module consisting of one or more paths. For deep prediction tasks, such a design may not be optimal. Moreover, the feature fusion process may not be sufficient for a limited computational budget, as most parameters and computational budget have been allocated to the classification network as backbone.To solve the problems, the present disclosure proposes an iterative architecture for obtaining multi-scale feature maps. The main idea of the iterative architecture may be to split a single backbone into a plurality of sub-backbone and cascade them, each of the sub-backbone, except for the most recent sub-backbone, being followed by a top-down feedback path. In this way, the extraction of multi-scale feature maps and the merging of multi-scale feature maps may be performed in an iterative manner, which may increase the performance of the multi-scale process without causing additional computational effort. For example, the proposed iterative architecture may maintain a similar model size and computational complexity as an FPN while increasing performance for deep prediction.FIG. 2 illustrates a schematic diagram of an FPN 200 for obtaining multi-scale feature maps. The FPN 200 may have an architecture (e.g., the architecture 140 of FIG. 1D ) including a backbone 210 and a relatively light feature fusion network 220, where nodes denotes C2, C3, C4, or C5 feature maps with different strides (e.g., denotes C2, C3, C4, or C5 feature maps with strides of 4, 8, 16, and 32, respectively), and prefixes C and P are used to distinguish the bottom-up path of the backbone 210 and the top-down path of the relatively light feature fusion network 220, respectively, and where edges denote an information flow between feature maps. An RGB input image 230 may be first fed to a lightweight main block 240 to generate first feature maps with a spatial resolution of H / 4×W / 4, where H and W denote the height and width of the input image 230. Subsequently, backbone 210 is used to extract multi-scale features, i.e., C2, C3, C4, and C5. Backbone 210 may be a strong classification network consisting of network blocks, e.g., more than one hundred layers in total. Feature maps of a particular scale in backbone 210 are transformed by multiple layered network blocks from the previous low-level features. Subsequently, the relatively light feature fusion network 220 is used to fuse the generated multi-scale feature maps. Finally, the output feature maps are fed into certain task headers.FIG. 3 illustrates a schematic diagram of an iterative architecture 300 for obtaining multi-scale feature maps according to one or more aspects of the present disclosure, wherein node T 2, T 3, T 4, or T 5 denotes one or more feature maps having a scale (i.e., spatial size) that matches that of node C 2, C 3, C 4, or C 5, respectively. The iterative architecture 300 may include a plurality of sub-networks (e.g., sub-network 310- aand sub-network 310- b) connected in cascade. Each of the plurality of sub-networks (e.g., sub-network 310- aor sub-network 310- b) may include a bottom-up sub-backbone (e.g., sub-backbone 311- aor sub-backbone 311- b) for feature extraction and a top-down feedback path (e.g., feedback path 312- aor feedback path 312- b) for feature fusion. Furthermore, the iterative architecture 300 may include an additional sub-backbone 311-c at the last stage. These included sub-backbone (e.g., sub-backbone 311-a, 311-b, and 311-c) may be split from a single backbone (e.g., backbone 210), and each of the sub-backbone may have a reduced number of network blocks, e.g., total dozens of layers. The overall size of these enclosed sub-backbone (e.g., sub-backbone 311-a, 311-b, and 311-c) may correspond to the size of the single backbone (e.g., backbone 210).In an aspect of the present disclosure, the iterative architecture 300 may use a main block 330 to first process an input image 320. For example, the input image 320 may be a digital image (e.g., a video, a radar, a LiDAR, an ultrasound, a motion, or a thermal image, and the like) of a particular scale captured by a sensor or camera or created by a computer. The output feature maps of a previous sub-network (e.g., sub-network 310- a) may be used as input to the next sub-network (e.g., sub-network 310- b). Furthermore, skip connections may be added between adjacent feature maps in each row, which may involve a copy operation.In another aspect of the present disclosure, the iterative architecture 300 may be configured to iteratively perform the feature extraction and the feature fusion, which may be optimal for the density prediction. In addition, since the overall size of the sub-backbone (e.g., sub-backbone 311-a, 311-b, and 311-c) included in the iterative architecture 300 may correspond to the size of the single backbone (e.g., backbone 210), no additional computational effort may arise.For example, the single backbone may be general CNNs (e.g., AlexNet, VGGNet, ResNet, ConvNeXt), transformer-based networks (e.g., Swin Transformer, Focal Transformer, PoolFormer), or variants thereof.It will be appreciated by those skilled in the art that for sub-backbone partitioning for the proposed iterative architecture, other backbone than the aforementioned networks may be applicable without causing any departure from the present disclosure. It should be appreciated that the number of sub-networks (i.e., 2) and the number of feature levels (i.e., 4) of FIG. 3 are merely illustrative and the available implementations of the iterative architecture 300 may not be limited to this particular example; instead, other numbers of sub-networks (e.g., 3, 4, 5, 6, 7, 8, 9, and the like) and / or other numbers of feature levels (e.g., 2, 3, 5, 6, and the like) may be possible without causing departing from the present disclosure.FIGS. 4A, 4B, and 4C illustrate schematic diagrams of three alternatives for a subnetwork architecture that may be implemented in the iterative architecture 300 according to one or more aspects of the present disclosure. An architecture 410 of FIG. 4A may have a symmetric structure in the bottom-up and top-down paths, i.e., its bottom-up path and top-down path share a similar structure and have similar computational complexity. For example, the architecture 410 may be a UNet style network slice.An architecture 420 of FIG. 4B may have an asymmetric structure in the bottom-up and top-down paths. The top-down path may be relatively light compared to the bottom-up path. For example, between adjacent nodes (e.g., P2, P3, P4, and P5) in the top-down path, a 1×1 convolution unit may be used first to reduce a channel dimension of high-level features, followed by a bilinear interpolation layer to adjust the spatial size of feature maps. Moreover, a 7×7 depth-wise separable convolution unit may be used to fuse the element-wise summed features in each plane before being fed to the next plane. For example, the architecture 420 may have a structure in an FPN style (i.e., with a similar structure of the top-down feedback path of the FPN as shown in FIGS. 1D and 2 ).An architecture 430 of FIG. 4C may also have an asymmetric structure in the bottom-up and top-down paths. The top-down path may be extremely light and include only 1×1 convolution units between adjacent nodes (e.g., Q2, Q3, Q4, and Q5) that need learning. Such a design may make the information flow between different feature planes (in the vertical direction of FIG. 4C ) and different sub-networks (in the horizontal direction of FIG. 4C ) more direct and effective, which may facilitate multi-scale feature fusion. Meanwhile, the architecture 430 of FIG. 4C may consume much less computational resources due to the extremely lightweight structure of the top-down path. In addition, the architecture 430 may include an interpolation unit or an upsampling unit between adjacent nodes (e.g., Q2, Q3, Q4, and Q5) in the top-down path to adjust the spatial size of feature maps.In one or more aspects of the present disclosure, architectures 410, 420, and 430 may all have a similar structure in the bottom-up path as shown in FIG. 5. Given the features from the main block (e.g., main block 330), a first feature level 510 consisting of a number (N 1) of layered network blocks (e.g., network block 501) may be used to transform these features. Within each subsequent feature level, a 2×2 convolution unit with a pitch of 2 502 may be used to reduce the spatial size of feature maps. In other words, feature maps within each feature level (denoted by the number after each letter in FIGS. 2, 3, 4A-4C ) may have the same spatial size, i.e., the same scale, while feature maps on different feature levels may have different spatial sizes, i.e., multiple scales. For example, as shown in FIG. 5, the spatial sizes of the feature maps across the different feature levels are decreased gradually (e.g., H / 4×W / 4 in the first feature level 510 (e.g., corresponding to nodes C 2, T 2, P 2, or Q 2), H / 8×W / 8 in a second feature level 520 (e.g., corresponding to nodes C 3, T 3, P 3, or Q 3), H / 16×W / 16 in a third feature level 530 (e.g., corresponding to nodes C 4, T 4, P4 or Q4) and H / 32×W / 32 in a fourth feature level 540 (e.g., corresponding to node C5, T5, P5, or Q5) in this example). For brevity and generality, the feature map output by the latest layer of a node may be passed as an edge input to the next node by default, and the feature map may be a multi-channel feature map.In one or more aspects of the present disclosure, if a feature level in the bottom-up path has more than one input, e.g., at a feature level of sub-backbone 311- bor 311- cof FIG. 3 (characterized by node C 2, C 3, C 4, or C 5), a 1×1 convolution unit may be introduced to fuse the inputs.Those skilled in the art will appreciate that kernel sizes and stride values other than the above example may be applicable to the sub-backbone of architectures 410, 420, 430, and 300 without causing departure from the present disclosure.FIG. 6 illustrates a schematic diagram of a design for the network block 600 that may be adopted in the sub-backbone, in accordance with one or more aspects of the present disclosure. The design of the network block 600 may be adopted to implement the network block 501 of FIG. 5. The design of the network block 600 may include three convolution units 610, 620, and 630, and a skip connection 601 between an input feature and an adder to output a feature of a network block. The design of the network block 600 may also include some other layers or units between the neighboring convolutional units, such as normalization layers or activation units (e.g., ReLU or GELU), which are omitted for brevity. For example, the convolution unit 610 may include a 1×1 convolution unit with n channels, the convolution unit 620 may include a 3×3 convolution unit with n channels, and the convolution unit 630 may include a 1×1 convolution unit with 4 n channels, and the input feature of the network block has a channel dimension of 4 n, which may be referred to as a residual neural network (ResNet) block. As another example, the convolution unit 610 may include a depth-by-depth 7×7 convolution unit (or a 3×3 convolution unit) with n channels, the convolution unit 620 may include a 1×1 convolution unit with 4 n channels, and the convolution unit 630 may include a 1×1 convolution unit with n channels, and the input feature of the network block has a channel dimension of n, which may be referred to as an iterative network block.FIG. 7 illustrates a schematic diagram of another design for a central network block 700 that may be adopted in the sub-backbone, in accordance with one or more aspects of the present disclosure. The central network block 700 may be adopted to implement the network block 501 of FIG. 5. The central network block 700 may include four convolution units 710, 720, 730 and 715 and three skip connections 701, 702 and 703. Central network block 700 may also include some other layers or units between the neighboring convolutional units, such as normalization layers or activation units (e.g., ReLU or GELU), which are omitted for brevity. The convolution units 710, 720, and 730 may be similar to the convolution units 610, 620, and 630, e.g., including a 7×7 deep convolution unit (or a 3×3 convolution unit) with n channels, a 1×1 convolution unit with 4 n channels, and a 1×1 convolution unit with n channels, respectively. The additional convolution unit 715 may include a depth-by-depth 7×7 convolution unit with an expansion rate of 3 and n channels. The two successive convolution units 710 and 715 may be used to detect locally dense features and globally sparse features, respectively. The architecture of the central network block 700 may include both fine grain local and coarse grain global interactions to achieve a more effective convolution mechanism.In one aspect of the present disclosure, the design of the central network block 700 may be applied only to the top feature level (e.g., feature level 540 corresponding to C5) of each sub-network to achieve a better trade-off between performance and efficiency. By increasing the receptive field of high-level neurons by applying the central network block to the feature-level network block 501 540 corresponding to C5, learning of subsequent low-level features may be improved, e.g., to improve semantics of low-level features.To collect long range information, ConvNet generally introduces large convolution kernels to simulate a function of self-attention. However, performance tends to be saturated because the kernel size is greater than 7×7. This may be because the effective receptive field of neurons saturates as network depth increases. In one or more aspects of the present disclosure, the power may not be saturated as the kernel size becomes larger, i.e., greater than 7×7 when the sub-backbone (e.g., sub-backbone 311- a, 311- b, or 311- c) is much flatter than the single backbone (e.g., backbone 210). For example, the kernel size of the convolution unit 710 and / or the convolution unit 715 may be greater than 7×7 or the expansion rate may be greater than 3 when the sub-backbone has a small number of network blocks (e.g., N 3< 4 and / or N 4< 2).FIG. 8 illustrates an example workflow of a method 800 for deep prediction using a network of iterative architecture 300 with architecture 410, architecture 420, or architecture 430, in accordance with one or more aspects of the present disclosure. At block 810, input may be provided to the network consisting of a plurality of sub-networks, wherein the plurality of sub-networks are cascaded and each of the plurality of sub-networks may include a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion, and wherein the network may further include lateral skip connections between feature maps at respective feature levels in a bottom-up sub-backbone and in a top-down feedback path of a sub-network, and lateral skip connections across adjacent sub-networks. For example, the bottom-up sub-backbone and top-down feedback path may comprise an architecture of architecture 410, architecture 420, or architecture 430, combined with one or more of network block 600 and / or central network block 700. At block 820, output network feature maps may be provided to a task of the deep prediction. For example, the input may include a digital image of a scale, and the output feature maps may include different scales, and the task may include image classification, instance segmentation, object recognition, or semantic segmentation, or the like.FIG. 9 illustrates an example of a hardware implementation for a device 900, in accordance with one or more aspects of the present disclosure. The device 900 for deep prediction may include a memory 910 and at least one processor 920. The processor 920 may be coupled to the memory 910 and configured to perform the method 800 with the architecture 300 combined with the architecture 410, 420, or 430, and the network block 600 and / or the central network block 700, as described above with respect to FIGS. 8, 3, 4A- 4C, and 6 and / or 7. Processor 920 may be a general purpose processor or may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. The memory 910 may store the input data, output data, data generated by the processor 920, and / or instructions executed by the processor 920.The various operations, models, and networks described herein in connection with the disclosure may be implemented in hardware, software executed by a processor, firmware, a computer, or any combination thereof. According to one or more aspects of the disclosure, a computer program product for deep prediction may include processor executable computer code for performing method 800 with architecture 300 combined with architecture 410, 420, or 430 and network block 600 and / or central network block 700, as described above with respect to FIGS. 8, 3, 4A-4C, and 6 and / or 7. According to another embodiment of the disclosure, a computer readable medium may store computer code for deep prediction, wherein the computer code, when executed by a processor, may cause the processor to perform the method 800 with the architecture 300 combined with the architecture 410, 420, or 430, and the network block 600 and / or the central network block 700, as described above with respect to FIGS. 8, 3, 4A-4C, and 6 and / or 7. Computer readable media includes both non-transitory computer readable storage media and communication media including any media that supports the transfer of a computer program from one location to another. Each connection may be referred to as a computer readable medium. Other embodiments and implementations are within the scope of the disclosure.In an embodiment of the present disclosure, the iterative architecture may further include lateral skip connections across two adjacent sub-networks (e.g., sub-network 310- aand sub-network 310- b). For example, there may be lateral skip connections between feature maps at respective feature levels of a sub-network (e.g., in both a bottom-up sub-backbone 311- aand a top-down feedback path 312- aof sub-network 310- a) and feature maps at respective feature levels in a bottom-up sub-backbone (e.g., sub-backbone 311- b) of an adjacent sub-network (e.g., sub-network 310- b). In other words, a node at a particular feature level of a sub-backbone may use the output feature maps of its previous subnet as inputs at the same feature level, and the output feature maps may originate from both a sub-backbone and a top-down feedback path of the previous subnet, i.e., more than one input feature map. In this case, a 1×1 convolution unit may be introduced to fuse the input feature maps.In an embodiment of the present disclosure, the cascaded sub-networks may have a size corresponding to a conventional classification model or network, and the cascaded architecture may allow the network to be configured such that the multi-scale feature extraction and the multi-scale feature fusion are performed in an iterative manner, thereby coupling the feature extraction and the feature fusion in the multi-scale feature learning to achieve optimal performance for the density prediction.In an embodiment of the present disclosure, each of the plurality of sub-networks may have a symmetric structure, a feature pyramid network (FPN) style structure, or an extremely light feedback path structure. The extremely light feedback path structure may include convolution units of one multiplied by one that are the only trainable units in the extremely light feedback path structure.In an embodiment of the present disclosure, the bottom-up sub-backbone may comprise residual neural network (ResNet) blocks or iterative network blocks and / or central network blocks. For example, the bottom-up sub-backbone may include multiple feature levels, and each of the feature levels may consist of one or more network blocks. The network blocks may be, for example, the ResNet blocks, the iterative network blocks, or the central network blocks. In another example, the central network blocks may be used only at the top feature level (e.g., at the fourth feature level 540 of FIG. 5 ).In an embodiment of the present disclosure, the input to the network, which is made up of a plurality of sub-networks, may include a digital image of a scale captured by a sensor or a camera or generated by a computer, and the output feature maps of the network may have different scales and have strong semantics on all scales.The foregoing description of the disclosed embodiments is provided to enable those skilled in the art to make or use the various embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the scope of the various embodiments. Thus, the claims are not intended to be limited to the embodiments shown herein, but are to be accorded the broadest scope consistent with the following claims and the principles and novel features disclosed herein.

Claims

A computer-implemented network for deep prediction, comprising: a plurality of sub-networks, wherein the plurality of sub-networks are cascaded, and wherein each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion; lateral skip connections between feature maps at respective feature levels in a bottom-up sub-backbone and in a top-down feedback path of a sub-network; and lateral skip connections across adjacent sub-networks.The computer-implemented network of claim 1, wherein the lateral skip connections across the adjacent sub-networks comprise: lateral skip connections between feature maps at respective feature levels in a sub-network and in a bottom-up sub-backbone of an adjacent sub-network.The computer-implemented network of claim 2, wherein the network is configured to perform the multi-scale feature extraction and the multi-scale feature fusion in an iterative manner.The computer-implemented network of claim 1, wherein each of the plurality of sub-networks has a symmetric structure, a feature pyramid network (FPN) style structure, or an extremely light feedback path structure.The computer-implemented network of claim 4, wherein the extremely light feedback path structure comprises convolution units of one multiplied by one that are the only trainable units in the extremely light feedback path structure.The computer-implemented network of claim 1, wherein the bottom-up sub-backbone comprises residual neural network (ResNet) blocks or iterative network blocks.The computer-implemented network of claim 6, wherein the bottom-up sub-backbone further comprises one or more central blocks at an uppermost feature level.A computer-implemented method for deep prediction using a network, comprising: providing an input to the network consisting of a plurality of sub-networks, wherein the plurality of sub-networks are cascaded and each of the plurality of sub-networks comprises a bottom-up sub-backbone for multi-scale feature extraction and a top-down feedback path for multi-scale feature fusion, and wherein the network further comprises lateral skip connections between feature maps at corresponding feature levels in a bottom up sub-backbone and in a top down feedback path of a sub-network, and lateral skip connections across adjacent sub-networks; and providing output feature maps of the network in headers for a task; and wherein the input comprises a digital image of a scale and the output feature maps have different scales, and wherein the task comprises image classification, instance segmentation, object detection, or semantic segmentation.An apparatus for deep prediction, comprising: a memory; and at least one processor coupled to the memory and configured to perform the computer-implemented method of claim 8.A computer program product for deep prediction, comprising: processor executable computer code for performing the computer implemented method of claim 8.A computer readable medium storing computer code for density prediction, wherein the computer code, when executed by a processor, causes the processor to perform the computer implemented method of claim 8.