Training neural network with generative many-to-one feature distillation

Generative many-to-one feature distillation in neural networks addresses the limitations of one-to-one knowledge transfer by transforming student network feature maps into multiple segments that approximate teacher features, resulting in improved accuracy and efficiency.

WO2025107244A1PCT designated stage expired Publication Date: 2025-05-30INTEL CORP +5
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/133667
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing knowledge distillation methods in neural networks are limited by one-to-one feature distillation between pre-selected teacher-student layer pairs, which restricts knowledge transfer and reduces performance.

Method used

The proposed solution is to implement generative many-to-one feature distillation, where a student network's feature map is transformed into multiple feature segments using a generative model, allowing each segment to approximate a teacher feature map, thereby providing multiple knowledge transfer inlets.

Benefits of technology

This approach enhances model accuracy and training efficiency by preserving the information learned by the pre-trained teacher network and allowing for more effective knowledge transfer between teacher and student networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023133667_30052025_PF_FP_ABST
    Figure CN2023133667_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A student network may be trained through generative many-to-one feature distillation using a pre-trained teacher network. Training samples are provided to both the student network and the teacher network. A teacher feature map may be generated in the teacher network. A student feature map may be generated in the student network and then transformed to a feature map including feature segments, each of which may have the same number of channels as the teacher feature map. Internal parameters of the student network may be adjusted to make each feature segment approximate the teacher feature map by minimizing a feature distance between each feature segment and the teacher feature map. The feature distance may be used to define a feature distillation loss. Another loss may be determined based on outputs of the student network and ground-truth labels of the training samples. The student network may be trained using the two losses.
Need to check novelty before this filing date? Find Prior Art

Description

TRAINING NEURAL NETWORK WITH GENERATIVE MANY-TO-ONE FEATURE DISTILLATIONTechnical Field

[0001] This disclosure relates generally to neural networks (also referred to as “deep neural networks” or “DNNs” ) , and more specifically, to training DNNs with generative many-to-one feature distillation.Background

[0002] DNNs are used extensively for a variety of artificial intelligence applications ranging from computer vision to speech recognition and natural language processing due to their ability to achieve high accuracy. However, the high accuracy comes at the expense of significant computation cost. DNNs have extremely high computing demands as each inference can require hundreds of millions of MAC (multiply-accumulate) operations as well as hundreds of millions of weight operand weights to be stored for classification or detection. Therefore, techniques to improve efficiency of DNNs are needed.Brief Description of the Drawings

[0003] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements. Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.

[0004] FIG. 1 illustrates an example DNN, in accordance with various embodiments.

[0005] FIG. 2 is a block diagram of a DNN system, in accordance with various embodiments.

[0006] FIG. 3 is a block diagram of a training module, in accordance with various embodiments.

[0007] FIG. 4 illustrates an example process of training a student network with a teacher network through generative many-to-one feature distillation, in accordance with various embodiments.

[0008] FIG. 5 illustrates a generative network, in accordance with various embodiments.

[0009] FIG. 6 illustrates a deep learning (DL) environment, in accordance with various embodiments.

[0010] FIG. 7 is a flowchart showing a method of training a DNN through generative many- to-one feature distillation, in accordance with various embodiments.

[0011] FIG. 8 is a block diagram of an example computing device, in accordance with various embodiments.Detailed Description

[0012] Overview

[0013] The last decade has witnessed a rapid rise in AI (artificial intelligence) based data processing, particularly based on DNNs. DNNs are widely used in the domains of image recognition, video understanding, image or video generation, machine translation, mathematical reasoning, and so on. A DNN typically includes a sequence of layers. A DNN layer may include one or more DL operations (also referred to as “neural network operations” ) , such as convolution, pooling, elementwise operation, linear operation, nonlinear operation, and so on. A DL operation in a DNN may be performed on one or more internal parameters of the DNNs (e.g., weights) , which are determined during the training phase, and one or more activations. An activation may be a data point (also referred to as “data elements” or “elements” ) . Activations or weights of a DNN layer may be elements of a tensor of the DNN layer. A tensor is a data structure having multiple elements across one or more dimensions. Example tensors include a vector, which is a one-dimensional tensor, and a matrix, which is a two-dimensional tensor. There can also be three-dimensional tensors and even higher dimensional tensors. A DNN layer may have an input tensor (also referred to as “input feature map (IFM) ” ) including one or more input activations (also referred to as “input elements” ) and a weight tensor including one or more weights. A weight is an element in the weight tensor. A weight tensor of a convolution may be a kernel, a filter, or a group of filters. The output data of the DNN layer may be an output tensor (also referred to as “output feature map (OFM) ” ) that includes one or more output activations (also referred to as “output elements” ) .

[0014] DNN architectures usually have large numbers of learnable parameters (including weights described above) stacked with complicated network topologies. This makes DNN models have much better ability in fitting training data than traditional machine learning methods. However, this also leads to challenges. One challenge is posing intensive memory, computation, and power costs during inference. Another challenge is increasing the difficulty of model training. Knowledge distillation techniques are used to strike accuracy- efficiency tradeoff for training high performance DNN models, specially tailored to low-cost AI model deployment on diverse edge devices or client devices.

[0015] Knowledge distillation usually uses a trained large model (called teacher) to a smaller target model (called student) . Knowledge distillation can improve the performance of the student model significantly. Vanilla Knowledge distillation approaches typically use the logits output from the last layer of the teacher network as soft knowledge. Other approaches show that the feature maps from intermediate layers can also be used as hints to improve distillation performance. Spatial attention maps are used sometimes instead of original feature maps. One-stage knowledge distillation methods can adopt an online training framework to collaboratively train the teacher and student models from the scratch. Progressive or contrastive or masked learning methods are also adopted to boost the performance of knowledge distillation solutions.

[0016] However, existing knowledge distillation solutions (in both two-stage and one-stage families) commonly adopt the one-to-one feature distillation between every pre-selected teacher-student layer pair. They are limited to one knowledge transfer inlet for any teacher-student layer pair. This can suppress knowledge distillation performance to a large extent.

[0017] Embodiments of the present disclosure may improve on at least some of the challenges and issues described above by training DNNs with generative many-to-one feature distillation. A feature map generated in a student can be converted to a number of feature segments using a generative model. Each of the feature segments can be forced to approximate a feature map generated in a teacher to get the knowledge learnt by the teacher, providing multiple knowledge transfer inlets.

[0018] In various embodiments of the, a teacher network is pre-trained for training one or more student networks. The teacher network and the one or more student networks may be DNNs, such as DNNs that can be used for performing the same type of AI tasks. In an example training process, a teacher feature map (also referred to as “teacher feature” ) generated in the teacher network may be identified. Also, a student feature map (also referred to as “student feature” ) generated in the student network is identified. The teacher network layer (also referred to as “teacher layer” ) generating the teacher feature map may correspond to the student network layer (also referred to as “student layer” ) generating the student feature map. For example, the teacher layer may be the last layer in the backbone of the teacher network, and the student layer may be the last layer in the backbone of the  student network. As another example, the teacher layer may be placed right before a fully-connected layer in the teacher network, and the student layer may be placed right before a fully-connected layer in the student network.

[0019] The student feature map may be transformed to an expanded feature map that has more channels than the student feature map. The transformation may be done in a linear layer inserted into the student network. The expanded feature map may be input into a generative model, which outputs a new feature map comprising a number of feature segments. A feature segment may have the same number of channels as the teacher feature map. Internal parameters of the student network may be adjusted to make each feature segment approximate the teacher feature map. The feature distillation from the teacher feature map to each feature segment may be done by minimizing a feature distance between each feature segment and the teacher feature map. The feature distance may be used to define a feature distillation loss. Another loss may be determined based on outputs of the student network and ground-truth labels of the training samples. The student network may be trained using the two losses.

[0020] Another linear layer may also be inserted into the student network to transform the expanded feature map to a feature map having the same spatial size as the student feature map. The transformation (e.g., contraction) in the second linear layer can facilitate the DL operation in the next layer of the student network by making sure that the feature map input to the next layer has the proper spatial size. The next layer may be a fully-connected layer. After the student network is trained, the two linear layers may be merged into the fully-connected layer. For instance, the internal parameters of the two linear layers may be incorporated into the fully-connected layer for inference of the student network.

[0021] With generative many-to-one feature distillation, intact information learnt by the pre-trained teacher network can be preserved. Generative many-to-one feature distillation can be used to train DNNs used to perform AI tasks in various applications, such as image classification, face recognition, action recognition, person re-identification, image or video generation machine translation, speech recognition, and so on. Compared with currently available knowledge distillation approaches, generative many-to-one feature distillation can provide better model accuracy and training efficiency.

[0022] For purposes of explanation, specific numbers, materials and configurations are set forth in order to provide a thorough understanding of the illustrative implementations.  However, it will be apparent to one skilled in the art that the present disclosure may be practiced without the specific details or / and that the present disclosure may be practiced with only some of the described aspects. In other instances, well known features are omitted or simplified in order not to obscure the illustrative implementations.

[0023] Further, references are made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, embodiments that may be practiced. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.

[0024] Various operations may be described as multiple discrete actions or operations in turn, in a manner that is most helpful in understanding the claimed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations may not be performed in the order of presentation. Operations described may be performed in a different order from the described embodiment. Various additional operations may be performed or described operations may be omitted in additional embodiments.

[0025] For the purposes of the present disclosure, the phrase “A or B” or the phrase "A and / or B" means (A) , (B) , or (A and B) . For the purposes of the present disclosure, the phrase “A, B, or C” or the phrase "A, B, and / or C" means (A) , (B) , (C) , (A and B) , (A and C) , (B and C) , or (A, B, and C) . The term "between, " when used with reference to measurement ranges, is inclusive of the ends of the measurement ranges.

[0026] The description uses the phrases "in an embodiment" or "in embodiments, " which may each refer to one or more of the same or different embodiments. The terms "comprising, " "including, " "having, " and the like, as used with respect to embodiments of the present disclosure, are synonymous. The disclosure may use perspective-based descriptions such as "above, " "below, " "top, " "bottom, " and "side" to explain various features of the drawings, but these terms are simply for ease of discussion, and do not imply a desired or required orientation. The accompanying drawings are not necessarily drawn to scale. Unless otherwise specified, the use of the ordinal adjectives “first, ” “second, ” and “third, ” etc., to describe a common object, merely indicates that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking or in any other manner.

[0027] In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.

[0028] The terms “substantially, ” “close, ” “approximately, ” “near, ” and “about, ” generally refer to being within + / -20%of a target value as described herein or as known in the art. Similarly, terms indicating orientation of various elements, e.g., “coplanar, ” “perpendicular, ” “orthogonal, ” “parallel, ” or any other angle between the elements, generally refer to being within + / -5-20%of a target value as described herein or as known in the art.

[0029] In addition, the terms “comprise, ” “comprising, ” “include, ” “including, ” “have, ” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a method, process, device, or DNN accelerator that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, device, or DNN accelerators. Also, the term “or” refers to an inclusive “or” and not to an exclusive “or. ”

[0030] The systems, methods and devices of this disclosure each have several innovative aspects, no single one of which is solely responsible for all desirable attributes disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the description below and the accompanying drawings.

[0031] Example DNN

[0032] FIG. 1 illustrates an example DNN 100, in accordance with various embodiments. For the purpose of illustration, the DNN 100 in FIG. 1 is a CNN. In other embodiments, the DNN 100 may be other types of DNNs. The DNN 100 is trained to receive images and output classifications of objects in the images. In the embodiments of FIG. 1, the DNN 100 receives an input image 105 that includes objects 115, 125, and 135. The DNN 100 includes a sequence of layers comprising a plurality of convolutional layers 110 (individually referred to as “convolutional layer 110” ) , a plurality of pooling layers 120 (individually referred to as “pooling layer 120” ) , and a plurality of fully-connected layers 130 (individually referred to as “fully-connected layer 130” ) . In other embodiments, the DNN 100 may include fewer, more, or different layers. In an inference of the DNN 100, the layers of the DNN 100 execute tensor computation that includes many tensor operations, such as convolution (e.g., multiply-accumulate (MAC) operations, etc. ) , pooling operations, elementwise operations (e.g., elementwise addition, elementwise multiplication, etc. ) , other types of tensor  operations, or some combination thereof.

[0033] The convolutional layers 110 summarize the presence of features in the input image 105. The convolutional layers 110 function as feature extractors. The first layer of the DNN 100 is a convolutional layer 110. In an example, a convolutional layer 110 performs a convolution on an input tensor 140 (also referred to as IFM 140) and a filter 150. As shown in FIG. 1, the IFM 140 is represented by a 7×7×3 three-dimensional (3D) matrix. The IFM 140 includes 3 input channels, each of which is represented by a 7×7 two-dimensional (2D) matrix. The 7×7 2D matrix includes 7 input elements (also referred to as input points) in each row and 7 input elements in each column. The filter 150 is represented by a 3×3×3 3D matrix. The filter 150 includes 3 kernels, each of which may correspond to a different input channel of the IFM 140. A kernel is a 2D matrix of weights, where the weights are arranged in columns and rows. A kernel can be smaller than the IFM. In the embodiments of FIG. 1, each kernel is represented by a 3×3 2D matrix. The 3×3 kernel includes 3 weights in each row and 3 weights in each column. Weights can be initialized and updated by backpropagation using gradient descent. The magnitudes of the weights can indicate importance of the filter 150 in extracting features from the IFM 140.

[0034] The convolution includes MAC operations with the input elements in the IFM 140 and the weights in the filter 150. The convolution may be a standard convolution 163 or a depthwise convolution 183. In the standard convolution 163, the whole filter 150 slides across the IFM 140. All the input channels are combined to produce an output tensor 160 (also referred to as output feature map (OFM) 160) . The OFM 160 is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements (also referred to as output points) in each row and 5 output elements in each column. For the purpose of illustration, the standard convolution includes one filter in the embodiments of FIG. 1. In embodiments where there are multiple filters, the standard convolution may produce multiple output channels in the OFM 160.

[0035] The multiplication applied between a kernel-sized patch of the IFM 140 and a kernel may be a dot product. A dot product is the elementwise multiplication between the kernel-sized patch of the IFM 140 and the corresponding kernel, which is then summed, always resulting in a single value. Because it results in a single value, the operation is often referred to as the “scalar product. ” Using a kernel smaller than the IFM 140 is intentional as it allows the same kernel (set of weights) to be multiplied by the IFM 140 multiple times at different  points on the IFM 140. Specifically, the kernel is applied systematically to each overlapping part or kernel-sized patch of the IFM 140, left to right, top to bottom. The result from multiplying the kernel with the IFM 140 one time is a single value. As the kernel is applied multiple times to the IFM 140, the multiplication result is a 2D matrix of output elements. As such, the 2D output matrix (i.e., the OFM 160) from the standard convolution 163 is referred to as an OFM.

[0036] In the depthwise convolution 183, the input channels are not combined. Rather, MAC operations are performed on an individual input channel and an individual kernel and produce an output channel. As shown in FIG. 1, the depthwise convolution 183 produces a depthwise output tensor 180. The depthwise output tensor 180 is represented by a 5×5×3 3D matrix. The depthwise output tensor 180 includes 3 output channels, each of which is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements in each row and 5 output elements in each column. Each output channel is a result of MAC operations of an input channel of the IFM 140 and a kernel of the filter 150. For instance, the first output channel (patterned with dots) is a result of MAC operations of the first input channel (patterned with dots) and the first kernel (patterned with dots) , the second output channel (patterned with horizontal strips) is a result of MAC operations of the second input channel (patterned with horizontal strips) and the second kernel (patterned with horizontal strips) , and the third output channel (patterned with diagonal stripes) is a result of MAC operations of the third input channel (patterned with diagonal stripes) and the third kernel (patterned with diagonal stripes) . In such a depthwise convolution, the number of input channels equals the number of output channels, and each output channel corresponds to a different input channel. The input channels and output channels are referred to collectively as depthwise channels. After the depthwise convolution, a pointwise convolution 193 is then performed on the depthwise output tensor 180 and a 1×1×3 tensor 190 to produce the OFM 160.

[0037] The OFM 160 is then passed to the next layer in the sequence. In some embodiments, the OFM 160 is passed through an activation function. An example activation function is rectified linear unit (ReLU) . ReLU is a calculation that returns the value provided as input directly, or the value zero if the input is zero or less. The convolutional layer 110 may receive several images as input and calculate the convolution of each of them with each of the kernels. This process can be repeated several times. For instance, the OFM 160  is passed to the subsequent convolutional layer 110 (i.e., the convolutional layer 110 following the convolutional layer 110 generating the OFM 160 in the sequence) . The subsequent convolutional layers 110 perform a convolution on the OFM 160 with new kernels and generate a new feature map. The new feature map may also be normalized and resized. The new feature map can be kernelled again by a further subsequent convolutional layer 110, and so on.

[0038] In some embodiments, a convolutional layer 110 has four hyperparameters: the number of kernels, the size F kernels (e.g., a kernel is of dimensions F×F×D pixels) , the S step with which the window corresponding to the kernel is dragged on the image (e.g., a step of one means moving the window one pixel at a time) , and the zero-padding P (e.g., adding a black contour of P pixels thickness to the input image of the convolutional layer 110) . The convolutional layers 110 may perform various types of convolutions, such as 2-dimensional convolution, dilated or atrous convolution, spatial separable convolution, depthwise separable convolution, transposed convolution, and so on. The DNN 100 includes 16 convolutional layers 110. In other embodiments, the DNN 100 may include a different number of convolutional layers.

[0039] The pooling layers 120 down-sample feature maps generated by the convolutional layers, e.g., by summarizing the presence of features in the patches of the feature maps. A pooling layer 120 is placed between two convolution layers 110: a preceding convolutional layer 110 (the convolution layer 110 preceding the pooling layer 120 in the sequence of layers) and a subsequent convolutional layer 110 (the convolution layer 110 subsequent to the pooling layer 120 in the sequence of layers) . In some embodiments, a pooling layer 120 is added after a convolutional layer 110, e.g., after an activation function (e.g., ReLU, etc. ) has been applied to the OFM 160.

[0040] A pooling layer 120 receives feature maps generated by the preceding convolution layer 110 and applies a pooling operation to the feature maps. The pooling operation reduces the size of the feature maps while preserving their important characteristics. Accordingly, the pooling operation improves the efficiency of the DNN and avoids over-learning. The pooling layers 120 may perform the pooling operation through average pooling (calculating the average value for each patch on the feature map) , max pooling (calculating the maximum value for each patch of the feature map) , or a combination of both. The size of the pooling operation is smaller than the size of the feature maps. In  various embodiments, the pooling operation is 2×2 pixels applied with a stride of two pixels, so that the pooling operation reduces the size of a feature map by a factor of 2, e.g., the number of pixels or values in the feature map is reduced to one quarter the size. In an example, a pooling layer 120 applied to a feature map of 6×6 results in an output pooled feature map of 3×3. The output of the pooling layer 120 is inputted into the subsequent convolution layer 110 for further feature extraction. In some embodiments, the pooling layer 120 operates upon each feature map separately to create a new set of the same number of pooled feature maps.

[0041] The fully-connected layers 130 are the last layers of the DNN. The fully-connected layers 130 may be convolutional or not. A fully-connected layer 130 may be a linear layer. In some embodiments, a fully-connected layer 130 (e.g., the first fully-connected layer in the DNN 100) may receive an input operand. The input operand may define the output of the convolutional layers 110 and pooling layers 120 and includes the values of the last feature map generated by the last pooling layer 120 in the sequence. The fully-connected layer 130 may apply a linear transformation to the input operand through a weight matrix. The weight matrix may be a kernel of the fully-connected layer 130. The linear transformation may include a tensor multiplication between the input operand and the weight matrix. The result of the linear transformation may be an output operand. In some embodiments, the fully-connected layer may further apply a non-linear transformation (e.g., by using a non-linear activation function) on the result of the linear transformation to generate an output operand. The output operand may contain as many elements as there are classes: element i represents the probability that the image belongs to class i. Each element is therefore between 0 and 1, and the sum of all is worth one. These probabilities are calculated by the last fully-connected layer 130 by using a logistic function (binary classification) or a SoftMax function (multi-class classification) as an activation function.

[0042] In some embodiments, the fully-connected layers 130 classify the input image 105 and return an operand of size N, where N is the number of classes in the image classification problem. In the embodiments of FIG. 1, N equals 3, as there are 3 objects 115, 125, and 135 in the input image. Each element of the operand indicates the probability for the input image 105 to belong to a class. To calculate the probabilities, the fully-connected layers 130 multiply each input element by weight, make the sum, and then apply an activation function (e.g., logistic if N=2, SoftMax if N>2) . This is equivalent to multiplying the input operand by  the matrix containing the weights. In an example, the vector includes 3 probabilities: a first probability indicating the object 115 being a tree, a second probability indicating the object 125 being a car, and a third probability indicating the object 135 being a person. In other embodiments where the input image 105 includes different objects or a different number of objects, the individual values can be different.

[0043] Example DNN System

[0044] FIG. 2 is a block diagram of a DNN system 200, in accordance with various embodiments. The DNN system 200 trains DNNs by using generative many-to-one feature distillation, e.g., knowledge distillation between a number of feature segments generated by a generative network taking a feature map from a student network as the input and a feature from a teacher network. A DNN can be used to perform one or more machine learning tasks. A machine learning task is a task of making an inference. The inference is a process of running available data into the DNN to generate an output, and the output provides a solution to a problem or question that is being asked. An example of the output is one or more numerical scores that can indicate a probability of an object in an image belonging to a category. The DNN system 200 can train DNNs that can be used to solve various problems, such as image classification, learning relationships between biological cells (e.g., DNA, proteins, etc. ) , control behaviors for devices (e.g., robots, machines, etc. ) , and so on.

[0045] The DNN system 200 includes an interface module 210, a training set generator 220, a student network generator 230, a teacher network generator 240, a training module 250, a validation module 260, and a datastore 270. In other embodiments, alternative configurations, different or additional components may be included in the DNN system 200. Further, functionality attributed to a component of the DNN system 200 may be accomplished by a different component included in the DNN system 200 or by a different system.

[0046] The interface module 210 facilitates communications of the DNN system 200 with other systems. For example, the interface module 210 establishes communications between the DNN system 200 with an external database to receive data that can be used to train DNNs or data that can be input into DNNs to perform machine learning tasks. As another example, the interface module 210 supports the DNN system 200 to distribute DNNs to other systems, e.g., computing devices configured to apply DNNs to perform tasks. The  computing devices may be an edge device, a client device, and so on.

[0047] The training set generator 220 forms training datasets that will be used to train DNNs. A training dataset includes training samples and ground-truth labels. The training dataset may include one or more ground-truth labels for each training sample. A ground-truth label of a training sample may be a known or verified label that answers the problem or question that the DNN will be used to answer. In an example where a DNN is trained to recognize objects in images, the training dataset includes training images and ground-truth labels that indicate classifications of objects in the training images. A ground-truth label in the example may be a number that indicates a probability that an object belongs to a class. The object may be associated with other ground-truth labels that indicate probabilities that the object belongs to other classes.

[0048] In some embodiments, the training set generator 220 may also form validation datasets for validating performance of trained DNNs by the validation module 260. A validation dataset may include validation samples and ground-truth labels of the validation samples. The validation dataset for a DNN may include different samples from the training dataset used for training the DNN. In an embodiment, a part of a training dataset may be used to initially train a DNN, and the rest of the training dataset may be held back as a validation subset used by the validation module 260 to validate performance of the trained DNN. The portion of the training dataset not including the validation subset may be used to train the DNN.

[0049] The student network generator 230 generates student networks. A student network is a DNN that after being trained, can be used to perform machine learning tasks. The student network generator 230 may generate a student network based on parameters that define the architecture of a DNN. Examples of the parameters include the number of layers, types of layers, sequence of layers, number of processing elements (PEs) in a layer, types of PEs, arrangement of PEs (e.g., interconnections between PEs, number of columns in a PE array, number of rows in a PE array, etc. ) in a layer, activation function, pooling function, or other types of parameters. A processing element performs MAC operations.

[0050] In some embodiments, the student network generator 230 determines some or all of the parameters, e.g., based on the problem or question to be answered by the DNN, resource available for training, resources available for inference, some other factors that may be critical to the architecture of the DNN, or some combination thereof. In other  embodiments, the student network generator 230 may receive some or all of the parameters from a different system (e.g., from a computing device that will run the DNN for inference, a system managing such computing devices, etc. ) or from a user (e.g., through a user interface that allows the user to provide information of the DNN) .

[0051] The architecture of a DNN includes an input layer, an output layer, and a plurality of hidden layers. The input layer of a DNN may include tensors (e.g., a multidimensional array) specifying attributes of the input image, such as the height of the input image, the width of the input image, and the depth of the input image (e.g., the number of bits specifying the color of a pixel in the input image) . The output layer includes labels of objects in the input layer. The hidden layers are layers between the input layer and output layer. The hidden layers include one or more convolutional layers and one or more other types of layers, such as ReLU layers, pooling layers, fully-connected layers, normalization layers, softmax or logistic layers, and so on. The convolutional layers of the DNN abstract the input image to a feature map that is represented by a tensor specifying the feature map height, the feature map width, and the feature map channels (e.g., red, green, blue images include three channels) . A pooling layer is used to reduce the spatial volume of input image after convolution. It is used between two convolutional layers. A fully-connected layer involves weights, biases, and neurons. It connects neurons in one layer to neurons in another layer. It is used to classify images between different category by training. An example DNN is the DNN 100 described above in conjunction with FIG. 1.

[0052] The teacher network generator 240 generates teacher networks to be used to train student networks through knowledge distillation. In some embodiments, the teacher network generator 240 may generate a single teacher network for training a single student network or multiple student networks or may generate multiple teacher networks to train a single student network. A teacher network may be a DNN. A teacher network may have a different architecture from a student network trained with the teacher network.

[0053] In some embodiments, the teacher network generator 240 determines a structure of a teacher network based on the structure of a student network. For instance, the teacher network generator 240 may generate a teacher network including the same types of layers as the student network. The arrangement of the layers in the teacher network ( “teacher layers” ) can be the same as the arrangement of the layers in the student network ( “student layers” ) . Also, for an individual teacher layer, the teacher network generator 240 may design  the teacher layer based on a corresponding student layer. The teacher network generator 240 may make the teacher layer mirror the student layer. For instance, the teacher layer can have the same DL operation (s) as the student layer. In some embodiments, the teacher network generator 240 may generate a teacher network that is larger than the student network. For instance, the teacher network may include more layers than the student network. Additionally or alternatively, a teacher layer may have more internal parameters than the corresponding student layer.

[0054] The teacher network generator 240 also generates internal connections within the teacher network. An internal connection may connect two teacher layers, e.g., from a first teacher layer to a second teacher layer. The second teacher layer may be arranged in after the first layer in the teacher network. The internal connection facilitates data transfer between the two teacher layers. For instance, the first layer can send features (e.g., OFM 160) to the second layer through the internal connection. The second layer receives the features and can aggregate the features from the first layer with features generated in the second layer to output aggregated features. An internal connection may be bidirectional, e.g., the second layer can also send data to the first layer.

[0055] The teacher network generator 240 may also train teacher networks. For instance, the teacher network generator 240 may input training samples from the training set generator 220 into a teacher network and update internal parameters of the teacher network to minimize a loss. The teacher network generator 240 may determine the loss based on the difference between the outputs of the teacher network and the ground-truth labels of the training samples. In some embodiments, the teacher network generator 240 may train the teacher network for tasks to be performed by one or more student networks that will be trained using the teacher network. For instance, the teacher network generator 240 may train the teacher network with the same types of training samples that the student network will receive. In an example where the student network will be used to classify objects in images, the teacher network generator 240 may train the teacher network using training samples including images.

[0056] The training module 250 trains DNNs. For instance, the training module 250 may train a teacher network first, then train a student network using the teacher network through generative many-to-one feature distillation. The student network may be generated by the student network generator 230. The teacher network may be generated  by the teacher network generator 240. In some embodiments, the training module 250 may train the teacher network for one or more AI tasks to be performed by the student network. The teacher network may be larger than the student network. For instance, the teacher network may have more layers or more internal parameters than the student network. In some embodiments, the training module 250 may use one teacher network to train multiple student networks. In other embodiments, the training module 250 may use multiple teacher networks to train one student network.

[0057] To train a student network, the training module 250 may provide the same training set to the teacher network and student network. The internal parameters of the teacher network are fixed during the process of the training the student network. Layers in the two networks may process the training samples in the training set and generate feature maps. The training module 250 may extract a teacher feature generated by a layer identified in the teacher network.

[0058] The training module 250 may insert additional layers into the student network. In an example, the training module 250 may identify a layer in the student network. The identified layer may be the layer that outputs a student feature map to be used for feature distillation. The training module 250 may insert two transformation modules right after the identified layer. The OFM of the identified layer would be the IFM of the first transformation module, and the second transformation module may receive the OFM of the first transformation module and would generate an OFM, which may be used as the IFM of the next layer in the student network. The two transformation modules may each be a linear layer that applies a linear transformation on the input of the transformation module. In some embodiments, the first transformation module may expand feature maps in the channel dimension, and the second transformation module may contract feature maps in the channel dimension. The training module 250 may also associate a generative model with the two transformation modules. The generative model may receive the OFM of the first transformation module and generate a new feature map.

[0059] The training module 250 may perform generative many-to-one feature distillation using the new feature map from the generative model and a teacher feature. In some embodiments, the training module 250 may determine a feature distillation loss based on a number of feature distances. A feature distance may be a feature distance between a segment in the new feature map and the teacher feature. In some embodiments, the  training module 250 may determine other types of losses. For instance, the training module 250 may further determine a task loss that measures the difference between the outputs of the student network with ground-truth labels of the training samples. The training module 250 may use a combination of the feature distillation loss and the task loss to adjust the internal parameters of the student network.

[0060] The training module 250 may determine hyperparameters for training a student network. Hyperparameters may be different from parameters inside the network (e.g., weights) . In some embodiments, the hyperparameters include variables which determine how the DNN is trained, such as batch size, number of epochs, etc. A batch size defines the number of training samples to work through before updating the parameters of the DNN. The batch size is the same as or smaller than the number of samples in the training dataset. The training dataset can be divided into one or more batches. The number of epochs defines how many times the entire training dataset is passed forward and backwards through the entire network. The number of epochs defines the number of times that the DL algorithm works through the entire training dataset. One epoch means that each training sample in the training dataset has had an opportunity to update the parameters inside the network. An epoch may include one or more batches. The number of epochs may be 10, 100, 500, 1000, or even larger. Certain aspects of the training module 250 are described below in conjunction with FIG. 3.

[0061] The validation module 260 verifies performance (e.g., accuracy) of trained DNNs, such as trained student networks that are separated from their corresponding teacher networks. The validation module 260 may determine the accuracy of a trained student network and determine whether the accuracy meets a threshold (e.g., a requirement for model accuracy) . In response to determining that the accuracy of the student network meets the threshold, the validation module 260 may deploy the student network to another system or device, e.g., through the interface module 210.

[0062] In some embodiments, the validation module 260 inputs samples in a validation dataset into the DNN and uses the outputs of the DNN to determine the model accuracy. In some embodiments, a validation dataset may be formed of some or all the samples in the training dataset. Additionally or alternatively, the validation dataset includes additional samples, other than those in the training sets. In some embodiments, the validation module 260 determines may determine an accuracy score measuring the precision, recall, or a  combination of precision and recall of the DNN. The validation module 260 may use the following metrics to determine the accuracy score: Precision = TP  /  (TP + FP) and Recall = TP  /  (TP + FN) , where precision may be how many the DNN correctly predicted (TP or true positives) out of the total it predicted (TP + FP or false positives) , and recall may be how many the DNN correctly predicted (TP) out of the total number of samples that did have the property in question (TP + FN or false negatives) . The F-score (F-score = 2 *PR  /  (P + R) ) unifies precision and recall into a single measure.

[0063] The datastore 270 stores data associated with the DNN system 200, such as data received, generated, or used by the DNN system 200. For instance, the datastore 270 may store parameters (e.g., internal parameters, hyperparameters, etc. ) of student networks or teacher networks generated by the student network generator 230, the teacher network generator 240, or the training module 250. The datastore 270 may also store training sets and validation sets used to train networks and validate networks. The datastore 270 may further store contrasting pairs formed by the training module 250. In some embodiments, the DNN system 200 may include or be associated with more than one datastore. The datastore 270 may be implemented as a random-access memory (RAM) , such as a static RAM (SRAM) , disk storage, nearline storage, online storage, offline storage, and so on.

[0064] FIG. 3 is a block diagram of the training module 250, in accordance with various embodiments. As described above, the training module 250 may train a student network based on a pre-trained teacher network using generative many-to-one feature distillation. As shown in FIG. 3, the training module 250 includes an input module 310, a layer identification module 320, an extraction module 330, an expansion module 340, a contraction module 350, a generative module 360, and a loss module 370. In other embodiments, alternative configurations, different or additional components may be included in the training module 250. For instance, the training module 250 may include multiple student transformation modules. Further, functionality attributed to a component of the training module 250 may be accomplished by a different component included in the training module 250 or by a different system.

[0065] The input module 310 inputs one or more training samples into the teacher network and the student network. The input module 310 may input the same training samples into the teacher network and the student network. The DL operations in the teacher network and the student network may be performed based on the training samples to generate  feature maps. The input module 310 may receive the training samples from the training set generator 220. A training sample may be associated with one or more ground-truth labels. The input module 310 may also prevent the internal parameters of the teacher network from being changed.

[0066] The layer identification module 320 identifies one or more layers in the teacher network and one or more in the student network that will be used to train the student network using generative many-to-one feature distillation. For instance, the training module 250 may identify a layer in the teacher network and a layer in the student network. The two identified layers may be of the same type. For instance, the two identified layers may include the same type of DL operation. Additionally or alternatively, the two identified layers may have the same position index in the two networks. In an example, the identified layers may each be the last layer in the backbone of the corresponding network or the layer immediately before the head of the corresponding network. In other examples, the identified layers may be at other positions in the networks.

[0067] In some embodiments, the layer identification module 320 may identify multiple layers in the student network. For example, the layer identification module 320 may identify multiple stages in the student network and select one or more layers from each of the identified stages. As another example, the layer identification module 320 may identify layers of different types from the student network. For each identified student layer, the layer identification module 320 may identify a corresponding layer in the teacher network. The layer identification module 320 may pair a student layer with the corresponding teacher layer for generative many-to-one feature distillation from the teacher layer to the student layer.

[0068] The extraction module 330 extracts feature maps from layers identified by the layer identification module 320. For instance, the extraction module 330 may extract the feature map generated by each identified teacher layer. This feature map is to be used as a teacher feature map. In an example, the teacher feature map may be a 3D tensor denoted as where H denotes the height of the teacher feature map, W denotes the width of the teacher feature map, and Ct denotes the depth of the teacher feature map. The height of the teacher feature map may be indicated by the number of activations in a row of the teacher feature map. The width of the teacher feature map may be indicated by the number of activations in a column of the teacher feature map. The depth of the teacher feature map  may be indicated by the number of channels in the teacher feature map.

[0069] The extraction module 330 may extract the feature map generated by each identified student layer. In an example, the student feature map may be a 3D tensor denoted as where H denotes the height of the student feature map, W denotes the width of the student feature map, and Cs denotes the depth of the student feature map. The height of the student feature map may be indicated by the number of activations in a row of the student feature map. The width of the student feature map may be indicated by the number of activations in a column of the student feature map. The depth of the student feature map may be indicated by the number of channels in the student feature map.

[0070] In some embodiments, the feature map extracted from a student layer may have one or more different dimensions from the feature map extracted from the corresponding teacher layer. For instance, the height, width, or depth of the feature map extracted from the student layer may be different from the height, width, or depth of the feature map extracted from the teacher layer.

[0071] The expansion module 340 is inserted into the student network before or during the training of the student network. The expansion module 340 may be one or more linear layers. In some embodiments, the expansion module 340 is inserted right after the layer identified by the layer identification module 320, i.e., between the identified layer and the next layer in the student network. The expansion module 340 may have one or more learnable parameters, such as weights whose values may be determined during the process of training the student network through generative many-to-one feature distillation. In some embodiments, the expansion module 340 may receive a student feature map generated by the identified layer (e.g., the student feature map extracted by the extraction module 330) and transform the student feature map to an expanded feature map.

[0072] In an example, the expansion module 340 may apply a linear transformation on the student feature using a weight tensor The expansion module 340 may multiply the weight tensor and the student feature map to compute the expanded feature map. The expansion module 340 may expand the depth of student feature map (i.e., Cs) to a larger depth NCt, which is N times the depth of the teacher feature map extracted by the extraction module 330. The expanded feature map can be divided into feature segments, each of which have a depth of Ct. The channel expansion by the expansion module 340 may  be defined as Fsc=Wsc*Fse, where *denotes linear transformation operation.

[0073] The contraction module 350 is inserted into the student network before or during the training of the student network. The contraction module 350 may be one or more linear layers. In some embodiments, the contraction module 350 is inserted right after the expansion module 340. The contraction module 350 may have one or more learnable parameters, such as weights whose values may be determined during the process of training the student network through generative many-to-one feature distillation.

[0074] In some embodiments, the contraction module 350 may contract the expanded feature map in the channel dimension to compute a contracted feature map that has the same spatial size as the student feature map so that the contracted feature map can be processed by the next layer in the student network. The next layer is a layer that is right after the identified layer in the student network before the insertion of the expansion module 340 and the contraction module 350. After the insertion of the expansion module 340 and the contraction module 350, the next layer is right after the contract module 350. In an example, the contraction module 350 may apply a linear transformation on the expanded feature using a weight tensor The contraction module 350 may multiply the weight tensor and the expanded feature map to compute the contracted feature map. The contraction module 350 may reduce the depth of expanded feature map (i.e., NCt) to a smaller depth Cs, i.e., the depth of the student feature map extracted by the extraction module 330. The channel contraction by the contraction module 350 may be defined as Fse=Wse*Fs, respectively, where *denotes linear transformation operation.

[0075] The generative module 360 is also inserted into the student network before or during the training to enable multiple knowledge transfer inlets with the teacher feature map during the training of the student network. The generative module 360 may facilitate semantic alignment over multiple feature segments to mimic the teacher feature map. The generative module 360 may be a network including a plurality of layers. An example of the generative module 360 is the generative network 500 in FIG. 5. The generative module 360 may apply a mask on the expanded feature map generated by the expansion module 340. The mask may zero out some of the channels in the expanded feature map. The number of the zeroed out channels may be rNCt, where r has a value between zero and one, such as 0.3, 0.5, 0.7, and so on.

[0076] The remaining (1-r) NCt channels in the expanded feature map may constitute a  subtensor, which may be further processed in one or more DL operations, such as convolutions, batch normalization, activation function, other DL operation, or some combination thereof. The one or more DL operations may compute a tensor that has rNCt channels. The output tensor of the one or more DL operations may be combined with the input tensor of the one or more DL operations to generate a feature map having NCt channels. This feature map may be divided to feature segments each having Ct channels. A feature distillation inlet may be built between an individual feature map and the teacher feature map. As there are N such feature segments, there can be N feature distillation inlets enabled using one teacher feature map.

[0077] The loss module 370 determines one or more losses and adjusts internal parameters of the student network based on the one or more losses. In some embodiments, the loss module 370 determines a feature distillation loss for the generative many-to-one feature distillation. The loss module 370 may determine the feature distillation loss using a loss function denoted as:

[0078] where Lfd is the feature distilaltion loss,  is the i-th feature segment with i indicating the index of the feature segment, N is the total number of feature segments, and Ft is the teacher feature map.  denotes the feature distance between the i-th feature segment and the teacher feature map. The feature distance may be a Euclidean distance such as an L2-norm distance. The feature distance may be determined using a feature distance function denoted as:

[0079] In some embodiments, the loss module 370 may also determine one or more other losses. For instance, the loss module 370 may determine a task loss of the student network 410 supervised with the ground-truth labels associated with the training samples. In some embodiments, the task loss may be a cross-entropy loss between probability distributions determined by the student network 410 using the training samples and the probability distributions in the ground-truth labels. The loss module 370 may determine the task loss using a cross-entropy loss function denoted as:

[0080] where LCE is the task loss (i.e., cross-entropy loss) , CE standards for cross-entropy, Y denotes the ground-truth labels, Fs denotes the student feature map, the denotes the  weight tensor of a fully-connected layer in the student network. The fully-connected layer may be the layer into which the expansion module 340 and contraction module 350 are merged.

[0081] The loss module 370 may determine a total loss by aggregating the feature distillation loss and the task loss. In some embodiments, the total loss L may be defined as: L= α·LCE+ β·Lfd,

[0082] where α and β are the weights of the task loss and the feature distillation loss, respectively. In some embodiments, the values of α and β may be predetermined, e.g., determined before the training of the student network starts. In an example, α=1 and β=8. In other examples, α or β may have different values.

[0083] The loss module 370 may adjust internal parameters of the student network to reduce or minimize the total loss. The internal parameters include weights in layers of the student network, including the weight tensor of the expansion module 340, contraction module 350, or generative module 360. The loss module 370 may adjust the internal parameters of the student network till one or more criteria are met. Examples of the criteria may include a threshold loss, a predetermined number of epochs, a target performance (e.g., an accuracy) of the student network, a threshold duration of time, other types of criteria, or some combination thereof.

[0084] After the training is done, the loss module 370 may merge the expansion module 340 and the contraction module 350 into a fully-connected layer after the identified layer in the student network. For instance, the loss module may modify a weight tensor of the fully-connected layer to incorporate the weight tensors of the expansion module 340 and the weight tensor of the contraction module 350. The merge may be represented by: Wfc=Wfc (WscWse+I) ,

[0085] where is an identity matrix. By merging the expansion module 340 and the contraction module 350 into an existing layer in the student network, no extra parameters or architectural modifications are introduced to the student network. The trained student network can be used to handle machine learning tasks. In some embodiments, the student network, or parameters of the student network, may be sent to another system or device (e.g., an edge device, a client device, etc. ) for inference.

[0086] Example Processes of Training Student Network

[0087] FIG. 4 illustrates an example process of training a student network 410 with a  teacher network 420 through generative many-to-one feature distillation, in accordance with various embodiments. The student network 410, after being trained, may be used to perform AI tasks. The teacher network 420 may have been trained to perform the same type of tasks. As shown in FIG. 4, the student network 410 includes layers 413 (individually referred to as “layer 413” ) , a layer 417, and a layer 419. The teacher network 420 includes layers 423 (individually referred to as “layer 423” ) , a layer 427, and a layer 429. In some embodiments, the teacher network 420 may include more layers than the student network 410. Additionally or alternatively, one or more layers 423 may include more parameters and have larger sizes than one or more layers 413. In other embodiments, the student network 410 or the teacher network 420 may include different, fewer, or more components. Examples of the layers 413, 417, 419, 423, 427, and 429 may include convolutional layers, self-attention layers, fully-connected layers, pooling layers, other types of neural network layers, or some combination thereof. An example of the student network 410 or an example of the teacher network 420 may be the DNN 100 in FIG. 1.

[0088] In some embodiments, the layer 413 may constitute the backbone of the student network 410. The layers 413 may be arranged in a sequence. A layer 413 may receive OFM of the previous layer 413 in the sequence as an IFM and generate a new OFM using one or more DL operations in the layer 413. The OFM of the layer 413 may be provided to the next layer 413 in the sequence as an IFM. The OFM of the last layer 413 in the sequence may be provided to the layer 417. The layer 417 and the layer 419 may constitute the head of the student network 410. The layer 417 may receive and output a vector. The layer 419 may apply an activation function on the vector to generate a new vector. In an example, the layer 417 may be a fully-connected layer, such as a fully-connected layer 130 in FIG. 1. The layer 419 may include a softmax activation function that can convert a vector to probabilities. The output of the layer 419 may be the output of the student network 410. Even though not shown in FIG. 4, the student network 410 may include an additional layer (e.g., a fully-connected layer) after the layer 419 in some embodiments, and the output of the additional layer may be the output of the student network 410.

[0089] In some embodiments, the layer 423 may constitute the backbone of the teacher network 420. The layers 423 may be arranged in a sequence. A layer 423 may receive OFM of the previous layer 423 in the sequence as an IFM and generate a new OFM using one or more DL operations in the layer 423. The OFM of the layer 423 may be provided to the next  layer 423 in the sequence as an IFM. The OFM of the last layer 423 in the sequence may be provided to the layer 427. The layer 427 and the layer 429 may constitute the head of the teacher network 420. The layer 427 may receive and output a vector. The layer 429 may apply an activation function on the vector to generate a new vector. In an example, the layer 427 may be a fully-connected layer, such as a fully-connected layer 130 in FIG. 1. The layer 429 may include a softmax activation function that can convert a vector to probabilities. The output of the layer 429 may be the output of the teacher network 420. Even though not shown in FIG. 4, the teacher network 420 may include an additional layer (e.g., a fully-connected layer) after the layer 429 in some embodiments, and the output of the additional layer may be the output of the teacher network 420.

[0090] To train the student network 410, an input 430 is provided to the student network 410 and the teacher network 420. The input 430 may include one or more training samples. The training samples may be generated by the training set generator 220. The layers 413 may extract features from the input 430 and generate a feature map 415. In some embodiments, the feature map 415 may be the OFM of the last layer 413 in the backbone of the student network 410. In other embodiments, the feature map 415 may be the OFM of a different layer 413 in the student network 410. The feature map 415 may be a 3D tensor and includes data elements arranged in three dimensions. One of the three dimensions represents the channels in the feature map 415. Each channel may be represented by a 2D tensor. The feature map 415 may include a plurality of pixels. Each pixel may be represented by a vector that includes a sequence of data elements in the channel dimension. The data elements in the vector may encode the pixel in different channels. In an example, the feature map 415 may be a 3D tensor denoted as H may denote the height of the feature map 415 along the Y axis and may equal the number of data elements in a column of the feature map 415. W may denote the width of the feature map 415 along the X axis and may equal the number of data elements in a row of the feature map 415. Cs may denote the depth of the feature map 415 along the Z axis and may equal the number of channels in the feature map 415.

[0091] The layers 423 in the teacher network 420 may extract features from the input 430 and generate a feature map 425. In some embodiments, the feature map 425 may be the OFM of the last layer 423 in the backbone of the teacher network 420. In other embodiments, the feature map 425 may be the OFM of a different layer 423 in the teacher  network 420. In some embodiments, the feature map 425 is a 3D tensor and includes data elements arranged in three dimensions. One of the three dimensions represents the channels in the feature map 425. Each channel may be represented by a 2D tensor. The feature map 425 may include a plurality of pixels. Each pixel may be represented by a vector that includes a sequence of data elements. The data elements in the vector may encode the pixel in different channels. The feature map 425 may have a different spatial size from the feature map 415. In an example, the feature map 415 may be a 3D tensor denoted as H may denote the height of the feature map 425 along the Y axis and may equal the number of data elements in a column of the feature map 425. W may denote the width of the feature map 425 along the X axis and may equal the number of data elements in a row of the feature map 425. Ct may denote the depth of the feature map 425 along the Z axis and may equal the number of channels in the feature map 425.

[0092] To facilitate the generative many-to-one feature distillation, two transformation modules 440 and 450 are inserted into the student network 410. An example of the transformation module 440 may be the expansion module 340 in FIG. 3. An example of the transformation module 450 may be the contraction module 350 in FIG. 3. In some embodiments, the two transformation modules 440 and 450 are inserted right after the layer 413 that generates the feature map 415. In an embodiment where the layer 413 generating the feature map 415 is the last layer 413, the two transformation modules 440 and 450 are inserted between the last layer 413 and the layer 417. A generative network 460 is also added to facilitate the generative many-to-one feature distillation. The generative network 460 may be added between the two transformation modules 440 and 450. An example of the generative network 460 may be the generative module 360 in FIG. 3.

[0093] In some embodiments, the transformation module 440 may be a linear layer that can expand the feature map 415 in the channel dimension to produce a feature map 470. The feature map 470 is an expanded feature map and has more channels than the feature map 415. A linear transformation in the linear layer may be applied on the feature map 415 to increase the number of channels of the feature map 415 to a multiple of the number of channels in the feature map 425. In an embodiment, the linear layer may have a weight tensor that may be denoted as where Cs denotes the number of channels in the feature map 415, Ct denotes the number of channels in the feature map 425, and N is an integer. N may have a value greater than one. The feature map 470 may be  denoted as The feature map 470 has N feature segments 475, individually referred to as “feature segment 475. ” The linear transformation may be denoted as Fse= Wse*Fs. A feature segment 475 may have a spatial size of H×W×Ct. The feature segments 475 may have the same spatial size, which may be the same as the spatial size of the feature map 425.

[0094] The transformation module 450 may be a liner layer that can contract the feature map 470 in the channel dimension, i.e., along the Z axis. In some embodiments, the linear layer may have a weight tensor denoted as to project each pixel in Fseback to the original channel dimension Cs, producing With the transformation module 450, the output of the transformation modules 440 and 450 may be a tensor having the proper spatial size that can be processed by the layer 417.

[0095] In some embodiments, the layer 417 may be a fully-connected layer. The layer 417 may process the output of the transformation module 450 and output a probability distribution. The layer 417 may have a weight tensor including weights that can be learnt during the process of training the student network 410. The weight tensors of the transformation modules 440 and 450 may also be learnt during the training process. After the training process, the transformation modules 440 and 450 may be merged into the layer 417, e.g., after the student network 410 is trained. In some embodiments, the trained weight tensor of the layer 417 may be modified based on the trained weight tensors of the transformation modules 440 and 450 during the merge.

[0096] The generative network 460 may receive the feature map 470 generated by the transformation module 440 and output a feature map 480. The generative network 460 may be used for semantic alignment of the feature map 470 to the feature map 425, which may be denoted as Fsg=G (Fse, r) , where Fsg denotes the feature map 480, G denotes the generative network 460, and r denotes a mask ratio in a range from zero to one. In some examples, r may be 0.3.0.4, 0.5, 0.6, or other values. In some embodiments, r indicates a percentage of channels in the feature map 470. The percentage of channels may be selected, e.g., randomly, and be masked. For instance, the mask may be applied to the selected channels to zero out the data elements in the selected channels. The number of the selected channels may be rNCt. The other channels of the feature map 470 may be input into the generative network 460. The input to the generative network 460 may be a  3D tensor denoted as The generative network 460 produces an output from the other channels of the feature map 470. The output of the generative network 460 may have rNCt channels. In an example, the output of the generative network 460 may be denoted as The output of the generative network 460 may be combined with the input of the generative network 460 to produce the feature map 480.

[0097] The feature map 480 may be denoted as The feature map 480 includes N feature segments 485, individually referred to as “feature segment 485. ” A feature segment 485 may be denoted as The spatial size of a feature segment 485 may be the same as the spatial size of the feature map 425. The feature segments 485 may have the same spatial size as each other. To train the student network 410, each feature segment 485 is a segmented student representation and may be forced to approximate the feature map 425, e.g., by minimizing a L2-norm feature distance function, increasing the number of knowledge transfer paths / inlets from one-to-one to N-to-one.

[0098] FIG. 5 illustrates a generative network 500, in accordance with various embodiments. The generative network 500 may be an example of the generative network 460 in FIG. 4. The generative network 500 includes layers 510, 520, 530, 540, and 550 that are arranged in a sequence. The layer 510 may receive feature maps as inputs, e.g., expanded feature maps generated by the transformation module 440. The layer 510 may include one or more DL operations that would be performed on a received feature map to generate an OFM. In some embodiments, the layer 510 is a convolutional layer. The layer 510 may include a convolution, which may be a standard convolution, depthwise convolution, etc. The convolution may have a weight tensor. The weight tensor may be a 3D tensor with a plurality of channels. Each channel may correspond to a 2D tensor, which may be a kernel including weights arranged in a 2D matrix. In some embodiments, the convolution may be accompanied by one or more padding operations or other types of operations.

[0099] The layer 520 receives OFMs of the layer 510 as inputs and normalizes values of the activations in the inputs. In some embodiments, the layer 520 is a batch normalization layer. The layer 520 may be used to make the training of a student network (e.g., the student network 410) faster and more stable. The layer 520 may normalize activations (e.g.,  activations output from the layer 510) . For instance, the layer 520 may determine the mean and variable of the activation values across the batch and normalize the activations based on the mean and variables. The layer 520 may include an operation to normalize activations by subtracting the batch mean and dividing by the batch standard deviation.

[0100] In some embodiments, the layer 520 may have one or more internal parameters that may be determined by training the generative network 500. In an example, the layer 520 may have a first parameter to adjust the standard deviation and a second parameter to adjust the bias. The outputs from the layer 520 may follow a standard normal distribution across the batch.

[0101] In some embodiments, the layer 520 may mitigate internal covariate shift. The internal covariate shift may be the change in the distribution of layer inputs during training. This shift can make the training process slower and more difficult. Batch normalization can mitigate this issue by ensuring that the IFMs have zero mean and unit variance, which can help to stabilize and speed up the training process. The layer 520 may also prevent overfitting by adding a small amount of noise to the IFMs, which may act as a form of regularization. This noise may help to reduce the reliance of the model on specific input values and encourages it to learn more robust and generalizable features.

[0102] The layer 530 receives outputs of the layer 520 as IFMs and generates new OFMs. In some embodiments, the layer 530 may be an activation function layer. The layer 530 may apply an activation function on an IFM to generate an OFM. An example of the activation function may be ReLU.

[0103] The layer 540 receives OFMs of the layer 530 as inputs. In some embodiments, the layer 540 may include one or more DL operations that would be performed on a received feature map to generate an OFM. In some embodiments, the layer 540 is a convolutional layer, such as one of the convolutional layers described above. In some embodiments, the convolution may be accompanied by one or more padding operations or other types of operations. In some embodiments, a weight tensor of the layer 540 may have the same spatial size (e.g., height, weight, and depth) as a weight tensor of the layer 510. Also, the weights in the layer 540 may be the same as the weights in the layer 510. In an example, the layers 510 and 540 may use the same weight tensor.

[0104] The layer 550 receives OFMs of the layer 540 as inputs. The layer 550 may normalize values of activations in the input and outputs activations having values that follow a  standard normal distribution. In some embodiments, the layer 550 is a batch normalization layer, such as one of the batch normalization layers described above.

[0105] The architecture of the generative network 500 shown in FIG. 5 is an example. The generative network 500 may have different architectures. For example, the layers 510, 520, 530, 540, and 550 in the generative network 500 may be arranged in different orders. As another example, the generative model 500 may include fewer, more, or different layers. In some embodiments, the generative model 500 may include one or more fully-connected layers in lieu of convolutional layers. Each fully-connected layer may be followed by a normalization layer and an activation function layer (e.g., a ReLU layer) .

[0106] Example DL Environment

[0107] FIG. 6 illustrates a DL environment 600, in accordance with various embodiments. The DL environment 600 includes a DL server 610 and a plurality of client devices 620 (individually referred to as client device 620) . The DL server 610 is connected to the client devices 620 through a network 630. In other embodiments, the DL environment 600 may include fewer, more, or different components.

[0108] The DL server 610 trains DL models using neural networks. A neural network is structured like the human brain and consists of artificial neurons, also known as nodes. These nodes are stacked next to each other in three types of layers: input layer, hidden layer (s) , and output layer. Data provides each node with information in the form of inputs. The node multiplies the inputs with random weights, calculates them, and adds a bias. Finally, nonlinear functions, also known as activation functions, are applied to determine which neuron to fire. The DL server 610 can use various types of neural networks, such as DNN, recurrent neural network (RNN) , generative adversarial network (GAN) , long short-term memory network (LSTMN) , and so on. During the process of training the DL models, the neural networks use unknown elements in the input distribution to extract features, group objects, and discover useful data patterns. The DL models can be used to solve various problems, e.g., making predictions, classifying images, and so on. The DL server 610 may build DL models specific to particular types of problems that need to be solved. A DL model is trained to receive an input and outputs the solution to the particular problem.

[0109] In FIG. 6, the DL server 610 includes a DNN system 640, a database 650, and a distributer 660. The DNN system 640 trains DNNs. The DNNs can be used to process images, e.g., images captured by autonomous vehicles, medical devices, satellites, and so on. In an  embodiment, a DNN receives an input image and outputs classifications of objects in the input image. An example of the DNNs is the DNN 100 described above in conjunction with FIG. 1 or the student network 410 described above in conjunction with FIG. 4. In some embodiments, the DNN system 640 trains DNNs through knowledge distillation, e.g., dense-connection based knowledge distillation. The trained DNNs may be used on low memory systems, like mobile phones, IOT edge devices, and so on. An embodiment of the DNN system 640 is the DNN system 200 described above in conjunction with FIG. 2.

[0110] The database 650 stores data received, used, generated, or otherwise associated with the DL server 610. For example, the database 650 stores a training dataset that the DNN system 640 uses to train DNNs. In an embodiment, the training dataset is an image gallery that can be used to train a DNN for classifying images. The training dataset may include data received from the client devices 620. As another example, the database 650 stores hyperparameters of the neural networks built by the DL server 610.

[0111] The distributer 660 distributes DL models generated by the DL server 610 to the client devices 620. In some embodiments, the distributer 660 receives a request for a DNN from a client device 620 through the network 630. The request may include a description of a problem that the client device 620 needs to solve. The request may also include information of the client device 620, such as information describing available computing resource on the client device. The information describing available computing resource on the client device 620 can be information indicating network bandwidth, information indicating available memory size, information indicating processing power of the client device 620, and so on. In an embodiment, the distributer may instruct the DNN system 640 to generate a DNN in accordance with the request. The DNN system 640 may generate a DNN based on the information in the request. For instance, the DNN system 640 can determine the structure of the DNN and / or train the DNN in accordance with the request.

[0112] In another embodiment, the distributer 660 may select the DNN from a group of pre-existing DNNs based on the request. The distributer 660 may select a DNN for a particular client device 620 based on the size of the DNN and available resources of the client device 620. In embodiments where the distributer 660 determines that the client device 620 has limited memory or processing power, the distributer 660 may select a compressed DNN for the client device 620, as opposed to an uncompressed DNN that has a larger size. The distributer 660 then transmits the DNN generated or selected for the client device 620 to  the client device 620.

[0113] In some embodiments, the distributer 660 may receive feedback from the client device 620. For example, the distributer 660 receives new training data from the client device 620 and may send the new training data to the DNN system 640 for further training the DNN. As another example, the feedback includes an update of the available computer resource on the client device 620. The distributer 660 may send a different DNN to the client device 620 based on the update. For instance, after receiving the feedback indicating that the computing resources of the client device 620 have been reduced, the distributer 660 sends a DNN of a smaller size to the client device 620.

[0114] The client devices 620 receive DNNs from the distributer 660 and applies the DNNs to perform machine learning tasks, e.g., to solve problems or answer questions. In various embodiments, the client devices 620 input images into the DNNs and uses the output of the DNNs for various applications, e.g., visual reconstruction, augmented reality, robot localization and navigation, medical diagnosis, weather prediction, and so on. A client device 620 may be one or more computing devices capable of receiving user input as well as transmitting and / or receiving data via the network 630. In one embodiment, a client device 620 is a conventional computer system, such as a desktop or a laptop computer. Alternatively, a client device 620 may be a device having computer functionality, such as a personal digital assistant (PDA) , a mobile telephone, a smartphone, an autonomous vehicle, or another suitable device. A client device 620 is configured to communicate via the network 630. In one embodiment, a client device 620 executes an application allowing a user of the client device 620 to interact with the DL server 610 (e.g., the distributer 660 of the DL server 610) . The client device 620 may request DNNs or send feedback to the distributer 660 through the application. For example, a client device 620 executes a browser application to enable interaction between the client device 620 and the DL server 610 via the network 630. In another embodiment, a client device 620 interacts with the DL server 610 through an application programming interface (API) running on a native operating system of the client device 620, such as or ANDROIDTM.

[0115] In an embodiment, a client device 620 is an integrated computing device that operates as a standalone network-enabled device. For example, the client device 620 includes display, speakers, microphone, camera, and input device. In another embodiment, a client device 620 is a computing device for coupling to an external media device such as a  television or other external display and / or audio output system. In this embodiment, the client device 620 may couple to the external media device via a wireless interface or wired interface (e.g., an HDMI cable) and may utilize various functions of the external media device such as its display, speakers, microphone, camera, and input devices. Here, the client device 620 may be configured to be compatible with a generic external media device that does not have specialized software, firmware, or hardware specifically for interacting with the client device 620.

[0116] The network 630 supports communications between the DL server 610 and client devices 620. The network 630 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 630 may use standard communications technologies and / or protocols. For example, the network 630 may include communication links using technologies such as Ethernet, 8010.11, worldwide interoperability for microwave access (WiMAX) , 3G, 4G, code division multiple access (CDMA) , digital subscriber line (DSL) , etc. Examples of networking protocols used for communicating via the network 630 may include multiprotocol label switching (MPLS) , transmission control protocol / Internet protocol (TCP / IP) , hypertext transport protocol (HTTP) , simple mail transfer protocol (SMTP) , and file transfer protocol (FTP) . Data exchanged over the network 630 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML) . In some embodiments, all or some of the communication links of the network 630 may be encrypted using any suitable technique or techniques.

[0117] Example Method of Training DNN

[0118] FIG. 7 is a flowchart showing a method 700 of training a DNN through generative many-to-one feature distillation, in accordance with various embodiments. The method 700 may be performed by the training module 250 in FIG. 2. Although the method 700 is described with reference to the flowchart illustrated in FIG. 7, many other methods for training a DNN through dense-connection based knowledge distillation may alternatively be used. For example, the order of execution of the steps in FIG. 7 may be changed. As another example, some of the steps may be changed, eliminated, or combined.

[0119] The training module 250 identifies 710 identifying an OFM generated in a layer of the first neural network. In some embodiments, the first neural network is a student network. In some embodiments, the layer is the last layer before a fully-connected layer in  the first neural network. In some embodiments, the layer is a convolutional layer.

[0120] The training module 250 generates 720 an expanded feature map from the OFM by expanding the OFM in a channel dimension of the OFM. In some embodiments, the training module 250 generates the expanded feature map by inputting the OFM into a linear layer comprising a linear transformation operation. The linear layer outputs the expanded feature map.

[0121] The training module 250 generates 730, using a generative network, a first feature map from the expanded feature map. The first feature map has more channels than the OFM and comprises a plurality of feature segments. In some embodiments, the training module 250 applies a mask on the expanded feature map to convert data elements in a first portion of the expanded feature map to zeros. The training module 250 inputs a second portion of the expanded feature map into the generative network. The generative network outputs an intermediate feature map. The intermediate feature map has a same number of channels as the first portion of the expanded feature map. The training module 250 generates the first feature map by combining the intermediate feature map with the second portion of the expanded feature map.

[0122] In some embodiments, the training module 250 also generates a contracted feature map from the expanded feature map. The contracted feature map has the same number of channels as the OFM and is input into a next layer in the first neural network. The next layer is after the layer of the first neural network. In some embodiments, the next layer is a fully-connected layer. In some embodiments, the training module 250 generates a contracted feature map by inputting the expanded feature map into a linear layer comprising a linear transformation operation. The linear layer outputs the contracted feature map.

[0123] The training module 250 identifies 740 a second feature map generated in a layer of a second neural network. The second feature map has the same number of channels as a feature segment in the first feature map. In some embodiments, the second neural network is a teacher network. In some embodiments, the layer is the last layer before a fully-connected layer in the second neural network. In some embodiments, the layer is a convolutional layer.

[0124] The training module 250 adjusts 750 one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the  second feature map.

[0125] In some embodiments, the training module 250 also adjusts one or more parameters in a linear layer based on the plurality of feature distillation operations. The linear layer is inserted into the first neural network at a position after the layer generating the OFM. The linear layer is used for generating the expanded feature map or the contracted feature map. After adjusting the one or more parameters in the linear layer, the training module 250 merges the linear layer into a next layer in the first neural network. The next layer is after the layer of the first neural network. In some embodiments, the next layer is a fully-connected layer.

[0126] Example Computing Device

[0127] FIG. 8 is a block diagram of an example computing device 800, in accordance with various embodiments. A number of components are illustrated in FIG. 8 as included in the computing device 800, but any one or more of these components may be omitted or duplicated, as suitable for the application. In some embodiments, some or all of the components included in the computing device 800 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 800 may not include one or more of the components illustrated in FIG. 8, but the computing device 800 may include interface circuitry for coupling to the one or more components. For example, the computing device 800 may not include a display device 806, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 806 may be coupled. In another set of examples, the computing device 800 may not include an audio input device 818 or an audio output device 808, but may include audio input or output device interface circuitry (e.g., connectors and supporting circuitry) to which an audio input device 818 or audio output device 808 may be coupled.

[0128] The computing device 800 may include a processing device 802 (e.g., one or more processing devices) . As used herein, the term "processing device" or "processor" may refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. The processing device 802 may include one or more digital signal processors (DSPs) , application-specific ICs (ASICs) , CPUs, GPUs, cryptoprocessors (specialized processors that execute cryptographic algorithms within hardware) , server processors, or  any other suitable processing devices. The computing device 800 may include a memory 804, which may itself include one or more memory devices such as volatile memory (e.g., DRAM) , nonvolatile memory (e.g., read-only memory (ROM) ) , flash memory, solid state memory, and / or a hard drive. In some embodiments, the memory 804 may include memory that shares a die with the processing device 802. In some embodiments, the memory 804 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for training DNNs, e.g., the method 700 described above in conjunction with FIG. 7 or the operations performed by the DNN system 200 described above in conjunction with FIG. 2 (e.g., operations performed by the training module 250) . The instructions stored in the one or more non-transitory computer-readable media may be executed by the processing device 802.

[0129] In some embodiments, the computing device 800 may include a communication chip 812 (e.g., one or more communication chips) . For example, the communication chip 812 may be configured for managing wireless communications for the transfer of data to and from the computing device 800. The term "wireless" and its derivatives may be used to describe circuits, devices, DNN accelerators, methods, techniques, communications channels, etc., that may communicate data through the use of modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not.

[0130] The communication chip 812 may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.13 family) , IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment) , Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as "3GPP2" ) , etc. ) . IEEE 802.16 compatible Broadband Wireless Access (BWA) networks are generally referred to as WiMAX networks, an acronym that stands for worldwide interoperability for microwave access, which is a certification mark for products that pass conformity and interoperability tests for the IEEE 802.16 standards. The communication chip 812 may operate in accordance with a Global system for Mobile Communication (GSM) , General Packet Radio Service (GPRS) , Universal Mobile Telecommunications system (UMTS) , High Speed Packet Access (HSPA) , Evolved HSPA (E-HSPA) , or LTE network. The communication chip 812 may operate in accordance with  Enhanced Data for GSM Evolution (EDGE) , GSM EDGE Radio Access Network (GERAN) , Universal Terrestrial Radio Access Network (UTRAN) , or Evolved UTRAN (E-UTRAN) . The communication chip 812 may operate in accordance with CDMA, Time Division Multiple Access (TDMA) , Digital Enhanced Cordless Telecommunications (DECT) , Evolution-Data Optimized (EV-DO) , and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. The communication chip 812 may operate in accordance with other wireless protocols in other embodiments. The computing device 800 may include an antenna 822 to facilitate wireless communications and / or to receive other wireless communications (such as AM or FM radio transmissions) .

[0131] In some embodiments, the communication chip 812 may manage wired communications, such as electrical, optical, or any other suitable communication protocols (e.g., the Ethernet) . As noted above, the communication chip 812 may include multiple communication chips. For instance, a first communication chip 812 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second communication chip 812 may be dedicated to longer-range wireless communications such as global positioning system (GPS) , EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication chip 812 may be dedicated to wireless communications, and a second communication chip 812 may be dedicated to wired communications.

[0132] The computing device 800 may include battery / power circuitry 814. The battery / power circuitry 814 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 800 to an energy source separate from the computing device 800 (e.g., AC line power) .

[0133] The computing device 800 may include a display device 806 (or corresponding interface circuitry, as discussed above) . The display device 806 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD) , a light-emitting diode display, or a flat panel display, for example.

[0134] The computing device 800 may include an audio output device 808 (or corresponding interface circuitry, as discussed above) . The audio output device 808 may include any device that generates an audible indicator, such as speakers, headsets, or earbuds, for example.

[0135] The computing device 800 may include an audio input device 818 (or corresponding interface circuitry, as discussed above) . The audio input device 818 may include any device that generates a signal representative of a sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments having a musical instrument digital interface (MIDI) output) .

[0136] The computing device 800 may include a GPS device 816 (or corresponding interface circuitry, as discussed above) . The GPS device 816 may be in communication with a satellite-based system and may receive a location of the computing device 800, as known in the art.

[0137] The computing device 800 may include an other output device 813 (or corresponding interface circuitry, as discussed above) . Examples of the other output device 813 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.

[0138] The computing device 800 may include an other input device 820 (or corresponding interface circuitry, as discussed above) . Examples of the other input device 820 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.

[0139] The computing device 800 may have any desired form factor, such as a handheld or mobile computing system (e.g., a cell phone, a smart phone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a PDA, an ultramobile personal computer, etc. ) , a desktop computing system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computing system. In some embodiments, the computing device 800 may be any other electronic device that processes data.

[0140] Select Examples

[0141] The following paragraphs provide various examples of the embodiments disclosed herein.

[0142] Example 1 provides a method for training a first neural network, the method including identifying an OFM generated in a layer of the first neural network; generating an expanded feature map from the OFM by expanding the OFM in a channel dimension of the OFM; generating, using a generative network, a first feature map from the expanded  feature map, the first feature map having more channels than the OFM and including a plurality of feature segments; identifying a second feature map generated in a layer of a second neural network, the second feature map having a same number of channels as a feature segment in the first feature map; and adjusting one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the second feature map.

[0143] Example 2 provides the method of example 1, in which generating the expanded feature map from the OFM includes inputting the OFM into a linear layer including a linear transformation operation, the linear layer outputting the expanded feature map.

[0144] Example 3 provides the method of example 2, further including adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; and after adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, in which the next layer is after the layer of the first neural network.

[0145] Example 4 provides the method of example 3, in which the next layer is a fully-connected layer.

[0146] Example 5 provides the method of any one of examples 1-4, further including generating a contracted feature map from the expanded feature map, in which the contracted feature map has a same number of channels as the OFM and is input into a next layer in the first neural network, and the next layer is after the layer of the first neural network.

[0147] Example 6 provides the method of example 5, in which generating the contracted feature map from the expanded feature map includes inputting the expanded feature map into a linear layer including a linear transformation operation, the linear layer outputting the contracted feature map.

[0148] Example 7 provides the method of example 6, further including adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; and after adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, in which the next layer is after the layer of the first neural network.

[0149] Example 8 provides the method of any one of examples 1-7, in which generating the  first feature map including applying a mask on the expanded feature map to convert data elements in a first portion of the expanded feature map to zeros; inputting a second portion of the expanded feature map into the generative network, the generative network outputting an intermediate feature map, the intermediate feature map having a same number of channels as the first portion of the expanded feature map; and generating the first feature map by combining the intermediate feature map with the second portion of the expanded feature map.

[0150] Example 9 provides the method of any one of examples 1-8, in which the generative network includes a first convolutional layer and a second convolutional layer.

[0151] Example 10 provides the method of example 9, in which: the generative network further includes a first batch normalization layer, an activation function layer, and a second batch normalization layer, the first batch normalization layer and the activation function layer are between the first convolutional layer and the second convolutional layer, and the second batch normalization layer is after the second convolutional layer.

[0152] Example 11 provides one or more non-transitory computer-readable media storing instructions executable to perform operations for training a first neural network, the operations including identifying an OFM generated in a layer of the first neural network; generating an expanded feature map from the OFM by expanding the OFM in a channel dimension of the OFM; generating, using a generative network, a first feature map from the expanded feature map, the first feature map having more channels than the OFM and including a plurality of feature segments; identifying a second feature map generated in a layer of a second neural network, the second feature map having a same number of channels as a feature segment in the first feature map; and adjusting one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the second feature map.

[0153] Example 12 provides the one or more non-transitory computer-readable media of example 11, in which generating the expanded feature map from the OFM includes inputting the OFM into a linear layer including a linear transformation operation, the linear layer outputting the expanded feature map.

[0154] Example 13 provides the one or more non-transitory computer-readable media of example 12, in which the operations further include adjusting one or more parameters in  the linear layer based on the plurality of feature distillation operations; and after adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, in which the next layer is after the layer of the first neural network.

[0155] Example 14 provides the one or more non-transitory computer-readable media of any one of examples 11-13, in which the operations further include generating a contracted feature map from the expanded feature map by inputting the expanded feature map into a linear layer including a linear transformation operation, the linear layer outputting the contracted feature map, in which the contracted feature map has a same number of channels as the OFM and is input into a next layer in the first neural network, and the next layer is after the layer of the first neural network.

[0156] Example 15 provides the one or more non-transitory computer-readable media of example 14, in which the operations further include adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; and after adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, in which the next layer is after the layer of the first neural network.

[0157] Example 16 provides the one or more non-transitory computer-readable media of any one of examples 11-15, in which generating the first feature map including applying a mask on the expanded feature map to convert data elements in a first portion of the expanded feature map to zeros; inputting a second portion of the expanded feature map into the generative network, the generative network outputting an intermediate feature map, the intermediate feature map having a same number of channels as the first portion of the expanded feature map; and generating the first feature map by combining the intermediate feature map with the second portion of the expanded feature map.

[0158] Example 17 provides the one or more non-transitory computer-readable media of any one of examples 11-16, in which: the generative network includes a first convolutional layer, a second convolutional layer, a first batch normalization layer, an activation function layer, and a second batch normalization layer, the first batch normalization layer and the activation function layer are between the first convolutional layer and the second convolutional layer, and the second batch normalization layer is after the second convolutional layer.

[0159] Example 18 provides an apparatus, including a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing  computer program instructions executable by the computer processor to perform operations including identifying an OFM generated in a layer of the first neural network, generating an expanded feature map from the OFM by expanding the OFM in a channel dimension of the OFM, generating, using a generative network, a first feature map from the expanded feature map, the first feature map having more channels than the OFM and including a plurality of feature segments, identifying a second feature map generated in a layer of a second neural network, the second feature map having a same number of channels as a feature segment in the first feature map, and adjusting one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the second feature map.

[0160] Example 19 provides the apparatus of example 18, in which generating the expanded feature map from the OFM includes inputting the OFM into a linear layer including a linear transformation operation, the linear layer outputting the expanded feature map.

[0161] Example 20 provides the apparatus of example 19, in which the operations further include adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; and after adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, in which the next layer is after the layer of the first neural network.

[0162] The above description of illustrated implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. These modifications may be made to the disclosure in light of the above detailed description.

Claims

1.A method for training a first neural network, the method comprising:identifying an output feature map generated in a layer of the first neural network;generating an expanded feature map from the output feature map by expanding the output feature map in a channel dimension of the output feature map;generating, using a generative network, a first feature map from the expanded feature map, the first feature map having more channels than the output feature map and comprising a plurality of feature segments;identifying a second feature map generated in a layer of a second neural network, the second feature map having a same number of channels as a feature segment in the first feature map; andadjusting one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the second feature map.2.The method of claim 1, wherein generating the expanded feature map from the output feature map comprises:inputting the output feature map into a linear layer comprising a linear transformation operation, the linear layer outputting the expanded feature map.3.The method of claim 2, further comprising:adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; andafter adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, wherein the next layer is after the layer of the first neural network.4.The method of claim 3, wherein the next layer is a fully-connected layer.5.The method of claim 1, further comprising:generating a contracted feature map from the expanded feature map,wherein the contracted feature map has a same number of channels as the output feature map and is input into a next layer in the first neural network, and the next layer is after the layer of the first neural network.6.The method of claim 5, wherein generating the contracted feature map from the expanded feature map comprises:inputting the expanded feature map into a linear layer comprising a linear transformation operation, the linear layer outputting the contracted feature map.7.The method of claim 6, further comprising:adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; andafter adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, wherein the next layer is after the layer of the first neural network.8.The method of claim 1, wherein generating the first feature map comprising:applying a mask on the expanded feature map to convert data elements in a first portion of the expanded feature map to zeros;inputting a second portion of the expanded feature map into the generative network, the generative network outputting an intermediate feature map, the intermediate feature map having a same number of channels as the first portion of the expanded feature map; andgenerating the first feature map by combining the intermediate feature map with the second portion of the expanded feature map.9.The method of claim 1, wherein the generative network comprises a first convolutional layer and a second convolutional layer.10.The method of claim 9, wherein:the generative network further comprises a first batch normalization layer, an activation function layer, and a second batch normalization layer,the first batch normalization layer and the activation function layer are between the first convolutional layer and the second convolutional layer, andthe second batch normalization layer is after the second convolutional layer.11.One or more non-transitory computer-readable media storing instructions executable to perform operations for training a first neural network, the operations comprising:identifying an output feature map generated in a layer of the first neural network;generating an expanded feature map from the output feature map by expanding the output feature map in a channel dimension of the output feature map;generating, using a generative network, a first feature map from the expanded feature map, the first feature map having more channels than the output feature map and comprising a plurality of feature segments;identifying a second feature map generated in a layer of a second neural network, the second feature map having a same number of channels as a feature segment in the first feature map; andadjusting one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the second feature map.12.The one or more non-transitory computer-readable media of claim 11, wherein generating the expanded feature map from the output feature map comprises:inputting the output feature map into a linear layer comprising a linear transformation operation, the linear layer outputting the expanded feature map.13.The one or more non-transitory computer-readable media of claim 12, wherein the operations further comprise:adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; andafter adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, wherein the next layer is after the layer of the first neural network.14.The one or more non-transitory computer-readable media of claim 11, wherein the  operations further comprise:generating a contracted feature map from the expanded feature map by inputting the expanded feature map into a linear layer comprising a linear transformation operation, the linear layer outputting the contracted feature map,wherein the contracted feature map has a same number of channels as the output feature map and is input into a next layer in the first neural network, and the next layer is after the layer of the first neural network.15.The one or more non-transitory computer-readable media of claim 14, wherein the operations further comprise:adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; andafter adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, wherein the next layer is after the layer of the first neural network.16.The one or more non-transitory computer-readable media of claim 11, wherein generating the first feature map comprising:applying a mask on the expanded feature map to convert data elements in a first portion of the expanded feature map to zeros;inputting a second portion of the expanded feature map into the generative network, the generative network outputting an intermediate feature map, the intermediate feature map having a same number of channels as the first portion of the expanded feature map; andgenerating the first feature map by combining the intermediate feature map with the second portion of the expanded feature map.17.The one or more non-transitory computer-readable media of claim 11, wherein:the generative network comprises a first convolutional layer, a second convolutional layer, a first batch normalization layer, an activation function layer, and a second batch normalization layer,the first batch normalization layer and the activation function layer are between the first convolutional layer and the second convolutional layer, andthe second batch normalization layer is after the second convolutional layer.18.An apparatus, comprising:a computer processor for executing computer program instructions; anda non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising:identifying an output feature map generated in a layer of the first neural network,generating an expanded feature map from the output feature map by expanding the output feature map in a channel dimension of the output feature map,generating, using a generative network, a first feature map from the expanded feature map, the first feature map having more channels than the output feature map and comprising a plurality of feature segments,identifying a second feature map generated in a layer of a second neural network, the second feature map having a same number of channels as a feature segment in the first feature map, andadjusting one or more parameters in the first neural network based on a plurality of feature distillation operations, a feature distillation operation performed on one feature segment of the plurality of feature segments and the second feature map.19.The apparatus of claim 18, wherein generating the expanded feature map from the output feature map comprises:inputting the output feature map into a linear layer comprising a linear transformation operation, the linear layer outputting the expanded feature map.20.The apparatus of claim 19, wherein the operations further comprise:adjusting one or more parameters in the linear layer based on the plurality of feature distillation operations; andafter adjusting the one or more parameters in the linear layer, merging the linear layer into a next layer in the first neural network, wherein the next layer is after the layer of the first neural network.

Citation Information

Patent Citations

  • Industrial image defect detection method based on knowledge distillation

    CN114663392A

  • Neural network training method and related device

    CN115640831A

  • Multi-scale distillation for low-resolution detection

    US20230153943A1