Methods and devices for performing many-to-one feature distillation in neural networks

The knowledge distillation architecture addresses the challenges of high computational requirements and latency in training student networks by simultaneously training teacher and student networks on edge devices using a many-to-one feature distillation method, resulting in improved efficiency and accuracy.

DE112022007693T5Pending Publication Date: 2025-06-26INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112022007693
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current neural network architectures face challenges in efficiently training student networks on edge devices due to high computational requirements and latency issues, particularly in the context of knowledge distillation processes.

Method used

The proposed solution involves a knowledge distillation architecture that trains both the teacher and student networks simultaneously on an edge device, utilizing a many-to-one feature distillation method to reduce computational demands and latency.

Benefits of technology

This approach reduces the computational burden on edge devices, decreases latency, and enhances the accuracy of the knowledge distillation process by enabling simultaneous training of teacher and student networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods, apparatus, systems, and articles of manufacture are disclosed for training a teacher network and a student network simultaneously without a pre-trained teacher model. An apparatus comprising at least one memory, machine-readable instructions, and processor circuitry to at least one of instantiate and execute the machine-readable instructions to generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; generate a second feature map for a student network based on the query, the second feature map having a second channel dimension, the second channel dimension being different from the first channel dimension; divide the second feature map into segments having the first channel dimension;training the teacher network using the segments of the second feature map and training the student network by applying a total loss value to the student network, the total loss value being based on a loss function, wherein the teacher network and the student network are trained on an edge device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELDThis disclosure relates generally to neural networks, and more particularly to methods and apparatus for performing a many-to-one feature distillation in neural networks.BACKGROUNDNeural networks are a subset of machine learning and form the heart of deep learning algorithms. Neural networks process input data and generate an output based on an analysis of the input data using a training model to generate the best (e.g., most accurate) response to a given input. Training models are constantly retraining to generate more accurate responses and optimize the neural network.BRIEF DESCRIPTION OF THE DRAWINGSFIG. 1 is a schematic illustration of a knowledge distillation architecture. FIG. 2 is a block diagram of an example knowledge distillation circuit of FIG. 1. FIG. 3 is a block diagram of the example teacher training circuit and student training circuit of FIG. 2. FIG. 4 is an operational illustration of the knowledge distillation architecture of FIG. 1. FIG. 5 is a flow diagram illustrating example machine readable instructions and / or example operations executable by example processor circuitry for implementing the knowledge distillation architecture of FIG. 1. FIG. 6 is a flow diagram illustrating example machine readable instructions and / or example operations executable by example processor circuitry for performing the knowledge distillation of FIG. 5. FIG. 7 is a flow diagram illustrating example machine readable instructions and / or example operations executable by example processor circuitry for implementing training of the teacher and student network of FIG. 6. FIG. 8 is a flowchart illustrating example machine readable instructions and / or example operations executable by an example overall loss generation processor circuit of FIG. 6. FIG. 9 is a block diagram of an example processing platform including processor circuitry structured to execute the example machine readable instructions and / or example operations of FIG. 3 to implement the knowledge distillation circuitry of FIG. 2. FIG. 10 is a block diagram of an example implementation of the processor circuit of FIG. 9. FIG. 11 is a block diagram of another example implementation of the processor circuit of FIG. 9. FIG. 12 is a block diagram of an example software distribution platform (e.g., one or more servers) for distributing software (e.g., software corresponding to the example machine readable instructions of FIGS. 5, 6, 7, and / or 8) to client devices, end users and / or consumers (e.g., for licenseing, sale, and / or use), retailers (e.g., for sale, resale, license, and / or underlication), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products, which are assigned, for example, to retailers and / or to other end users such as direct customers).In general, throughout the one or more drawings and the accompanying written description, the same reference numerals are used to refer to the same or similar parts. The figures are not to scale. Rather, the thickness of the layers or regions in the drawings may be increased. Although the figures show layers and regions with clear lines and boundaries, some or all of these lines and / or boundaries may be idealized. In reality, the boundaries and / or lines may not be perceptible, overlaid, and / or irregular.Unless expressly stated otherwise, descriptors such as "first / r / s", "second / r / s", "third / r / s", etc. are used herein without in any way placed under or otherwise indicating any meaning of priority, physical order, arrangement in a list and / or order, but are used merely as labels and / or arbitrary names to distinguish elements for ease of understanding of the disclosed examples. In some examples, the term "first" may be used to refer to an element in the detailed description, while the same element in a claim may be referred to with a different term, such as "second" or "third.". In such cases, it will be appreciated that such descriptors are used merely to uniquely identify those elements that might otherwise share a same name, for example.As used herein, "about" and "about" modify their subjects / values to account for the potential presence of variations that occur in real applications. For example, "about" and "about" may modify dimensions that may not be exact due to manufacturing tolerances and / or other real imperfections present, as understood by those of ordinary skill in the art. For example, "about" and "about" may indicate that such dimensions may be within a tolerance range of + / - 10% unless otherwise indicated in the description below. As used herein, "substantially real-time" refers to near instantaneous occurrence, recognizing that there may be delays in real world for computing time, transmission, etc. Thus, unless otherwise stated, "substantially real time" refers to real time + / - 1 second.As used herein, the term "in communication," including variations thereof, encompasses direct communication and / or indirect communication through one or more intermediary components, and does not require direct physical (e.g., wired) communication and / or constant communication, but additionally encompasses selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.As used herein, the term "processor circuit" is defined to include: (i) one or more special purpose electrical circuits structured to perform particular operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors); and / or (ii) one or more general purpose electrical semiconductor-based circuits programmed with instructions to perform particular operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuits include programmable microprocessors, field programmable gate arrays (FPGAs) that can instantiate instructions, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), XPUs, or microcontrollers, and integrated circuits such as application specific integrated circuits (ASICs). For example, an XPU may be implemented by a heterogeneous system that includes multiple types of processor circuitry (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc., and / or a combination thereof) and application programming interface(s) (API(s)) that may / may assign the computing task(s) to the / those of the multiple types of processor circuitry that is / are best suited for executing the computing task(s).DETAILED DESCRIPTION. . Neural networks implementing knowledge distillation typically train a student network based on a pre-trained teacher network. Feature maps of the teacher network are created to be used as knowledge points for the student network for learning. This methodology is known as two-step knowledge distillation. Alternatively, some neural networks utilize an online training framework that does not require pretraining the teacher model. This methodology is known as single stage knowledge distillation.Knowledge distillation typically requires a pre-trained teacher network in order to be able to train a student network. The teacher network is typically trained with a machine having high computing resources and / or capabilities (e.g., high-performance CPUs, GPUs, RAM, etc.). Such computing resources are typically not implemented in local devices (e.g., edge devices, user devices, etc.) to train the teacher network due to the required computing power.Since student networks are typically operated on edge devices and do not have the necessary computing power to train the teacher network, the student network must rely on the pre-trained teacher network to learn. This is undesirable for edge devices because they must retrieve data from the teacher network feature maps to determine an output (e.g., a response to a question, a result, etc.), and the neural network must train the teacher model in advance before the student network can be trained.To train the student network, neural networks generate feature maps from the teacher network, and transformation modules are generated to adapt the feature maps to the student network. Once the teacher network is trained, the transformation modules are discarded and each feature from the feature map is associated with the student network in a one-to-one representation. Latency in neural networks is increased by this process because the steps are required to train the teacher network and create the one-to-one representations. Such a process is known as two-stage knowledge distillation.Alternatively, in the single stage knowledge distillation, the neural network retrieves data from an online teacher network, which is a network without any training. Because the neural network does not train the teacher model itself, the computational requirements at the edge device level are reduced. However, latency continues to be increased because data (e.g., feature maps, feature representations, etc.) must be retrieved from online sources (e.g., from sources external to the device performing training of the student network). In addition, these processes still use a one-to-one representation of the features in the teacher network feature map and are less accurate than the two-stage knowledge distillation because the neural network does not train the teacher network.Therefore, there is a need for an online knowledge distillation process (e.g., single-stage knowledge distillation) that trains the teacher network and the student network simultaneously while reducing computing requirements at the edge device level, reducing latency, and simultaneously increasing accuracy.FIG. 1 is a schematic illustration of a knowledge distillation architecture 100. The knowledge distillation architecture 100 of the illustrated example of FIG. 1 includes an edge device 110 and a data host 120. In some examples, the edge device 110 accesses data from the data host 120 via a network 125. In other examples, the edge device 110 accesses the data from the data host 120 by direct communication with the data host 120 (e.g., Ethernet, Universal Serial Bus (USB), Peripheral Component Interconnect (PCI), etc.).The edge device 110 includes a data fetch circuit 130, a knowledge distillation circuit 140, and an output circuit 150. In some examples, the edge device 110 includes an edge memory 160 to store data directly on the edge device 110.The data fetch circuit 130 communicates with the data host 120 to fetch data to train a teacher and student network (e.g., a teacher network 410 and a student network 420). In some examples, the data fetch circuit 130 communicates with the data host 120 via the network 125. In other examples, data fetch circuit 130 directly accesses data from data host 120. In some examples, the data fetcher circuit 130 provides a data source before the data is fetched. In such an example, the data fetch circuit 130 determines whether to fetch data from the data host 120 via the network 125 (e.g., to fetch data from online sources, big data sources, etc.) or to fetch data from local storage devices (e.g., a memory 180 on the data host 120).The knowledge distillation circuit 140 performs a knowledge distillation process to train the teacher network 410 and the student network 420 from the data retrieved from the data retrieval circuit 130. In the examples described herein, the knowledge distillation circuit 140 communicates with the data fetch circuit 130. However, the edge device 110 could omit the data fetch circuit 130, and the knowledge distillation circuit 140 may communicate directly with the data host 120.The output circuit 150 outputs a result of the trained teacher and student networks 410, 420. In some examples, the result could include a response to a query (such as a natural language response), an image, an audio clip, a command, etc. In some examples, the output circuit 150 communicates the result to the data host 120 via the network 125. In other examples, the output circuit 150 stores the result in the data host 120 by communicating directly with the data host 120. In some examples, the output circuit 150 directly stores the result in the edge memory 160.The data host 120 includes external data 170. In some examples, the external data 170 includes data from big data sources and may be retrieved by the data host 120 when the data retrieval circuit 130 has retrieved the data from the data host 120 via the network 125. In some examples, the data host 120 includes the memory 180 that contains data that can be retrieved / accessed by the data fetcher circuit 130 over the network 125 and / or directly.FIG. 2 is a block diagram of an example knowledge distillation circuit 140 for training teacher network 410 and student network 420. The knowledge distillation circuit 140 of FIG. 2 may be instantiated by a processor circuit, such as a central processing unit executing instructions (e.g., creating an instance, realizing for any amount of time, realizing, implementing, etc.). Additionally or alternatively, the knowledge distillation circuit 140 of FIG. 2 may be instantiated by an ASIC or an FPGA structured to perform operations according to the instructions (e.g., generating an instance, realizing for any amount of time, realizing, implementing, etc.). It will be appreciated that some or all of the circuits of FIG. 2 may thus be instantiated at the same time or at different times. For example, a portion or all of the circuitry may be instantiated in one or more threads executing concurrently on hardware and / or in series on hardware. Moreover, in some examples, a portion or all of the circuit of FIG. 2 may be implemented by microprocessor circuitry executing instructions to implement one or more virtual machines and / or containers.The knowledge distillation circuit 140 includes a feature map circuit 220, a teacher training circuit 230, a student training circuit 240, a loss circuit 250, and a result generation circuit 260.The feature map circuit 220 generates teacher and student feature maps (e.g., a teacher feature map 440 and a student feature map 450) from the data retrieved by the data retrieval circuit 130 based on the teacher network 410 and the student network 420, respectively. In some examples, the feature maps 440, 450 are generated based on a query received by the data fetch circuit 130. In some examples, the feature map circuit 220 communicates with the data fetch circuit 130, the output circuit 150, and / or the edge memory 160 via an I / O interface 210. In some examples, the output circuit 220 is instantiated by a processor circuit executing output instructions and / or configured to perform operations such as those illustrated by the flowchart of FIG. 6.In some examples, the knowledge distillation circuit 140 includes means for generating teacher and student feature maps 440, 450. For example, the means for generating may be implemented by the feature map circuit 220. In some examples, the feature map circuit 220 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the feature map circuit 220 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in blocks 610, 620, and 630 of FIG. 6. In some examples, the feature map circuit 220 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the feature map circuit 220 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the feature map circuit 220 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.Teacher training circuit 230 trains teacher network 410. In some examples, teacher training circuit 230 communicates with student training circuit 240 while teacher network 410 is being trained. In some examples, teacher training circuit 230 is instantiated by processor circuitry executing teacher training instructions and / or configured to perform operations as illustrated by the flowchart of FIGS. 6 and / or 7.In some examples, the knowledge distillation circuit 140 includes means for training the teacher network 410. For example, the means for training teacher network 410 may be implemented by teacher training circuit 230. In some examples, teacher training circuit 230 may be instantiated by processor circuitry, such as example processor circuitry 912 of FIG. 9. For example, teacher training circuit 230 may be instantiated by example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in blocks 640 of FIG. 6 and / or 730 and 740 of FIG. 7. In some examples, teacher training circuit 230 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, teacher training circuit 230 may be instantiated by any other combination of hardware, software, and / or firmware. For example, teacher training circuit 230 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without execution of software or firmware, but other structures are also suitable.The student training circuit 240 trains the student network 420. In some examples, student training circuit 240 communicates with teacher training circuit 230 while training student network 420. In some examples, the student training circuit 240 is instantiated by processor circuitry executing student training instructions and / or configured to perform operations as illustrated by the flowchart of FIGS. 6 and / or 7.In some examples, the knowledge distillation circuit 140 includes means for training the student network 420. For example, the means for training the student network 420 may be implemented by the student training circuit 240. In some examples, the student training circuit 240 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the student training circuit 240 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in blocks 640 of FIG. 6 and / or 710, 720, 750, and 760 of FIG. 7. In some examples, the student training circuit 240 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, student training circuit 240 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the student training circuit 240 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.Loss circuit 250 calculates a total loss value of teacher and student networks 410, 420. In some examples, loss circuit 250 is instantiated by a processor circuit executing loss instructions and / or configured to perform operations such as those illustrated by the flowchart of FIGS. 6 and / or 8.In some examples, the knowledge distillation circuit 140 includes means for calculating a total loss value. For example, the means for computing may be implemented by the loss circuit 250. In some examples, loss circuit 250 may be instantiated by processor circuitry, such as example processor circuitry 912 of FIG. 9. For example, loss circuit 250 may be instantiated by example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in blocks 650 of FIG. 6 and / or 810, 820, 830, 840, and / or 850 of FIG. 8. In some examples, loss circuit 250 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, loss circuit 250 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the loss circuit 250 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.The result generator circuit 260 generates a result based on the training of the teacher and student networks 410, 420. In some examples, the result generator circuit 260 communicates with the data fetch circuit 130, the output circuit 150, and / or the edge memory 160 via the I / O interface 210. In some examples, the result generator circuit 260 is instantiated by a processor circuit executing result instructions and / or configured to perform operations like those illustrated by the flowchart of FIG. 6.In some examples, the knowledge distillation circuit 140 includes means for generating a result. For example, the means for generating may be implemented by the result generator circuit 260. In some examples, the result generator circuit 260 may be instantiated by processor circuitry, such as the example processor circuit 912 of FIG. 9. For example, the result generator circuit 260 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in the block 660 of FIG. 6. In some examples, the result generator circuit 260 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the result generator circuit 260 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the result generator circuit 260 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.FIG. 3 is a block diagram of the example teacher training circuit 230 and the example student training circuit 240 for training the teacher network 410 and the student network 420, respectively. Teacher training circuit 230 and student training circuit 240 of FIG. 3 may be instantiated by processor circuitry, such as a central processing unit executing instructions (e.g., creating an instance, realizing for any amount of time, realizing, implementing, etc.). Additionally or alternatively, teacher training circuit 230 and student training circuit 240 of FIG. 3 may be instantiated by an ASIC or an FPGA structured to perform operations corresponding to the instructions (e.g., creating an instance, realizing for any amount of time, realizing, implementing, etc.). It will be appreciated that some or all of the circuits of FIG. 3 may thus be instantiated at the same time or at different times. For example, a portion or all of the circuitry may be instantiated in one or more threads executing concurrently on hardware and / or in series on hardware. Moreover, in some examples, a portion or all of the circuit of FIG. 3 may be implemented by microprocessor circuitry executing instructions to implement one or more virtual machines and / or containers.Teacher training circuit 230 and student training circuit 240 include feature map projection circuit 310, feature division circuit 320, matching circuit 330, reparametrization circuit 340, and matching application circuit 350. In some examples, student training circuit 240 includes feature map projection circuit 310, feature splitting circuit 320, and matching application circuit 350. In some examples, teacher training circuit 230 includes matching circuit 330 and reparametrizing circuit 340.The feature map projection circuit 310 projects the student feature map 450 to a desired channel dimension. In some examples, the feature map projection circuit 310 projects the student feature map 450 into a predefined channel dimension prior to training the student network 420, where the predefined channel dimension is used for a many-to-one representation of the teacher network 410. In some examples, the feature map projection circuit 310 projects the student feature map 450 back to its original channel dimension after the student network 420 has been trained. In some examples, the feature map projection circuit 310 is instantiated by a processor circuit executing feature map projection instructions and / or configured to perform operations as illustrated by the flowchart of FIG. 7.In some examples, the student training circuit 240 includes means for projecting the student feature map 450 to a desired channel dimension. For example, the means for projecting may be implemented by the feature map projection circuit 310. In some examples, the feature map projection circuit 310 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the feature map projection circuit 310 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in blocks 710 and 760 of FIG. 7. In some examples, the feature map projection circuit 310 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the feature map projection circuit 310 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the feature map projection circuit 310 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.The feature splitting circuit 320 splits the student feature map 450 into segments having the same channel dimension as the teacher feature map 440. In some examples, the feature splitting circuit 320 is instantiated by a processor circuit executing feature splitting instructions and / or configured to perform operations as illustrated by the flowchart of FIG. 7.In some examples, the student training circuit 240 includes means for dividing the student feature map 450 into segments having the same channel dimension as the teacher feature map 440. For example, the means for partitioning may be implemented by the feature partitioning circuit 320. In some examples, the feature splitting circuit 320 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the feature splitting circuit 320 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in the block 720 of FIG. 7. In some examples, the feature splitting circuit 320 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the feature separation circuit 320 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the feature splitting circuit 320 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.Matching circuit 330 matches segments of student feature map 450 generated by feature splitting circuit 320 with teacher feature map 440. In some examples, the alignment circuit 330 is instantiated by a processor circuit executing alignment instructions and / or configured to perform operations such as those illustrated by the flowchart of FIG. 7.In some examples, teacher training circuit 230 includes means for matching segments of student feature map 450 with teacher feature map 440. For example, the means for matching may be implemented by the matching circuit 330. In some examples, the alignment circuit 330 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the alignment circuit 330 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in the block 730 of FIG. 7. In some examples, the alignment circuit 330 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the alignment circuit 330 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the alignment circuit 330 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.The reparametrizing circuit 340 reparameterizes the teacher network 410 based on a combination (e.g., a summation) of the segments of the student feature map 450. In some examples, the reparametrizing circuit 340 is instantiated by a processor circuit executing reparametrizing instructions and / or configured to perform operations as illustrated by the flowchart of FIG. 7.In some examples, teacher training circuit 230 includes means for reparametry teacher network 410 based on a combination of the segments of student feature map 450. For example, the means for reparametry may be implemented by the reparametry circuit 340. In some examples, the reparametrizing circuit 340 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the reparametrizing circuit 340 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in the block 740 of FIG. 7. In some examples, the reparametrizing circuit 340 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the reparametrizing circuit 340 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the reparametrization circuit 340 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.The matching application circuit 350 applies a result of matching the segments of the student feature map 450 with the teacher network 410 to the student network 420. In some examples, the matching application circuit 350 is instantiated by a processor circuit executing matching application instructions and / or configured to perform operations as illustrated by the flowchart of FIG. 7.In some examples, the student training circuit 240 includes means for applying the matching of the segments of the student feature map 450 with the teacher feature map 410 to the student network 420. For example, the means for applying the matching may be implemented by the matching application circuit 350. In some examples, the alignment application circuit 350 may be instantiated by processor circuitry, such as the example processor circuitry 912 of FIG. 9. For example, the alignment application circuit 350 may be instantiated by the example microprocessor 1000 of FIG. 10 executing machine-executable instructions as implemented at least in the block 750 of FIG. 7. In some examples, the alignment application circuit 350 may be instantiated by hardware logic circuitry, which may be implemented by an ASIC, an XPU, or the FPGA circuit 1100 of FIG. 11 structured to perform operations corresponding to the machine readable instructions. Additionally or alternatively, the matching application circuit 350 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the alignment application circuit 350 may be implemented by at least one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.) structured to execute some or all of the machine readable instructions and / or perform some or all of the operations corresponding to the machine readable instructions without execution of software or firmware, but other structures are also suitable.In some examples, the reparametrizing circuit 240 and the matching application circuit 250 are executed in parallel to simultaneously train the teacher network 410 and the student network 420. In other examples, the reparametrizing circuit 240 and the matching application circuit 250 are executed in series. The order in which the reparametrizing circuit 240 and the matching application circuit 250 are executed is interchangeable and / or modifiable. In other words, the execution order of both the reparametrizing circuit 240 and the matching application circuit 250 is not limited to the specific examples described herein.While an example manner of implementing the knowledge distillation circuit 140 of FIG. 1 is illustrated in FIGS. 2 and / or 3, one or more of the elements, processes, and / or devices illustrated in FIGS. 2 and / or 3 may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other manner. Further, the example feature map circuit 220, the example teacher training circuit 230, the example student training circuit 240, the example loss circuit 250, the example result generator circuit 260, the example feature map projection circuit 310, the example feature partitioning circuit 320, the example matching circuit 330, the example reparametrization circuit 340, the example matching application circuit 350, and / or, more generally, the example knowledge distillation circuit 140 of FIG. 1 may be implemented by hardware alone or by hardware in combination with software and / or firmware. Thus, for example, any of the example feature map circuit 220, the example teacher training circuit 230, the example student training circuit 240, the example loss circuit 250, the example result generator circuit 260, the example feature map projection circuit 310, the example feature partitioning circuit 320, the example matching circuit 330, the example reparametrizing circuit 340, the example matching application circuit 350, and / or, more generally, the example knowledge distillation circuit 140 may be processed circuits, (an) analog circuit(s), (a) digital circuit(s), (a) logic circuit(s), (a) programmable processor(s), (a) programmable microcontroller(s), (a) graphics processing unit(s) (GPU(s)), (a) digital signal processor(s) (DSP(s)), (An) Application Specific Integrated Circuit(s) (ASIC(s)), (Programmable Logic Device(s) (PLD(s)), and / or (Field Programmable Logic Device(s) (FPLD(s)) such as Field Programmable Gate Arrays (FPGAs). Further, the example knowledge distillation circuit 140 of FIG. 1 may include one or more elements, processes, and / or devices in addition to or in place of those illustrated in FIGS. 2 and / or 3, and / or may include more than one or all of the illustrated elements, processes, and devices.Flowcharts illustrating example machine readable instructions executable to configure processor circuitry to implement the knowledge distillation circuitry 140 of FIG. 2 are shown in FIGS. 5, 6, 7, and / or 8. The machine readable instructions may be one or more executable programs or portions of an executable program for execution by the processor circuitry, such as the processor circuitry 912 shown in the example processor platform 900 discussed below in connection with FIG. 9 and / or the example processor circuitry discussed below in connection with FIGS. 10 and / or 11. The program may be embodied in software stored on one or more non-transitory computer readable storage media, such as a compact disk (CD), a floppy disk, a hard disk drive (HDD), a solid state drive (SSD), a digital versatile disk (DVD), a blu-ray disk, a volatile memory (e.g., random access memory (RAM) of any type, etc.), or a non-volatile memory (e.g., electrically erasable programmable read only memory (EEPROM), flash memory, a HDD, an SSD, etc.), associated with a processor circuit located in one or more hardware devices; Alternatively, however, the entire program and / or portions thereof may be executed by one or more hardware devices other than the processor circuit and / or embodied in firmware or dedicated hardware. The machine readable instructions may be distributed among multiple hardware devices and / or executed by two or more hardware devices (e.g., a server and a client hardware device). For example, the client hardware device may be implemented by an endpoint client hardware device (e.g., a hardware device associated with a user) or a gateway of an intermediate client hardware device (e.g., a radio access network (RAN)), which may facilitate communication between a server and an endpoint client hardware device. Similarly, the non-transitory computer readable storage media may include one or more media residing in one or more hardware devices. Further, although the example program is described with reference to the one or more flowcharts illustrated in FIGS. 5, 6, 7, and / or 8, many other methods for implementing the example knowledge distillation circuit 140 may alternatively be used. For example, the order of execution of the blocks may be changed and / or some of the blocks described may be changed, removed, or combined. Additionally or alternatively, some or all of the blocks may be implemented by one or more hardware circuits (e.g., a processor circuit, a discrete and / or integrated analog and / or digital circuit, an FPGA, an ASIC, a comparator, an operational amplifier (op-amp), or a logic circuit, etc.) structured to perform the corresponding operation without executing software or firmware. The processor circuitry may be distributed to different network locations and / or locally to one or more hardware devices (e.g., a single core processor (e.g., a single core central processing unit CPU)), a multi-core processor (e.g., a multi-core CPU, an XPU, etc.) in a single machine, multiple processors distributed across multiple servers of a server rack, multiple processors distributed across one or more server racks, a CPU and / or an FPGA placed in the same package (e.g., the same integrated circuit (IC) package, or in two or more separate packages, etc.).The machine readable instructions described herein may be stored in a compressed format and / or an encrypted format and / or a fragmented format and / or a compiled format and / or an executable format and / or a packaged format, etc. Machine-readable instructions as described herein may be stored as data or data structure (e.g., as portions of instructions, code, representations of code, etc.) that may be used to generate, produce, and / or produce machine-executable instructions. For example, the machine readable instructions may be fragmented and stored on one or more storage devices and / or computing devices (e.g., servers) located at the same or different locations of a network or collection of networks (e.g., in the cloud, edge devices, etc.). The machine readable instructions may require one or more of the following operations: installation, modification, adjustment, updating, combining, complementing, configuring, decryption, decompression, unpacking, distribution, reallocation, compilation, etc., to render them directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, the machine-readable instructions may be stored in multiple portions that are individually compressed, encrypted, and / or stored on separate computing devices, where the portions, when decrypted, decompressed, and / or combined, form a set of machine-executable instructions that implement one or more operations that together may form a program such as that described herein.In another example, the machine readable instructions may be stored in a state where they may be read by a processor circuit, but may require addition to a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., to execute the machine readable instructions on a particular computing device or other device. In another example, the machine-readable instructions may need to be configured (e.g., settings stored, data input, network addresses recorded, etc.) before the machine-readable instructions and / or the one or more corresponding programs may be executed in whole or in part. Thus, machine-readable media as used herein may include machine-readable instructions and / or one or more programs, regardless of the particular format or state of the machine-readable instructions and / or the one or more programs when stored or otherwise in the sleep or transition state.The machine readable instructions described herein may be represented by any previous, current, or future instruction language, scripting language, programming language, etc. For example, the machine readable instructions may be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.As mentioned above, the example operations of FIGS. 5, 6, 7, and / or 8 may be implemented using executable instructions (e.g., computer and / or machine readable instructions) stored on one or more non-transitory computer and / or machine readable media, such as optical storage devices, magnetic storage devices, an HDD, a flash memory, a read only memory (ROM), a CD, a DVD, a cache, a RAM of any type, a register, and / or any other storage device or storage disk on which information is stored for any duration (e.g., for extended periods of time, permanently, for brief instances, for temporarily buffering, and / or buffering the information). As used herein, the terms non-transitory computer readable medium, non-transitory computer readable storage medium, non-transitory machine readable medium, and / or non-transitory machine readable storage medium are expressly defined to include any type of computer readable storage device and / or storage disk and to exclude propagating signals and to exclude transmission media. The terms "computer readable storage medium" and "machine readable storage medium" are defined herein to include any physical (mechanical and / or electrical) structure for storing information, but not the transmission of signals and transmission media. Examples of computer readable storage devices and / or machine readable storage devices include random access memory (RAM) of any type, read only memory (ROM) of any type, solid state memory, flash memory, optical disks, magnetic disks, disk drives, and / or redundant array of independent disks (RAID) systems. As used herein, the term "device" refers to a tangible structure such as mechanical and / or electrical equipment, hardware, and / or circuitry that may or may not be configured by computer readable instructions, machine readable instructions, etc. and / or manufactured to execute computer readable instructions, machine readable instructions, etc."Including" and "comprising" (and all forms and time forms thereof) are used herein as open terms. Thus, when any form of "include" or "comprise" (e.g., comprises, includes, comprises, including, having, etc.) is used in a claim as a preamble or in any claim formulation, it is understood that additional elements, labels, etc. may be present without departing from the scope of the corresponding claim or formulation. Likewise, when the phrase "at least" ("at least") as used herein is used as a transitional phrase in, for example, a preamble of a claim, it is an open phrase such as the terms "comprising" and "including" are open terms. The term "and / or," when used in a form such as A, B, and / or C, for example, refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A with B, (5) A with C, (6) B with C, or (7) A with B, and with C. As used herein in connection with describing structures, components, objects, and / or things, the term "at least one of A and B" is intended to refer to implementations including any of (1) at least one of A, (2) at least one of B, or (3) at least one of A and at least one of B. Similarly, the term "at least one of A or B" as used herein in the context of describing structures, components, objects, and / or things, is intended to refer to implementations including any of (1) at least one of A, (2) at least one of B, or (3) at least one of A and at least one of B. As used herein in connection with describing the performance or execution of processes, instructions, actions, activities, and / or steps, the term "at least one of A and B" is intended to refer to implementations including any of (1) at least one of A, (2) at least one of B, or (3) at least one of A and at least one of B. Likewise, the term "at least one of A or B," as used herein in the context of describing the performance or execution of processes, instructions, actions, activities, and / or steps, is intended to refer to implementations including any of (1) at least one of A, (2) at least one of B, or (3) at least one of A and at least one of B.As used herein, references in the singular (e.g., "a", "an", "a", "first", "second", etc.) do not exclude a plurality. The term "a" object, as used herein, refers to one or more of that object. The terms "a / e", "one / e or more", and "at least one / e" are used interchangeably herein. Moreover, a plurality of means, elements or method actions, even if individually listed, may be performed, for example, by the same entity or object. Additionally, although individual features may be included in different examples or claims, these may possibly be combined, and inclusion in different examples or claims does not mean that a combination of features is not feasible and / or advantageous.FIG. 4 is an operational representation of the knowledge distillation architecture 100 of FIG. 1. The example illustrated operation 400 includes the data host 120, the loss circuit 250, the teacher network 410, the student network 420, and a plurality of convolutional layers 430. Teacher network 410 includes teacher feature map 440. The student network 420 includes the student feature map 450. The illustrated operation 400 may be inserted at any convolution layer of the plurality of convolution layers 430. Depending on the neural network, the illustrated operation 400 of FIG. 4 may be inserted at more than one convolutional layer.Teacher feature map 440 (also shown as F t) includes teacher channel dimension C t. The student feature map 450 (also shown as F s) includes a student channel dimension C s. In examples C t described herein, the teacher channel dimension is different from the student channel dimension C s. Therefore, the student feature map 450 should be projected onto a channel dimension that is equivalent to the teacher channel dimension C t.The student network 420 is used by a Schülerkanaldimensionserweiterer 455 (also referred to as W se) to project the student feature map 450 into a channel dimension that is the same as the teacher channel dimension C t. The student network 420 then divides the projected student feature map into an extended student feature map 460 (also referred to as F se) having N non-overlapping segments. Therefore, the extended student feature map 460 is represented as N x C t because each non-overlapping segment N has the same channel dimension as the teacher feature map 440.Teacher network 410 is used to perform a many-to-one feature representation using the N non-overlapping segments of extended student feature map 460 by matching features on extended student feature map 460 with features on teacher feature map 440.The teacher network 410 is used by a reparametrizer 470 to reparameterize the teacher feature map 440, wherein the reparametrizer 470 uses a combination of the N non-overlapping segments of the extended student feature map 460. In some examples, the combination may be a summation of the N non-overlapping segments. In other examples, the combination may be a difference, a multiplication, a division, etcThe student network 420 is used to apply a result of matching the N non-overlapping segments to the enhanced student feature map 460. As disclosed above, the reparametrization of teacher feature map 440 may be performed in parallel using reparametrizer 470 and applying the result of matching the N non-overlapping segments.The student network 420 is used by the student channel dimension assembler 475 (also referred to as W sc) to create an output student feature map. Such generation of the output student feature may be accomplished by projecting the extended student feature map 560 back to its original channel dimension C s. The output student feature map, also referred to as F so, is then weighted to represent a fully connected student network layer 480, also referred to as (e.g., the convolutional layer has been fully learned).Similarly, teacher network 410 is used to weight the reparametrized teacher feature map, also referred to as F tr to represent a fully connected teacher network layer 490, also referred to as (e.g., the convolutional layer has been fully learned)The fully connected student layer 480 and the fully connected teacher layer 490 are passed to the loss circuit 250 to calculate a total loss value. The calculation of the total loss value will be further described with reference to the flowchart of FIG. 8 below.FIG. 5 is a flow diagram illustrating example machine readable instructions and / or example operations that may be executed and / or instantiated by example processor circuitry to implement the knowledge distillation architecture 100 of FIG. 1. The knowledge distillation process 500 of FIG. 5 begins at block 510, at which the data fetch circuit 130 provides a data source for the fetched data. In some examples, as described above, the data may be stored in external data 170 (e.g., online sources, big data sources, etc.) that may be retrieved via the network 125, or in memory 180 directly accessible by the data retrieval circuit 170. In such an example, the data fetch circuit 130 may specify the source of the data regardless of whether it originates from the external data 170 or the memory 180.The knowledge distillation process 500 then continues at the location where the data fetch circuit 130 fetches the data from the established data source. (Block 520). In some examples, the data may be stored on internal memory buffers within the edge device 110. In other examples, the data may be temporarily stored in the edge memory 160. In still other examples, the data may not be stored and may be passed directly to the knowledge distillation circuit 140, where the data is deleted once the knowledge distillation process 500 is completed.Next, the knowledge distillation circuit 140 performs the knowledge distillation on the data fetched by the data fetch circuit 130. (Block 530). In some examples, the knowledge distillation circuit 140 accesses data from the data fetch circuit 130. In other examples, the knowledge distillation circuit 140 accesses data stored on the edge memory 160.Once the knowledge distillation circuit 140 performs the knowledge distillation on the fetched data, the output circuit 150 outputs a result of the knowledge distillation. (Block 540). In some examples, the result is forwarded to the data host 120 via the network 125 to output the result to a graphical user interface (GUI) to allow an operator to display the result. In other examples, the result is stored in the edge memory 160 and / or the memory 180 on the data host 120.FIG. 6 is a flow diagram illustrating example machine readable instructions and / or example operations that may be executed and / or instantiated by the processor circuitry to perform the knowledge distillation of the data retrieved by the data retrieval circuit 130. Performing the knowledge distillation of FIG. 6 begins at block 610, where the feature map circuit 220 generates the teacher feature map 440 and the student feature map 450.Once teacher feature map 440 and student feature map 450 have been generated, feature map circuit 220 identifies the channel dimensions for teacher feature map 440 and student feature map 450. (Block 620). In the examples described herein and as indicated above, teacher feature map 440 has a different channel dimension than student feature map 450. In some examples, the channel dimension includes a channel height, width, and number representing a three-dimensional feature map. In other examples, the feature maps 440, 450 may have additional dimensions or fewer dimensions.The feature map circuit 220 then determines whether the teacher feature map 440 and the student feature map 450 are in the same feature space. (Block 625). In some examples, teacher feature map 440 is located in a different feature space (e.g., the dimensions in which the features of the teacher feature map are stored) than the feature space of student feature map 450. In such an example, once the feature map circuit 220 identifies the channel dimensions of the feature map 440, 450, the feature map circuit 220 should convert the feature maps 440, 450 to the same feature space.If the feature map circuit 220 has determined that the teacher and student feature maps 440, 450 are not in the same feature space (e.g., block 625 returns the result to NO), then the feature map circuit 220 may perform feature conversion on the teacher feature map 440 and / or the student feature map 450 to project the feature maps 440, 450 into the same feature space. (Block 630). In some examples, both feature maps 440, 450 may be converted to the same feature space. Such an example may be desired to reduce computational complications by reducing channel dimensions, modifying channel dimensions, etc. In other examples, only one of the feature maps 440, 450 may be converted.Once the feature maps 440, 450 are converted by the feature map circuit 220, or if the feature map circuit 220 has determined that the feature maps 440, 450 are in the same feature space (e.g., block 625 returns the result YES), then the teacher training circuit 230 and the student training circuit 240 train the teacher network 410 and the student network 420, respectively. (Block 640). In the examples described herein, teacher network 410 and student network 420 are trained simultaneously (e.g., in parallel). However, in other examples, teacher network 410 and student network 420 may be trained in series.Once teacher training circuit 230 and student training circuit 240 train teacher network 410 and student network 420, loss circuit 250 generates a total loss value based on trained teacher and student networks 410, 420. (Block 650). In some examples, the total loss value is a result of teacher network loss and student network loss. In other examples, the total loss value may include additional loss characteristics. Further information for generating the total loss value is disclosed with reference to FIG. 8.Once the total loss value is generated by the loss circuit 250, the result generation circuit 260 generates a result to be sent to the output circuit 150. (Block 660). In some examples, the result is a report of the success, failure, efficiency, etc. of the knowledge distillation process being performed. In some examples, the result is generated in a format for a user to visualize the result of training the teacher and student networks 410, 420. In other examples, the result may be stored in the edge memory 160 and / or the memory 180. Once the result is generated by the result generator circuit 260, the example operations of the flowchart of FIG. 6 are complete.FIG. 7 is a flowchart illustrating example machine readable instructions and / or example operations that may be executed and / or instantiated by example processor circuitry for training teacher network 410 and student network 420. Training the teacher network 410 and the student network 420 of FIG. 7 begins at block 710, where the feature map projection circuit 310 projects the student feature map 450 to a predefined channel dimension. In some examples, the predefined channel dimension is equal to the channel dimension of the teacher feature map 440.Once the student feature map 440 is projected by the feature map projection circuit 310 to a predefined channel dimension, the feature splitting circuit 320 then splits the projected student feature map into a predefined number of non-overlapping segments (e.g., the extended student feature map 460) having the same channel dimension as the teacher feature map 440. (Block 720). Any number of non-overlapping segments may be used to generate the many-to-one feature representation, as described above. In some examples, fewer (e.g., eight or less) non-overlapping segments may be desired to increase compute time. In other examples, more than eight non-overlapping segments may be desired to increase the accuracy of the learning process and decrease the overall loss value.Once the student feature map 450 has been divided into a predefined number of non-overlapping segments, the matching circuit 330 then begins training the teacher by matching the non-overlapping segments of the extended student feature map 460 with features on the teacher feature map 440. (Block 730). In some examples, matching circuit 330 matches the non-overlapping segments to obtain the information that is transmitted from student network 420 to teacher network 410 and then back to student network 420. Matching the non-overlapping segments allows teacher network 410 to learn from many features of extended student feature map 460, while conventional neural networks use a one-to-one feature representation that increases computational time to achieve the same result as the many-to-one feature representation described herein.Once the matching circuit 330 matches the non-overlapping segments with the teacher feature map 440, the reparametrizing circuit 340 reparameterizes the teacher feature map 440 using a combination of the non-overlapping student segments. (Block 740). In some examples, the reparametrization circuit 330 enables backward propagation through the teacher network 410 and the student network 420. The reparametrizing circuit 340 promotes the learning process for the teacher network 450.Alternatively or simultaneously, the matching application circuit 350 applies a result of matching the non-overlapping segments to the enhanced student feature map 460. (Block 750). In some examples, application of the result of matching the non-overlapping segments trains the student network based on the output of matching circuit 330. In the examples described herein, the reparametrizing circuit 340 and the matching application circuit 350 may be executed in parallel to train the teacher network 440 and the student network 450 simultaneously.Once teacher feature map 440 has been reparameterized and matching of non-overlapping segments has been applied to expanded student feature map 460, feature map projection circuit 310 projects expanded student feature map 460 back to the original channel dimension (e.g., the channel dimension of student feature map 450). (Block 760). In some examples, the projection back to the original channel dimension is to allow the loss calculation to be performed on the output student feature map F so. Once the extended student feature map 460 has been projected back to its original channel dimension, training of the teacher and student networks 410, 420 is completed.FIG. 8 is a flowchart illustrating example machine readable instructions and / or example operations that may be executed and / or instantiated by a processor circuit to generate the total loss. Generating the total loss value of FIG. 8 begins at block 810, where loss circuit 250 generates a feature matching loss. In some examples, the feature matching loss is represented by Equation 1 below:As shown in Equation 1, N represents the number of non-overlapping segments (e.g., divided by the student feature map 450), F se represents the extended student feature map 460, and F t represents the teacher feature map 440.Loss circuit 250 generates teacher network loss after the teacher is reparameterized (e.g., according to the instructions of block 740). (Block 820). The generation of teacher network loss is represented by Equation 2 below:As represented by equation 2, the weight of the fully connected teacher network layer (e.g., fully learned layer) is. CE is a cross entropy loss (also known as protocol loss, for example) that measures a performance of a classification model whose output is a probability value between 0 and 1. Softmax is a normalized exponential function that converts a real number vector into a probability distribution of possible results. Finally, GT is ground truth (e.g., a goal of training or validating the neural network with a labeled dataset).Loss circuit 250 generates a student network loss after the student is projected back to its original channel dimension (e.g., according to the instructions of block 760). (Block 830). The creation of the student network loss is represented by equation 3 below:As represented by equation 3, the weight of the fully connected student network layer (e.g., fully learned layer) is.Corresponding to each individual loss contribution (e.g., the feature matching loss L mofd, the teacher network loss L t and the student network loss L s) the loss circuit 250 generates loss coefficients for each loss contribution. (Block 840). In some examples, the loss coefficients, represented as a student network loss coefficient β, a teacher network loss coefficient α, and a feature matching loss coefficient γ, are generated by experimental analysis (e.g., updating the loss coefficients based on the result of the previous loss calculation). In other examples, the loss coefficients are constant.Once the loss coefficients are generated, the loss circuit 250 calculates the total loss value. (Block 850). Calculating the total loss value is represented by Equation 4 below:As represented by Equation 4 above, L represents the total loss value as a combination of the feature matching loss L mofd with the feature matching loss coefficient γ, the teacher network loss L t with the teacher network loss coefficient β, and the student network loss L s with the student network loss coefficient α. The total loss value is used in the output of the result of the learning process for teacher network 410 and student network 420. During each learning cycle (e.g., the instructions of FIG. 5 ), the total loss value is recomputed and the result is therefore different, where the neural network learns from the result to output a better (e.g., more accurate and with less data containing) result at subsequent learning cycles.FIG. 9 is a block diagram of an example processor platform 900 structured to execute and / or instantiate the machine readable instructions and / or operations of FIGS. 5, 6, 7, and / or 8 to implement the knowledge distillation circuit 140 of FIG. 2. The processor platform 900 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smart phone, a tablet such as a iPad™ ), a personal digital assistant (PDA), an Internet appliance, a game console, a personal video recorder, a set-top box, a headset (e.g., an augmented reality (AR) headset, a virtual reality (VR) headset, etc.), or another wearable device, or any other type of computing device and / or electronic device.The processor platform 900 of the illustrated example includes a processor circuit 912. The processor circuit 912 of the illustrated example is hardware. The processor circuit 912 may be implemented, for example, by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers of any desired family or manufacturer. The processor circuit 912 may be implemented by one or more semiconductor-based (e.g., silicon-based) devices. In this example, the processor circuit 912 implements the feature map circuit 220, teacher training circuit 230, student training circuit 240, loss circuit 250, result generation circuit 260, feature map projection circuit 310, feature division circuit 320, matching circuit 330, reparametrization circuit 340, and matching application circuit 350.The processor circuit 912 of the illustrated example includes a local memory 913 (e.g., a cache, registers, etc.). The processor circuit 912 of the illustrated example is in communication with a main memory including a volatile memory 914 and a nonvolatile memory 916 via a bus 918. The volatile memory 914 may be implemented by a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), a RAMBUS® dynamic random access memory (RDRAM®), and / or another type of RAM device. The non-volatile memory 916 may be implemented by flash memory and / or any other desired type of storage device. Access to main memory 914, 916 of the illustrated example is controlled by a memory controller 917.The processor platform 900 of the illustrated example also includes an interface circuit 920. The interface circuit 920 may be implemented by hardware according to any type of interface standard, such as an Ethernet interface, a universal serial bus (USB) interface, a Bluetooth® interface, a near field communication (NFC) interface, a peripheral component interconnect (PCI) interface, and / or a peripheral component interconnect express (PCIe) interface.In the illustrated example, one or more input devices 922 are connected to the interface circuit 920. The one or more input devices 922 enable a user to input data and / or commands to the processor circuit 912. The input devices 922 may be implemented, for example, by an audio sensor, a microphone, a camera (photo or video camera), a keyboard, a key, a mouse, a touch screen, a track pad, a track ball, an isopoint device, and / or a voice recognition system.One or more output devices 924 are also connected to the interface circuit 920 of the illustrated example. The one or more output devices 924 may be implemented, for example, by display devices (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube (CRT) display, an in-place switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. The interface circuit 920 of the illustrated example therefore typically includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit such as a GPU.The interface circuit 920 of the illustrated example also includes a communication device such as a transmitter, a receiver, a transceiver, a modem, a residential gateway, a wireless access point, and / or a network interface to enable data exchange with external machines (e.g., computing devices of any type) via a network 926. The communication may be through, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a visual connection system, a cellular telephone system, an optical connection, etc.The processor platform 900 of the illustrated example also includes one or more mass storage disks or devices 928 for storing software and / or data. Examples of such mass storage devices 928 are magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disk drives, redundant array of independent disks (RAID) systems, solid state storage devices such as flash memory devices and / or SSDs and DVD drives.The machine readable instructions 932, which may be implemented by the machine readable instructions of FIGS. 5, 6, 7, and / or 8, may be stored in the mass storage device 928, the volatile memory 914, the nonvolatile memory 916, and / or on at least one removable nonvolatile computer readable storage medium such as a CD or DVD.FIG. 10 is a block diagram of an example implementation of the processor circuit 912 of FIG. 9. in this example, the processor circuit 912 of FIG. 9 is implemented by a microprocessor 1000. Microprocessor 1000 may be, for example, a general purpose microprocessor (e.g., a general purpose microprocessor circuit). The microprocessor 1000 executes some or all of the machine readable instructions of the flowcharts in FIGS. 5, 6, 7, and / or 8 to instantiate the knowledge distillation circuit 140 of FIG. 2 as logic circuitry and perform the operations corresponding to those machine readable instructions. In some such examples, the knowledge distillation circuit 140 of FIG. 2 is instantiated by the hardware circuits of the microprocessor 1000 in combination with the instructions. The microprocessor 1000 may be implemented by, for example, a multi-core hardware circuit such as a CPU, a DSP, a GPU, an XPU, etc. Although it may include any number of example cores 1002 (e.g., 1 core), the microprocessor 1000 of this example is a multi-core semiconductor device with N cores. Cores 1002 of microprocessor 1000 may operate or cooperate independently to execute machine readable instructions. For example, machine code corresponding to a firmware program, an embedded software program, or a software program may be executed by one of the cores 1002, or may be executed by multiple ones of the cores 1002 at the same time or at different times. In some examples, machine code corresponding to the firmware program, embedded software program, or software program is broken into threads and executed in parallel by two or more of the cores 1002. The software program may correspond to a portion or all of the machine readable instructions and / or operations depicted in the flowcharts of FIGS. 5, 6, 7, and / or 8.Cores 1002 may communicate via a first example bus 1004. In some examples, the first bus 1004 may be implemented by a communication bus to cause communication associated with one or more of the cores 1002. The first bus 1004 may be implemented by, for example, at least one of an inter-integrated circuit (I2C) bus, a serial peripheral interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the first bus 1004 may be implemented by any other type of computer bus or an electrical bus. Cores 1002 may receive data, instructions, and / or signals from one or more external devices using, for example, example, an example interface circuit 1006. The cores 1002 may output data, instructions, and / or signals to the one or more external devices through the interface circuit 1006. Although the cores 1002 of this example include the example local memory 1020 (e.g., level 1 (L1) cache, which may be split into an L1 data cache and an L1 instruction cache), the microprocessor 1000 also includes an example shared memory 1010 that may be shared by the cores (e.g., level 2 (L2) cache) for high speed access to data and / or instructions. The data and / or instructions may be transferred (e.g., shared) by writing to and / or reading from the shared memory 1010. The local memory 1020 of each of the cores 1002 and the shared memory 1010 may be part of a hierarchy of data storage devices and may include multiple levels of cache and main memory (e.g., main memory 914, 916 of FIG. 9 ). Typically, higher levels of memory in the hierarchy have a lower access time and a lower storage capacity than lower levels of memory. Changes at the various levels of the cache hierarchy are managed (e.g., coordinated) by a cache coherency policy.Each core 1002 may be referred to as a CPU, a DSP, a GPU, etc., or any other type of hardware circuit. Each core 1002 includes a controller circuit 1014, an arithmetic and logic (AL) circuit (sometimes also referred to as an ALU) 1016, a plurality of registers 1018, the local memory 1020, and a second example bus 1022. Other structures may be present. For example, each core 1002 may include a vector unit circuit, a single instruction multiple data (SIMD) unit circuit, a load / store unit (LSU) circuit, a branch / jump unit circuit, a floating point unit (FPU) circuit, etc. The controller circuit 1014 includes semiconductor-based circuits structured to control (e.g., coordinate) data movement within the corresponding core 1002. AL circuit 1016 includes semiconductor-based circuitry structured to perform one or more mathematical and / or logical operations on the data in the corresponding core 1002. The AL circuit 1016 of some examples performs integer-based operations. In other examples, AL circuit 1016 also performs floating point operations. In still further examples, AL circuit 1016 may include a first AL circuit that performs integer-based operations and a second AL circuit that performs floating point operations. In some examples, AL circuit 1016 may be referred to as an arithmetic logic unit (ALU). The registers 1018 are semiconductor-based structures for storing data and / or instructions, such as results of one or more of the operations performed by the AL circuit 1016 of the corresponding core 1002. The registers 1018 may include, for example, vector registers, SIMD registers, general purpose registers, flag registers, segment registers, machine specific registers, instruction pointer registers, control registers, debug registers, memory management registers, machine check registers, etc. The registers 1018 may be arranged in a bank as shown in Fig. 10. Alternatively, the registers 1018 may be organized in any other arrangement, format, or structure, including as distributed throughout the core 1002, to reduce access time. The second bus 1022 may be implemented by an I2C bus and / or an SPI bus and / or a PCI bus and / or a PCIe bus.Each core 1002 and / or, more generally, the microprocessor 1000 may include structures that are additional and / or alternative to those shown and described above. For example, one or more clock circuits, one or more power supplies, one or more power gates, one or more cache home agents (CHA), one or more converged / common mesh stops (CMS), one or more shifters (e.g., barrel shifters), and / or other circuitry may be present. The microprocessor 1000 is a semiconductor device manufactured to include many interconnected transistors to implement the structures described above in one or more ICs (integrated circuits) included in one or more packaging. The processor circuitry may include and / or cooperate with one or more accelerators. In some examples, the accelerators are implemented by logic circuitry to perform certain tasks more quickly and / or efficiently than is possible with a general purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. A GPU or other programmable device may also be accelerators. Accelerators may be included in the processor circuit, in the same die package as the processor circuit, and / or in one or more packages separate from the processor circuit.FIG. 11 is a block diagram of another example implementation of the processor circuit 912 of FIG. 9. For example, the FPGA circuit 1100 may be implemented by an FPGA. For example, the FPGA circuit 1100 may be used to perform operations that could otherwise be performed by the example microprocessor 1000 of FIG. 10 executing corresponding machine readable instructions. However, once configured, the FPGA circuit 1100 instantiates the machine readable instructions in hardware and, accordingly, can often perform the operations faster than they could be performed by a general purpose microprocessor executing the corresponding software.More specifically, unlike the microprocessor 1000 of FIG. 10 described above (which is a general purpose device that can be programmed to execute some or all of the machine readable instructions represented by the flowcharts of FIGS. 5, 6, 7, and / or 8, but whose connections and logic circuits are fixed after manufacture), the FPGA circuit 1100 of the example of FIG. 11 includes connections and logic circuits that can be configured and / or connected in various ways after manufacture, for example, to instantiate some or all of the machine readable instructions represented by the flowcharts of FIGS. 5, 6, 7, and / or 8. In particular, the FPGA circuit 1100 may be thought of as an array of logic gates, links, and switches. The switches may be programmed to change the manner in which the logic gates are interconnected by the interconnects, thereby effectively forming one or more dedicated logic circuits (unless the FPGA circuit 1100 is reprogrammed). The configured logic circuits enable the logic gates to cooperate in different ways to perform different operations on the data received by input circuits. These operations may correspond to some or all of the software depicted in the flowcharts of FIGS. 5, 6, 7, and / or 8. Thus, the FPGA circuit 1100 may be structured to correspond to some or all of the machine readable instructions of the flowcharts of FIGS. 5, 6, 7, and / or 8, effectively instantiated as dedicated logic circuits to perform the operations corresponding to these software instructions in a dedicated manner analogous to an ASIC. Therefore, the FPGA circuit 1100 may perform the operations corresponding to some or all of the machine readable instructions of FIGS. 5, 6, 7, and / or 8 faster than the general purpose microprocessor may perform them.In the example of FIG. 11, the FPGA circuit 1100 is structured to be programmed (and / or reprogrammed once or more times) by an end user through a hardware description language (HDL) such as Verilog. The FPGA circuit 1100 of FIG. 11 includes an example input / output (I / O) circuit 1102 to receive and / or output data from and to the example configuration circuit 1104 and / or external hardware 1106. The configuration circuit 1104 may be implemented, for example, by an interface circuit that may receive machine readable instructions to configure the FPGA circuit 1100, or portions thereof. In some such examples, the configuration circuit 1104 may receive the machine readable instructions from a user, a machine (e.g., hardware circuit (e.g., programmable or dedicated circuit) that may implement an artificial intelligence / machine learning (AI / ML) model to generate the instructions), etc. In some examples, external hardware 1106 may be implemented by an external hardware circuit. External hardware 1106 may be implemented by microprocessor 1000 of FIG. 10, for example. The FPGA circuit 1100 also includes an array of an example logic gate circuit 1108, a plurality of example configurable interconnects 1110, and an example memory circuit 1112. Logic gate circuit 1108 and configurable interconnects 1110 are configurable to instantiate one or more operations corresponding to at least some of the machine readable instructions of FIGS. 5, 6, 7, and / or 8, and / or other desired operations. The logic gate circuit 1108 shown in FIG. 11 is fabricated in groups or blocks. Each block includes semiconductor-based electrical structures configurable into logic circuits. In some examples, the electrical structures include logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that represent basic logic circuit building blocks. Electrically controllable switches (e.g., transistors) are provided within each of the logic gate circuits 1108 to enable configuration of the electrical structures and / or logic gates to form circuits for performing desired operations. Logic gate circuit 1108 may include other electrical structures such as look-up tables (LUTs), registers (e.g., flip-flop memories or latches), multiplexers, etc.The configurable interconnects 1110 of the illustrated example are conductive paths, traces, vias, or the like, which may include electrically controllable switches (e.g., transistors), the state of which may be changed by programming (e.g., using an HDL command language) to enable or disable one or more interconnects between one or more of the logic gate circuits 1108 to program desired logic circuits.The memory circuit 1112 of the example shown is structured to store the result or results of one or more of the operations performed by the respective logic gates. The memory circuit 1112 may be implemented by registers or the like. In the illustrated example, the memory circuit 1112 is distributed among the logic gate circuit 1108 to facilitate access and increase execution speed.The example FPGA circuit 1100 of FIG. 11 also includes an example dedicated operation circuit 1114. In this example, dedicated operation circuit 1114 includes special purpose circuit 1116 that can be invoked to implement frequently used functions to avoid the need to program these functions in practical use. Examples of such special purpose circuitry 1116 include memory (e.g., DRAM) control circuitry, PCIe control circuitry, clock circuitry, transceiver circuitry, memory, and multiplier-accumulator circuitry. Other types of special purpose circuits may be present. In some examples, the FPGA circuit 1100 may also include an example programmable general purpose circuit 1118, such as an example CPU 1120 and / or an example DSP 1122. Additionally or alternatively, there may also be another programmable general purpose circuit 1118, such as a GPU, an XPU, etc., that may be programmed to perform further operations.Although FIGS. 10 and 11 illustrate two example implementations of the processor circuit 912 of FIG. 9, many other approaches are contemplated. For example, as mentioned above, an FPGA circuit may include an integrated CPU, such as one or more of the example CPUs 1120 of FIG. 11 Therefore, the processor circuit 912 of FIG. 9 may additionally be implemented by combining the example microprocessor 1000 of FIG. 10 and the example FPGA circuit 1100 of FIG. 11. In some such hybrid examples, a first portion of the machine readable instructions represented by the flowcharts of FIGS. 5, 6, 7, and / or 8 may be executed by one or more of the cores 1002 of FIG. 10, a second portion of the machine readable instructions represented by the flowcharts of FIGS. 5, 6, 7, and / or 8 may be executed by the FPGA circuit 1100 of FIG. 11, and / or a third portion of the machine readable instructions represented by the flowcharts of FIGS. 5, 6, 7, and / or 8 may be executed by an ASIC. It should be appreciated that some or all of the elements of the knowledge distillation circuit 140 of FIG. 2 may thus be instantiated at the same time or at different times. For example, a portion or all of the circuitry may be instantiated in one or more threads executing simultaneously and / or in series. Moreover, in some examples, some or all of the elements of the knowledge distillation circuit 140 of FIG. 2 may be implemented in one or more virtual machines and / or containers executing on the microprocessor.In some examples, the processor circuit 912 of FIG. 9 may be located in one or more housings. For example, the microprocessor 1000 of FIG. 10 and / or the FPGA circuit 1100 of FIG. 11 may be housed in one or more housings. In some examples, an XPU may be implemented by the processor circuit 912 of FIG. 9, which may be located in one or more housings. For example, the XPU may include a CPU in one package, a DSP in another package, a GPU in another package, and an FPGA in yet another package.A block diagram illustrating an example software distribution platform 1205 for executing software, such as the example machine readable instructions 932 of FIG. 9, on third party owned and / or operated hardware devices is shown in FIG. 12. The example software distribution platform 1205 may be implemented by any computer server, data device, cloud service, etc., capable of storing software and transmitting it to other computing devices. The third party may be customers of the enterprise that is owners and / or operators of the software distribution platform 1205. The entity owning and / or operating the software distribution platform 1205 may be, for example, a developer, a seller, and / or a licenser of software, such as the example machine readable instructions 932 of FIG. 9. The third party may be consumers, users, retailers, OEMs, etc., that purchase and / or license the software for use and / or resale and / or for sub-licenses. In the illustrated example, the software distribution platform 1205 includes one or more servers and one or more storage devices. The memory devices store the machine readable instructions 932, which may correspond to the machine readable instruction examples of FIGS. 5, 6, 7, and / or 8 described above. The one or more servers of the example software distribution platform 1205 are in communication with an example network 1210, which may correspond to one or more of the Internet and / or one of the example networks 926 described above. In some examples, the one or more servers are responsive to requests to send the software to a requesting party as part of a commercial transaction. The payment for the delivery, sale and / or license of the software may be handled via the one or more servers of the software distribution platform and / or via a payment entity of a third party. The servers allow buyers and / or licensees to download the machine readable instructions 932 from the software distribution platform 1205. For example, software that may correspond to the example machine readable instructions of FIGS. 5, 6, 7, and / or 8 may be downloaded to the example processor platform 900 that is to execute the machine readable instructions 932 to implement the knowledge distillation circuit 140. In some examples, one or more servers of the software distribution platform 1205 periodically provide, transmit, and / or force updates for the software (e.g., the machine readable instructions 932 of FIG. 9 ) to ensure that improvements, patches, updates, etc., are distributed and applied to the software at the end-user devices.From the foregoing, it will be appreciated that example systems, methods, apparatus, and articles of manufacture have been disclosed that train a teacher network and a student network simultaneously without a pre-trained teacher model. The disclosed systems, methods, devices, and articles of manufacture improve the efficiency of use of a computing device by using an online data source to train the teacher network and the student network simultaneously with a many-to-one feature rendering approach. The disclosed systems, methods, devices, and articles of manufacture are accordingly directed to one or more improvements in the operation of a machine, such as a computer or other electronic and / or mechanical device.Example methods, apparatuses, systems, and articles of manufacture for simultaneously training a teacher network and a student network without a pre-trained teacher model are disclosed herein. Further examples and combinations thereof include the following:Example 1 includes an apparatus comprising at least one memory, machine readable instructions, and processor circuitry to at least one of instantiate and execute the machine readable instructions to generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension, generate a second feature map for a student network based on the query, the second feature map having a second channel dimension, the second channel dimension different from the first channel dimension, split the second feature map into segments having the first channel dimension, train the teacher network using the segments of the second feature map, and train the student network by applying a total loss value to the student network, the total loss value based on a loss function, wherein the teacher network and the student network are trained on an edge device.Example 2 includes the apparatus of example 1, wherein the processor circuitry is further to match the segments of the second feature map with the first feature map to train the teacher network.Example 3 includes the apparatus of example 1, wherein the processor circuit is further to reparameterize the teacher network based on the combination of the segments of the second feature map.Example 4 includes the apparatus of example 1, wherein the processor circuit is further to retrieve data from a data source, the data including information regarding the query.Example 5 includes the apparatus of example 4, wherein the processor circuit is further to provide the data source before retrieving the data.Example 6 includes the apparatus of any of Examples 1 to 5, wherein the processor circuitry is further to project the second feature map onto a predefined channel dimension before performing the partitioning, wherein the predefined channel dimension is different than the second channel dimension.Example 7 includes the apparatus of example 6, wherein the processor circuitry is further to project the second feature map back to the second channel dimension after training the teacher network.Example 8 includes the apparatus of example 1, wherein the segments of the second feature map are non-overlapping segments.Example 9 includes the apparatus of example 1, wherein the processor circuit is further to generate the total loss value.Example 10 includes the apparatus of example 9, wherein the total loss value includes a combination of a matching loss corresponding to matching the segments of the second feature map with the first feature map, a teacher network loss, and a student network loss.Example 11 includes the apparatus of example 1, wherein the processor circuitry is further to generate a result corresponding to training of the teacher network and the student network.Example 12 includes the apparatus of example 1, wherein training of the teacher network and the student network is performed simultaneously.Example 13 includes an apparatus for performing feature distillation in neural networks, comprising means for generating a first feature map for a teacher network, the first feature map having a first channel dimension, and a second feature map for a student network, the second feature map having a second channel dimension different from the first channel dimension, means for dividing the second feature map into segments having the first channel dimension, first means for training the teacher network using the segments of the second feature map, and second means for training the student network using an output of a loss function, wherein the first and second means for training are implemented on an edge device.Example 14 includes the apparatus of example 13, wherein the first means for training the teacher network further includes means for matching the segments of the second feature map with the first feature map.Example 15 includes the apparatus of example 14, wherein the second means for training the student network further includes means for applying a result of the matching means to the student network.Example 16 includes the apparatus of example 13, wherein the first means for training the teacher network further includes means for reparametry the teacher network based on the combination of the segments of the second feature map.Example 17 includes the apparatus of example 13, wherein the second means for training the student network further includes means for projecting the second feature map onto a predefined channel dimension prior to splitting the second feature map, wherein the predefined channel dimension is different than the second channel dimension.Example 18 includes the apparatus of example 17, wherein the means for projecting is a first means for projecting, and the second means for training the student network further includes a second means for projecting the second feature map back to the second channel dimension after training the teacher network.Example 19 includes the apparatus of example 13, further including means for calculating a total loss value, wherein the total loss value corresponds to the output of the loss function.Example 20 includes the apparatus of example 13, further including means for generating a result of the first and second training means.Example 21 includes a non-transitory machine-readable storage medium comprising instructions that, when executed, cause the processor circuitry to generate at least a first feature map for a teacher network based on a query, the first feature map having a first channel dimension, generate a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension, divide the second feature map into segments having the first channel dimension, train the teacher network using the segments of the second feature map, and train the student network by applying a total loss value to the student network, the total loss value based on a loss function, wherein the teacher network and the student network are trained on an edge device.Example 22 includes the non-transitory machine readable storage medium of example 21, wherein the instructions, when executed, further cause the processor circuitry to match the segments of the second feature map with the first feature map to train the teacher network.Example 23 includes the non-transitory machine readable storage medium of example 21, wherein the instructions, when executed, further cause the processor circuitry to reparameterize the teacher network based on the combination of the segments of the second feature map.Example 24 includes the non-transitory machine readable storage medium of example 21, wherein the instructions, when executed, further cause the processor circuit to retrieve data from a data source, the data including information regarding the query.Example 25 includes the non-transitory machine readable storage medium of example 24, wherein the instructions, when executed, further cause the processor circuit to provide the data source before the data is retrieved.Example 26 includes the non-transitory machine readable storage medium of example 21, wherein the instructions, when executed, further cause the processor circuitry to project the second feature map onto a predefined channel dimension before the partitioning is performed, wherein the predefined channel dimension is different than the second channel dimension.Example 27 includes the non-transitory machine readable storage medium of any of Examples 21-26, wherein the instructions, when executed, further cause the processor circuitry to project the second feature map back to the second channel dimension after training the teacher network.Example 28 includes the non-transitory machine readable storage medium of example 21, wherein the instructions, when executed, further cause the processor circuitry to generate the total loss value using a combination of a matching loss corresponding to a matching of the segments of the second feature map with the first feature map, a teacher network loss, and a student network loss.Example 29 includes the non-transitory machine readable storage medium of example 21, wherein the instructions, when executed, further cause the processor circuitry to generate a result corresponding to training of the teacher network and the student network.Example 30 includes a method for training a teacher network and a student network, comprising generating a first feature map for a teacher network based on a query, the first feature map having a first channel dimension, generating a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension, dividing the second feature map into segments having the first channel dimension, training the teacher network using the segments of the second feature map, and training the student network by applying a total loss value to the student network, wherein the training of the teacher network and the student network is performed on an edge device.Example 31 includes the method of example 30, further including matching the segments of the second feature map with the first feature map to train the teacher network.Example 32 includes the method of example 30, further including reparametry the teacher network based on the combination of the segments of the second feature map.Example 33 includes the method of example 30, further including retrieving data from a data source, the data including information regarding the query.Example 34 includes the method of example 33, further including providing the data source before retrieving the data.Example 35 includes the method of example 30, further including projecting the second feature map onto a predefined channel dimension before performing the partitioning, wherein the predefined channel dimension is different than the second channel dimension.Example 36 includes the method of any of Examples 30-35, further including projecting the second feature map back to the second channel dimension after training the teacher network.Example 37 includes the method of example 30, further including generating the total loss value by employing a loss function.Example 38 includes the method of example 30, wherein generating the total loss value further includes generating a matching loss corresponding to matching the segments of the second feature map with the first feature map, generating a teacher network loss, generating a student network loss, and calculating the total loss value by using a combination of the matching loss, the teacher network loss, and the student network loss.Example 39 includes the method of example 30, further including generating a result corresponding to training the teacher network and the student network.Example 40 includes the method of example 30, wherein training of the teacher network and the student network is performed simultaneously.The following claims are hereby incorporated by reference into this detailed description in their entirety. Although certain exemplary systems, methods, apparatus, and articles of manufacture have been disclosed herein, the scope of the invention of this patent is not limited thereto. Rather, this patent covers all systems, methods, apparatus and articles of manufacture that reasonably fall within the scope of the claims of this patent.

Claims

An apparatus, comprising: at least one memory; machine readable instructions; and processor circuitry to instantiate and / or execute the machine readable instructions to: generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; generate a second feature map for a student network based on the query, the second feature map having a second channel dimension, the second channel dimension different from the first channel dimension; split the second feature map into segments having the first channel dimension; train the teacher network using the segments of the second feature map; and train the student network by applying a total loss value to the student network, the total loss value based on a loss function; wherein the teacher network and the student network are trained on an edge device.The apparatus of claim 1, wherein the processor circuitry is further to match the segments of the second feature map with the first feature map to train the teacher network.The apparatus of claim 1, wherein the processor circuitry is further to reparameterize the teacher network based on the combination of the segments of the second feature map.The apparatus of claim 1, wherein the processor circuitry is further to retrieve data from a data source, the data including information regarding the query.The apparatus of claim 4, wherein the processor circuitry is further to provide the data source before the data is retrieved.The apparatus of any of claims 1 to 5, wherein the processor circuitry is further to project the second feature map onto a predefined channel dimension before performing the partitioning, wherein the predefined channel dimension is different from the second channel dimension.The apparatus of claim 6, wherein the processor circuitry is further to project the second feature map back onto the second channel dimension after training the teacher network.The apparatus of claim 1, wherein the segments of the second feature map are non-overlapping segments.The apparatus of claim 1, wherein the processor circuit is further to generate the total loss value.The apparatus of claim 9, wherein the total loss value includes a combination of a matching loss corresponding to matching the segments of the second feature map with the first feature map, a teacher network loss, and a student network loss.The apparatus of claim 1, wherein the processor circuitry is further to generate a result corresponding to training of the teacher network and the student network.The apparatus of claim 1, wherein training of the teacher network and the student network is performed simultaneously.An apparatus for performing feature distillation in neural networks, comprising: means for generating a first feature map for a teacher network, the first feature map having a first channel dimension and a second feature map for a student network, the second feature map having a second channel dimension different from the first channel dimension; means for dividing the second feature map into segments having the first channel dimension; a first means for training the teacher network using the segments of the second feature map; and a second means for training the student network using an output of a loss function; wherein the first and second means for training are implemented on an edge device.The apparatus of claim 13, wherein the first means for training the teacher network further includes means for matching the segments of the second feature map with the first feature map.The apparatus of claim 14, wherein the second means for training the student network further includes means for applying a result of the matching means to the student network.The apparatus of claim 13, wherein the first means for training the teacher network further includes means for reparametry the teacher network based on the combination of the segments of the second feature map.The apparatus of claim 13, wherein the second means for training the student network further includes means for projecting the second feature map onto a predefined channel dimension prior to splitting the second feature map, wherein the predefined channel dimension is different than the second channel dimension.The apparatus of claim 17, wherein the means for projecting is a first means for projecting, and the second means for training the student network further includes a second means for projecting the second feature map back to the second channel dimension after training the teacher network.The apparatus of claim 13, further comprising means for calculating a total loss value, wherein the total loss value corresponds to the output of the loss function.A non-transitory machine-readable storage medium comprising instructions that, when executed, cause a processor circuit to at least: generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; generate a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; divide the second feature map into segments having the first channel dimension; train the teacher network using the segments of the second feature map; and train the student network by applying a total loss value to the student network, the total loss value based on a loss function; wherein the teacher network and the student network are trained simultaneously on an edge device.A method for training a teacher network and a student network, comprising: generating a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; generating a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; dividing the second feature map into segments having the first channel dimension; training the teacher network using the segments of the second feature map; and training the student network by applying a total loss value to the student network; wherein the training of the teacher network and the student network is performed on an edge device.The method of claim 21, further comprising matching the segments of the second feature map with the first feature map to train the teacher network.The method of claim 21, further comprising reparametry the teacher network based on the combination of the segments of the second feature map.The method of claim 21, further comprising projecting the second feature map onto a predefined channel dimension before performing the partitioning, wherein the predefined channel dimension is different from the second channel dimension.The method of any of claims 21 to 24, further including projecting the second feature map back onto the second channel dimension after training the teacher network.