Method and apparatus for performing many-to-one feature distillation in neural network

Through the online knowledge distillation process, the teacher network and student network are trained simultaneously in the neural network, and the problems of insufficient computing power and increased delay of edge devices are solved, achieving efficient and accurate characteristic distillation effect.

CN120129909APending Publication Date: 2025-06-10INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101301.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Prior art When performing many-to-one feature distillation in neural networks, pre-training of the teacher network is required, resulting in insufficient computing power and increased latency of edge devices.

Method used

Using the online knowledge distillation process, data is retrieved from the data host through the data retrieval circuit module, and the knowledge distillation circuit module is used to train the teacher network and student network at the same time, reducing the computing power requirements of edge devices, and improving accuracy and reducing delays.

Benefits of technology

The simultaneous training of teacher and student networks on edge devices is achieved, reducing computing power requirements, improving accuracy, and reducing latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129909A_ABST
    Figure CN120129909A_ABST
Patent Text Reader

Abstract

Methods, apparatus, systems, and articles of manufacture to train a teacher network and a student network simultaneously without a pre-trained teacher model are disclosed. An apparatus comprising: at least one memory; a machine readable instruction; and a processor circuit module to at least one of instantiate or execute the machine-readable instructions to generate a first feature map for the teacher network based on the query, the first feature map having a first channel dimension; generating a second feature map for the student network based on the query, the second feature map having a second channel dimension, the second channel dimension being different from the first channel dimension; segmenting the second feature map into fragments with a first channel dimension; training a teacher network using the segments of the second feature map; and training the student network by applying a total loss value to the student network, the total loss value based on a loss function, where the teacher network and the student network are trained on the edge device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to neural networks, and more particularly to methods and apparatus for performing many-to-one feature distillation in neural networks. Background Art

[0002] Neural networks are a subset of machine learning and are the core of deep learning algorithms. Neural networks take input data and use a trained model to generate an output based on an analysis of the input data to produce the best (e.g., most accurate) answer for a given input. The trained model is continuously retrained to generate more accurate answers and optimize the neural network. Brief Description of the Drawings

[0003] Figure 1 is a schematic diagram of a knowledge distillation architecture.

[0004] Figure 2 is Figure 1 a block diagram of an example knowledge distillation circuit module of

[0005] Figure 3 is Figure 2 a block diagram of an example teacher training circuit module and student training circuit module of

[0006] Figure 4 is Figure 1 an operational illustration of the knowledge distillation architecture of

[0007] Figure 5 is a flowchart representing example machine-readable instructions and / or example operations that may be executed by an example processor circuit module to implement Figure 1 the knowledge distillation architecture of

[0008] Figure 6 is a flowchart representing example machine-readable instructions and / or example operations that may be executed by an example processor circuit module to implement Figure 5 the knowledge distillation of

[0009] Figure 7 is a flowchart representing example machine-readable instructions and / or example operations that may be executed by an example processor circuit module to implement Figure 6 the training of the teacher network and the training of the student network of

[0010] Figure 8 is a flowchart representing example machine-readable instructions and / or example operations that may be executed by an example processor circuit module to generate Figure 6 the total loss value of

[0011] Figure 9is a block diagram of an example processing platform that includes a processor circuit module configured to execute Figure 3 example machine-readable instructions and / or example operations to implement Figure 2 a knowledge distillation circuit module of

[0012] Figure 10 is Figure 9 a block diagram of an example implementation of the processor circuit module of

[0013] Figure 11 is Figure 9 a block diagram of another example implementation of the processor circuit module of

[0014] Figure 12 is a block diagram of an example software distribution platform (e.g., one or more servers) for distributing software (e.g., software corresponding to example machine-readable instructions of Figure 5 , Figure 6 , Figure 7 and / or Figure 8 to client devices associated with: end users and / or consumers (e.g., for licensing, selling, and / or using), retailers (e.g., for selling, reselling, licensing, and / or sublicensing), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products to be distributed to, e.g., retailers and / or other end users such as direct purchase customers).

[0015] Generally, the same reference numerals will be used throughout the drawings and the accompanying written description to refer to the same or like parts. The drawings are not drawn to scale. Instead, the thickness of layers or regions may be exaggerated in the drawings. Although the drawings show layers and regions with clean lines and boundaries, some or all of these lines and / or boundaries may be idealized. In reality, the boundaries and / or lines may be unobservable, blended, and / or irregular.

[0016] Unless otherwise clearly stated, descriptors such as "first", "second", "third", etc. are used herein and do not in any way imply any meaning of arrangement, priority, physical order, and / or sorting in a list, but are only used as labels and / or arbitrary names to distinguish elements for ease of understanding the disclosed examples. In some examples, the descriptor "first" may be used to refer to an element in a particular implementation, while the same element may be referred to in a claim with a different descriptor (such as "second" or "third"). In such instances, it should be understood that such descriptors are only used to clearly identify those elements that may otherwise share the same name.

[0017] As used herein, "approximately" and "about" modify the subject / value to recognize the potential presence of variations that occur in real-world applications. For example, as would be understood by one of ordinary skill in the art, "approximately" and "about" can modify dimensions that may be imprecise due to manufacturing tolerances and / or other real-world imperfections. For example, "approximately" and "about" can indicate that such dimensions can be within a tolerance range of + / - 10%, unless otherwise specified in the description below. As used herein, "substantially real-time" means occurring in a near-instantaneous manner that recognizes that there may be real-world delays in computing time, transmission, etc. Thus, unless otherwise stated, "substantially real-time" means real-time + / - 1 second.

[0018] As used herein, the phrase "in communication" (including its variants) encompasses direct and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but rather includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.

[0019] As used herein, "processor circuitry" is defined to include (i) one or more dedicated electrical circuits that are configured to perform specific operations and include one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors), and / or (ii) one or more general-purpose semiconductor-based electrical circuits that are programmable with instructions to perform specific operations and include one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuitry include programmable microprocessors, field-programmable gate arrays (FPGAs) that can instantiate instructions, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), XPUs, or microcontrollers, and integrated circuits such as application-specific integrated circuits (ASICs). For example, an XPU can be implemented by a heterogeneous computing system that includes multiple types of processor circuitry (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc. and / or combinations thereof) and (multiple) application programming interfaces (APIs) that can assign (multiple) computing tasks to any (multiple) processor circuitry of the multiple types that is most suitable for performing the (multiple) computing tasks. Detailed Description

[0020] Neural networks that implement knowledge distillation typically train a student network based on a pre-trained teacher network. Feature maps of the teacher network are created to be used as knowledge points for the student network to learn. This method is called two-step knowledge distillation. Alternatively, some neural networks utilize an online training framework that does not require a pre-trained teacher model. This method is called one-stage knowledge distillation.

[0021] Knowledge distillation typically requires a pre-trained teacher network to train a student network. The teacher network is usually trained using a machine that exhibits a large amount of computing resources and / or capabilities (e.g., high-end CPUs, GPUs, RAM, etc.). Due to the computing power required to train the teacher network, such computing resources are typically not implemented on local devices (e.g., edge devices, user devices, etc.).

[0022] Since the student network is usually housed on an edge device and does not have the ability to train the teacher network, the student network must rely on the pre-trained teacher network to teach the student network. This is undesirable for edge devices because they must retrieve data from the feature maps on the teacher network to determine the output (e.g., the answer to a question, the result, etc.), and the neural network must pre-train the teacher model before it can train the student network.

[0023] To train the student network, the neural network creates feature maps based on the teacher network and creates a transformation module to adapt the feature maps to the student network. Once the teacher is trained, the transformation module is discarded, and each feature from the feature map is matched with the student network in a one-to-one representation. Due to the steps required to train the teacher network and create the one-to-one representation, the latency in the neural network increases due to this process. This process is called two-step knowledge distillation.

[0024] Alternatively, in one-step knowledge distillation, the neural network retrieves data from an online teacher network, which is a network that has not been trained at all. Since the neural network does not train the teacher model itself, the computing power requirements are reduced at the edge device level, but the latency still increases due to the data (e.g., feature maps, feature representations, etc.) that needs to be retrieved from an online source (e.g., a source external to the device on which the student network is being trained). Additionally, these processes still utilize a one-to-one representation of the features in the feature maps on the teacher network and are less accurate compared to two-step knowledge distillation because the neural network does not train the teacher network.

[0025] Therefore, there is a need for an online knowledge distillation process (e.g., one-step knowledge distillation) that trains both the teacher network and the student network simultaneously, while reducing the computing power requirements at the edge device level and reducing latency while improving accuracy.

[0026] Figure 1It is a schematic diagram of the knowledge distillation architecture 100. Figure 1 The illustrated knowledge distillation architecture 100 of the example includes an edge device 110 and a data host 120. In some examples, the edge device 110 accesses data from the data host 120 via a network 125. In other examples, the edge device 110 accesses data from the data host 120 by directly communicating with the data host 120 (e.g., Ethernet, Universal Serial Bus (USB), Peripheral Component Interconnect (PCI), etc.).

[0027] The edge device 110 includes a data retriever circuit module 130, a knowledge distillation circuit module 140, and an output circuit module 150. In some examples, the edge device 110 includes an edge storage device 160 for directly storing data on the edge device 110.

[0028] The data retriever circuit module 130 communicates with the data host 120 to retrieve data for training a teacher network and a student network (e.g., teacher network 410 and student network 420). In some examples, the data retriever circuit module 130 communicates with the data host 120 via the network 125. In other instances, the data retriever circuit module 130 directly accesses data from the data host 120. In some examples, the data retriever circuit module 130 establishes a data source before retrieving data. In such examples, the data retriever circuit module 130 determines whether to retrieve data from the data host 120 via the network 125 (e.g., retrieve data from an online source, a big data source, etc.) or from a local storage device (e.g., storage device 180 on the data host 120).

[0029] The knowledge distillation circuit module 140 performs a knowledge distillation process to train the teacher network 410 and the student network 420 based on the data retrieved from the data retriever circuit module 130. In the examples described herein, the knowledge distillation circuit module 140 communicates with the data retriever circuit module 130. However, the edge device 110 may omit the data retriever circuit module 130, and the knowledge distillation circuit module 140 may communicate directly with the data host 120.

[0030] The output circuit module 150 outputs the results of the trained teacher network 410 and student network 420. In some examples, the results may include an answer to a query (such as a natural language answer), an image, an audio clip, a command, etc. In some examples, the output circuit module 150 transmits the results to the data host 120 via the network 125. In other instances, the output circuit module 150 stores the results in the data host 120 by directly communicating with the data host 120. In some examples, the output circuit module 150 directly stores the results in the edge storage device 160.

[0031] The data host 120 includes external data 170. In some examples, the external data 170 includes data from a big data source and can be retrieved by the data host 120 when the data retriever circuit module 130 retrieves data from the data host 120 via the network 125. In some examples, the data host 120 includes a storage device 180 that houses data that can be retrieved / accessed directly and / or by the data retriever circuit module 130 via the network 125.

[0032] Figure 2 is a block diagram of an example knowledge distillation circuit module 140 for training a teacher network 410 and a student network 420. Figure 2 The knowledge distillation circuit module 140 can be instantiated (e.g., create an instance thereof, bring it into existence for any length of time, materialize, implement, etc.) by a processor circuit module such as a central processing unit that executes instructions. Additionally or alternatively, Figure 2 The knowledge distillation circuit module 140 can be instantiated (e.g., create an instance thereof, bring it into existence for any length of time, materialize, implement, etc.) by an ASIC or FPGA configured to perform operations corresponding to the instructions. It should be understood that Figure 2 Some or all of the circuit modules can thus be instantiated at the same or different times. Some or all of the circuit modules can be instantiated, for example, in one or more threads that execute simultaneously on hardware and / or serially on hardware. Additionally, in some examples, Figure 2 Some or all of the circuit modules can be implemented by a microprocessor circuit module that executes instructions to implement one or more virtual machines and / or containers.

[0033] The knowledge distillation circuit module 140 includes a feature map circuit module 220, a teacher training circuit module 230, a student training circuit module 240, a loss circuit module 250, and a result generator circuit module 260.

[0034] The feature map circuit module 220 generates a teacher feature map and a student feature map (e.g., teacher feature map 440 and student feature map 450) based on the data retrieved by the data retriever circuit module 130 according to the teacher network 410 and the student network 420, respectively. In some examples, the feature maps 440, 450 are generated based on a query received by the data retriever circuit module 130. In some examples, the feature map circuit module 220 communicates with the data retriever circuit module 130, the output circuit module 150, and / or the edge storage device 160 via the I / O interface 210. In some examples, the feature map circuit module 220 is instantiated by a processor circuit module that executes feature map instructions and / or is configured to perform operations such as those represented by the Figure 6 flowchart shown.

[0035] In some examples, the knowledge distillation circuit module 140 includes units for generating a teacher feature map 440 and a student feature map 450. For example, the units for generating may be implemented by the feature map circuit module 220. In some examples, the feature map circuit module 220 may be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912). For example, the feature map circuit module 220 may be instantiated by Figure 10 the example microprocessor 1000 executing machine-executable instructions (such as those implemented by at least Figure 6 the blocks 610, 620, and 630). In some examples, the feature map circuit module 220 may be instantiated by a hardware logic circuit module, which may be implemented by an Figure 11 ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the feature map circuit module 220 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the feature map circuit module 220 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0036] The teacher training circuit module 230 trains the teacher network 410. In some examples, the teacher training circuit module 230 communicates with the student training circuit module 240 while training the teacher network 410. In some examples, the teacher training circuit module 230 is instantiated by a processor circuit module that executes teacher training instructions and / or is configured to perform operations such as those represented by the Figure 6 and / or Figure 7 flowcharts.

[0037] In some examples, the knowledge distillation circuit module 140 includes a unit for training the teacher network 410. For example, the unit for training the teacher network 410 may be implemented by the teacher training circuit module 230. In some examples, the teacher training circuit module 230 may be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912). For example, the teacher training circuit module 230 may be instantiated by Figure 10 the example microprocessor 1000 executing machine-executable instructions (such as those implemented by at least Figure 6 the block 640 and / or Figure 7instantiated by the frame 730 and / or 740 (and those implemented by the like). In some examples, the teacher training circuit module 230 may be instantiated by a hardware logic circuit module, which may be an ASIC, XPU or FPGA circuit module 1100 configured to perform operations corresponding to machine-readable instructions. Additionally or alternatively, the teacher training circuit module 230 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the teacher training circuit module 230 may be implemented by at least one or more hardware circuits (e.g., a processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuit, etc.), and the at least one or more hardware circuit modules are configured to execute some or all of the machine-readable instructions and / or implement some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable. Figure 11 implemented by the ASIC, XPU or FPGA circuit module 1100. Additionally or alternatively, the teacher training circuit module 230 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the teacher training circuit module 230 may be implemented by at least one or more hardware circuits (e.g., a processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuit, etc.), and the at least one or more hardware circuit modules are configured to execute some or all of the machine-readable instructions and / or implement some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0038] The student training circuit module 240 trains the student network 420. In some examples, the student training circuit module 240 communicates with the teacher training circuit module 230 while training the student network 420. In some examples, the student training circuit module 240 is instantiated by a processor circuit module that executes student training instructions and / or is configured to implement operations such as those represented by the flowchart of Figure 6 and / or Figure 7 and the like.

[0039] In some examples, the knowledge distillation circuit module 140 includes a unit for training the student network 420. For example, the unit for training the student network 420 may be implemented by the student training circuit module 240. In some examples, the student training circuit module 240 may be instantiated by a processor circuit module (such as the example processor circuit module 912 of Figure 9 ). For example, the student training circuit module 240 may execute machine-executable instructions (such as those implemented by at least the frame 640 of Figure 10 and / or Figure 6 the frame 710, 720, 750, and / or 760 of Figure 7 and the like) through the example microprocessor 1000. In some examples, the student training circuit module 240 may be instantiated by a hardware logic circuit module, which may be an ASIC, XPU or FPGA circuit module 1100 configured to perform operations corresponding to machine-readable instructions. Figure 11implemented by the ASIC, XPU, or FPGA circuit module 1100. Additionally or alternatively, the student training circuit module 240 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the student training circuit module 240 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.), the at least one or more hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0040] The loss circuit module 250 calculates the total loss value of the teacher network 410 and the student network 420. In some examples, the loss circuit module 250 is instantiated by a processor circuit module that executes loss instructions and / or is configured to perform operations such as those Figure 6 and / or Figure 8 represented by the flowchart of.

[0041] In some examples, the knowledge distillation circuit module 140 includes a unit for calculating the total loss value. For example, the unit for calculation may be implemented by the loss circuit module 250. In some examples, the loss circuit module 250 may be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912). For example, the loss circuit module 250 may be executed by Figure 10 the example microprocessor 1000 to execute machine-executable instructions (such as those implemented by at least Figure 6 the block 650 and / or Figure 8 the blocks 810, 820, 830, 840, and / or 850). In some examples, the loss circuit module 250 may be instantiated by a hardware logic circuit module, which may be implemented by the Figure 11 ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the loss circuit module 250 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the loss circuit module 250 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.), the at least one or more hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0042] The result generator circuit module 260 generates results based on the training of the teacher network 410 and the training of the student network 420. In some examples, the result generator circuit module 260 communicates with the data retriever circuit module 130, the output circuit module 150, and / or the edge storage device 160 via the I / O interface 210. In some examples, the result generator circuit module 260 is instantiated by a processor circuit module that executes result instructions and / or is configured to perform operations such as those represented by the Figure 6 flowchart shown.

[0043] In some examples, the knowledge distillation circuit module 140 includes a unit for generating results. For example, the unit for generating can be implemented by the result generator circuit module 260. In some examples, the result generator circuit module 260 can be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912). For example, the result generator circuit module 260 can be instantiated by Figure 10 the example microprocessor 1000 executing machine-executable instructions (such as those implemented by at least Figure 6 block 660). In some examples, the result generator circuit module 260 can be instantiated by a hardware logic circuit module, which can be implemented by an Figure 11 ASIC, XPU, or FPGA circuit module 1100 constructed to perform operations corresponding to machine-readable instructions. Additionally or alternatively, the result generator circuit module 260 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the result generator circuit module 260 can be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPUs, comparators, operational amplifiers (op-amps), logic circuits, etc.) that are constructed to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0044] Figure 3 are block diagrams of an example teacher training circuit module 230 and an example student training circuit module 240 for training the teacher network 410 and the student network 420, respectively. Figure 3 The teacher training circuit module 230 and the student training circuit module 240 of Figure 3The teacher training circuit module 230 and the student training circuit module 240 can be instantiated (e.g., create its instance, make it for any length of time, materialize, implement, etc.) by an ASIC or FPGA configured to perform operations corresponding to the instructions. It should be understood that Figure 3 Some or all of the circuit modules of can thus be instantiated at the same or different times. Some or all of the circuit modules can be instantiated, for example, in one or more threads that execute simultaneously on hardware and / or serially on hardware. Additionally, in some examples, Figure 3 Some or all of the circuit modules of can be implemented by a microprocessor circuit module that executes instructions to implement one or more virtual machines and / or containers.

[0045] The teacher training circuit module 230 and the student training circuit module 240 include a feature map projection circuit module 310, a feature segmentation circuit module 320, a matching circuit module 330, a reparameterization circuit module 340, and a matching application circuit module 350. In some examples, the student training circuit module 240 includes the feature map projection circuit module 310, the feature segmentation circuit module 320, and the matching application circuit module 350. In some examples, the teacher training circuit module 230 includes the matching circuit module 330 and the reparameterization circuit module 340.

[0046] The feature map projection circuit module 310 projects the student feature map 450 onto a desired channel dimension. In some examples, the feature map projection circuit module 310 projects the student feature map 450 into a predefined channel dimension before training the student network 420, where the predefined channel dimension is used for the many-to-one representation of the teacher network 410. In some examples, after the student network 420 has been trained, the feature map projection circuit module 310 projects the student feature map 450 back to its original channel dimension. In some examples, the feature map projection circuit module 310 is instantiated by a processor circuit module that executes feature map projection instructions and / or is configured to perform operations such as those represented by the Figure 7 flowchart shown.

[0047] In some examples, the student training circuit module 240 includes a unit for projecting the student feature map 450 onto a desired channel dimension. For example, the unit for projection can be implemented by the feature map projection circuit module 310. In some examples, the feature map projection circuit module 310 can be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912). For example, the feature map projection circuit module 310 can be executed by Figure 10 the example microprocessor 1000 to execute machine-executable instructions (such as at least by Figure 7instantiated by those implemented by the frames 710 and 760. In some examples, the feature map projection circuit module 310 may be instantiated by a hardware logic circuit module, which may be an ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to machine-readable instructions. Additionally or alternatively, the feature map projection circuit module 310 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the feature map projection circuit module 310 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPUs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other architectures are equally suitable. Figure 11 The feature segmentation circuit module 320 segments the student feature map 450 into segments having the same channel dimension as the teacher feature map 440. In some examples, the feature segmentation circuit module 320 is instantiated by a processor circuit module that executes feature segmentation instructions and / or is configured to perform operations such as those represented by the flowchart of

[0048] the example flowchart. Figure 7 In some examples, the student training circuit module 240 includes a unit for segmenting the student feature map 450 into segments having the same channel dimension as the teacher feature map 440. For example, the unit for segmentation may be implemented by the feature segmentation circuit module 320. In some examples, the feature segmentation circuit module 320 may be instantiated by a processor circuit module (such as the example processor circuit module 912 of

[0049] the example). For example, the feature segmentation circuit module 320 may be instantiated by the example microprocessor 1000 of Figure 9 executing machine-executable instructions (such as those implemented by at least the frame 720 of Figure 10 the example). In some examples, the feature segmentation circuit module 320 may be instantiated by a hardware logic circuit module, which may be an ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to machine-readable instructions. Figure 7 In some examples, the feature segmentation circuit module 320 may be instantiated by a hardware logic circuit module, which may be an ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to machine-readable instructions. Figure 11implemented by the ASIC, XPU, or FPGA circuit module 1100. Additionally or alternatively, the feature segmentation circuit module 320 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the feature segmentation circuit module 320 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.), the at least one or more hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0050] The matching circuit module 330 matches the segments of the student feature map 450 created by the feature segmentation circuit module 320 with the teacher feature map 440. In some examples, the matching circuit module 330 is instantiated by a processor circuit module that executes matching instructions and / or is configured to perform operations such as those represented by Figure 7 the flowchart shown.

[0051] In some examples, the teacher training circuit module 230 includes a unit for matching the segments of the student feature map 450 with the teacher feature map 440. For example, the unit for matching may be implemented by the matching circuit module 330. In some examples, the matching circuit module 330 may be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912 shown). For example, the matching circuit module 330 may be instantiated by Figure 10 the example microprocessor 1000 executing machine-executable instructions (such as those implemented by at least Figure 7 the block 730 shown). In some examples, the matching circuit module 330 may be instantiated by a hardware logic circuit module, which may be implemented by the Figure 11 ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the matching circuit module 330 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the matching circuit module 330 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.), the at least one or more hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0052] The reparameterization circuit module 340 reparameterizes the teacher network 410 based on a combination (e.g., summation) of segments of the student feature map 450. In some examples, the reparameterization circuit module 340 is instantiated by a processor circuit module that executes reparameterization instructions and / or is configured to perform operations such as those represented by the Figure 7 flowchart shown.

[0053] In some examples, the teacher training circuit module 230 includes units for reparameterizing the teacher network 410 based on a combination of segments of the student feature map 450. For example, the unit for reparameterization can be implemented by the reparameterization circuit module 340. In some examples, the reparameterization circuit module 340 can be instantiated by a processor circuit module (such as the Figure 9 example processor circuit module 912). For example, the reparameterization circuit module 340 can be instantiated by Figure 10 the example microprocessor 1000 executing machine-executable instructions (such as those implemented by at least the Figure 7 box 740). In some examples, the reparameterization circuit module 340 can be instantiated by a hardware logic circuit module, which can be implemented by an Figure 11 ASIC, XPU, or FPGA circuit module 1100 constructed to perform operations corresponding to machine-readable instructions. Additionally or alternatively, the reparameterization circuit module 340 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the reparameterization circuit module 340 can be implemented by at least one or more hardware circuits (e.g., processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuit, etc.) that are constructed to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0054] The match application circuit module 350 applies the match result of the segments of the student feature map 450 with the teacher network 410 to the student network 420. In some examples, the match application circuit module 350 is instantiated by a processor circuit module that executes match application instructions and / or is configured to perform operations such as those represented by the Figure 7 flowchart shown.

[0055] In some examples, the student training circuit module 240 includes a unit for applying the result of matching a segment of the student feature map 450 with the teacher network 410 to the student network 420. For example, the unit for application can be implemented by the matching application circuit module 350. In some examples, the matching application circuit module 350 can be instantiated by a processor circuit module (such as Figure 9 the example processor circuit module 912). For example, the matching application circuit module 350 can be instantiated by Figure 10 the example microprocessor 1000 executing machine-executable instructions (such as those implemented by at least Figure 7 block 750). In some examples, the matching application circuit module 350 can be instantiated by a hardware logic circuit module, which can be implemented by an Figure 11 ASIC, XPU, or FPGA circuit module 1100 configured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the matching application circuit module 350 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the matching application circuit module 350 can be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.), which are configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0056] In some examples, the reparameterization circuit module 240 and the matching application circuit module 250 execute in parallel to train the teacher network 410 and the student network 420 simultaneously. In other examples, the reparameterization circuit module 240 and the matching application circuit module 250 execute serially. The order of execution of the reparameterization circuit module 240 and the matching application circuit module 250 is interchangeable and / or modifiable. In other words, the order of execution of either the reparameterization circuit module 240 or the matching application circuit module 250 is not limited to the specific examples described herein.

[0057] Although examples of ways to implement Figure 2 and / or Figure 3 the knowledge distillation circuit module 140 are shown in Figure 1 , Figure 2 and / or Figure 3One or more of the elements, processes, and / or devices shown may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, example feature map circuit module 220, example teacher training circuit module 230, example student training circuit module 240, example loss circuit module 250, example result generator circuit module 260, example feature map projection circuit module 310, example feature segmentation circuit module 320, example matching circuit module 330, example reparameterization circuit module 340, example matching application circuit module 350, and / or more generally Figure 1 Example knowledge distillation circuit module 140 may be implemented by hardware alone or by hardware in combination with software and / or firmware. Thus, for example, any one of example feature map circuit module 220, example teacher training circuit module 230, example student training circuit module 240, example loss circuit module 250, example result generator circuit module 260, example feature map projection circuit module 310, example feature segmentation circuit module 320, example matching circuit module 330, example reparameterization circuit module 340, example matching application circuit module 350, and / or more generally example knowledge distillation circuit module 140 may be implemented by a processor circuit module, (one or more) analog circuits, (one or more) digital circuits, (one or more) logic circuits, (one or more) programmable processors, (one or more) programmable microcontrollers, (one or more) graphics processing units (GPUs), (one or more) digital signal processors (DSPs), (one or more) application specific integrated circuits (ASICs), (one or more) programmable logic devices (PLDs), and / or (one or more) field programmable logic devices (FPLDs) (such as field programmable gate arrays (FPGAs)). Additionally, Figure 1 Example knowledge distillation circuit module 140 may include one or more elements, processes, and / or devices in addition to or instead of those shown in Figure 2 and / or Figure 3 those shown in Figure 2 and / or Figure 3 and / or may include more than one of any or all of the elements, processes, and devices shown.

[0058] In Figure 5 , Figure 6 , Figure 7 and / or Figure 8 are shown flowcharts representing example machine-readable instructions that may be executed to configure a processor circuit module to implement Figure 2 knowledge distillation circuit module 140. The machine-readable instructions may be one or more executable programs or (one or more) portions of an executable program for execution by a processor circuit module, such as the one described below in connection with Figure 9The processor circuit module 912 shown in the example processor platform 900 under discussion and / or below in connection with Figure 10 and / or Figure 11 the example processor circuit module under discussion. The program can be embodied in software stored on one or more non-transitory computer-readable storage media, such as a compact disc (CD), floppy disk, hard disk drive (HDD), solid state drive (SSD), digital versatile disc (DVD), Blu-ray disc, volatile memory (e.g., any type of random access memory (RAM), etc.) or non-volatile memory associated with a processor circuit module located in one or more hardware devices (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, HDD, SSD, etc.), but the entire program and / or portions thereof can alternatively be executed by one or more hardware devices other than the processor circuit module and / or embodied in firmware or dedicated hardware. The machine-readable instructions can be distributed across multiple hardware devices and / or executed by two or more hardware devices (e.g., server and client hardware devices). For example, the client hardware device can be implemented by an endpoint client hardware device (e.g., a hardware device associated with a user) or an intermediate client hardware device (e.g., a radio access network (RAN) gateway that can facilitate communication between a server and an endpoint client hardware device). Similarly, the non-transitory computer-readable storage media can include one or more media located in one or more hardware devices. Additionally, although reference is made to Figure 5 、 Figure 6 、 Figure 7 and / or Figure 8 the flowcharts shown describe an example program, many other methods of implementing the example knowledge distillation circuit module 140 can alternatively be used. For example, the order of execution of the blocks can be changed, and / or some of the blocks described can be changed, eliminated, or combined. Additionally or alternatively, any one or all of the blocks can be implemented by one or more hardware circuits (e.g., processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, comparator, operational amplifier (op-amp), logic circuits, etc.), the one or more hardware circuits being constructed to perform the corresponding operations without executing software or firmware. The processor circuit module can be distributed across different network locations and / or be local to one or more hardware devices (e.g., a single-core processor (e.g., a single-core central processing unit (CPU)), a multi-core processor in a single machine (e.g., a multi-core CPU, XPU, etc.), multiple processors distributed across multiple servers in a server rack, multiple processors distributed across one or more server racks, a CPU and / or FPGA located in the same package (e.g., the same integrated circuit (IC) package or two or more separate enclosures, etc.).

[0059] The machine-readable instructions described herein may be stored in one or more of a compressed format, an encrypted format, a segmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions as described herein may be stored as data or data structures (e.g., as part of instructions, code, code representations, etc.) that can be used to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be segmented and stored on one or more storage devices and / or computing devices (e.g., servers) located at the same or different locations of a network or network collection (e.g., in the cloud, at an edge device, etc.). The machine-readable instructions may need to be installed, modified, adapted, updated, combined, supplemented, configured, decrypted, decompressed, unpackaged, distributed, redistributed, compiled, etc. in order to make them directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, the machine-readable instructions may be stored in multiple parts that are individually compressed, encrypted, and / or stored on separate computing devices, where the parts form a set of machine-executable instructions when decrypted, decompressed, and / or combined, and the set of machine-executable instructions implements one or more operations that can together form a program such as described herein.

[0060] In another example, the machine-readable instructions may be stored in a state where they can be read by a processor circuit module, but libraries (e.g., dynamic link libraries (DLLs)), software development kits (SDKs), application programming interfaces (APIs), etc. need to be added in order to execute the machine-readable instructions on a particular computing device or other device. In another example, the machine-readable instructions may need to be configured (e.g., stored settings, data inputs, recorded network addresses, etc.) before the machine-readable instructions and / or corresponding program can be executed in whole or in part. Thus, as used herein, a machine-readable medium may include machine-readable instructions and / or programs regardless of the particular format or state of the machine-readable instructions and / or programs when stored or otherwise at rest or in transit.

[0061] The machine-readable instructions described herein may be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented using any one of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0062] As described above, Figure 5 、 Figure 6 、 Figure 7 and / or Figure 8Example operations can be implemented using executable instructions (e.g., computer and / or machine-readable instructions) stored on one or more non-transitory computer and / or machine-readable media, such as optical storage devices, magnetic storage devices, HDDs, flash memories, read-only memories (ROMs), CDs, DVDs, caches, any type of RAM, registers, and / or any other storage device or storage disk where information is stored for any duration (e.g., an extended period of time, permanently, briefly, temporarily buffered, and / or cached). As used herein, the terms non-transitory computer-readable medium, non-transitory computer-readable storage medium, non-transitory machine-readable medium, and non-transitory machine-readable storage medium are expressly defined to include any type of computer-readable storage device and / or storage disk and to exclude propagated signals and to exclude transmission media. As used herein, the terms "computer-readable storage device" and "machine-readable storage device" are defined to include any physical (mechanical and / or electrical) structure for storing information, but to exclude propagated signals and to exclude transmission media. Examples of computer-readable storage devices and machine-readable storage devices include any type of random access memory, any type of read-only memory, solid-state memory, flash memory, optical disks, magnetic disks, disk drives, and / or redundant array of independent disks (RAID) systems. As used herein, the term "device" refers to a physical structure that may or may not be configured and / or manufactured to execute computer-readable instructions, machine-readable instructions, etc., by computer-readable instructions, machine-readable instructions, etc., such as a mechanical and / or electrical device, a hardware and / or circuit module.

[0063] "Including" and "comprising" (and all of their forms and tenses) are used herein as open-ended terms. Thus, whenever a claim uses any form of "include" or "comprise" (e.g., comprises, includes, comprising, including, having, etc.) as a preamble or within any kind of claim recitation, it should be understood that additional elements, terms, etc. may exist without falling outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in, for example, the preamble of a claim, it is open-ended in the same manner as the terms "including" and "comprising". When used in the form of, for example, A, B, and / or C, the term "and / or" refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, or (7) A and B and C. As used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A and B" is intended to refer to an implementation that includes any one of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A or B" is intended to refer to an implementation that includes any one of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein in the context of describing the implementation or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A and B" is intended to refer to an implementation that includes any one of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing the implementation or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to refer to an implementation that includes any one of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.

[0064] As used herein, singular references (e.g., "a", "an", "first", "second", etc.) do not exclude a plurality. As used herein, the term "a" or "an" object refers to one or more of such objects. The terms "a" (or "an"), "one or more", and "at least one" are used interchangeably herein. In addition, although listed separately, multiple units, elements, or method acts may be implemented by, for example, the same entity or object. Additionally, although the various features may be included in different examples or claims, these features may be combined, and inclusion in different examples or claims does not imply that the combination of features is not feasible and / or advantageous.

[0065] Figure 4 is Figure 1 an operational representation of the knowledge distillation architecture 100. The example operation 400 shown includes a data host 120, a loss circuit module 250, a teacher network 410, a student network 420, and a plurality of convolutional layers 430. The teacher network 410 includes teacher feature maps 440. The student network 420 includes student feature maps 450. The shown operation 400 can be inserted at any convolutional layer among the plurality of convolutional layers 430. Depending on the neural network, the shown operation 400 can be inserted at more than one convolutional layer Figure 4 of.

[0066] The teacher feature maps 440 (also denoted as F t ) include a teacher channel dimension C t . The student feature maps 450 (also denoted as F s ) include a student channel dimension C s . In the example described herein, the teacher channel dimension C t is different from the student channel dimension C s . Thus, the student feature maps 450 should be projected to a channel dimension equal to the teacher channel dimension C t .

[0067] The student network 420 is extended by a student channel dimension expander 455 (also referred to as W se ) to project the student feature maps 450 into the same channel dimension as the teacher channel dimension C t . Then, the student network 420 segments the projected student feature maps into an extended student feature map 460 (also referred to as F se ) having N non-overlapping segments. Thus, the extended student feature map 460 is denoted as N×C t , since each non-overlapping segment N is the same channel dimension as the teacher feature maps 440.

[0068] The teacher network 410 is used to perform many-to-one feature representation using N non-overlapping segments of the expanded student feature map 460 by matching the features on the expanded student feature map 460 with the features on the teacher feature map 440.

[0069] The teacher network 410 is used by the reparameterizer 470 to reparameterize the teacher feature map 440, where the reparameterizer 470 uses a combination of N non-overlapping segments of the expanded student feature map 460. In some examples, the combination can be the sum of the N non-overlapping segments. In other examples, the combination can be a difference, multiplication, division, etc.

[0070] The student network 420 is used to apply the result of the matching of the N non-overlapping segments to the expanded student feature map 460. As described above, the reparameterization of the teacher feature map 440 using the reparameterizer 470 and the application of the result of the matching from the N non-overlapping segments can be performed in parallel.

[0071] The student channel dimension contractor 475 (also known as Wsc) uses the student network 420 to create an output student feature map. This creation of the output student features can be done by projecting the expanded student feature map 560 back to its original channel dimension Cs. The output student feature map (also known as F so ) is then weighted to represent the fully connected student network layer 480 (e.g., the convolutional layer has been fully learned), also known as

[0072] Similarly, the teacher network 410 is used to weight the reparameterized teacher feature map (also known as F tr ) to represent the fully connected teacher network layer 490 (e.g., the convolutional layer has been fully learned), also known as

[0073] The fully connected student layer 480 and the fully connected teacher layer 490 are passed to the loss circuit module 250 to calculate the total loss value. The calculation of the total loss value is further described below with reference to Figure 8 the flowchart of.

[0074] Figure 5 is a flowchart representing example machine-readable instructions and / or example operations that can be executed and / or instantiated by a processor circuit module to implement Figure 1 the knowledge distillation architecture 100. Figure 5The knowledge distillation process 500 begins at block 510, where the data retriever circuit module 130 establishes a data source for the data to be retrieved. In some examples, as described above, the data can be contained in external data 170 (e.g., online sources, big data sources, etc.) accessible via network 125, or in a storage device 180 directly accessible by the data retriever circuit module 170. In such examples, the data retriever circuit module 130 can establish the source of the data, whether from external data 170 or the storage device 180.

[0075] The knowledge distillation process 500 then continues, where the data retriever circuit module 130 retrieves data from the established data source (block 520). In some instances, the data can be stored on an internal storage buffer within the edge device 110. In other examples, the data can be temporarily stored in the edge storage device 160. In other examples, the data can not be stored and can be directly passed to the knowledge distillation circuit module 140, where the data is deleted once the knowledge distillation process 500 is complete.

[0076] Next, the knowledge distillation circuit module 140 performs knowledge distillation on the data retrieved by the data retriever circuit module 130 (block 530). In some examples, the knowledge distillation circuit module 140 accesses the data from the data retriever circuit module 130. In other examples, the knowledge distillation circuit module 140 accesses the data stored on the edge storage device 160.

[0077] Once the knowledge distillation circuit module 140 performs knowledge distillation on the retrieved data, the output circuit module 150 outputs the result of the knowledge distillation (block 540). In some examples, the result is passed via network 125 to the data host 120 to output the result to a graphical user interface (GUI) for an operator to view the result. In other examples, the result is stored to the edge storage device 160 and / or the storage device 180 on the data host 120.

[0078] Figure 6 is a flowchart representing example machine-readable instructions and / or example operations that can be executed and / or instantiated by a processor circuit module to perform knowledge distillation on data retrieved by the data retriever circuit module 130. Execution Figure 6 of the knowledge distillation begins at block 610, where the feature map circuit module 220 generates a teacher feature map 440 and a student feature map 450.

[0079] Once the teacher feature map 440 and the student feature map 450 are generated, the feature map circuit module 220 identifies the channel dimensions for the teacher feature map 440 and the student feature map 450 (block 620). In the examples described herein and as mentioned above, the teacher feature map 440 has different channel dimensions from the student feature map 450. In some examples, the channel dimensions include the channel height, width, and number representing a three-dimensional feature map. In other examples, the feature maps 440, 450 may have additional dimensions or smaller dimensions.

[0080] Then, the feature map circuit module 220 determines whether the teacher feature map 440 and the student feature map 450 are in the same feature space (block 625). In some examples, the teacher feature map 440 is in a feature space different from the feature space of the student feature map 450 (e.g., the dimension storing the features of the teacher feature map). In such examples, once the feature map circuit module 220 identifies the channel dimensions of the feature maps 440, 450, the feature map circuit module 220 should transform the feature maps 440, 450 into the same feature space.

[0081] When the feature map circuit module 220 determines that the teacher map 440 and the student map 450 are not in the same feature space (e.g., the result of block 625 returns "no"), the feature map circuit module 220 can then perform a feature transformation on the teacher feature map 440 and / or the student feature map 450 to project the feature maps 440, 450 into the same feature space (block 630). In some examples, both feature maps 440, 450 can be transformed into the same feature space. Such examples may be desirable to reduce computational complexity by reducing the channel dimensions, modifying the channel dimensions, etc. In other examples, only one of the feature maps 440, 450 can be transformed.

[0082] Once the feature maps 440, 450 are transformed by the feature map circuit module 220, or when the feature map circuit module 220 determines that the feature maps 440, 450 are in the same feature space (e.g., the result of block 625 returns "yes"), the teacher training circuit module 230 and the student training circuit module 240 train the teacher network 410 and the student network 420 respectively (block 640). In the examples described herein, the teacher network 410 and the student network 420 are trained simultaneously (e.g., in parallel). However, in other examples, the teacher network 410 and the student network 420 can be trained serially.

[0083] Once the teacher training circuit module 230 and the student training circuit module 240 train the teacher network 410 and the student network 420, the loss circuit module 250 generates a total loss value based on the trained teacher network 410 and the trained student network 420 (block 650). In some examples, the total loss value is the result of the teacher network loss and the student network loss. In other examples, the total loss value may include additional loss characteristics. Refer to Figure 8 Further information regarding the generation of the total loss value is disclosed.

[0084] Once the loss circuit module 250 generates the total loss value, the result generator circuit module 260 generates a result to be sent to the output circuit module 150 (block 660). In some examples, the result is a report of the success, failure, efficiency, etc. of the knowledge distillation process performed. In some examples, the result is generated in a format that allows a user to visualize the training results of the teacher network 410 and the student network 420. In other examples, the result may be stored in the edge storage device 160 and / or the storage device 180. Once the result is generated by the result generator circuit module 260, Figure 6 the example operations of the flowchart are complete.

[0085] Figure 7 is a flowchart depicting example machine-readable instructions and / or example operations that may be executed and / or instantiated by the processor circuit module to train the teacher network 410 and the student network 420. The training Figure 7 of the teacher network 410 and the student network 420 begins at block 710, where the feature map projection circuit module 310 projects the student feature map 450 onto a predefined channel dimension. In some examples, the predefined channel dimension is equal to the channel dimension of the teacher feature map 440.

[0086] Once the student feature map 440 is projected by the feature map projection circuit module 310 onto the predefined channel dimension, the feature segmentation circuit module 320 then segments the projected student feature map into a predefined number of non-overlapping segments having the same channel dimension as the teacher feature map 440 (e.g., the expanded student feature map 460) (block 720). Any number of non-overlapping segments may be used to create the many-to-one feature representation described above. In some examples, it may be desirable to have fewer (e.g., eight or fewer) non-overlapping segments to increase computational time. In other examples, more than eight non-overlapping segments may be required to improve the accuracy of the learning process and reduce the total loss value.

[0087] Once the student feature map 450 has been segmented into a predefined number of non-overlapping segments, the matching circuit module 330 then begins training the teacher by matching the non-overlapping segments of the extended student feature map 460 with the features on the teacher feature map 440 (block 730). In some examples, the matching circuit module 330 matches the non-overlapping segments to preserve information passed from the student network 420 to the teacher network 410 and then back to the student network 420. The matching of non-overlapping segments allows the teacher network 410 to learn from many features of the extended student feature map 460, while traditional neural networks use one-to-one feature representations, which increases the computational time for producing the same results as the many-to-one feature representations described herein.

[0088] Once the matching circuit module 330 matches the non-overlapping segments with the teacher feature map 440, the reparameterization circuit module 340 reparameterizes the teacher feature map 440 using a combination of the non-overlapping student segments (block 740). In some examples, the reparameterization circuit module 330 allows backpropagation through the teacher network 410 and the student network 420. The reparameterization circuit module 340 facilitates the learning process of the teacher network 450.

[0089] Alternatively or concurrently, the matching application circuit module 350 applies the matching results of the non-overlapping segments to the extended student feature map 460 (block 750). In some examples, the application of the matching results of the non-overlapping segments trains the student network based on the output of the matching circuit module 330. In the examples described herein, the reparameterization circuit module 340 and the matching application circuit module 350 can be executed in parallel to train the teacher network 440 and the student network 450 simultaneously.

[0090] Once the teacher feature map 440 has been reparameterized and the matching of the non-overlapping segments has been applied to the extended student feature map 460, the feature map projection circuit module 310 projects the extended student feature map 460 back to the original channel dimension (e.g., the channel dimension of the student feature map 450) (block 760). In some examples, projecting back to the original channel dimension is used to allow for the execution of loss calculation on the output student feature map F so Once the extended student feature map 460 has been projected back to its original channel dimension, the training of the teacher network 410 and the training of the student network 420 are complete.

[0091] Figure 8 is a flowchart of example machine-readable instructions and / or example operations that may be executed and / or instantiated by a processor circuit module to generate a total loss value. Generating Figure 8 the total loss value begins at block 810, where the loss circuit module 250 generates a feature matching loss. In some examples, the feature matching loss is represented by Equation 1 below:

[0092]

[0093] As shown in Equation 1, N represents the number of non-overlapping segments (e.g., segmented from the student feature map 450), and F se represents the expanded student feature map 460, and F t represents the teacher feature map 440.

[0094] The loss circuit module 250 generates a teacher network loss (block 820) after the teacher has been re-parameterized (e.g., instructions corresponding to block 740). The generation of the teacher network loss is represented by Equation 2 below:

[0095]

[0096] As shown in Equation 2, are the weights of the fully connected teacher network layer (e.g., fully learned layer). CE is the cross-entropy loss (e.g., also known as log loss), which measures the performance of a classification model whose output is a probability value between 0 and 1. Softmax is a normalized exponential function that converts a vector of real numbers into a probability distribution of possible outcomes. Finally, GT is the ground truth (e.g., the target for training or validating a neural network with a labeled dataset).

[0097] After the student has been projected back to its original channel dimension (e.g., instructions corresponding to block 760), the loss circuit module 250 generates a student network loss (block 830). If represented by Equation 3 below, the generation of the student network loss is:

[0098]

[0099] As shown in Equation 3, are the weights of the fully connected student network layer (e.g., fully learned layer).

[0100] Corresponding to each individual loss contribution (e.g., the feature matching loss L mofd , the teacher network loss L t and the student network loss L s ), the loss circuit module 250 generates loss coefficients for each loss contribution (block 840). In some examples, the loss coefficients represented as the student network loss coefficient α, the teacher network loss coefficient β, and the feature matching loss coefficient γ are generated through experimental analysis (e.g., updating the loss coefficients based on the results of previous loss calculations). In other examples, the loss coefficients are constant.

[0101] Once the loss coefficients are generated, the loss circuit module 250 calculates the total loss value (block 850). The calculation of the total loss value is represented by Equation 4 below:

[0102] L = α·L s + β·L t + γ·L mofd Equation 4

[0103] As shown in Equation 4 above, L represents the total loss value as the feature matching loss L mofd with the feature matching loss coefficient γ, the teacher network loss L t with the teacher network loss coefficient β, and the student network loss L s and the student network loss coefficient α. The total loss value is used for the output of the results of the learning processes of the teacher network 410 and the student network 420. During each learning cycle (e.g., Figure 5 instructions), the total loss value is recalculated, and thus the results are different, where the neural network learns from the results to output better (e.g., more accurate and less data) results on subsequent learning cycles.

[0104] Figure 9 is constructed to execute and / or instantiate Figure 5 , Figure 6 , Figure 7 and / or Figure 8 machine-readable instructions and / or operations to implement Figure 2 of the knowledge distillation circuit module 140. The example processor platform 900 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cellular phone, a smart phone, a tablet such as an iPad TM tablet), a personal digital assistant (PDA), an Internet device, a game console, headphones (e.g., augmented reality (AR) headphones, virtual reality (VR) headphones, etc.) or other wearable devices or any other type of computing device.

[0105] The example processor platform 900 shown includes a processor circuit module 912. The example processor circuit module 912 shown is hardware. For example, the processor circuit module 912 can be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. The processor circuit module 912 can be implemented by one or more semiconductor (e.g., silicon-based) devices. In this example, the processor circuit module 912 implements the feature map circuit module 220, the teacher training circuit module 230, the student training circuit module 240, the loss circuit module 250, the result generator circuit module 260, the feature map projection circuit module 310, the feature segmentation circuit module 320, the matching circuit module 330, the reparameterization circuit module 340, and the matching application circuit module 350.

[0106] The processor circuit module 912 of the illustrated example includes a local memory 913 (e.g., cache, registers, etc.). The processor circuit module 912 of the illustrated example communicates with a main memory including a volatile memory 914 and a non-volatile memory 916 via a bus 918. The volatile memory 914 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), dynamic random access memory and / or any other type of RAM device. The non-volatile memory 916 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memories 914, 916 of the illustrated example is controlled by a memory controller 917.

[0107] The processor platform 900 of the illustrated example further includes an interface circuit module 920. The interface circuit module 920 can be implemented in hardware according to any type of interface standard (such as an Ethernet interface, a universal serial bus (USB) interface, interface, a near field communication (NFC) interface, a peripheral component interconnect (PCI) interface, and / or a peripheral component interconnect express (PCIe) interface).

[0108] In the illustrated example, one or more input devices 922 are connected to the interface circuit module 920. The input devices 922 allow a user to input data and / or commands into the processor circuit module 912. The input devices 922 can be implemented by, for example, audio sensors, microphones, cameras (static or video), keyboards, buttons, mice, touchscreens, touchpads, trackballs, point devices such as these, and / or a voice recognition system.

[0109] One or more output devices 924 are also connected to the interface circuit module 920 of the illustrated example. The output devices 924 can be implemented by, for example, display devices (e.g., light emitting diodes (LEDs), organic light emitting diodes (OLEDs), liquid crystal displays (LCDs), cathode ray tube (CRT) displays, in-plane switching (IPS) displays, touchscreens, etc.), haptic output devices, printers, and / or speakers. Thus, the interface circuit module 920 of the illustrated example generally includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit module, such as a GPU.

[0110] The interface circuit module 920 of the illustrated example also includes a communication device, such as a transmitter, receiver, transceiver, modem, residential gateway, wireless access point, and / or network interface, for facilitating data exchange with an external machine (e.g., any kind of computing device) via network 926. The communication can be carried out, for example, via an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a site-line wireless system, a cellular phone system, an optical connection, etc.

[0111] The processor platform 900 of the illustrated example also includes one or more mass storage devices 928 for storing software and / or data. Examples of such mass storage devices 928 include magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disc drives, redundant arrays of independent disks (RAID) systems, solid-state storage devices such as flash devices and / or SSDs, and DVD drives.

[0112] Can be implemented by Figure 5 、 Figure 6 、 Figure 7 and / or Figure 8 The machine-readable instructions 932 that can be implemented by the machine-readable instructions of

[0113] Figure 10 are Figure 9 A block diagram of an example implementation of the processor circuit module 912 of Figure 9 In this example, Figure 5 、 Figure 6 、 Figure 7 and / or Figure 8 Some or all of the machine-readable instructions of the flowchart of Figure 2 to effectively instantiate the knowledge distillation circuit module 140 of Figure 2The knowledge distillation circuit module 140 is instantiated by the hardware circuit of the microprocessor 1000 in combination with instructions. For example, the microprocessor 1000 can be implemented by a multi-core hardware circuit module (e.g., CPU, DSP, GPU, XPU, etc.). Although it can include any number of example cores 1002 (e.g., 1 core), the example microprocessor 1000 is a multi-core semiconductor device including N cores. The cores 1002 of the microprocessor 1000 can operate independently or can cooperate to execute machine-readable instructions. For example, the machine code corresponding to a firmware program, an embedded software program, or a software program can be executed by one of the cores 1002, or can be executed by multiple cores among the cores 1002 at the same or different times. In some examples, the machine code corresponding to a firmware program, an embedded software program, or a software program is divided into threads and executed in parallel by two or more of the cores 1002. The software program can correspond to part or all of the machine-readable instructions and / or operations represented by the Figure 5 , Figure 6 , Figure 7 and / or Figure 8 flowchart.

[0114] The cores 1002 can communicate via a first example bus 1004. In some examples, the first bus 1004 can be implemented by a communication bus to enable communication associated with one (or more) of the cores 1002. For example, the first bus 1004 can be implemented by at least one of an integrated circuit (I2C) bus, a serial peripheral interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the first bus 1004 can be implemented by any other type of computing bus or electrical bus. The cores 1002 can obtain data, instructions, and / or signals from one or more external devices via an example interface circuit module 1006. The cores 1002 can output data, instructions, and / or signals to one or more external devices via the interface circuit module 1006. Although the example cores 1002 include an example local memory 1020 (e.g., a level 1 (L1) cache that can be divided into an L1 data cache and an L1 instruction cache), the microprocessor 1000 also includes an example shared memory 1010 (e.g., a level 2 (L2) cache) that can be shared by the cores for high-speed access to data and / or instructions. Data and / or instructions can be transferred (e.g., shared) by writing to and / or reading from the shared memory 1010. The local memory 1020 and the shared memory 1010 of each core among the cores 1002 can be a memory including multiple levels of cache memory and a main memory (e.g., Figure 9is part of the storage device hierarchy of the main memories 914, 916). Generally, higher-level memories in the hierarchy exhibit lower access times and have smaller storage capacities compared to lower-level memories. Changes to the various levels of the cache hierarchy are managed (e.g., coordinated) by a cache coherence policy.

[0115] Each core 1002 can be referred to as a CPU, DSP, GPU, etc., or any other type of hardware circuit module. Each core 1002 includes a control unit circuit module 1014, an arithmetic and logic (AL) circuit module (sometimes referred to as an ALU) 1016, multiple registers 1018, a local memory 1020, and a second example bus 1022. Other structures may exist. For example, each core 1002 can include a vector unit circuit module, a single instruction multiple data (SIMD) unit circuit module, a load / store unit (LSU) circuit module, a branch / jump unit circuit module, a floating-point unit (FPU) circuit module, etc. The control unit circuit module 1014 includes semiconductor-based circuitry configured to control (e.g., coordinate) data movement within the corresponding core 1002. The AL circuit module 1016 includes semiconductor-based circuitry configured to perform one or more mathematical and / or logical operations on data within the corresponding core 1002. Some example AL circuit modules 1016 perform integer-based operations. In other examples, the AL circuit module 1016 also performs floating-point operations. In other examples, the AL circuit module 1016 can include a first AL circuit module that performs integer-based operations and a second AL circuit module that performs floating-point operations. In some examples, the AL circuit module 1016 can be referred to as an arithmetic logic unit (ALU). The registers 1018 are semiconductor-based structures for storing data and / or instructions, such as the results of one or more operations performed by the AL circuit module 1016 of the corresponding core 1002. For example, the registers 1018 can include vector registers, SIMD registers, general-purpose registers, flag registers, segment registers, machine-specific registers, instruction pointer registers, control registers, debug registers, memory management registers, machine check registers, etc. The registers 1018 can be arranged in banks as Figure 10 shown. Alternatively, the registers 1018 can be organized in any other arrangement, format, or structure including being distributed throughout the core 1002 to shorten access times. The second bus 1022 can be implemented by at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.

[0116] Each core 1002 and / or, more generally, microprocessor 1000 may include additional and / or alternative structures to those shown and described above. For example, there may be one or more clock circuits, one or more power supplies, one or more power gates, one or more cache coherence agents (CHA), one or more aggregation / common mesh stops (CMS), one or more shifters (e.g., (multiple) barrel shifters), and / or other circuit modules. Microprocessor 1000 is a semiconductor device fabricated to include a number of transistors that are interconnected to implement the above structures in one or more integrated circuits (ICs) contained within one or more packages. The processor circuit module may include one or more accelerators and / or cooperate with one or more accelerators. In some examples, the accelerator is implemented by a logic circuit module to perform certain tasks faster and / or more efficiently than a general-purpose processor can. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. A GPU or other programmable device may also be an accelerator. The accelerator may be on the processor circuit module, in the same chip package as the processor circuit module, and / or in one or more packages separate from the processor circuit module.

[0117] Figure 11 is Figure 9 A block diagram of another example implementation of processor circuit module 912. In this example, processor circuit module 912 is implemented by FPGA circuit module 1100. For example, FPGA circuit module 1100 may be implemented by an FPGA. FPGA circuit module 1100 can be used, for example, to perform operations that might otherwise be performed by Figure 10 example microprocessor 1000. However, once configured, FPGA circuit module 1100 instantiates machine-readable instructions in hardware and can thus generally perform operations faster than a general-purpose microprocessor executing the corresponding software.

[0118] More specifically, compared to the above Figure 10 microprocessor 1000 (which is a general-purpose device that can be programmed to execute some or all of the machine-readable instructions represented by the Figure 5 , Figure 6 , Figure 7 and / or Figure 8 flowcharts, but whose interconnections and logic circuit modules are fixed once fabricated), Figure 11 example FPGA circuit module 1100 of Figure 5 , Figure 6 , Figure 7 and / or Figure 8Some or all of the machine-readable instructions represented by the flowchart of. In particular, the FPGA circuit module 1100 can be considered an array of interconnects, switches, and logic gates. The switches can be programmed to change how the logic gates are interconnected via the interconnects, effectively forming one or more dedicated logic circuits (unless and until the FPGA circuit module 1100 is reprogrammed). The configured logic circuits enable the logic gates to cooperate in different ways to perform different operations on the data received by the input circuit module. These operations can correspond to Figure 5 , Figure 6 , Figure 7 and / or Figure 8 some or all of the software represented by the flowchart of. Thus, the FPGA circuit module 1100 can be constructed to effectively instantiate Figure 5 , Figure 6 , Figure 7 and / or Figure 8 some or all of the instances of the machine-readable instructions of the flowchart of into dedicated logic circuits to perform operations corresponding to those software instructions in a dedicated manner similar to an ASIC. Therefore, the FPGA circuit module 1100 can perform operations corresponding to some or all of Figure 5 , Figure 6 , Figure 7 and / or Figure 8 of the machine-readable instructions faster than a general-purpose microprocessor can perform operations corresponding to some or all of Figure 5 , Figure 6 , Figure 7 and / or Figure 8 of the machine-readable instructions.

[0119] In an example of Figure 11 , the FPGA circuit module 1100 is constructed to be programmed (and / or reprogrammed one or more times) by an end user via a hardware description language (HDL) such as Verilog. Figure 11 The FPGA circuit module 1100 of Figure 11 includes an example input / output (I / O) circuit module 1102 for obtaining data and / or outputting data to / from an example configuration circuit module 1104 and / or external hardware 1106. For example, the configuration circuit module 1104 can be implemented by an interface circuit module that can obtain machine-readable instructions to configure the FPGA circuit module 1100 or its part(s). In some such examples, the configuration circuit module 1104 can obtain machine-readable instructions from a user, a machine (e.g., a hardware circuit module that can implement an artificial intelligence / machine learning (AI / ML) model to generate instructions (e.g., a programming or dedicated circuit module)), etc. In some examples, the external hardware 1106 can be implemented by an external hardware circuit module. For example, the external hardware 1106 can be implemented by Figure 10implemented by the microprocessor 1000. The FPGA circuit module 1100 also includes an example logic gate circuit module 1108, an array of multiple example configurable interconnections 1110, and an example storage circuit module 1112. The logic gate circuit module 1108 and the configurable interconnections 1110 can be configured to instantiate one or more operations of at least some of the machine-readable instructions that can correspond to Figure 5 , Figure 6 , Figure 7 and / or Figure 8 and / or other desired operation instances. Figure 11 The logic gate circuit module 1108 shown in is fabricated in groups or blocks. Each block includes a semiconductor-based electrical structure that can be configured into a logic circuit. In some examples, the electrical structure includes logic gates (e.g., AND gates, OR gates, NAND gates, etc.) that provide basic building blocks for the logic circuit. Electro-controllable switches (e.g., transistors) are present within each of the logic gate circuit modules 1108 to enable the configuration of the electrical structure and / or the logic gates to form a circuit that performs the desired operation. The logic gate circuit module 1108 can include other electrical structures such as look-up tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.

[0120] The configurable interconnections 1110 of the example shown in are conductive circuit paths, traces, vias, etc., which can include electro-controllable switches (e.g., transistors) whose states can be changed by programming (e.g., using an HDL instruction language) to activate or deactivate one or more connections between one or more of the logic gate circuit modules 1108 to program the desired logic circuit.

[0121] The storage circuit module 1112 of the example shown in is constructed to store the result of one or more operations performed by the corresponding logic gates. The storage circuit module 1112 can be implemented by registers, etc. In the example shown, the storage circuit module 1112 is distributed among the logic gate circuit modules 1108 to facilitate access and improve execution speed.

[0122] Figure 11The example FPGA circuit module 1100 also includes an example dedicated operation circuit module 1114. In this example, the dedicated operation circuit module 1114 includes a specific-purpose circuit module 1116, which can be invoked to implement common functions, so as to avoid the need to program these functions on-site. Examples of such specific-purpose circuit modules 1116 include memory (e.g., DRAM) controller circuit modules, PCIe controller circuit modules, clock circuit modules, transceiver circuit modules, memory, and multiplier-accumulator circuit modules. There can be other types of specific-purpose circuit modules. In some examples, the FPGA circuit module 1100 can also include an example general-purpose programmable circuit module 1118, such as the example CPU 1120 and / or the example DSP 1122. There can additionally or alternatively be other general-purpose programmable circuit modules 1118 that can be programmed to perform other operations, such as GPUs, XPU, etc.

[0123] Although Figure 10 and Figure 11 show Figure 9 two example implementations of the processor circuit module 912, many other methods can be envisioned. For example, as described above, modern FPGA circuit modules can include on-board CPUs, such as Figure 11 one or more of the example CPUs 1120 in Figure 9 Therefore, Figure 10 the processor circuit module 912 in Figure 11 can alternatively be implemented by combining Figure 5 the example microprocessor 1000 in Figure 6 and Figure 7 the example FPGA circuit module 1100 in Figure 8 . In some such hybrid examples, the first part of the machine-readable instructions represented by the flowcharts in Figure 10 can be executed by one or more cores 1002 in the core 1002 in Figure 5 , Figure 6 , Figure 7 and / or Figure 8 , the second part of the machine-readable instructions represented by the flowcharts in Figure 11 can be executed by the FPGA circuit module 1100 in Figure 5 , Figure 6 , Figure 7 and / or Figure 8 , and the third part of the machine-readable instructions represented by the flowcharts in Figure 2 can be executed by the ASIC. It should be understood that Figure 2 some or all of the knowledge distillation circuit modules 140 in can thus be instantiated at the same or different times. Some or all of the circuit modules can be instantiated, for example, in one or more threads executing simultaneously and / or serially. Additionally, in some examples,Figure 2 Some or all of the knowledge distillation circuit module 140 can be implemented within one or more virtual machines and / or containers that can be executed on a microprocessor.

[0124] In some examples, Figure 9 the processor circuit module 912 can be in one or more packages. For example, Figure 10 the microprocessor 1000 and / or Figure 11 the FPGA circuit module 1100 can be in one or more packages. In some examples, the XPU can be implemented by Figure 9 the processor circuit module 912, which can be in one or more packages. For example, the XPU can include a CPU in one package, a DSP in another package, a GPU in yet another package, and an FPGA in still another package.

[0125] Figure 12 illustrates a block diagram of an example software distribution platform 1205 for distributing software, such as example machine-readable instructions 932, to a hardware device owned and / or operated by a third party. The example software distribution platform 1205 can be implemented by any computer server, data facility, cloud service, etc. that can store software and transfer the software to other computing devices. The third party can be a customer of the entity that owns and / or operates the software distribution platform 1205. For example, the entity that owns and / or operates the software distribution platform 1205 can be a developer, seller, and / or licensor of software (such as example machine-readable instructions 932). The third party can be a consumer, user, retailer, OEM, etc. that purchases and / or licenses the software for use and / or resale and / or sub-licensing. In the illustrated example, the software distribution platform 1205 includes one or more servers and one or more storage devices. The storage device stores the machine-readable instructions 932, which can correspond to the example machine-readable instructions of Figure 9 as described above, Figure 9 The one or more servers of the example software distribution platform 1205 communicate with an example network 1210, which can correspond to the Internet and / or any one or more of the example networks 926 described above. In some examples, the one or more servers respond to a request to send the software to a requester as part of a commercial transaction. Payment for the delivery, sale, and / or license of the software can be handled by one or more servers of the software distribution platform and / or by a third-party payment entity. The server enables a purchaser and / or licensee to download the machine-readable instructions 932 from the software distribution platform 1205. For example, it can correspond to Figure 5 , Figure 6 , Figure 7 and / or Figure 8 of the example machine-readable instructions. Figure 5 , Figure 6 ,Figure 7 and / or Figure 8 The software of example machine-readable instructions of Figure 7 can be downloaded to an example processor platform 900 that executes the machine-readable instructions 932 to implement the knowledge distillation circuit module 140. In some examples, one or more servers of the software distribution platform 1205 periodically provide, send, and / or enforce updates to the software (e.g., Figure 9 example machine-readable instructions 932 of Figure 7 ) to ensure that improvements, patches, updates, etc. are distributed and applied to the software at the end-user device.

[0126] In view of the foregoing, it should be understood that example systems, methods, apparatuses, and articles for training a teacher network and a student network simultaneously without a pre-trained teacher model have been disclosed. The disclosed systems, methods, apparatuses, and articles improve the efficiency of using a computing device by simultaneously training a teacher network and a student network using an online data source with a many-to-one feature representation method. Accordingly, the disclosed systems, methods, apparatuses, and articles relate to one or more improvements in the operation of machines such as computers or other electronic and / or mechanical devices.

[0127] Example methods, apparatuses, systems, and articles for training a teacher network and a student network simultaneously without a pre-trained teacher model are disclosed herein. Other examples and combinations thereof include the following:

[0128] Example 1 includes an apparatus that includes at least one memory, machine-readable instructions, and a processor circuit module that is configured to perform at least one of instantiating or executing the machine-readable instructions for: generating, based on a query, a first feature map for a teacher network, the first feature map having a first channel dimension; generating, based on the query, a second feature map for a student network, the second feature map having a second channel dimension that is different from the first channel dimension; partitioning the second feature map into segments having the first channel dimension; using the segments of the second feature map to train the teacher network; and training the student network by applying a total loss value to the student network, the total loss value being based on a loss function, wherein the teacher network and the student network are trained on an edge device.

[0129] Example 2 includes the apparatus of example 1, wherein the processor circuit module is further configured to match the segments of the second feature map with the first feature map to train the teacher network.

[0130] Example 3 includes the apparatus as described in example 1, wherein the processor circuit module is further configured to reparameterize the teacher network based on a combination of the segments of the second feature map.

[0131] Example 4 includes the apparatus as described in Example 1, wherein the processor circuit module is further configured to retrieve data from a data source, the data including information related to the query.

[0132] Example 5 includes the apparatus of Example 4, wherein the processor circuit module is further configured to establish a data source before retrieving the data.

[0133] Example 6 includes the apparatus as described in any one of Examples 1-5, wherein the processor circuit module is further configured to project the second feature map onto a predefined channel dimension before performing the segmentation, the predefined channel dimension being different from the second channel dimension.

[0134] Example 7 includes the apparatus of Example 6, wherein the processor circuit module is further configured to project the second feature map back to the second channel dimension after training the teacher network.

[0135] Example 8 includes the apparatus of Example 1, wherein the segments of the second feature map are non-overlapping segments.

[0136] Example 9 includes the apparatus of Example 1, wherein the processor circuit module is further configured to generate a total loss value.

[0137] Example 10 includes the apparatus of Example 9, wherein the total loss value includes a combination of a matching loss corresponding to the matching of the segments of the second feature map with the first feature map, a teacher network loss, and a student network loss.

[0138] Example 11 includes the apparatus of Example 1, wherein the processor circuit module is further configured to produce results corresponding to the training of the teacher network and the training of the student network.

[0139] Example 12 includes the apparatus of Example 1, wherein the training of the teacher network and the training of the student network are performed simultaneously.

[0140] Example 13 includes an apparatus for performing feature distillation in a neural network, including: a unit for generating a first feature map for a teacher network and a second feature map for a student network, the first feature map having a first channel dimension and the second feature map having a second channel dimension different from the first channel dimension; a unit for segmenting the second feature map into segments having the first channel dimension; a first unit for training the teacher network using the segments of the second feature map; and a second unit for training the student network using the output of a loss function, wherein the first unit and the second unit for training are implemented on an edge device.

[0141] Example 14 includes the apparatus of Example 13, wherein the first unit for training the teacher network further includes a unit for matching the segments of the second feature map with the first feature map.

[0142] Example 15 includes the apparatus of Example 14, wherein the second unit for training the student network further includes a unit for applying the result of the matching unit to the student network.

[0143] Example 16 includes the apparatus of Example 13, wherein the first unit for training the teacher network further includes a unit for reparameterizing the teacher network based on a combination of segments of the second feature map.

[0144] Example 17 includes the apparatus as described in Example 13, wherein the second unit for training the student network further includes a unit for projecting the second feature map onto a predefined channel dimension before segmenting the second feature map, the predefined channel dimension being different from the second channel dimension.

[0145] Example 18 includes the apparatus of Example 17, wherein the projection unit is the first unit for projection, and the second unit for training the student network further includes a second unit for projecting the second feature map back to the second channel dimension after training the teacher network.

[0146] Example 19 includes the device of Example 13, further including a unit for calculating a total loss value corresponding to the output of a loss function.

[0147] Example 20 includes the apparatus of Example 13, further including a unit for generating the results of the first training unit and the second training unit.

[0148] Example 21 includes a non-transitory machine-readable storage medium including instructions that, when executed, cause a processor circuit module to at least: generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; generate a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; segment the second feature map into segments having the first channel dimension; train the teacher network using the segments of the second feature map; and train the student network by applying a total loss value to the student network, the total loss value being based on a loss function, wherein the teacher network and the student network are trained on an edge device.

[0149] Example 22 includes the non-transitory machine-readable storage medium of Example 21, wherein the instructions, when executed, further cause the processor circuit module to match the segments of the second feature map with the first feature map to train the teacher network.

[0150] Example 23 includes the non-transitory machine-readable storage medium of Example 21, wherein the instructions, when executed, further cause the processor circuit module to reparameterize the teacher network based on a combination of the segments of the second feature map.

[0151] Example 24 includes the non-transitory machine-readable storage medium described in Example 21, wherein the instructions, when executed, further cause the processor circuitry to retrieve data from a data source, the data including information related to the query.

[0152] Example 25 includes the non-transitory machine-readable storage medium described in Example 24, wherein the instructions, when executed, further cause the processor circuitry to establish a data source before retrieving the data.

[0153] Example 26 includes the non-transitory machine-readable storage medium described in Example 21, wherein the instructions, when executed, further cause the processor circuitry to project the second feature map onto a predefined channel dimension before performing the segmentation, the predefined channel dimension being different from the second channel dimension.

[0154] Example 27 includes the non-transitory machine-readable storage medium described in any one of Examples 21-26, wherein the instructions, when executed, further cause the processor circuitry to project the second feature map back to the second channel dimension after training the teacher network.

[0155] Example 28 includes the non-transitory machine-readable storage medium described in Example 21, wherein the instructions, when executed, further cause the processor circuitry to generate the total loss value using a combination of a matching loss, a teacher network loss, and a student network loss corresponding to the matching of the segments of the second feature map with the first feature map.

[0156] Example 29 includes the non-transitory machine-readable storage medium described in Example 21, wherein the instructions, when executed, further cause the processor circuitry to produce results corresponding to the training of the teacher network and the training of the student network.

[0157] Example 30 includes a method for training a teacher network and a student network, including: generating a first feature map for the teacher network based on a query, the first feature map having a first channel dimension; generating a second feature map for the student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; segmenting the second feature map into segments having the first channel dimension; training the teacher network using the segments of the second feature map; and training the student network by applying a total loss value to the student network, wherein the training of the teacher network and the training of the student network are performed on an edge device.

[0158] Example 31 includes the method of Example 30, further including matching segments of the second feature map with the first feature map to train the teacher network.

[0159] Example 32 includes the method of Example 30 and further includes reparameterizing the teacher network based on a combination of segments of the second feature map.

[0160] Example 33 includes the method of Example 30 and further includes retrieving data from a data source, the data containing information related to the query.

[0161] Example 34 includes the method of Example 33 and further includes establishing the data source before retrieving the data.

[0162] Example 35 includes the method of Example 30 and further includes projecting the second feature map onto a predefined channel dimension before performing the segmentation, the predefined channel dimension being different from the second channel dimension.

[0163] Example 36 includes the method of any one of Examples 30 - 35 and further includes projecting the second feature map back to the second channel dimension after training the teacher network.

[0164] Example 37 includes the method of Example 30 and further includes generating a total loss value by employing a loss function.

[0165] Example 38 includes the method of Example 30, wherein generating the total loss value further includes: generating a matching loss corresponding to the matching of segments of the second feature map with the first feature map; generating a teacher network loss; generating a student network loss; and calculating the total loss value by using a combination of the matching loss, the teacher network loss, and the student network loss.

[0166] Example 39 includes the method of Example 30 and further includes producing results corresponding to the training of the teacher network and the training of the student network.

[0167] Example 40 includes the method of Example 30, wherein the training of the teacher network and the training of the student network are performed simultaneously.

[0168] The appended claims are hereby incorporated by reference into this detailed description. Although certain example systems, methods, devices, and articles have been disclosed herein, the scope of this patent is not limited thereto. On the contrary, this patent covers all systems, methods, devices, and articles that fall entirely within the scope of the claims of this patent.

Claims

1. A device, comprising: at least one memory; machine-readable instructions; and a processor circuit module for instantiating or executing at least one of the machine-readable instructions for: generating a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; generating a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; segmenting the second feature map into segments having the first channel dimension; using the segments of the second feature map to train the teacher network; and training the student network by applying a total loss value to the student network, the total loss value being based on a loss function; wherein the teacher network and the student network are trained on an edge device.

2. The device according to claim 1, wherein the processor circuit module is further configured to match the segments of the second feature map with the first feature map to train the teacher network.

3. The device according to claim 1, wherein the processor circuit module is further configured to reparameterize the teacher network based on a combination of the segments of the second feature map.

4. The device according to claim 1, wherein the processor circuit module is further configured to retrieve data from a data source, the data containing information related to the query.

5. The device according to claim 4, wherein the processor circuit module is further configured to establish the data source before retrieving the data.

6. The device according to any one of claims 1-5, wherein the processor circuit module is further configured to project the second feature map to a predefined channel dimension different from the second channel dimension before performing the segmentation.

7. The device according to claim 6, wherein the processor circuit module is further configured to project the second feature map back to the second channel dimension after training the teacher network.

8. The device according to claim 1, wherein the segments of the second feature map are non-overlapping segments.

9. The device according to claim 1, wherein the processor circuit module is further configured to generate the total loss value.

10. The device according to claim 9, wherein the total loss value includes a combination of: a matching loss corresponding to the matching of the segments of the second feature map with the first feature map, a teacher network loss, and a student network loss.

11. The device according to claim 1, wherein the processor circuit module is further configured to produce results corresponding to the training of the teacher network and the training of the student network.

12. The device according to claim 1, wherein the training of the teacher network and the training of the student network are performed simultaneously.

13. A device for performing feature distillation in a neural network, comprising: A unit for generating a first feature map for a teacher network and a second feature map for a student network, the first feature map having a first channel dimension and the second feature map having a second channel dimension different from the first channel dimension; A unit for splitting the second feature map into segments having the first channel dimension; A first unit for training the teacher network using the segments of the second feature map; And A second unit for training the student network using the output of a loss function; Wherein, the first unit and the second unit for training are implemented on an edge device.

14. The apparatus according to claim 13, Wherein, The first unit for training the teacher network further includes: a unit for matching the segments of the second feature map with the first feature map.

15. The apparatus according to claim 14, Wherein, The second unit for training the student network further includes a unit for applying the result of the matching unit to the student network.

16. The apparatus according to claim 13, Wherein, The first unit for training the teacher network further includes a unit for reparameterizing the teacher network based on a combination of the segments of the second feature map.

17. The apparatus according to claim 13, Wherein, The second unit for training the student network further includes a unit for projecting the second feature map onto a predefined channel dimension before splitting the second feature map, the predefined channel dimension being different from the second channel dimension.

18. The apparatus according to claim 17, Wherein, The projection unit is a first unit for projection, and the second unit for training the student network further includes a unit for projecting the second feature map back to the second channel dimension after training the teacher network.

19. The apparatus according to claim 13, further includes a unit for calculating a total loss value corresponding to the output of the loss function.

20. A non-transitory machine-readable storage medium including instructions that, when executed, cause a processor circuit module to at least: Generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; Generate a second feature map for a student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; Split the second feature map into segments having the first channel dimension; Train the teacher network using the segments of the second feature map; And Train the student network by applying a total loss value to the student network, the total loss value being based on a loss function; Wherein, the teacher network and the student network are trained simultaneously on an edge device.

21. A method for training a teacher network and a student network, Including: Generate a first feature map for a teacher network based on a query, the first feature map having a first channel dimension; Generate a second feature map for the student network based on the query, the second feature map having a second channel dimension different from the first channel dimension; Segment the second feature map into segments having the first channel dimension; Use the segments of the second feature map to train the teacher network; and Train the student network by applying a total loss value to the student network; wherein the training of the teacher network and the training of the student network are performed on an edge device.

22. The method according to claim 21, further comprising matching the segments of the second feature map with the first feature map to train the teacher network.

23. The method according to claim 21, further comprising reparameterizing the teacher network based on a combination of the segments of the second feature map.

24. The method according to claim 21, further comprising projecting the second feature map to a predefined channel dimension before performing the segmentation, the predefined channel dimension being different from the second channel dimension.

25. The method according to any one of claims 21-24, further comprising projecting the second feature map back to the second channel dimension after training the teacher network.