Artificial neural network processing method and system
By training students' ANN module on edge devices and using composite loss functions for knowledge refining, the deployment problem of complex neural networks on edge devices is solved, and efficient and highly adaptable model compression and performance retention are achieved.
Patent Information
- Application Number
- CN202510047280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-17
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively deploy complex artificial neural network models on edge devices with limited computing resources and memory resources, and the existing knowledge refining methods have problems of poor performance and limited adaptability.
By providing a large teacher ANN module and a relatively simple student ANN module, the student module is trained using the composite loss function to reduce errors, realize knowledge refinement, reduce computational complexity, and deploy compressed models on edge devices.
Efficient deployment of complex neural networks is realized on edge devices, maintaining the performance of the model, adapting to different complex backbones, and reducing computing complexity and storage requirements.
Smart Images

Figure CN120337992A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of Italian Patent Application No. 102024000000861, filed on January 18, 2024, which is hereby incorporated herein by reference. Technical field
[0003] This specification relates to a method and system for processing artificial neural networks.
[0004] One or more embodiments may relate to a processing device configured to perform neural network processing operations, such as an edge computing processing device. Background art
[0005] Complex artificial neural network processing models (currently referred to as "backbones" or "machine learning") may require computing resources and / or data storage resources that exceed the capabilities of edge processing devices (e.g., such as microcontrollers).
[0006] One of the problems of adapting large - scale machine learning models and applications to edge computing is the limited computing resources of the latter.
[0007] Existing solutions to this problem include attempts to "distill" (or compress) the knowledge obtained from large models into smaller models with reduced computational usage.
[0008] For example, the following literature discusses existing solutions:
[0009] Hinton, G.E., Vinyals, O., & Dean, J. (2015): "Distilling the Knowledge in a Neural Network", ArXiv, abs / 1503.02531 discusses a method of compressing the knowledge in an ensemble into a single model, which is more easily deployable by introducing a new type of ensemble consisting of one or more full models and many specialist models that learn to distinguish the confusions of the full models at a fine - grained class level;
[0010] Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., & Bengio, Y. (2014): "FitNets: Hints for Thin Deep Nets", CoRR, abs / 1412.6550 discussed knowledge distillation, which is used to allow training of a student node that is deeper and thinner than the teacher node, thus using the intermediate representations learned by the teacher as hints to improve the training process and final performance of the student node; and
[0011] Yim, J., Joo, D., Bae, J., & Kim, J. (2017): "A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning", 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 7130 - 7138 discussed a novel knowledge transfer technique where knowledge from a pre-trained deep neural network (DNN) is distilled and transferred to another DNN, showing that the student DNN learning the distilled knowledge is optimized much faster than the original model and outperforms the original DNN.
[0012] Existing solutions have one or more of the following drawbacks:
[0013] Poor performance and automation;
[0014] Limited ability to adapt to different complex backbones, especially for embedded solutions, and
[0015] Reduced distillation ability for large models. Summary of the Invention
[0016] An object of one or more embodiments is to contribute to overcoming the foregoing drawbacks.
[0017] According to one or more embodiments, this object can be achieved by a method having the features set forth in the appended claims.
[0018] A computer-implemented method can be an example of such a method.
[0019] One or more embodiments can relate to a corresponding processing device.
[0020] One or more embodiments may include a computer program product that is loadable in a memory of at least one processing circuit (e.g., a computer) and includes software code portions for performing steps of a method when the product is run on the at least one processing circuit. As used herein, a reference to such a computer program product should be understood as equivalent to a reference to a computer-readable medium containing instructions for controlling a processing system to coordinate the implementation of a method according to one or more embodiments. A reference to “at least one computer” is intended to emphasize the possibility of implementing one or more embodiments in a modular and / or distributed form.
[0021] The claims are an integral part of the technical teachings provided herein with reference to the embodiments.
[0022] One or more embodiments facilitate the deployment of complex machine learning methods on relatively simple devices such as microcontrollers.
[0023] One or more embodiments may be deployed on a group of microcontrollers arranged in a joint configuration.
[0024] One or more embodiments may utilize the addition of neuron tags with embedded encryption codes. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] One or more embodiments will now be described by way of non-limiting example only with reference to the drawings, in which:
[0026] Figure 1 is a schematic example of a deep neural network (DNN) topology;
[0027] Figure 2 is a schematic example of a first stage of a method according to the present disclosure;
[0028] Figure 3 is a schematic example of a second stage of a method according to the present disclosure;
[0029] Figure 4 is a schematic example of a signal processing pipeline according to the present disclosure;
[0030] Figure 5 and Figure 6 is a schematic example of a performance benchmark of one or more embodiments;
[0031] Figure 7 and Figure 8 is a schematic example of an alternative performance benchmark of one or more embodiments; and
[0032] Figure 9 is a schematic example of a processing device according to the present disclosure.
[0033] Unless otherwise noted, corresponding reference numerals and symbols in different figures generally denote corresponding parts.
[0034] The figures are drawn to clearly illustrate relevant aspects of the embodiments, and the figures are not necessarily drawn to scale.
[0035] The edges of the features drawn in the figures do not necessarily indicate the termination of the scope of the feature. Detailed Description
[0036] In the following description, one or more specific details are illustrated, aiming to provide an in - depth understanding of examples of the embodiments of this specification. Embodiments can be obtained without one or more specific details, or by using other methods, components, materials, etc. In other cases, well - known structures, materials, or operations are not illustrated or described in detail so as not to obscure certain aspects of the embodiments.
[0037] References to "an embodiment" or "one embodiment" in the context of this specification are intended to indicate that a particular configuration, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, phrases such as "in an embodiment" or "in one embodiment" that may appear in one or more places in this specification do not necessarily refer to one and the same embodiment.
[0038] Furthermore, a particular construction, structure, or characteristic may be combined in any suitable manner in one or more embodiments.
[0039] As used herein, the term "or" is an inclusive "or" operator and is equivalent to the phrase "A or B, or both" or "A or B or C, or any combination thereof", and is treated similarly for lists with additional elements. The term "based on" is not exclusive and allows for additional features, functions, aspects, or limitations not described, unless the context otherwise indicates. Also, throughout the specification, the meanings of "a" and "the" include both singular and plural references.
[0040] The reference numerals used herein are provided for convenience only and thus do not limit the scope of protection or the scope of the embodiments.
[0041] For simplicity, in the following detailed description, the same reference numeral symbols may be used to denote nodes / lines in a circuit and the signals that may appear at that node or line.
[0042] The term "processing device" may be used interchangeably hereinafter to refer to a "processing system" and is intended to represent a computing device / system that can easily process data signals.
[0043] The term "data set" may be used hereinafter to refer to a collection of like or unlike signals that may be stored in at least one data storage unit (or memory), such as a database accessible via an Internet connection.
[0044] A variety of technical fields, such as, for example, computer vision, speech recognition, and / or signal processing applications, may benefit from the use of an artificial neural network ANN, which is a processing method that can quickly apply hundreds, thousands, or even millions of parallel processing operations to data signals. As discussed in this disclosure, ANN methods may fall under the technical categories of learning / inference machines, machine learning, artificial intelligence, artificial neural networks, probabilistic inference engines, backbones, etc.
[0045] Such a learning / inference machine may have an underlying topology or architecture currently known as a deep convolutional neural network (DCNN).
[0046] A DCNN is a computer-based tool that applies data processing to large amounts of data and adaptively "learns" to perform pattern recognition on the data by combining adjacent related features within the data, thereby making broad predictions and refining the predictions based on reliable conclusions and new combinations.
[0047] For example, a convolutional neural network (CNN) is a type of DCNN.
[0048] As Figure 1 illustrated, the CNN pipeline 100 includes multiple "layers" 12, 13, 14, 16, 18, and different types of data processing operations, such as feature extraction 11 and / or classification 15, are performed at each layer.
[0049] The most commonly used layer types are convolutional layers 13, fully connected or dense layers 16, and pooling layers 14 (max pooling, average pooling, etc.). The data exchanged between layers is referred to as features.
[0050] As can be understood by those skilled in the art, each layer of the CNN 10 includes multiple computational units, which are currently represented as perceptrons, and the description is performed via a parameter tuple. These parameters may include, for example:
[0051] A set of learnable parameters, commonly referred to as weights W, and
[0052] Other parameters P, such as activation function type, padding, stride, etc., depending on the type of ANN processing layer.
[0053] A processing layer configured to apply ANN processing (e.g., convolution) to the input data provided at the input layer, thereby providing processed data at the output layer is currently referred to as a "hidden layer".
[0054] CNNs are particularly suitable for recognition tasks, such as recognizing numbers or objects in images, and can provide highly accurate results.
[0055] As can be understood by those skilled in the art, the computations performed by a CNN or by other neural networks typically involve repetitive computations on large amounts of data. Thus, such "large" models can be executed on computer devices with hardware acceleration subsystems or on a wide network including computing resources and data storage resources, such as the computing resources and data storage resources of a server.
[0056] The inventors have observed that in order to perform operations similar to those available under large machine learning in an environment with limited computing resources and memory resources, the "large" ANN stage can teach the "smaller" ANN stage how to process data, thereby contributing to an almost lossless compression of the machine learning model in terms of its performance.
[0057] For simplicity, one or more embodiments are mainly discussed herein with reference to a convolutional neural network (CNN), which is a deep neural network (DNN) topology for a large or "teacher" ANN network, but it should be understood that one or more embodiments can theoretically be applied to any complex ANN topology or pipeline.
[0058] As Figure 2 and Figure 3 illustrated in
[0059] providing a first "teacher" ANN module 20, such as a large CNN processing pipeline 10 having dozens of layers of various different types;
[0060] providing a second "student" ANN module 30, such as a smaller ANN with a complexity at least one order of magnitude lower;
[0061] providing a training data set TD (e.g., a set of labeled images) that includes calibration data of known ground truth;
[0062] training the teacher ANN module 20 to perform artificial neural network (ANN) processing (e.g., classifying images in the training data set), while training the student ANN module 30 to perform the same operation as the teacher by using a composite loss function that takes into account the performance of both ANN processing pipelines 20, 30 in reducing the error regarding the known classification.
[0063] Using Figure 2In the method illustrated in [reference], a trained teacher ANN module 20T (with fixed weight values) and a partially trained student ANN module 30' can be obtained, where the weight values of the student ANN module 30' are based on the "observation" of the learning process of the teacher ANN module 20.
[0064] As Figure 3 Illustrated in [reference], the "knowledge distillation" method for reducing ANN processing complexity includes, in the second stage (currently also referred to as the "inference stage"):
[0065] Providing an additional dataset UD (currently also referred to as the "unlabeled dataset") for which the ground truth cannot be obtained a priori;
[0066] Applying ANN processing to the unlabeled dataset UD using both the trained teacher ANN module 20T and the partially trained student ANN module 30';
[0067] Minimizing the loss function of the student by reducing the error of the output provided by the student ANN module 30' relative to the output provided by the trained teacher ANN module 20T.
[0068] Figure 2 And Figure 3 The training operation illustrated in [reference] includes minimizing at least one loss function LOSS based on the mean squared error (MSE) between the logits z of the teacher 20 and the student 30.
[0069] For example, the loss function L can be expressed as:
[0070] L(z s (τ), z t (τ)) = ‖z S (τ) - z t (τ)‖ 2 2
[0071] Where
[0072] z s(τ) Represents the logits of the teacher ANN module 20;
[0073] z T(τ) Represents the logits of the student ANN module 30.
[0074] The logit function Z is mathematically defined as the logarithm of the odds of the probability p of an event occurring, which can be expressed as:
[0075] z(p) = log(p / (1 - p))
[0076] Where p represents the probability of the event, and log represents the natural logarithm.
[0077] As exemplified herein, the logit function Z is used as a link function to map a probability (between 0 and 1) to a real number, which can then be used to express a linear relationship.
[0078] In one or more embodiments, known ANN processing pipelines are suitable for use as teacher ANN modules 20, 20T, such as those discussed in the following:
[0079] Bellitto, G., Proietto Salanitri, F., Palazzo, S. et al.: “Hierarchical Domain-Adapted Feature Learning for Video Saliency Prediction”, Int J Comput Vis 129, 3216-3232 (2021), doi:10.1007 / s11263-021-01519-y;
[0080] “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, ArXiv (2020), abs / 2010.11929.
[0081] For example, the teacher ANN module includes a CNN processing stage or a transformer network processing stage.
[0082] For example, one of these computer program products may use 100GB of RAM, and thanks to the method according to the present disclosure, a student ANN module can be trained, which can reproduce its performance on a processing device limited to 4MB of data storage space.
[0083] Figure 4 is a schematic example of the “knowledge distillation” pipeline according to the present disclosure, which can be used in the first stage exemplified in Figure 2 and / or in the second stage exemplified in Figure 3 according to the method of the present disclosure.
[0084] As Figure 4 exemplified in
[0085] the teacher ANN module 20T includes a plurality of ANN processing layers 22, 23, 24, 26, 28, the plurality of ANN processing layers including an input layer 22, a convolutional layer 23, a pooling layer 24, a fully connected layer 26 and an output layer 28, and
[0086] The student ANN module 30 includes an input layer 32, a general hidden layer 35, and an output layer 38.
[0087] Thus, the topologies of the student ANN modules 30, 30' are significantly simpler than the structures of the teacher ANN modules 20, 20T (e.g., one-third smaller in the example of Figure 4 ).
[0088] For simplicity, this configuration of the teacher and student ANN modules 20, 20T, 30, 30' is illustrated, but it should be understood that this configuration is merely exemplary and in no way limiting.
[0089] In one or more embodiments, the topologies of the student ANN modules 30, 30' can consider the processing capabilities of edge devices (e.g., microcontroller devices) in a heuristic manner at design time, e.g., to find a trade-off between application and computational performance.
[0090] As Figure 4 illustrated, the unlabeled dataset UD includes a set of images. Similarly, Figure 4 this type of training data is illustrated only for simplicity, but it should be understood that in theory any type of unlabeled data can be used to perform knowledge distillation as illustrated herein.
[0091] In one or more embodiments, known datasets such as the publicly available Cifar-100 and / or ImageNette datasets can be advantageously used. The Canadian Institute for Advanced Research, the CIFAR-100 dataset is a collection of images commonly used for training machine learning and computer vision algorithms. ImageNette is a subset of ten "easy-to-classify" classes from the ImageNette dataset. It was originally prepared by Jeremy Howard of FastAI.
[0092] As Figure 4 illustrated:
[0093] The images of the datasets TD, UD are processed by the networks of the teachers 20, 20T and the students 30, 30';
[0094] The output data of the teacher ANN modules 20, 20T and the output data of the student ANN modules 30, 30' are compared;
[0095] A distillation loss L is calculated based on the output data of the teacher ANN modules 20, 20T and the output data of the student ANN modules 30, 30' D ( Figure 4 box 40 in
[0096] Based on the output data of the student ANN modules 30, 30', calculate the classification loss L of the student ANN modules 30, 30'. CE ( Figure 4 box 42 in
[0097] Based on the distillation loss L D and the classification loss L CE calculate the total loss L( Figure 4 box 44 in
[0098] Backpropagate the calculated total loss value L to the student ANN modules 30, 30' (preferably also to the teacher ANN module 20), and adjust the parameters of the student ANN modules 30, 30' (such as the weights Ws of each layer and / or other parameters Ps) until the total loss value L reaches a relative minimum.
[0099] In one or more embodiments, a selection can be made among different metric functions to measure the loss function used in the pipeline exemplified in Figure 4 .
[0100] As Figure 4 exemplified in, the classification loss LCE can be expressed as:
[0101]
[0102] where
[0103] p j is the probability distribution of the log-odds error,
[0104] y j is the target label to be assigned,
[0105] τ is the time variable.
[0106] For example, the KLD or MSE function can be used for the distillation loss function LD.
[0107] For example, using the KLD function, the distillation function can be expressed as
[0108]
[0109] where
[0110] p j is the probability distribution of the log-odds error,
[0111] y j is the target label to be assigned,
[0112] τ is the time variable.
[0113] For example, the softmax function can be used to calculate the probability distribution of the log-odds z, which can be expressed as:
[0114]
[0115]
[0116] where
[0117] z k represents the k-th log-odds or error value, which indicates the difference between the ground truth table and the label assigned by the corresponding processing stage.
[0118] For example, Figure 4 box 44 of
[0119]
[0120] where
[0121] α is a value in the range [0, 1], preferably closer to the upper limit value (one) of this range, so as to better propagate to the student ANN module 30.
[0122] As Figure 4 illustrated in
[0123] As illustrated herein, the "teacher" networks 20, 20T pre-trained on large datasets are used as a guide for developing the compressed networks 30, 30', and the operational functions of the teacher ANN modules 20, 20' are transferred to the compressed networks 30, 30' without repeating the same computational complexity.
[0124] Compared with the "teacher", the compressed models 30, 30' have a reduced number of ANN parameters and / or a simpler topology, thus resulting in compatibility with edge devices with limited processing capabilities. For example, the STM32 cube device can be equipped with the compressed models 30, 30' to perform compressed ANN processing.
[0125] One or more embodiments use a total loss function L, which includes a first loss function L CE of the student 30, 30' and a distilled loss function L D based on the result comparison of the student 30, 30' relative to the teacher 20, 20T The weighted sum. For example, the total loss L can be expressed as:
[0126] For example, the Kullback-Leibler divergence or the mean squared error can be used as the distillation loss L D .
[0127] As Figure 4 illustrated in, the method according to the present disclosure includes:
[0128] Providing a first artificial neural network ANN processing stage 20, 20T, the processing stage including a first set of ANN processing layers 22, 23, 24, 26, 28; and
[0129] Providing a second ANN processing stage 30, 30', the processing stage including a second set of ANN processing layers 32, 35, 38, the second set of ANN processing layers having a set of processing layer parameters Ws, Ps (including at least a set of ANN processing weights W s ).
[0130] As illustrated herein, the number of processing layers in the first set of ANN processing layers is greater than the number of processing layers in the second set of ANN processing layers.
[0131] As Figure 4 illustrated in, the method further includes:
[0132] Applying the first ANN processing to at least one input data set TD; UD via the first ANN processing stage 20T, generating a first set of output values z (T) as a result;
[0133] Applying the second ANN processing to at least one input data set via the second ANN processing stage 30', generating a second set of output values z (S) as a result;
[0134] Calculating 40 a first loss value L based on the first set of output values and based on the second set of output values D ;
[0135] Calculating 42 a second loss value L based on the second set of output values CE ;
[0136] Based on the first loss value L D and the second loss value L CE Calculating 44 the total loss L, and
[0137] Adjusting 46 the values of the processing layer parameters in the set of processing layer parameters Ws, Ps (including at least a set of weight values Ws) of the second ANN processing stage 30' based on the total loss value.
[0138] As Figure 4 illustrated in, the method further includes:
[0139] Apply normalization processing to the first set of output values z provided by the first ANN processing stage (T) , to obtain a first set of probability values;
[0140] Apply normalization processing to the second set of output values (z (S) ) provided by the second ANN processing stage, to obtain a second set of probability values;
[0141] Calculate a first loss value based on the first set of probability values and based on the second set of probability values, and
[0142] calculate a second loss value based on the second set of probability values.
[0143] For example, applying the normalization processing exemplified in Figure 4 includes applying the softmax function to the corresponding set of output values.
[0144] As exemplified in Figure 4 , calculating the first loss value includes calculating the mean squared error MSE between the first set of output values (z (T) ) and the second set of output values, or calculating the Kullback-Leibler divergence KLDiv between the first set of output values and the second set of output values.
[0145] As exemplified in Figure 4 , calculating the total loss based on the first loss value and based on the second loss value includes calculating a linear combination of the first loss value and the second loss value.
[0146] For example, the total loss L is expressed as:
[0147] L = (1 - α)·L CE + α·L D
[0148] where,
[0149] α is a parameter within the range of values from 0 to 1, preferably within the range of values from 0.5 to 0.9;
[0150] L CE is the first loss value, and
[0151] L D is the second loss value.
[0152] As exemplified in Figure 4 , providing the first artificial neural network ANN processing stage includes providing a convolutional neural network CNN processing stage or a transformer network processing stage.
[0153] As exemplified in Figure 4As illustrated, the number of processing layers in the first group of ANN processing layers (22, 23, 24, 26, 28) is at least three times the corresponding number of processing layers in the second group of ANN processing layers (32, 35, 38).
[0154] Figures 5 to 8 It is an exemplary diagram of the performance of students 30, 30' relative to teachers 20, 20T for the same dataset for different datasets and teacher / student ANN module topologies.
[0155] Figure 5 and Figure 6 relates to the benchmark of student ANN module VGG 11 and teacher ANN module Vision Transformer (ViT)-16, where the student ANN module includes a CNN containing a first plurality (about 107 million) of parameters, and the teacher ANN module contains a second plurality (about 307 million, e.g., about three times the first plurality) of parameters.
[0156] As Figure 5 and Figure 6 illustrated, both student VGG 11 and teacher ViT-16 receive the ImageNette dataset containing N = 10 classes as input data TD, UD.
[0157] Figure 5 is a plot of the accuracy (ordinate, in percentage) of student ANN module VGG 11 and teacher ANN module ViT-16 in the ImageNette scenario over time (abscissa, in epochs).
[0158] Figure 6 is a plot of the loss function L (ordinate, in percentage) of student ANN module VGG 11 and teacher ANN module ViT-16 in the ImageNette scenario over time (abscissa, in epochs).
[0159] As Figure 5 and Figure 6 illustrated, in the ImageNette scenario, student model VGG 11 achieves an accuracy of approximately 77%, and when calculating the distillation loss L using the KLDiv expression D the accuracy can be improved to 79.65%.
[0160] Figure 7 and Figure 8 relates to student ANN module VGG 11The benchmark of the student ANN module Vision Transformer (ViT)-16 and the teacher ANN module, where the student ANN module includes a CNN containing a first plurality (about 107 million) of parameters, and the teacher ANN module contains a second plurality (about 307 million, e.g., about three times the first plurality) of parameters.
[0161] As Figure 7 and Figure 8 illustrated in, the student VGG 11 and the teacher ViT-16 both receive CIFAR100 with N = 100 classes as input data TD, UD.
[0162] Figure 7 is the graph of the evolution of the accuracy (ordinate, in percentage) of the student ANN module VGG 11 and the teacher ANN module ViT-16 over time (abscissa, in epochs) in the CIFAR100 scenario.
[0163] Figure 8 is the graph of the evolution of the loss function L (ordinate, in percentage) of the student ANN module VGG 11 and the teacher ANN module ViT-16 over time (abscissa, in epochs) in the CIFAR100 scenario.
[0164] As Figure 7 and Figure 8 illustrated in, in the CIFAR100 scenario, the student model VGG 11 achieves an accuracy of about 62%, and when using the KLDiv expression to calculate the distillation loss L D the accuracy can be increased to almost 65%.
[0165] Figure 9 is a block diagram of a processing device or system 90 suitable for executing the instructions of the student ANN modules 30, 30'.
[0166] As Figure 9 illustrated in, the system 90 includes:
[0167] One or more processing cores or circuits 92, which are configured to control the overall operation of the system 90, the execution of the system 90 on application programs (e.g., programs for classifying images using CNN), etc.;
[0168] One or more memories 94, such as one or more volatile and / or non-volatile memories, which may store all or part of the instructions and data related to, for example, the control of system 90, the applications and operations executed by system 90, etc.; for example, the weight values Ws and ANN parameters Ps of the student ANN modules 30, 30' may be stored in the memory 94 of system 90;
[0169] One or more sensors 96 (e.g., image sensors, audio sensors, accelerometers, pressure sensors, temperature sensors, etc.);
[0170] One or more interfaces 97 (e.g., wireless communication interfaces, wired communication interfaces, etc.), and
[0171] Other circuitry 98, which may include antennas, power supplies, one or more built-in self-tests (abbreviated, BIST) circuits, etc., and a main bus system 99.
[0172] For example, the processing core 92 may include one or more processors, state machines, microprocessors, programmable logic circuits, discrete circuitry, logic gates, registers, etc. and / or various combinations thereof.
[0173] For example, one or more of the memories in memory 94 may include a memory array, which may be shared by one or more processes executed by system 90 during operation.
[0174] For example, the main bus system 99 may include one or more data, address, power, and / or control buses coupled to various components of system 90.
[0175] As Figure 9 illustrated, system 90 preferably further includes one or more hardware accelerators 100, which accelerate the execution of one or more operations associated with implementing a CNN during operation. The illustrated hardware accelerator 100 includes one or more convolution accelerators to facilitate, for example, the efficient execution of convolutions associated with the convolutional layers of a CNN.
[0176] As illustrated herein, a computer program product includes instructions that, when executed by a computer, cause the computer to perform Figure 4 the method illustrated in
[0177] As illustrated herein, a computer-readable medium has stored therein the values of the set of processing layer parameters Ws, Ps obtained using Figure 4 the method illustrated in
[0178] As Figure 4 and Figure 9As illustrated herein, a method of a processing device 90 whose operation is configured to perform artificial neural network (ANN) processing according to a set of processing layer parameters Ws, Ps includes:
[0179] Accessing 94 the values of the sets of processing layer parameters Ws, Ps obtained using the Figure 4 method illustrated herein,
[0180] and performing artificial neural network (ANN) processing 30, 30' according to the values of the sets of processing layer parameters.
[0181] As illustrated herein, a computer program product includes instructions that, when executed by a processing device 90, cause the processing device to perform ANN processing according to the method of the present disclosure.
[0182] As illustrated herein, a computer-readable medium includes instructions that, when executed by a processing device 90, cause the processing device to perform ANN processing according to the method illustrated herein.
[0183] As Figure 9 illustrated herein, the processing device 90 includes a storage circuitry 94 in which there has been stored:
[0184] The adjusted values of the set of processing layer parameters Ws, Ps obtained using Figure 4 the method illustrated herein, and
[0185] instructions that, when executed in the processing device, cause the processing device to:
[0186] Access 94 the adjusted values of the set of processing layer parameters, and
[0187] perform ANN processing according to the adjusted values of the set of processing layer parameters.
[0188] For example, the processing device includes a microcontroller device.
[0189] It will be understood that the various individual implementation options illustrated in the drawings throughout this specification are not necessarily intended to be employed in the same combinations as illustrated in each of the figures. Thus, one or more embodiments may employ these (otherwise non-mandatory) options individually and / or in combinations different from those illustrated in the figures.
[0190] Without departing from the underlying principles, details and embodiments may vary, even significantly, from what has been described only by way of example, without departing from the scope of protection. The scope of protection is defined by the appended claims.
Claims
1. A computer-implemented method, comprising: Providing a first artificial neural network (ANN) processing stage, the first artificial neural network (ANN) processing stage including a first set of ANN processing layers; Providing a second ANN processing stage, the second ANN processing stage including a second set of ANN processing layers, the second set of ANN processing layers having a set of processing layer parameters, the set of processing layer parameters including at least a set of ANN processing weights, a first number of processing layers in the first set of ANN processing layers being greater than a second number of processing layers in the second set of ANN processing layers; Applying, via the first ANN processing stage, a first ANN processing to at least one input data set to produce a first set of output values; Applying, via the second ANN processing stage, a second ANN processing to the at least one input data set to produce a second set of output values; Calculating, via a processing core, a first loss value based on the first set of output values and based on the second set of output values; Calculating, via the processing core, a second loss value based on the second set of output values; Calculating, via the processing core, a total loss value based on the first loss value and based on the second loss value; And Adjusting, via the processing core, a first value of the processing layer parameters of the second ANN processing stage based on the total loss value.
2. The method according to claim 1, comprising: Applying a normalization process to the first set of output values provided by the first ANN processing stage to provide a first set of probability values; Applying a normalization process to the second set of output values provided by the second ANN processing stage to provide a second set of probability values; Calculating the first loss value based on the first set of probability values and based on the second set of probability values; And Calculating the second loss value based on the second set of probability values.
3. The method according to claim 2, wherein applying the normalization process includes applying a softmax function to a corresponding set of output values.
4. The method according to claim 1, wherein calculating the first loss value includes calculating: The mean squared error (MSE) between the first set of output values and the second set of output values; or The Kullback-Leibler divergence (KLDiv) between the first set of output values and the second set of output values.
5. The method according to claim 1, wherein calculating the total loss value includes calculating a linear combination of the first loss value and the second loss value.
6. The method according to claim 5, wherein the total loss value is L and is expressed as: L = (1 - α)·L CE + α·L D Where α is a parameter between 0 and 1, including 0 and 1, L CE is the first loss value, and L D is the second loss value.
7. The method according to claim 6, wherein α is between 0.5 and 0.9, including 0.5 and 0.
9.
8. The method according to claim 1, wherein providing the ANN processing stage includes providing: A convolutional neural network (CNN) processing stage, or A transformer network processing stage.
9. The method according to claim 1, wherein the first number of processing layers in the first set of ANN processing layers is at least three times as large as the second number of processing layers in the second set of ANN processing layers.
10. The method according to claim 1, further comprising: Access the adjusted first value of the processing layer parameters for the second ANN processing stage; And Perform ANN processing based on at least the first value of the processing layer parameters for the second ANN processing stage.
11. The method according to claim 1, wherein the data in the at least one input dataset is obtained from one or more sensors, and the one or more sensors include at least one of the following: an image sensor, an audio sensor, an accelerometer, a pressure sensor, a temperature sensor.
12. A non-transitory computer-readable medium storing computer instructions that, when executed by a processing device, cause the processing device to perform the following steps: Provide a first artificial neural network (ANN) processing stage that includes a first set of ANN processing layers; Provide a second ANN processing stage that includes a second set of ANN processing layers, the second set of ANN processing layers having a set of processing layer parameters that includes at least a set of ANN processing weights, and a first number of processing layers in the first set of ANN processing layers is greater than a second number of processing layers in the second set of ANN processing layers; Apply a first ANN processing to at least one input dataset via the first ANN processing stage to generate a first set of output values; Apply a second ANN processing to the at least one input dataset via the second ANN processing stage to generate a second set of output values; Calculate a first loss value based on the first set of output values and based on the second set of output values; Calculate a second loss value based on the second set of output values; Calculate a total loss value based on the first loss value and based on the second loss value; And Adjust a first value of the processing layer parameters for the second ANN processing stage based on the total loss value.
13. The non-transitory computer-readable medium according to claim 12, including further instructions that, when executed by a processing device, cause the processing device to perform the following steps: Apply a normalization process to the first set of output values provided by the first ANN processing stage to provide a first set of probability values; Apply a normalization process to the second set of output values provided by the second ANN processing stage to provide a second set of probability values; Calculate the first loss value based on the first set of probability values and based on the second set of probability values; And Calculate the second loss value based on the second set of probability values.
14. The non-transitory computer-readable medium according to claim 12, wherein the instructions that cause the processing device to calculate the first loss value include instructions that cause the processing device to calculate one of the following: The mean squared error (MSE) between the first set of output values and the second set of output values; or The Kullback-Leibler divergence (KLDiv) between the first set of output values and the second set of output values.
15. The non-transitory computer-readable medium according to claim 12, wherein the instructions that cause the processing device to calculate the total loss value include instructions that cause the processing device to calculate a linear combination of the first loss value and the second loss value.
16. The non-transitory computer-readable medium of claim 12, wherein the instructions that cause the processing device to provide the ANN processing phase include instructions that cause the processing device to provide: a convolutional neural network (CNN) processing phase, or a transformer network processing phase.
17. A processing device, comprising: a non-transitory memory circuitry containing instructions; and a processing core in communication with the memory circuitry, wherein the processing core executes the instructions to: provide a first ANN processing phase including a first set of artificial neural network (ANN) processing layers; provide a second ANN processing phase including a second set of ANN processing layers, the second set of ANN processing layers having a set of processing layer parameters including at least a set of ANN processing weights, a first number of processing layers in the first set of ANN processing layers being greater than a second number of processing layers in the second set of ANN processing layers; apply a first ANN processing to at least one input data set via the first ANN processing phase to produce a first set of output values; apply a second ANN processing to the at least one input data set via the second ANN processing phase to produce a second set of output values; calculate a first loss value based on the first set of output values and based on the second set of output values; calculate a second loss value based on the second set of output values; calculate a total loss value based on the first loss value and based on the second loss value; and adjust a first value of the processing layer parameters of the second ANN processing phase based on the total loss value.
18. The processing device of claim 17, wherein the processing core executes further instructions to: apply a normalization process to the first set of output values provided by the first ANN processing phase to provide a first set of probability values; apply a normalization process to the second set of output values provided by the second ANN processing phase to provide a second set of probability values; calculate the first loss value based on the first set of probability values and based on the second set of probability values; and calculate the second loss value based on the second set of probability values.
19. The processing device of claim 17, wherein the processing core executing the instructions to calculate the first loss value includes the processing core executing the instructions to calculate: the mean squared error (MSE) of the first set of output values and the second set of output values; or the Kullback-Leibler divergence (KLDiv) of the first set of output values and the second set of output values.
20. The processing device of claim 17, wherein the processing core executing the instructions to calculate the total loss value includes the processing core executing the instructions to calculate a linear combination of the first loss value and the second loss value.
21. The processing device of claim 17, wherein the processing core executing the instructions to provide the ANN processing phase includes the processing core executing the instructions to provide: a convolutional neural network (CNN) processing phase, or a transformer network processing phase.