A Quantification Method for Deep Convolutional Neural Networks
Patent Information
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- ASELSAN ELEKTRONIK SANAYI & TICARET ANONIM SIRKETI
- Filing Date
- 2024-12-16
- Publication Date
- 2026-06-22
Abstract
Description
A QUANTIFICATION METHOD FOR DEEP CONVENTIONAL NEURAL NETWORKS Technical Area This breakthrough improves quantization in deep convolutional neural networks compared to the current state of the art. According to the methods, deep processing involves less power and resource usage in the processor. It relates to a quantization method for convolutional neural networks. Previous Technique Convolutional Neural Networks (CNNs) are used in image processing and commonly used in problems that work with high-dimensional data, such as video analysis. These are deep learning models used. CNNs detect local patterns in images and Features such as classification and object recognition are learned through learning abilities. They provide high accuracy in tasks. CNNs are generally used to process large amounts of data and It learns by processing many parameters. This process is quite intensive during training. computationally intensive operation and inference phase, especially with limited hardware. This can create challenges in devices with limited capacity. This is where quantization comes into play. (quantization) comes into play. Quantization is the process of converting the weights and activations of neural networks into integers. represented by lower precision data types (typically 8-bit) This is the process. This method reduces the memory usage and computational speed of neural networks. It is used to optimize. Through quantification, neural network models become smaller. They are made more efficient and can be operated with fewer resources. This is especially true for mobile devices. It is particularly useful in embedded systems and hardware with energy constraints. Quantization in CNNs involves weighting and activating the model using lower bit rates. It is the process of representing things at different levels. Generally, this happens after neural network training is complete. then full precision (32-bit floating point) weights, to lower bit levels are reduced. Weights and activations are re-expressed as 8-bit integers, Thus, the model uses less memory and performs calculations faster. Quantization The process is generally carried out in the following steps: 1 - Model Training: CNNs initially require high precision (typically 32-bit floating-point). It is trained with (dotted) weights. During training, the model's parameters are optimized. By doing this, important features in the input data are learned. Quantification Application: The pre-trained network is quantized and retrained for the desired bit width. In this process, each floating-point weight and activation has a specific scaling. It is rescaled by a factor and represented with a limited depth. This, This makes the model lighter. Also, the weights and activations... Representing the model at a lower resolution reduces the computational burden. and enables it to operate with lower hardware requirements. - Inference Phase: After quantification, CNN now uses sliding scales when making inferences. It works with integer arithmetic instead of decimal operations. This improves the model's performance. It accelerates and increases energy consumption, especially in embedded systems or mobile devices. It saves money. The main advantage of quantification is that it reduces the size of the model and the computational process. It speeds things up, especially on mobile devices or those with limited processing power. Quantization is an ideal solution for improving the performance of CNNs in systems. It also provides energy efficiency, which is important for devices with limited battery life. However, quantification can negatively impact model accuracy in some cases. When weights and activations are expressed in lower bits, the original model Its sensitivity may decrease somewhat, which can lead to losses in accuracy. Brief Description of the Invention The aim of this invention is to improve deep convolutional neural networks in the current state of the art. According to quantization methods, power and resource usage in the processor is lower. a quantization method for deep convolutional neural networks that enables this to accomplish. The aim of this invention is to improve deep convolutional neural networks in the current state of the art. According to quantization methods, the load on the processor and memory unit is reduced. The aim is to implement a quantization method for convolutional neural networks. 2 The first step taken to achieve the purpose of this invention, and the steps associated with that step... to operate within a processor and a memory unit defined in the requirements adapted, local patterns in the deep image located within the memory unit and through their learning abilities, classification and object recognition. Determining the weight value in convolutional neural networks used in the processes. It is a quantization method and its characteristic is that it is entered into the memory unit by the user. positioned, layers and pre-trained floating point weights The processor of any layer selected from a model that has the specified values the weight histogram of the layer read by the processor extracting the values and saving them into the memory unit, memory unit The processor calculates the average (μ) of the weight histogram values positioned within it. calculated by and saved to the memory unit, into the memory unit The standard deviation (σ) of the positioned weighted histogram values is calculated by the processor. the calculation and saving of the data to the memory unit by the processor The cropping limit is calculated using the formula = − 2 and Saving to the memory unit, the upper clipping limit by the processor, = ü Calculation using + 2 formulas and saving to memory unit, The weight histogram values calculated by the processor represent the absolute value of the layer. normalized value obtained by dividing by the maximum histogram values. ( ) finding and saving to the memory unit, the operation by the user The processor is provided with information on how many bits (bits) will be used, and the processor then processes this information. ü Data retrieved from memory unit is processed using the formula = Calculation of the SW weighting scale factor and the calculated weighting scale Saving the factor (SW) to the memory unit is quantized by the processor. Calculation of the weight value = ğ ğ and saving to the memory unit, for each layer located within the memory unit the repetition and memory of the quantized weight value obtained for each layer. The procedure for recording it in the unit includes the steps involved. The known technique... In this case, multiple users are used in convolutional neural networks. 3. Previously saved into the memory unit and processed by the processor Because the weight values used during the weighting process are numerous. It places a load on the processor and memory unit. Processor and memory unit To reduce the load on it, the quantification method of the invention is weighted processing. By first applying it, the convolutional neural network will try to reach the correct result. It reduces the number of weight values. The subject of the invention is the quantification method and the technique. according to the quantization methods in deep convolutional neural networks in their known state This reduces the load on the processor and memory unit. Detailed Description of the Invention In order to make the invention more understandable, explanations will be given. The symbols used and their descriptions are as follows: μ: Average σ: Standard deviation : Lower cropping limit : Upper cropping limit ü Norm: Normalize : Normalized value Q: scale factor SW: Weight scale factor Bit: Bit value Memory is a component adapted to operate within a processor and a memory unit. Learning local patterns and features in the deep image within the unit Thanks to its capabilities, it is used in classification and object recognition processes. A quantification method for determining the weight value in convolutional neural networks. its characteristic is, - layers positioned within the memory unit by the user and a model with pre-trained floating-point weight values The processor reads any selected layer from within it, 4 - Extracting the weight histogram values of the layer read by the processor. and being saved into the memory unit, - weight histogram values located within the memory unit the average (μ) is calculated by the processor and stored in the memory unit. recording, - Standard weight histogram values located within the memory unit the deviation (σ) is calculated by the processor and stored in the memory unit. recording, - The lower trimming limit set by the processor is calculated using the formula = − 2 Calculation using and saving to memory unit, - The upper clipping limit set by the processor is = + 2 formula ü Calculation using and saving to memory unit, - the weight histogram values calculated by the processor, absolute value of the layer normalized value obtained by dividing by the maximum histogram values. ( ) finding and saving to the memory unit, - information provided by the user regarding the number of bits (bits) required for the operation entering the processor, ü - data retrieved from the memory unit by the processor = Calculation of the SW weighting scale factor using the formula and Saving the calculated weight scale factor (SW) to the memory unit, - the weight value quantized by the processor is ğ ğ = calculation and saving to memory unit, - to be repeated for each layer located within the memory unit and each the quantized weight value obtained for the layer is stored in the memory unit. The recording process includes the steps involved. In the current state of the technique, multiple methods are used in convolutional neural networks. data previously saved by the user into the memory unit and processed by the processor The weight value is determined by the processor and memory units due to their large number. It creates a load on the processor and memory unit. Reducing the load on the processor and memory unit. For the invention, the subject of the quantification method is the process of determining the weight value. By applying this, the weights that the convolutional neural network will try to reach the correct result It reduces the number of values. The subject of the invention is quantified using a quantitative method from the beginning. The processor and memory units of a convolutional neural network are determined by assigning a small number of weight values. This also helps to reduce the load on it. In one application of the invention, it is designed to operate within a processor and a memory unit. local in the deep image located within an adapted memory unit Through their ability to learn patterns and characteristics, they can classify and analyze objects. Convolutional neural networks, used in processes such as recognition, are a known technique. by reducing the load on the processor and memory unit according to the methods in this situation It is a quantification method that works efficiently, and its characteristic is... - the memory unit contains different layers previously defined by the user. from within the model that has trained floating-point weight values the layer being read by the processor, - Extracting the weight histogram values of the layer read by the processor and being saved into the memory unit, - weight histogram values located within the memory unit Calculation of the average (μ) by the processor, - Standard weight histogram values located within the memory unit the calculation of the deviation (σ) by the processor, - The lower trimming limit set by the processor is calculated using the formula = − 2 Calculation using, - The upper clipping limit set by the processor is calculated using the formula = + 2. ü Calculation using, - the weighted histogram values by the processor, absolute maximum histogram Dividing by its values to find the normalized value ( ) and memory to be recorded in the unit, - information provided by the user regarding the number of bits (bits) required for the operation entering the processor, 6 ü - by the processor using the = formula sw weight Calculation of the scale factor (SW) and weighting of the calculated scale factor. (SW) saving to the memory unit, - the weight value quantized by the processor is ğ ğ = calculation and saving to memory unit, - repeating the first step for each layer within the memory unit and The quantized weight value obtained for each layer is stored in the memory unit. recording, - input pixel values entered into the memory unit by the user read by the processor, - the histogram of the pixel values of the read image is generated by the processor created and saved to the memory unit, - the processor calculates the average (μ) of the generated histogram, - the standard deviation (σ) value of the generated histogram is determined by the processor. calculation, - upper clipping limit set by the processor = +2 formula ü Calculation using, - input pixels to find the normalized (norm) value by the processor normalized by dividing the values by the absolute maximum input pixel value. Finding the value of the value obtained ( ) and saving it to the memory unit, - retrieving a user-defined bit value from the memory unit, ü - scale factor (S) by the processor using the formula = Calculation and saving of the calculated s scale factor to the memory unit, - activation value by the processor = calculation and saving to memory unit, - Normalized weight values and activation obtained by the processor. the use of values in convolutional neural networks, - Presenting the data generated by the convolutional neural network to the user for approval, 7 - If the user indicates that the output data is incorrect, the weight in the first step will be adjusted. Updating the values and restarting the process, - The process stops if the user indicates that the output data is correct. It includes the methodological steps. A histogram of the weight values of a convolutional neural network layer, with a mean of zero. It has a shape similar to the Gaussian distribution. Histogram of activation values. It also has a Gaussian distribution, but the negative values of the Gaussian distribution It is equal to zero. Based on this information, the most meaningful values are around zero. It is observed that this is because the highest weight or activation values are zero. Its value is around zero. As it moves away from zero, the frequency of the value increases. is decreasing. Based on this information, the minimum and maximum values are determined for each We limited it to an adaptive number for the layer or channel and made it easier in the quantization process. Frequently used coefficients can be quantified in more detail. In this way, in-depth analysis is possible. We have observed that the performance rate increases when we train convolutional cipher networks. To achieve appropriate accuracy levels, a fixed limit value is not sufficient, and It should be adjustable to suit each layer or channel. High precision (32-bit) values to achieve maximum accuracy. A model trained using the Quantization-Sensitive Training (QAT) method is selected. The model is retrained using this method, taking into account the effects of quantification, and much more. a model / network that can operate with low bandwidth weighting and activation values This results in a quantified model that provides efficient inference. It is optimized for the process. This enables real-time applications on edge devices. a rapid and resource-efficient extraction process that is vital for This enables the process. The resulting final model is then used to determine the model's suitability for its task. The inference process is carried out as follows. The algorithm integrates quantification into the training process, thereby measuring the parameters of the model. It ensures that performance is optimized with the least possible compromise. Each During each epoch, the algorithm assesses the effects of quantification during inference. To minimize this, forward and backward convolutions are quantized in each layer. It carries out backpropagation iteratively. In quantized convolutional flow. First, the weight and activation values are subjected to a quantification process. After that... quantified weight and activation values are then scaled using a scale factor. The inverse of the quantization process is performed. After this step, the 32-bit convolution operation is carried out. The process continues. This algorithm quantizes the deep convolutional neural network during training. by taking into account its effects in a way that minimizes the reduction in the success rate We are able to train them. As a result, we obtain quantized weight tensors and thus It reduces the memory and resource requirements we need for the inference process. The algorithm uses mean and standard deviation of weight and activation tensors. It starts the process by calculating statistical data. This is followed by weighting and activation. It trims the extreme values of the tensors. Thus, those that have a small contribution to the performance rate By trimming the extreme values, the more frequently used coefficients achieve higher accuracy. This allows for quantification. The normalization step involves ensuring that the values fit within the desired range. It scales. Then, the range to be quantified is divided by the number of steps to be quantified. The scale factor is calculated. The normalized weight and activation values are used to calculate the scale. The value is divided by the factor and rounded to the nearest integer. In this way... The quantification step is complete. The pre-trained weight values and instantaneous values are used before the convolution process. Convolution is performed using the quantified activation values. The values... Despite the loss of precision, the quantized convolution during inference is similar to the original convolution. It preserves the essence of the process. By including quantification in the inference process, it becomes convolutional. Processes can be executed efficiently on edge devices without compromising accuracy. and the limited computational capabilities of deep convolutional neural network models It becomes possible to use it on devices that have it. 9
Claims
1. Adapted to operate within a processor and a memory unit, local patterns in the deep image located within the memory unit and Through their ability to learn features, they can classify and recognize objects. the weight value in convolutional neural networks used in the processes It is a quantification method for determining its characteristic, - located within the memory unit by the user, with layers and pre-trained floating point weight values any layer selected from within a model is processed by the processor. reading, - the weight histogram values of the layer read by the processor extracting and saving into the memory unit, - weight histogram values located within the memory unit the average (μ) is calculated by the processor and stored in the memory unit. recording, - weight histogram values located within the memory unit the standard deviation (σ) is calculated by the processor and stored in memory. to be recorded in the unit, - The lower clipping limit set by the processor is = − 2 formula Calculation using and saving to memory unit, - The upper clipping limit set by the processor is = + 2 formula ü Calculation using and saving to memory unit, - the weight histogram values calculated by the processor, layer normalized by dividing by absolute maximum histogram values Finding the value ( ) and saving it to the memory unit, - information provided by the user regarding the number of bits (bits) required for the operation entering the processor, - data retrieved from the memory unit by the processor = ü using the formula, the sw weight scale factor Calculation of (SW) and memory of the calculated weight scale factor (SW). to be recorded in the unit, - the weight value quantized by the processor calculation and saving to the memory unit, - to be repeated for each layer located within the memory unit and memory of the quantized weight value obtained for each layer It includes the steps for recording it in the unit.
2. A device adapted to operate within a processor and a memory unit. local patterns in the deep image located within the memory unit and Through their ability to learn features, they can classify and recognize objects. Convolutional neural networks are known techniques used in processes such as these. the load on the processor and memory unit according to the methods in this situation It is a quantification method to work efficiently by reducing the number of items. feature, - the memory unit contains different layers created by the user a model with pre-trained floating point weight values a layer within it being read by the processor, - the weight histogram values of the layer read by the processor extracting and saving into the memory unit, - weight histogram values located within the memory unit Calculation of the average (μ) by the processor, - weight histogram values located within the memory unit the calculation of the standard deviation (σ) by the processor, - The lower clipping limit set by the processor is = − 2 formula Calculation using, - The upper clipping limit set by the processor is = + 2 formula ü Calculation using, 11 - absolute maximum of weighted histogram values by the processor normalized value obtained by dividing by histogram values ( ) finding and saving to the memory unit, - information provided by the user regarding the number of bits (bits) required for the operation entering the processor, ü - sw by the processor using the = formula Calculation of the weighting scale factor (SW) and the calculated weighting scale Saving the factor (SW) to the memory unit, - the weight value quantized by the processor calculation and saving to the memory unit, - the first step for each layer located within the memory unit repetition and quantified weight obtained for each layer saving its value to the memory unit, - input pixels entered into the memory unit by the user the processor reading the values, - the histogram of the pixel values of the read image is generated by the processor created and saved to the memory unit, - the average (μ) of the generated histogram is calculated by the processor. calculation, - the standard deviation (σ) value of the generated histogram is determined by the processor. calculation, - upper clipping limit set by the processor = +2 formula ü Calculation using, - input to find the normalized (norm) value by the processor dividing pixel values by the absolute maximum input pixel value Finding the value normalized with ( ) and storing it in the memory unit recording, - the bit value (bit) specified by the user is retrieved from the memory unit withdrawal, 12 ü - the scale factor is calculated by the processor using the formula = (S) calculation and the memory unit of the calculated s scale factor. recording, - activation value by the processor = calculation and saving to memory unit, - normalized weight values obtained by the processor and The use of activation values in convolutional neural networks, - Presenting the data generated by the convolutional neural network to the user for approval, - If the user indicates that the output data is incorrect, the first step... Restarting the process by updating the weight values. - If the user confirms that the output data is correct, the process will proceed. The stopping process involves procedural steps. 13