Neural network system and its learning method, and transfer learning method
By iterating through multi-level learning in the neural network processor and determining the layer of learning interrupts, the problem of long training time of neural network systems in the prior art is solved, and faster learning speed and efficiency are achieved.
Patent Information
- Application Number
- CN202010252735.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-08
- Filing Date
- 2020-04-01
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-04-01
AI Technical Summary
Existing neural network systems take longer to train, especially in transfer learning, resulting in slower learning.
By implementing multi-level learning iterations in a neural network processor, and by comparing the distribution of weight values between different learning iterations, the layer of learning interruptions is determined, thereby optimizing the learning process and reducing unnecessary iterations.
It effectively reduces the learning time of neural network systems, especially in transfer learning, and improves learning speed and efficiency.
Smart Images

Figure CN111914989B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This patent application claims the priority benefit of Korean Patent Application No. 10-2019-0053887 filed in the Korean Intellectual Property Office on May 8, 2019, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates to a neural network system, a learning method of the system, and a transfer learning method of a neural network processor. More specifically, the present disclosure relates to a neural network system for performing learning, a learning method of the system, and a transfer learning method of a neural network processor. Background Art
[0004] A neural network refers to a computing architecture that is a model of a biological brain. With the recent development of neural network technology, many studies have been conducted in various electronic systems on using a neural network device that uses at least one neural network model to analyze input data and obtain output information.
[0005] Neural networks can provide a mapping between input patterns and output patterns, which means that neural networks have learning capabilities. Neural networks have generalization capabilities, through which neural networks can provide relatively correct output patterns for input patterns that have not yet been used for learning based on the learning results of their learning capabilities.
[0006] Training a neural network to obtain learning results in a neural network system is important and takes a lot of time. Therefore, a technique for increasing the learning speed of a neural network is needed. Summary of the invention
[0007] According to various aspects of the present disclosure, a neural network system, a learning method of the system, and a transfer learning method of a neural network processor are provided. Through the system and method, the total time spent on learning is reduced, more specifically, the total time spent on transfer learning is reduced, and the transfer learning speed is improved.
[0008] According to one aspect of the present disclosure, a neural network system includes a neural network processor and a memory. The neural network processor is configured to perform learning including multiple learning iterations on multiple layers, and determine at least one layer where learning is interrupted among the multiple layers. The determination of at least one layer where learning is interrupted is based on a result of comparing, for each layer among the multiple layers, a distribution of first weight values obtained through a first learning iteration with a distribution of second weight values obtained through a second learning iteration after the first learning iteration. The neural network processor is also configured to perform a third learning iteration after the second learning iteration on multiple layers among the multiple layers except for at least one layer where learning has been determined to be interrupted. The memory is configured to store first distribution information about the distribution of the first weight value and second distribution information about the distribution of the second weight value when the second learning iteration is completed, and is configured to provide the first distribution information and the second distribution information to the neural network processor.
[0009] According to another aspect of the present disclosure, a learning method for a neural network system includes multiple learning iterations on multiple layers. The learning method includes: storing a first weight obtained by the Nth learning iteration in a memory, and determining at least one layer of the multiple layers where learning is interrupted based on a first weight value included in the first weight and a second weight value included in the second weight obtained by the (N-1)th learning iteration. The learning method for a neural network system also includes: performing an (N+1)th learning iteration on multiple layers other than at least one layer where learning has been determined to be interrupted.
[0010] According to another aspect of the present disclosure, a transfer learning method for a neural network processor includes multiple learning iterations on multiple layers. The transfer learning method includes: storing a first weight value obtained by a first learning iteration in a memory outside the neural network processor; and storing a second weight value obtained by a second learning iteration in the memory. The second learning iteration is after the first learning iteration. The transfer learning method also includes: receiving first distribution information of the first weight value and second distribution information of the second weight value from the memory; and determining at least one layer of the multiple layers where learning is interrupted based on the first distribution information and the second distribution information. The transfer learning method also includes: performing a third learning iteration on multiple layers other than at least one layer of the multiple layers where learning has been determined to be interrupted. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The embodiments of the inventive concept of the present disclosure will be more clearly understood according to the following detailed description in conjunction with the accompanying drawings, in which:
[0012] Figure 1 An electronic system according to an example embodiment is shown;
[0013] Figure 2Another electronic system according to an example embodiment is shown;
[0014] Figure 3 A neural network according to an example embodiment is shown;
[0015] Figure 4 A transfer learning method according to an example embodiment is shown;
[0016] Figure 5 A neural network system according to an example embodiment is shown;
[0017] Figure 6 shows result values obtained through a transfer learning process according to an example embodiment;
[0018] Figure 7 A flow chart showing a learning method of a neural network system according to an example embodiment;
[0019] Fig. 8A shows a change in the distribution of weight values according to an example embodiment;
[0020] Figure 8B Also shown is a change in the distribution of weight values according to an example embodiment;
[0021] Fig.9A shows distribution information according to an example embodiment;
[0022] Fig. 9B Distribution information according to an example embodiment is also shown;
[0023] Fig.10 Another flow chart showing a learning method of a neural network system according to an example embodiment;
[0024] Fig.11 Another flow chart showing a learning method of a neural network system according to an example embodiment;
[0025] Fig.12 Another neural network system according to an example embodiment is shown;
[0026] Fig.13 Another neural network system according to an example embodiment is shown;
[0027] Fig.14 Another neural network system according to an example embodiment is shown; and
[0028] Fig.15 A memory device according to an example embodiment is shown. DETAILED DESCRIPTION
[0029] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings.
[0030] Figure 1 An electronic system 10 according to an example embodiment is shown. The electronic system 10 can analyze input data in real time based on a neural network, obtain valid information, and identify the status or control element of an electronic device equipped with the electronic system 10 based on the valid information. The valid information can be reasonable, reliable, trustworthy and / or other useful / usable information obtained based on applying the neural network to the input data. The status can be identified by identifying the environment and / or characteristics where, when, how, why and / or by whom the input data is obtained. For example, the electronic system 10 can be implemented in a drone, a robotic device such as an advanced driver assistance system (ADAS), a smart TV (TV), a smart phone, a medical device, a mobile device, an image display, a measuring device, an Internet of Things (IoT) device or otherwise applied to these devices. The electronic system 10 can be installed on or incorporated into any of a variety of different types of electronic devices. The electronic system 10 using a neural network can be referred to as a neural network system.
[0031] The electronic system 10 may include at least one intellectual property (IP) block and a neural network processor 100. For example, the electronic system 10 may include a first IP block IP1, a second IP block IP2, a third IP block IP3, and a neural network processor 100. The acronym "IP" stands for the term "intellectual property", and the term "IP block" refers to unique circuits and circuit components, each of which may be individually protected by intellectual property rights. When used in the description herein, the term "IP block" may be synonymous with similar terms such as "IP circuit".
[0032] Electronic system 10 may include various IP blocks. For example, IP blocks may include processing units, multiple cores included in the processing units, multi-format codecs (MFCs), video modules, three-dimensional (3D) graphics cores, audio systems, display drivers, volatile memory, non-volatile memory, memory controllers, input and output interface blocks, cache memory, etc. For example, the processing unit as an IP block may be or include a processor or an application specific integrated circuit (ASIC) that executes software instructions. The video module as an IP block may be or include a camera interface, a joint photo expert group (JPEG) processor, a video processor, and / or a mixer. Each IP block in the first IP block IP1, the second IP block IP2, and the third IP block IP3 may include at least one IP block selected from various IP blocks.
[0033] The connection method based on the system bus can be used as a technology for connecting IP blocks. For example, the Advanced Microcontroller Bus Architecture (AMBA) protocol of Advanced RISC Machines (ARM) Co., Ltd. can be used as a standard bus protocol. The AMBA protocol can include bus types, such as Advanced High Performance Bus (AHB), Advanced Peripheral Bus (APB), Advanced Extensible Interface (AXI), AXI4 and AXI Consistency Extension (ACE). In these bus types, AXI can be used as an interface protocol between IP blocks and can provide multiple outstanding addressing functions and data interleaving functions. In addition to the above-mentioned protocols, other types of protocols can also be used, such as the open core protocol of uNetwork of SONICs Co., Ltd., CoreConnect of IBM and OCP-IP.
[0034] The neural network processor 100 can generate a neural network, train or implement learning of a neural network and / or learning performed by a neural network, perform operations based on input data and generate information signals based on the results of the operations, or retrain the neural network. The neural network can be modeled based on various neural network models, such as convolutional neural networks (CNNs) (e.g., GoogLeNet, AlexNet, or VGG networks), regions with CNNs (R-CNNs), region proposal networks (RPNs), recursive neural networks (RNNs), stack-based deep neural networks (S-DNNs), state-space dynamic neural networks (S-SDNNs), deconvolution networks, deep belief networks (DBNs), restricted Boltzmann machines (RBMs), fully convolutional networks, long short-term memory (LSTM) networks, and classification networks. However, the neural networks and neural network models described herein are not limited thereto. The neural network processor 100 may include at least one processor that performs operations according to one or more neural network models. The neural network processor 100 may include a separate memory that stores programs corresponding to corresponding neural network models. The neural network processor 100 may be referred to as a neural network processing device, a neural network integrated circuit, or a neural network processing unit (NPU).
[0035] The neural network processor 100 may receive various input data from at least one IP block through a system bus, and may generate an information signal based on the input data. For example, the neural network processor 100 may generate an information signal by performing a neural network operation on the input data, and the neural network operation may include a convolution operation. The information signal generated by the neural network processor 100 may include at least one recognition signal selected from various recognition signals such as a voice recognition signal, an object recognition signal, an image recognition signal, and a biometric recognition signal. For example, the neural network processor 100 may receive frame data included in a video stream as input data, and may generate a recognition signal for an object included in an image represented by the frame data based on the frame data. However, the embodiments are not limited to the video stream as input data or the above-mentioned type of recognition signal as the generated information signal. The neural network processor 100 may receive various types of input data and generate a recognition signal based on the input data. For this operation, the neural network processor 100 may implement learning of a neural network and / or learning by a neural network, which will be referred to in detail. Figure 3 Describe in detail.
[0036] According to example embodiments, the neural network processor 100 of the electronic system 10 may perform transfer learning. Transfer learning may refer to learning of a neural network applied to a second model using a neural network applied to a first model and / or learning performed by a neural network applied to a second model. For example, transfer learning may include learning of a neural network applied to a specific model using a neural network algorithm or weights and / or learning performed by a neural network applied to a specific model, wherein the neural network algorithm or weights are applied to an existing trained general model. Figure 4This is described in detail. In an embodiment, transfer learning may include multiple learning iterations in multiple layers. Since transfer learning is performed based on an existing trained model, the weights corresponding to some of the multiple layers may have values that are saturated after fewer learning iterations than the total number of learning iterations. As an illustration, if the output result from a certain layer is limited to a certain range, saturation may occur when the output result from the layer reaches the end value of the range, and will not change in subsequent iterations. A simple example of saturation of a single processing element in a certain layer is a process of comparing a first input value with a second hold value maintained by a single processing element in the layer. If the first input value is less than the second hold value, the process may discard the first input value and output the second hold value as the output result. If the first input value is greater than the second hold value, the process may replace the second hold value with the first input value as the new hold value, and output the first input value as the new hold value as the output result. If both the hold value and the output result are limited to between 0 and 100, then once the hold value is 100, subsequent processing of the processing element cannot increase the hold value and the output result to greater than 100, and the logic of the processing element requires that the hold value and the output result will not decrease in any iteration. Therefore, in this example, once the output result is equal to 100, the operation result from the processing element will be saturated. The above description of the logic that may cause the processing elements in the layer to saturate is only exemplary, and the output results of the processing elements in the layer may be saturated in other ways. For the purposes described herein, saturation can be considered to be a state in which further processing of a certain layer does not satisfactorily change the operation results from the layer from one iteration to the next.
[0037] Saturation is merely an example of why learning is considered interrupted. As explained below, learning can be considered interrupted when the results of an operation from a layer do not change, or even if the results of an operation from the layer change from one iteration to the next but do not change satisfactorily (such as when the change is below a threshold). That is, the interruption itself may be a waste of time and resources to perform iterations in a non-productive layer, causing the entire process to be interrupted in other ways while time and resources are wasted on iterations in the non-productive layer. In addition, the basis for determining that learning is interrupted is not limited to one aspect of the output from a layer (such as a single value) or the output from a single processing element in a layer. Instead, changes or relative lack of changes can be reflected in distribution statistics and distribution histograms with multiple values from multiple processing elements, which reflect the distribution of weight values output from multiple processing elements of the layer at each iteration.
[0038] According to an example embodiment, the neural network processor 100 of the electronic system 10 may determine a layer where learning is to be interrupted based on a weight value obtained through a current learning iteration and a weight value obtained through a previous learning iteration. The neural network processor 100 may perform subsequent learning iterations only in layers other than the layer where learning is determined to be interrupted. For example, the neural network processor 100 may determine a layer where learning is to be interrupted based on a result of comparing the distribution of weight values obtained through a current learning iteration with the distribution of weight values obtained through a previous learning iteration. The neural network processor 100 may perform subsequent learning iterations only in layers other than the layer where learning is determined to be interrupted. Therefore, the neural network processor 100 may perform subsequent learning iterations faster, and thus the total transfer learning time may be reduced. In addition, the transfer learning speed of the neural network processor 100 may be increased. Transfer learning of the neural network processor 100 and / or transfer learning performed by the neural network processor 100 will be described in more detail with reference to the following drawings.
[0039] Above Figure 1 The description of describes a comparison of weight values between two iterations. The current iteration and the previous iteration do not have to be in immediate order, but a comparison can be made between the weight values of the third iteration and the weight values of the first iteration, wherein the second iteration between the first iteration and the third iteration is not part of the comparison. In addition, more than one comparison can be performed between more than two layers. For example, a comparison can be made between the weight values from the second iteration and the weight values from the first iteration, a comparison can be made between the weight values from the third iteration and the weight values from the second iteration, and a comparison can be made between the weight values from the third iteration and the weight values from the first iteration. For example, when an initial comparison identifies a substantial change in the weight value, but over time, the substantial change is less significant or appears less significant so that more and more comparisons are considered necessary due to imminent saturation, a decision to perform additional comparisons can also be made dynamically.
[0040] Figure 2 Another electronic system 20 according to an example embodiment is shown. Specifically, Figure 2 Shown according to Figure 1 The specific embodiment of the electronic system 10 will be omitted. Figure 1 A redundant description of the electronic system 10 is provided.
[0041] The electronic system 20 may include an NPU 1001, a random access memory (RAM) 2001, a processor 300, a memory 400, and a sensor module 500. The NPU 1001 may correspond to Figure 1 The neural network processor 100 in.
[0042] The RAM 2001 may temporarily store programs, data, or instructions. The programs and / or data stored in the memory 400 may be temporarily loaded to the RAM 2001 according to the control or boot code of the processor 300. The RAM 2001 may be implemented using a memory such as a dynamic RAM (DRAM) or a static RAM (SRAM).
[0043] The processor 300 may control all operations of the electronic system 20. For example, the processor 300 may be implemented as a central processing unit (CPU). The processor 300 may include a single core or multiple cores. The processor 300 may process or execute programs and / or data stored in the RAM 2001 and the memory 400. For example, the processor 300 may control the functions of the electronic system 20 by executing a program stored in the memory 400.
[0044] The memory 400 is a storage device for storing data, and can store, for example, an operating system (OS), various programs and various data. The memory 400 may include DRAM, but is not limited thereto. The memory 400 may include at least one memory selected from a volatile memory and a non-volatile memory. The non-volatile memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a phase change RAM (PRAM), a magnetic RAM (MRAM), a resistive RAM (RRAM) and a ferroelectric RAM (FeRAM). The volatile memory may include DRAM, SRAM, synchronous DRAM (SDRAM), PRAM, MRAM, RRAM and FeRAM. In an embodiment, the memory 400 may include at least one selected from the following items: a hard disk drive (HDD), a solid state drive (SSD), a high density flash memory (CF) memory, a secure digital (SD) memory, a micro SD memory, a mini SD memory, an extreme digital (xD) memory and a memory stick.
[0045] The sensor module 500 can collect surrounding information of the electronic system 20. The sensor module 500 can sense or receive image signals from outside the electronic system 20, and can convert the image signals into image data, such as image frames. For this operation, the sensor module 500 may include at least one sensing device selected from various sensing devices, such as an image pickup device, an image sensor, a light detection and ranging (LIDAR) sensor, an ultrasonic sensor, an infrared sensor, a gyroscope, an accelerometer, a thermometer, and a compass. The sensor module 500 may also or alternatively receive a sensing signal from a sensing device. In an embodiment, the sensor module 500 may provide an image frame to the NPU 1001. For example, the sensor module 500 may include an image sensor and may generate a video stream by capturing the surrounding environment of the electronic system 20, and sequentially provide continuous image frames in the video stream to the NPU 1001.
[0046] According to an example embodiment, the NPU 1001 of the electronic system 20 may perform transfer learning. Transfer learning may refer to learning of a neural network applied to a second model using a neural network applied to a first model and / or learning performed by a neural network applied to a second model. For example, transfer learning may include learning of a neural network applied to a specific model using a neural network algorithm or weights and / or learning performed by a neural network applied to a specific model, wherein the neural network algorithm or weights are applied to an existing trained general model. Figure 4 This is described in detail. In an embodiment, transfer learning includes multiple learning iterations in multiple layers. Since transfer learning is performed based on an existing trained model, weights corresponding to some layers in the multiple layers may have values that are saturated after fewer learning iterations than the total number of learning iterations.
[0047] According to an example embodiment, the NPU 1001 of the electronic system 20 may determine the layer where learning is interrupted based on the weight value obtained by the current learning iteration and the weight value obtained by the previous learning iteration. The NPU 1001 may perform subsequent learning iterations only in layers other than the layer where learning is determined to be interrupted. For example, the NPU 1001 may determine the layer where learning will be interrupted based on the result of comparing the distribution of the weight value obtained by the current learning iteration with the distribution of the weight value obtained by the previous learning iteration. The NPU 1001 may perform subsequent learning iterations only in layers other than the layer where learning is determined to be interrupted. Therefore, the NPU 1001 may perform subsequent learning iterations faster, thereby reducing the total transfer learning time. In addition, the transfer learning speed of the NPU 1001 may be improved. The transfer learning of the NPU 1001 and / or the transfer learning performed by the NPU 1001 will be described in more detail with reference to the following drawings.
[0048] Figure 3 A neural network 1000 according to an example embodiment is shown. The neural network 1000 may include an input layer 1100, hidden layers 1220 and 1240, and an output layer 1300. The neural network 1000 may perform operations based on input data I1 and I2 and generate output data O1 and O2 based on the operation results. Figure 3 The neural network 1000 includes two hidden layers, but this is only an example and the embodiment is not limited thereto. For example, the neural network 1000 may include one or more hidden layers.
[0049] Each of the layers included in the neural network 1000 (i.e., the input layer 1100, the hidden layers 1220 and 1240, and the output layer 1300) may include a plurality of neurons. Neurons may refer to artificial nodes referred to as processing elements (PEs), processing units, or similar terms. For example, Figure 3 As shown in , the input layer 1100 may include two neurons, each of the hidden layers 1220 and 1240 may include three neurons, and the output layer 1300 may include two neurons. However, Figure 3 The configuration in is only an example, and each of the input layer 1100, the hidden layers 1220 and 1240, and the output layer 1300 in the neural network 1000 may include various numbers of neurons.
[0050] The neurons included in each of the input layer 1100, the hidden layers 1220 and 1240, and the output layer 1300 in the neural network 1000 can be connected to neurons in different layers and can exchange data with neurons in different layers. Neurons can receive data from other neurons, perform operations on the data, and output the operation results to other neurons.
[0051] The input to each neuron and the output from each neuron may be referred to as input activation and output activation, respectively. An activation may be a parameter that corresponds to both the output from a neuron in a layer and the input to one or more neurons included in a successive layer. In other words, the output activation from a neuron in one layer may be the input activation to one or more neurons in a successive layer. Each neuron may be based on the input activation received from a neuron included in a previous layer, Figure 3The weights in the weights shown as W11, W12, W13, W21, W22 and W23 in FIG. 1 and the bias determine (e.g., calculate, generate, etc.) the output activation. In an embodiment, each of the weights W11, W12, W13, W21, W22 and W23 itself can be a weight matrix including multiple weight values as components. The weights and biases are parameters used to calculate the output activation in the neuron. The weights can be values assigned to the connection relationship between neurons, and the biases can indicate the weights associated with a single neuron.
[0052] As described above, each layer in the input layer 1100, the hidden layers 1220 and 1240, and the output layer 1300 can perform at least one operation so that the neuron can determine (e.g., calculate, generate, etc.) an output activation. As described above, the output activation can be considered to be the output from one of the input layer 1100, the hidden layers 1220 and 1240, and the output layer 1300, which can also be the input to the successive layers in the hidden layers 1220 and 1240 and the output layer 1300 (if applicable).
[0053] General learning of the neural network 1000 and / or general learning performed by the neural network 1000 may include multiple learning iterations and initial weights. The initial weights W11, W12, W13, W21, W22, and W23 may have random values. The neural network 1000 may change the weights W11, W12, W13, W21, W22, and W23 by back propagation using the output data O1 and O2 at each learning iteration. The learning iteration may be referred to as an epoch. The neural network 1000 that has been trained by performing multiple learning iterations may include trained weights W11, W12, W13, W21, W22, and W23. For ease of description, the weights W11, W12, W13, W21, W22, and W23 generated after training may be described herein around training the weights W11, W12, W13, W21, W22, and W23 through the learning operation of the neural network 1000.
[0054] Transfer learning of and / or by the neural network 1000 may train weights W11, W12, W13, W21, W22, and W23 through a similar process. However, in the case of transfer learning, weights W11, W12, W13, W21, W22, and W23 may have initial values corresponding to other trained weights, rather than having random initial values as in general learning.
[0055] Figure 4A transfer learning method according to an example embodiment is shown. Transfer learning may refer to learning of a neural network applied to a second model using a neural network applied to a first model and / or learning performed by a neural network applied to a second model. For example, transfer learning may include learning of a neural network applied to a specific model using a neural network algorithm or weights and / or learning performed by a neural network applied to a specific model, wherein the neural network algorithm or weights are applied to an existing trained general model.
[0056] In other words, the process of obtaining weights W_S11 and W_S12 applied to a specific model using trained weights W_G11 and W_G12 that have been trained in a general model may be referred to as transfer learning. Transfer learning may indicate a process of obtaining weights W_S11 and W_S12 applied to a specific model by performing multiple learning iterations on multiple layers using trained weights W_G11 and W_G12 as initial weights. In other words, the trained weights W_G11 and W_G12 from a general model may be the initial weights for training a specific model.
[0057] To aid understanding, an exemplary general model may be a model trained to recognize images of dogs. When the general model trained to recognize images of dogs includes weights, the weights from the general model may be used to train a model that recognizes a specific dog breed. For example, to train a model that recognizes a Chihuahua breed, a neural network system may perform transfer learning using weights corresponding to a general model trained to recognize images of dogs. The weights from the general model may be initial values in the training of a specific model. Since the general model is an existing trained model, the weights from the general model as initial values applied to one or more of the layers of the specific model may have been saturated or may be close to saturation.
[0058] Figure 5 A neural network system 30 is shown according to an example embodiment. Figure 5 The neural network system 30 may correspond to Figure 1 The electronic system 10 and Figure 2 The neural network system 30 may include a neural network processor 100 and a memory 200. The neural network processor 100 may be Figure 1 The neural network processor 100 and / or Figure 2 Detailed example of NPU 1001 in FIG. Figure 5 The memory 200 in the embodiment may correspond to Figure 2 RAM 2001 and / or Figure 2 The memory 400 in the embodiment of the present invention is not limited thereto. For example, Figure 5 The memory 200 may include Figure 2RAM 2001 and memory 400 in.
[0059] Figure 5 The neural network processor 100 may include a transfer learning circuit 120 and a transfer learning controller 140.
[0060] The transfer learning circuit 120 may perform transfer learning of the neural network processor 100 and / or transfer learning performed by the neural network processor 100. For this operation, the transfer learning circuit 120 may perform multiple learning iterations in multiple layers. The transfer learning circuit 120 may update the weights each time a learning iteration is performed in each of the multiple layers. For example, the transfer learning circuit 120 may update the weights based on the output data obtained by the Nth learning iteration (where N is a natural number). The transfer learning circuit 120 may output the updated weights as the Nth learning result Res_N. The transfer learning circuit 120 may provide the Nth learning result Res_N to the memory 200 so that the memory 200 stores the Nth learning result Res_N.
[0061] The transfer learning of the transfer learning circuit 120 can be implemented in various forms and can be implemented by hardware and / or software according to various embodiments. For example, when the transfer learning is implemented by hardware, the transfer learning circuit 120 may include a circuit that performs the transfer learning. When the transfer learning is implemented by software, the transfer learning can be performed by executing a program (or, instructions) stored in the memory 200 using the neural network processor 100 or at least one other processor. However, the transfer learning of the transfer learning circuit 120 is not limited to these embodiments and can be implemented by a combination of software and hardware (e.g., firmware).
[0062] The transfer learning controller 140 may control the transfer learning of the transfer learning circuit 120 and / or the transfer learning performed by the transfer learning circuit 120. In an embodiment, the transfer learning controller 140 may determine the layer where learning will be interrupted in the subsequent learning iteration based on the result of comparing the distribution of the weight values obtained by the current learning iteration with the distribution of the weight values obtained by the previous learning iteration. The transfer learning controller 140 may control the transfer learning circuit 120 to perform learning only in the layers other than the layers where learning interruption is determined in the subsequent learning iteration. For example, after the Nth learning iteration ends, the transfer learning controller 140 may receive distribution information Info_DB_N-1 about the distribution of the weight values obtained by the (N-1)th learning iteration and distribution information Info_DB_N about the distribution of the weight values obtained by the Nth learning iteration from the memory 200. The transfer learning controller 140 may determine the layer where learning is interrupted based on the distribution information Info_DB_N-1 and the distribution information Info_DB_N. The method for determining learning interruption will be described in detail below with reference to the accompanying drawings.
[0063] The memory 200 may store the transfer learning result provided from the transfer learning circuit 120. For example, after the Nth learning iteration ends, the memory 200 may store the Nth learning result Res_N provided from the transfer learning circuit 120. In an embodiment, the memory 200 may include a DRAM.
[0064] The memory 200 may store distribution information Info_DB indicating information about the distribution of weight values. In an embodiment, when the weights are stored in the memory 200, the memory 200 may obtain the distribution information Info_DB using a logic circuit in the memory 200, and may store the distribution information Info_DB therein. In other words, the memory 200 may be implemented as a process in memory (PIM), and may directly obtain the distribution information Info_DB according to the weights. The memory 200 may provide the distribution information Info_DB to the neural network processor 100.
[0065] According to an exemplary embodiment, the neural network system 30 including the neural network processor 100 and the memory 200 excludes subsequent learning iterations in one or more layers where the weight values are saturated based on the result of comparing the distribution of previous weight values with the distribution of current weight values. As a result, the neural network system 30 reduces the total transfer learning time and ultimately improves the transfer learning speed.
[0066] Figure 6 Result values obtained through the transfer learning process according to an example embodiment are shown. Specifically, Figure 6The (N-1)th learning result Res_N-1 obtained after the (N-1)th learning iteration and the Nth learning result Res_N obtained after the Nth learning iteration (where N is a natural number of 2 or greater). For ease of description, it is assumed that Figure 6 Used in Figure 3 However, this is only an example, and the embodiments are not limited to Figure 3 The number of layers and neurons in . Figure 5 and Figure 6 Give a description.
[0067] The (N-1)th learning result Res_N-1 may include: the (N-1)th learning result in the first hidden layer, i.e., the (N-1)th learning result Res_N-1_L1 of the first hidden layer; and the (N-1)th learning result in the second hidden layer, i.e., the (N-1)th learning result Res_N-1_L2 of the second hidden layer. The (N-1)th learning result Res_N-1_L1 of the first hidden layer may include weights W11_N-1, W12_N-1, and W13_N-1. The (N-1)th learning result Res_N-1_L2 of the second hidden layer may include weights W21_N-1, W22_N-1, and W23_N-1. When the weight W11_N-1 is described as a representative of the weights W11_N-1, W12_N-1, W13_N-1, W21_N-1, W22_N-1, and W23_N-1, the weight W11_N-1 may include a plurality of weight values w111, w112, w113, w114, w115, w116, w117, w118, and w119. Here, the number of weight values is an example, and the number of weight values of the weight may be less than or more than nine. Here, the weight values w111, w112, w113, w114, w115, w116, w117, w118, and w119 may be referred to as weight components w111, w112, w113, w114, w115, w116, w117, w118, and w119.
[0068] Similarly, the Nth learning result Res_N may include: the Nth learning result in the first hidden layer, i.e., the Nth learning result Res_N_L1 of the first hidden layer; and the Nth learning result in the second hidden layer, i.e., the Nth learning result Res_N_L2 of the second hidden layer. The Nth learning result Res_N_L1 of the first hidden layer may include weights W11_N, W12_N, and W13_N. The Nth learning result Res_N_L2 of the second hidden layer may include weights W21_N, W22_N, and W23_N. When weight W11_N is described as a representative of weights W11_N, W12_N, W13_N, W21_N, W22_N, and W23_N, weight W11_N may include a plurality of weight values w111', w112', w113', w114', w115', w116', w117', w118', and w119'. Here, the number of weight values is also an example, and the number of weight values of the weight may be less than or more than nine. Here, the weight values w111', w112', w113', w114', w115', w116', w117', w118', and w119' may be referred to as weight components w111', w112', w113', w114', w115', w116', w117', w118', and w119'.
[0069] As described above, as the learning process proceeds, the weights included in the learning result and the weight values included in each weight may be updated.
[0070] However, in transfer learning using weight values of a trained neural network, weight values included in a specific layer may be saturated from the initial iteration or after only a small number of learning iterations. Performing learning in a layer where weight values are saturated may be inefficient. According to an example embodiment, as described below, a neural network system can reduce the time spent on transfer learning by excluding (e.g., skipping, bypassing, etc.) learning in a layer where weight values are saturated.
[0071] Figure 7 A flow chart of a learning method of a neural network system according to an example embodiment is shown. Specifically, Figure 7 It can be shown Figure 5 A flowchart of a transfer learning method of a neural network system 30. Figure 5 and Figure 7 Give a description.
[0072] In operation S120, when the Nth learning iteration ends, the neural network system 30 may store the weights obtained after the Nth learning iteration. For example, the neural network processor 100 may perform transfer learning including multiple learning iterations in multiple layers. After performing the Nth learning iteration among the multiple learning iterations (where N is a natural number), the neural network processor 100 may provide the memory 200 with the weights obtained through the Nth learning iteration, and the memory 200 may store therein the weights obtained through the Nth learning iteration.
[0073] The neural network system 30 may compare the first distribution information about the weight value obtained after the Nth learning iteration with the second distribution information about the weight value obtained after the (N-1)th learning iteration. In operation S140, the neural network system 30 may determine the layer where learning is interrupted. In an embodiment, the memory 200 may provide the neural network processor 100 with the first distribution information about the first weight value obtained through the Nth learning iteration and the second distribution information about the second weight value obtained through the (N-1)th learning iteration, and the neural network processor 100 may determine the layer where learning is interrupted based on the first distribution information and the second distribution information. Fig.10 and Fig.11 A method for determining a layer where learning is interrupted is described in detail. In an embodiment, the memory 200 may provide a separate processor with first distribution information about a first weight value obtained by the Nth learning iteration and second distribution information about a second weight value obtained by the (N-1)th learning iteration. The separate processor may determine the layer where learning is interrupted based on the first distribution information and the second distribution information. The separate processor may also provide the neural network processor 100 with information about the layer where learning has been determined to be interrupted. Such an embodiment will be described with reference to FIG. 14.
[0074] In an embodiment, when N is "1", that is, when the first learning iteration ends, the neural network system 30 may compare the first distribution information about the weight value obtained after the first learning iteration with the distribution information about the initial weight value (instead of the weight value obtained after the (N-1)th learning iteration). In other words, since there is no learning iteration like the zeroth learning iteration, there is no weight value obtained after the zeroth learning iteration, and the initial weight value may be used in operation S140 instead of the non-existent weight value obtained after the zeroth learning iteration.
[0075] In operation S160 , the neural network system 30 may perform the (N+1)th learning iteration in other layers except the layer where the learning interruption has been determined.
[0076] When using the transfer learning method according to the example embodiment, any further learning is excluded (eg, skipped, bypassed, etc.) in the layer where the weight value is saturated. Therefore, the time spent on transfer learning can be reduced and the overall transfer learning speed can be improved.
[0077] Fig. 8A and Figure 8B 8A and 8B show the change in the distribution of weight values according to an example embodiment. Figure 8B The distribution of weight values after the (N-1)th learning iteration and the distribution of weight values after the Nth learning iteration can be shown. Specifically, Fig. 8A and Figure 8B The distribution in a specific layer can be shown. Figure 5 and Figure 7 as well as Fig. 8A and Figure 8B Give a description.
[0078] refer to Fig. 8A , in the case of a particular hidden layer, the distribution of weight values may change after the Nth learning iteration. Fig. 8A As shown in , when the distribution of weight values significantly changes after learning, in operation S140, the neural network system 30 may determine that the learning for the hidden layer is not interrupted.
[0079] refer to Figure 8B , in the case of another specific hidden layer, the distribution of weight values may not change substantially after the Nth learning iteration. Figure 8B As shown in , when the distribution of weight values slightly changes or does not change at all after learning, in operation S140, the neural network system 30 may determine that the learning for the hidden layer is interrupted.
[0080] Fig.9A and Fig. 9B Distribution information Info_DB according to an example embodiment is shown. Fig.9A and Fig. 9B Distribution information Info_DB about weight values obtained after the Nth learning iteration may be shown. Fig.9A and Fig. 9B The case where the neural network includes K hidden layers (where K is a natural number) is shown. Figure 5 , Fig.9A and Fig. 9B Give a description.
[0081] refer to Fig.9A , the distribution information Info_DB may include histogram information about weight values corresponding to a plurality of layers.
[0082] like Fig.9AAs shown on the left side of , the histogram information corresponding to each layer may include information about the weight value counts corresponding to each weight value. Specifically, for example, the histogram information of the first layer may include counts c_11, c_12, ..., and c_1M corresponding to the weight values wv_1, wv_2, ..., and wv_M, respectively, where M is a natural number. In other words, the first count c_11 may be the number (total number, sum of the number, count of the number, etc.) of the first weight values wv_1 in the first layer, and the second count c_12 may be the number (total number, sum of the number, count of the number, etc.) of the second weight values wv_2 in the first layer.
[0083] Similarly, for example, the histogram information of the second layer may include counts c_21, c_22, ..., and c_2M corresponding to the weight values wv_1, wv_2, ..., and wv_M, respectively, where M is a natural number. In other words, the first count c_21 may be the number (total number, sum of the number, count of the number, etc.) of the first weight values wv_1 in the second layer, and the second count c_12 may be the number (total number, sum of the number, count of the number, etc.) of the second weight values wv_2 in the second layer.
[0084] Similarly, for example, the histogram information of the Kth layer may include counts c_K1, c_K2, ..., and c_KM corresponding to the weight values wv_1, wv_2, ..., and wv_M, respectively, where M is a natural number. In other words, the first count c_K1 may be the number (total number, sum of the number, count of the number, etc.) of the first weight values wv_1 in the Kth layer, and the second count c_K2 may be the number (total number, sum of the number, count of the number, etc.) of the second weight values wv_2 in the Kth layer. Figure 5 The neural network processor 100 in can determine the layer where the subsequent learning iteration is interrupted based on such histogram information. Fig.10 This is described in detail.
[0085] refer to Fig. 9B , the distribution information Info_DB may include statistical information about weight values corresponding to a plurality of layers. The statistical information may indicate a value obtained by statistical processing using the weight values. In an embodiment, the statistical information may include at least one statistical information selected from various statistical information, such as a mean, variance, standard deviation, maximum value, and minimum value of the weight values in each layer. Fig. 9B The case where the statistical information includes the mean and standard deviation is shown.
[0086] refer to Fig. 9B, for example, the mean of the weight values corresponding to the first layer may have a first mean m_1, and the standard deviation of the weight values corresponding to the first layer may have a first standard deviation value S_1. Similarly, the mean of the weight values corresponding to the second layer may have a second mean m_2, and the standard deviation of the weight values corresponding to the second layer may have a second standard deviation value S_2. Similarly, the mean of the weight values corresponding to the Kth layer may have a Kth mean m_K, and the standard deviation of the weight values corresponding to the Kth layer may have a Kth standard deviation value S_K. Figure 5 The neural network processor 100 in can determine the layer where the subsequent learning iteration is interrupted based on such statistical information. Fig.11 This is described in detail.
[0087] Fig.10 Another flow chart of a learning method of a neural network system according to an example embodiment is shown. Specifically, Fig.10 It may be a flow chart of a transfer learning method corresponding to a case where the distribution information includes histogram information (eg, the histogram information shown in FIG. 9A ). In addition, Fig.10 It can be shown Figure 7 Reference will be made to the detailed flowchart of operation S140 in FIG. Figure 5 and Fig.10 Give a description.
[0088] In operation S142, the neural network system 30 may subtract the second histogram information included in the second distribution information from the first histogram information included in the first distribution information. For example, the neural network system 30 may subtract the first count of the first histogram information from the second count of the second histogram information, wherein the first count and the second count correspond to the same weight value among all weight values in each layer. The neural network system 30 may obtain the subtracted histogram information for each layer by performing such subtraction on all weight values.
[0089] In operation S144, the neural network system 30 may obtain a difference indicator value by adding the absolute values of the values included in the histogram information obtained by the subtraction process, the histogram information obtained by the subtraction process being obtained by the subtraction step of operation S142. When the difference indicator value is high, this means that the difference between the weight value obtained by the (N-1)th learning iteration and the weight value obtained by the Nth learning iteration is large. When the difference indicator value is low, this means that the difference between the weight value obtained by the (N-1)th learning iteration and the weight value obtained by the Nth learning iteration is small. A smaller difference may be considered to indicate that learning is interrupted in the layer because few or no values are changing, as indicated by the smaller difference. A larger difference may be considered to indicate that learning is occurring because a large number of values are changing.
[0090] In operation S146, the neural network system 30 may determine a learning interruption for a layer having a difference indicator value equal to or less than a threshold value. Here, the threshold value may be predetermined and may be a constant value or a variable value.
[0091] Fig.11 Another flow chart of a learning method of a neural network system according to an example embodiment is shown. Specifically, Fig.11 It can be a flow chart of a transfer learning method corresponding to the case where the distribution information includes statistical information (for example, the statistical information shown in FIG. 9B ). In addition, Fig.11 It can be shown Figure 7 The detailed flow chart of operation S140 in FIG. Figure 5 and Fig.11 Give a description.
[0092] In operation S141, the neural network system 30 may compare the first statistical information included in the first distribution information with the second statistical information included in the second distribution information for (with respect to) each layer. Fig. 9B As described, the first statistical information and the second statistical information may include at least one selected from a mean value, a variance value, a standard deviation value, a maximum value, and a minimum value.
[0093] In operation S143, the neural network system 30 may determine a learning interruption for a layer where the difference between the first statistical information and the second statistical information is equal to or less than a threshold value. For example, the neural network system 30 may determine a learning interruption for a layer where the difference between the mean included in the first statistical information and the mean included in the second statistical information is equal to or less than a threshold value. In another example, the neural network system 30 may determine a learning interruption for a layer where the difference between the standard deviation value included in the first statistical information and the standard deviation value included in the second statistical information is equal to or less than a threshold value. Here, the threshold value may be predetermined and may be a constant value or a variable value. However, the embodiment is not limited thereto. For example, the neural network system 30 may determine a learning interruption for a layer where the difference between the means is equal to or less than a first threshold value and the difference between the standard deviation values is equal to or less than a second threshold value.
[0094] Fig.12 Another neural network system 40 according to an example embodiment is shown. Figures 1 to 11 Redundant description of neural network system 10 , neural network system 20 , and neural network system 30 .
[0095] The memory 200 may include a memory device 210 and a transfer learning manager 220. The memory device 210 may refer to a physically addressable memory storing data.
[0096] After the Nth learning iteration is finished, the neural network processor 100 may provide the Nth learning result Res_N to the memory 200. The memory 200 may store the Nth learning result Res_N in the memory device 210.
[0097] At this time, the transfer learning manager 220 may generate distribution information Info_DB of the weight value using the weight value included in the Nth learning result Res_N. The distribution information Info_DB may include: Fig.9A The histogram information shown in Fig. 9B The transfer learning manager 220 may store the distribution information Info_DB therein. After the Nth learning iteration is completed, the transfer learning manager 220 may provide the neural network processor 100 with the distribution information Info_DB_N-1 of the weight values obtained through the (N-1)th learning iteration and the distribution information Info_DB_N of the weight values obtained through the Nth learning iteration. The transfer learning manager 220 may be referred to as a transfer learning management circuit.
[0098] The transfer learning manager 220 may be implemented in the memory 200, for example, in a logic circuit area of the memory 200. Alternatively, as described below Fig.15 As shown in , the transfer learning manager 220 can be formed in the buffer die area. As described above, a configuration for performing processing operations can be included in the memory 200. For example, the memory 200 can be implemented as a memory processing (PIM) having a processor integrated in the memory 200 (e.g., on a single chip).
[0099] Fig.13 Another neural network system 50 according to an example embodiment is shown. Figures 1 to 12 A redundant description of neural network system 10, neural network system 20, neural network system 30, and neural network system 40. Specifically, the description will focus on Fig.13 The neural network system 50 and Fig.12 40 differences between neural network systems.
[0100] The memory 200 may include a first memory device 210_1, a first migration learning manager 220_1 corresponding to the first memory device 210_1, a second memory device 210_2, a second migration learning manager 220_2 corresponding to the second memory device 210_2, a third memory device 210_3, a third migration learning manager 220_3 corresponding to the third memory device 210_3, and a main processor 230. For ease of description, Fig.13 The number of memory devices in is only an example, and embodiments are not limited thereto.
[0101] The main processor 230 may control various operations of the memory 200 .
[0102] For ease of description, assume that Figure 3 The neural network 1000 of FIG. 1 is also assumed to store weights in the first hidden layer in the memory 200. In an embodiment, the memory 200 may store the weight W11 in the first memory device 210_1, the weight W12 in the second memory device 210_2, and the weight W13 in the third memory device 210_3. In the process of storing each weight, the memory 200 may also obtain distribution information of the weight value included in the weight from the transfer learning manager corresponding to each memory device.
[0103] In an embodiment, the weight W11 and the weight W12 may be grouped together and stored in the first memory device 210_1 , and the weight W13 may be stored in the second memory device 210_2 .
[0104] The addressing for each weight may be determined by an external processor (e.g., a CPU), determined under the control of the main processor 230, or determined based on a group ID stored in each of the first to third transfer learning managers 220_1 to 220_3. That is, the weight is stored at a memory address, and the memory address of the weight may be used to retrieve and update the weight in iterative transfer learning performed by the neural network 1000. Therefore, the memory address of the weight may be determined and managed by an external processor, determined and managed under the control of the main processor 230, or determined based on a group ID stored in each of the first to third transfer learning managers 220_1 to 220_3.
[0105] Fig.14 Another neural network system 60 according to an example embodiment is shown. Figures 1 to 13 A redundant description of neural network system 10, neural network system 20, neural network system 30, neural network system 40, and neural network system 50. Specifically, the description will focus on Fig.14 The neural network system 60 and Fig.12 40 differences between neural network systems.
[0106] and Fig.12 Differently, after the Nth learning iteration is completed, the transfer learning manager 220 may provide the distribution information Info_DB_N-1 and the distribution information Info_DB_N to the processor 300 instead of the neural network processor 100. Therefore, the processor 300 may perform the operation of determining the layer where the learning is interrupted as a Figure 7 The processor 300 may provide the neural network processor 100 with layer information Info_Layer about the layer for which learning interruption has been determined. The neural network processor 100 may perform the (N+1)th learning iteration on other layers except the layer for which learning interruption has been determined based on the layer information Info_Layer provided from the processor 300.
[0107] Fig.15 2 shows a memory 2200 according to an example embodiment. The memory 2200 may correspond to Figure 12 to Figure 14 The memory 200 in. Fig.15 Memory 2200 is shown implemented as a high bandwidth memory (HBM) having increased bandwidth by including multiple channels with independent interfaces.
[0108] The memory 2200 may include a plurality of layers. For example, the memory 2200 may include a buffer die 2210 and a structure in which at least one core die 2220 is stacked on the buffer die 2210. For example, the first core die 2221 may include a first channel CH1 and a third channel CH3, the second core die 2222 may include a second channel CH2 and a fourth channel CH4, the third core die 2223 may include a fifth channel CH5 and a seventh channel CH7, and the fourth core die 2224 may include a sixth channel CH6 and an eighth channel CH8.
[0109] The buffer die 2210 may communicate with the memory controller, receive commands, addresses, and data from the memory controller, and provide the commands, addresses, and data to the at least one core die 2220. The buffer die 2210 may communicate with the memory controller through a conductive member (e.g., a bump formed on an outer surface of the buffer die 2210). The buffer die 2210 may buffer commands, addresses, and data, so that the memory controller may interface with the at least one core die 2220 by driving only the load of the buffer die 2210.
[0110] The memory 2200 may also include a plurality of TSVs 2230 (Through Silicon Vias) penetrating these layers. The TSVs 2230 may be arranged corresponding to multiples of the first to eighth channels CH1 to CH8. When each of the first to eighth channels CH1 to CH8 has a bandwidth of 128 bits, the TSVs 2230 may include a component for inputting and outputting 1024 bits of data.
[0111] The buffer die 2210 may include a TSV region 2212, a physical region (PHY region) 2213, and a direct access region (DA region) 2214. The TSV region 2212 is a region in which a TSV 2230 is formed for communicating with at least one core die 2220. The PHY region 2213 may include a plurality of input / output circuits for communicating with an external memory controller. Various signals from the memory controller may be provided to the TSV region 2212 through the PHY region 2213 and provided to at least one core die 2220 through the TSV 2230.
[0112] According to an example embodiment, such as Figure 12 to Figure 14 The transfer learning manager (TLM) 2240 shown in FIG. 2240 may be implemented in the buffer die 2210. The transfer learning manager 2240 may correspond to the reference Figure 12 to Figure 14 The transfer learning manager 220 is described.
[0113] The DA region 2214 may directly communicate with an external tester through a conductive member provided on an outer surface of the memory 2200 in a test mode of the memory 2200. Various signals from the tester may be provided to at least one core die 2220 through the DA region 2214 and the TSV region 2212. In a modified embodiment, various signals from the tester may be provided to at least one core die 2220 through the DA region 2214, the PHY region 2213, and the TSV region 2212.
[0114] As described above, when performing learning, a neural network processor in a neural network system can reduce learning time, save power, reduce processing requirements, and otherwise reduce resource consumption. The processor can exclude one or more layers from the learning iteration based on determining that one or more layers of the neural network will not benefit from or will not greatly benefit from the learning iteration that excludes one or more layers. Comparison of weights (e.g., distribution of weight values) between iterations can be used as a basis for determining that one or more layers will not benefit from additional learning iterations or will not greatly benefit from additional learning iterations. As described for various embodiments, weights from different iterations can be compared based on histogram information, statistics, or raw data of weight values in the weights. As a particular benefit, when weights from a trained neural network are used as initial weights for another neural network to be trained, the neural network processor can reduce the transfer learning time because layers with saturated weights from the trained neural network can be quickly detected based on the teachings herein. Therefore, based on the techniques described herein for increasing the learning speed, the amount of time required to train a neural network to obtain a learning result can be reduced.
[0115] While the inventive concepts of the present disclosure have been particularly shown and described with reference to embodiments thereof, it will be understood that various changes in form and details may be made therein without departing from the spirit and scope of the appended claims.
Claims
1. A neural network system for image processing, comprising: A neural network processor configured to: perform learning including a plurality of learning iterations on a plurality of layers; determining at least one layer among the plurality of layers in which the learning is interrupted based on a result of comparing, for each layer among the plurality of layers, a distribution of first weight values obtained through a first learning iteration with a distribution of second weight values obtained through a second learning iteration after the first learning iteration; and performing a third learning iteration after the second learning iteration on a plurality of layers among the plurality of layers except for the at least one layer in which the learning is determined to be interrupted; as well as A memory configured to store first distribution information about the distribution of the first weight value and second distribution information about the distribution of the second weight value when the second learning iteration is completed, and configured to provide the first distribution information and the second distribution information to the neural network processor.
2. The neural network system according to claim 1, wherein the first distribution information includes first histogram information about the first weight value in each of the plurality of layers, the second distribution information includes second histogram information about the second weight value in each of the plurality of layers, and The neural network processor is configured to determine the at least one layer in the plurality of layers where the learning is interrupted based on a difference between the first histogram information and the second histogram information in each of the plurality of layers.
3. The neural network system according to claim 2, The neural network processor is configured to: subtract the second histogram information from the first histogram information; obtain a difference indicator value by adding the absolute values of the values included in the histogram information obtained by subtraction, wherein the histogram information obtained by subtraction is obtained by the subtraction step; and determine the interruption of the learning for a layer having a difference indicator value equal to or less than a threshold.
4. The neural network system according to claim 1, wherein the first distribution information includes first statistical information about the first weight value in each of the plurality of layers, the second distribution information includes second statistical information about the second weight value in each of the plurality of layers, and The neural network processor is configured to determine the at least one layer in the plurality of layers where the learning is interrupted based on a difference between the first statistical information and the second statistical information in each of the plurality of layers.
5. The neural network system according to claim 4, wherein the first statistical information includes at least one selected from a mean, a variance, a standard deviation, a maximum value, and a minimum value of the first weight values in each of the plurality of layers; the second statistical information includes at least one selected from a mean, a variance, a standard deviation, a maximum value, and a minimum value of the second weight values in each of the plurality of layers; and The neural network processor is configured to calculate a difference between the first statistic and the second statistic and determine an interruption of the learning for a layer having a difference equal to or less than a threshold.
6. The neural network system according to claim 1, The memory includes a transfer learning management circuit, which is configured to: obtain the first distribution information based on the first weight value and store the first distribution information after the first learning iteration is completed; and obtain the second distribution information based on the second weight value and store the second distribution information after the second learning iteration is completed.
7. A learning method for a neural network system for image processing, comprising multiple learning iterations on multiple layers, the learning method comprising: Storing the first weight obtained through the Nth learning iteration in a memory; determining at least one layer in which learning is interrupted among the plurality of layers based on a first weight value included in the first weight and a second weight value included in a second weight obtained through an (N-1)th learning iteration; as well as An (N+1)th learning iteration is performed on a plurality of layers among the plurality of layers except the at least one layer for which it has been determined that the learning is interrupted.
8. The learning method according to claim 7, Wherein determining the at least one layer where the learning is interrupted includes determining the at least one layer among the multiple layers where the learning is interrupted based on a result of comparing first distribution information of the first weight value in each of the multiple layers with second distribution information of the second weight value.
9. The learning method according to claim 8, wherein the first distribution information includes first histogram information about the first weight value in each of the plurality of layers, the second distribution information includes second histogram information about the second weight value in each of the plurality of layers, and Determining the at least one layer where the learning is interrupted further includes determining the at least one layer where the learning is interrupted among the multiple layers based on a difference between the first histogram information and the second histogram information in each of the multiple layers.
10. The learning method according to claim 9, The at least one layer determining that the learning is interrupted further comprises: subtracting the second histogram information from the first histogram information; Obtaining a difference indicator value by adding absolute values of values included in the subtracted histogram information obtained by the subtracting step; as well as An interruption of the learning is determined for a layer having a difference indicator value equal to or less than a threshold value.
11. The learning method according to claim 8, wherein the first distribution information includes first statistical information about the first weight value in each of the plurality of layers, the second distribution information includes second statistical information about the second weight value in each of the plurality of layers, and Determining the at least one layer where the learning is interrupted further comprises: The at least one layer in the plurality of layers in which the learning is interrupted is determined based on a difference between the first statistical information and the second statistical information in each of the plurality of layers.
12. The learning method according to claim 11, wherein the first statistical information includes at least one selected from the mean, variance, standard deviation, maximum value and minimum value of the first weight value in each of the plurality of layers; the second statistical information includes at least one selected from the mean, variance, standard deviation, maximum value and minimum value of the second weight value in each of the plurality of layers; and Determining the at least one layer where the learning is interrupted further comprises: calculating the difference between the first statistical information and the second statistical information, and Interruption of the learning is determined for layers having a difference equal to or less than a threshold value.
13. The learning method according to claim 8, further comprising: After the Nth learning iteration is completed, using a transfer learning management circuit included in the memory to generate the first distribution information based on the first weight value; as well as The first distribution information is stored in the transfer learning management circuit.
14. The learning method according to claim 13, wherein the memory comprises at least one memory device, In the structure of the at least one memory device, at least one core die is stacked on a buffer die, and The transfer learning management circuit is implemented in the buffer die.
15. The learning method according to claim 13, Also includes: After the Nth learning iteration is completed, using the memory to provide the first distribution information and the second distribution information to the neural network processor, wherein the first distribution information and the second distribution information are stored in the transfer learning management circuit, The at least one layer that determines that the learning is interrupted is executed by the neural network processor.
16. The learning method according to claim 13, Also includes: After the Nth learning iteration is completed, using the memory to provide the first distribution information and the second distribution information to the processor, wherein the first distribution information and the second distribution information have been stored in the transfer learning management circuit; performing, using the processor, determining the at least one layer for which the learning is interrupted; as well as Using the processor, information is provided to a neural network processor regarding the at least one layer for which it is determined that the learning is interrupted.
17. A transfer learning method for a neural network processor for image processing, comprising multiple learning iterations on multiple layers, the transfer learning method comprising: Storing a first weight value obtained through a first learning iteration in a memory outside the neural network processor; storing a second weight value obtained by a second learning iteration in the memory, the second learning iteration being after the first learning iteration; receiving first distribution information of the first weight value and second distribution information of the second weight value from the memory, and determining at least one layer of the plurality of layers in which the learning is interrupted based on the first distribution information and the second distribution information; as well as A third learning iteration is performed on a plurality of layers among the plurality of layers except the at least one layer for which it has been determined that the learning is interrupted.
18. The transfer learning method according to claim 17, wherein the first distribution information includes first histogram information about the first weight value in each of the plurality of layers, the second distribution information includes second histogram information about the second weight value in each of the plurality of layers, and Determining the at least one layer where the learning is interrupted further comprises: The at least one layer in which the learning is interrupted among the plurality of layers is determined based on a difference between the first histogram information and the second histogram information in each of the plurality of layers.
19. The transfer learning method according to claim 18, The at least one layer determining that the learning is interrupted further comprises: subtracting the second histogram information from the first histogram information; Obtaining a difference indicator value by adding absolute values of values included in the subtracted histogram information obtained by the subtracting step; as well as An interruption of the learning is determined for a layer having a difference indicator value equal to or less than a threshold value.
20. The transfer learning method according to claim 17, wherein the first distribution information includes first statistical information about the first weight value in each of the plurality of layers, the second distribution information includes second statistical information about the second weight value in each of the plurality of layers, and Determining the at least one layer where the learning is interrupted includes determining the at least one layer where the learning is interrupted among the plurality of layers based on a difference between the first statistical information and the second statistical information in each of the plurality of layers.
Citation Information
Patent Citations
How the safety device works
KR1020190053887A
Method for analyzing and processing experimental data based on artificial neural network
CN101814158A
Method for self-adaptively adjusting learning rate by tracking and controlling neural network
CN103926832A