Electronic device and method for performing operation of the electronic device

By selectively adjusting the operation data volume of the neural network model according to the hardware resource information and the importance of data, the problem of difficult reduction in the data capacity of the artificial intelligence model under the limitation of hardware resources is solved, and performance is maintained and equipment operation time is extended.

CN112052943BActive Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010493206.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-05
Filing Date
2020-06-03
Publication Date
2025-05-13
Estimated Expiration
2040-06-03

AI Technical Summary

Technical Problem

When hardware resources are limited, the data capacity of the artificial intelligence model is difficult to reduce without affecting performance, resulting in reduced inference accuracy or greater latency.

Method used

By acquiring the hardware resource information of the electronic device, the operation data used for the neural network model is selectively acquired and used according to the importance of the data, and the amount of data is flexibly adjusted to adapt to hardware conditions.

Benefits of technology

It realizes that under the constraints of hardware resources, the data capacity of the artificial intelligence model is reduced without significantly reducing performance, extending the operating time of the device and maintaining high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112052943B_ABST
    Figure CN112052943B_ABST
Patent Text Reader

Abstract

A method for an electronic device to perform operations of an artificial intelligence model includes: when multiple data for the operation of a neural network model are stored in a memory, obtaining resource information about the hardware of the electronic device, the multiple data respectively having different levels of importance from each other; based on the obtained resource information, obtaining data to be used for the operation of the neural network model among the multiple data according to the level of importance of each of the multiple data; and performing the operation of the neural network model by using the obtained data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on and claims the benefit of priority from Korean Patent Application No. 10-2019-0066396 filed in the Korean Intellectual Property Office on June 5, 2019, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates to an electronic device and a method for performing operations of the electronic device, and more particularly, to a method for performing operations of an artificial neural network model. Background Art

[0004] Recently, research on implementing artificial intelligence models (such as deep learning models) by using hardware is ongoing. In the case of implementing artificial intelligence models by using hardware, the operation speed of artificial intelligence models can be greatly improved, and the use of various deep learning models that were previously difficult to use due to memory size or restrictions on response time becomes possible.

[0005] Algorithms are proposed for continuously improving the performance of artificial intelligence models from the perspective of hardware implementation, such as data quantization technology that reduces the amount of operating data to reduce operating delays and power consumption.

[0006] Data quantization is a method of reducing the amount of information representing matrix parameters, for example, and can decompose real data into binary data and scaling factors, and represent the data as approximate values. Since quantized data cannot achieve the accuracy of the original data, the accuracy of reasoning of the artificial intelligence model using quantized data may be lower than the accuracy of the original reasoning of the artificial intelligence model. However, considering the limited situation of hardware, quantization can save the usage of memory or the consumption of computing resources to a certain extent, so it is actively being studied in the field of on-device artificial intelligence. Summary of the invention

[0007] Embodiments of the present disclosure provide an electronic device that reduces the data capacity of an artificial intelligence model while minimizing degradation in the performance of the artificial intelligence model, and a method for performing operations of the artificial intelligence model.

[0008] According to one aspect of the present disclosure, a method for an electronic device to perform operations of an artificial intelligence model includes the following operations: when a plurality of data having different degrees of importance from each other are stored in a memory, obtaining resource information about hardware of the electronic device, the plurality of data being used for the operation of a neural network model; based on the obtained resource information, obtaining some data among the plurality of data to be used for the operation of the neural network model according to the degree of importance of each of the plurality of data; and performing the operation of the neural network model by using some of the obtained data.

[0009] According to one aspect of the present disclosure, an electronic device includes: a memory storing a plurality of data having different degrees of importance from each other; and a processor configured to obtain some data to be used for the operation of a neural network model among the plurality of data stored in the memory according to the degree of importance of each of the plurality of data based on resource information of hardware for the electronic device, and to perform the operation of the neural network model by using the obtained some data.

[0010] According to an embodiment, the amount of data used for the neural network model can be flexibly adjusted according to the requirements of the hardware. For example, an improved effect can be expected in at least one of delay, power consumption, or user aspects.

[0011] In terms of latency, the time for requesting to run the neural network model can be taken into account while excluding binary data with low importance and selectively using only binary data with high importance in the operation of the neural network model, thereby meeting the requirements with minimal accuracy reduction.

[0012] In terms of power consumption, in an example where it is determined that the remaining battery level of an electronic device is low, the amount of data can be controlled so that the neural network model operates at minimum performance taking into account the conditions of the hardware (such as low battery level), thereby extending the operating time of the electronic device.

[0013] On the user (or developer) side, in the related art, it is difficult for the user to judge the optimal amount of data to be used for the neural network model, considering the amount of operation of the artificial intelligence application installed on the electronic device and other limitations. However, according to an embodiment, the amount of data can be automatically and appropriately adjusted based on the conditions of the hardware in consideration of latency and power consumption, so the accuracy of the reasoning of the artificial intelligence model can be maintained above a certain level while effectively operating the neural network model.

[0014] According to the embodiment, the problem that the neural network model does not operate or the delay becomes large in the case of limited hardware resources can be overcome. That is, the amount of data used for the operation of the neural network model can be flexibly adjusted in consideration of the currently available resources of the hardware, so that even if the amount of data increases, the delay remains below a certain level, and the neural network model can operate without interruption and above a certain accuracy threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram showing a configuration of an electronic device according to an embodiment;

[0016] Figure 2 is a block diagram illustrating components for neural network operations including a processing unit according to an embodiment;

[0017] Figure 3 An example of a scheduling syntax according to an embodiment is shown;

[0018] Figure 4 is a block diagram illustrating components for neural network operations including a plurality of processing units according to an embodiment;

[0019] Figure 5 is a diagram showing a process of storing quantized parameter values ​​in a memory for each bit sequence according to an embodiment;

[0020] Figure 6 is a diagram showing a process of storing quantized parameter values ​​in a memory for each bit sequence according to an embodiment;

[0021] Figure 7 is a flow chart illustrating a method for an electronic device to perform an operation according to an embodiment; and

[0022] Figure 8 is a block diagram showing a detailed configuration of an electronic device according to an embodiment. DETAILED DESCRIPTION

[0023] The embodiments of the present disclosure may be modified in various ways. Therefore, specific embodiments are shown in the drawings and described in detail in the detailed description. However, it will be understood that the present disclosure is not limited to specific embodiments, but includes all modifications, equivalents and substitutions without departing from the scope and spirit of the present disclosure. In addition, because known functions or configurations may make the present disclosure unclear with unnecessary details, they are not described in detail.

[0024] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0025] Figure 1 1 is a block diagram showing a configuration of an electronic device 100 according to an embodiment. Figure 1 As shown, the electronic device 100 includes a memory 110 and a processor 120 .

[0026] The electronic device 100 may be a server, a desktop PC, a laptop computer, a smart phone, a tablet PC, etc. Alternatively, the electronic device 100 is a device using an artificial intelligence model, and it may be a cleaning robot, a wearable device, a home appliance, a medical device, an Internet of Things (IoT) device, or an autonomous vehicle.

[0027] Figure 1 The memory 110 in the electronic device 100 may store a plurality of data having different degrees of importance from each other. The plurality of data having different degrees of importance from each other may include, for example, parameter values ​​of quantized matrices used for the operation of the neural network model. In this case, if there are a plurality of matrices used for the operation of the neural network model, the electronic device 100 may include parameter values ​​having different degrees of importance from each other for each of the plurality of quantized matrices.

[0028] As another example, the plurality of data respectively having different degrees of importance from each other may be a plurality of neural network layers respectively having different degrees of importance from each other for operation of a neural network model.

[0029] As another example, the plurality of data having different importance levels may be parameter values ​​of a matrix before being quantized. In this case, the parameter values ​​of the matrix may be composed of binary data having different importance levels, and, for example, the importance levels may increase according to the bit order.

[0030] In the case where the plurality of data are parameter values ​​of a quantized matrix, the quantization process of the matrix performed to obtain the parameter values ​​may be performed by the electronic device 100. Alternatively, the quantization process may be performed at an external device, and the parameter values ​​of the quantized matrix may be pre-stored in the memory 110.

[0031] In the case where the parameter values ​​of the matrix are quantized, the parameter values ​​of the matrix of full-precision values ​​can be converted into k numbers of binary data (e.g., +1 and -1) (or quantized bits) bi values ​​and scaling coefficient factors ai values. In the case where the operation of the neural network model is performed by using the parameter values ​​of the quantized matrix, the usage of memory and the usage of the computer during the reasoning between the neural network layers can be reduced, but the accuracy of the reasoning may be deteriorated.

[0032] Therefore, various quantization algorithms for improving inference accuracy can be used.

[0033] For example, in the case of quantizing the parameter value of the w matrix to have a number of bits of k, various algorithms that satisfy the conditions of [Formula 1] can be used.

[0034] [Formula 1]

[0035] where b i ∈{-1,+1} n ,

[0036] In order to satisfy the conditions of [Formula 1], as an example, an alternating algorithm (e.g., an alternating multi-bit algorithm) or the like may be used. The alternating algorithm is an algorithm that repeatedly updates binary data and coefficient factors and finds a value that minimizes [Formula 1]. For example, in the alternating algorithm, the binary data may be calculated and updated again based on the updated coefficient factors, and the coefficient factors may be calculated and updated again based on the updated binary data. This process may be repeated until the error value becomes less than or equal to a specific value.

[0037] The alternating algorithm can guarantee high accuracy, but may require a large amount of computing resources and operation time for updating binary data and coefficient factors. In particular, in the alternating algorithm, when the parameter value is quantized into k number of bits, all k bits have similar importance, so accurate reasoning is possible only when the operation of the neural network model is performed by using all k number of bits.

[0038] In other words, in the case where the operation of the neural network model is performed while omitting some bits, the accuracy of the neural network operation may be degraded. As an example, in a resource-constrained environment of the electronic device 100 (e.g., an on-device artificial intelligence chip environment), in the case where the operation of the neural network model is performed by using only some bits in consideration of the resources of the electronic device 100, the accuracy of the neural network operation may be degraded.

[0039] Therefore, an algorithm that can flexibly respond to limited hardware resources may be required.As an example, for quantization of parameter values ​​of a matrix, a greedy algorithm may be used that quantizes the parameter values ​​so that each bit of binary data has a different degree of importance from each other.

[0040] In the case where the parameter value of the matrix is ​​quantized into k number of bits by using a greedy algorithm, the first binary data and coefficient factors of k number of bits in [Formula 1] may be calculated by using [Formula 2].

[0041] [Formula 2]

[0042]

[0043] Next, the i-th bit (1 < i ≤ k) can repeat the same calculation as in [Formula 3] for r, where r is the difference between the original parameter value and the first quantized value. That is, by calculating the i-th bit by virtue of using the residual remaining after calculating the (i - 1)-th bit, the parameter value of the quantized matrix with k bits can be obtained.

[0044] [Formula 3]

[0045] where

[0046] Therefore, the parameter value of the quantized matrix with k bits can be obtained.

[0047] In addition to the above, in order to further minimize the error between the original parameter value of the matrix and the quantized parameter value, a fine-grained greedy algorithm based on the greedy algorithm can be used. The fine-grained greedy algorithm can update the coefficient factor in [Formula 4] by using the vector b determined via the greedy algorithm.

[0048] [Formula 4]

[0049] where B j = [b 1 ,..., b j ,

[0050] In the case of using the greedy algorithm (or the fine-grained greedy algorithm), as the order of the bits becomes higher, the value of the coefficient factor becomes smaller, and thus the importance of the bits decreases. Therefore, even if the bit omission operation for bits with a high bit order is performed, it hardly affects the inference of the neural network. In the neural network model, a method of deliberately applying noise to about 10% of the parameter values can be used to improve the inference accuracy. In this case, even if an operation is performed while omitting some bits of the binary data quantized by the greedy algorithm, it is difficult to consider that the inference accuracy of the neural network deteriorates overall. Instead, there may be a case where the inference accuracy of the neural network is considerably improved.

[0051] When the parameter value of the matrix is quantized to include bits with different importance levels, it becomes possible to perform an adaptive operation of the neural network model in consideration of the requirements of given computing resources (such as power consumption, operation time). That is, it becomes possible to adjust the performance of the neural network model according to the importance level of each of various neural network models.

[0052] In addition, without having to laboriously consider the optimal number of quantized bits for each neural network model, it becomes possible to flexibly adjust the number of quantized bits to be applied to the neural network model according to conditions or limited conditions after a matrix quantized to a certain extent is loaded. In addition, the development cost required to find the optimal operating conditions for each neural network model can be saved.

[0053] Figure 1 The processor 120 in the electronic device 100 can control the overall operation of the electronic device 100. The processor 120 can be a general-purpose processor (such as a central processing unit (CPU) or an application processor), a graphics-specific processor (such as a GPU), or a system on a chip (SoC) that performs processing (such as an artificial intelligence (AI) chip on a device), a large-scale integration (LSI), or a field programmable gate array (FPGA). The processor 120 may include one or more of a CPU, a microcontroller unit (MCU), a microprocessing unit (MPU), a controller, an application processor (AP) or a communication processor (CP), and an ARM processor, or may be defined by the term.

[0054] Although the processor 120 stores a plurality of data having different degrees of importance from each other in the memory 110, the processor 120 may obtain some data to be used for the operation of the neural network model among the plurality of data according to the degree of importance of each of the plurality of data stored in the memory 110 based on resource information for the hardware of the electronic device 100. As an example, in the case where the plurality of data includes binary data as parameter values ​​of a quantized matrix, the processor 120 may obtain the number of binary data to be used for the operation of the neural network model among the plurality of data. For example, as the bit order of the binary data increases, the degree of importance may decrease.

[0055] When some data to be used for the operation of the neural network model is obtained, the processor 120 can perform the operation of the neural network model by using the obtained data. As an example, the processor 120 can perform matrix operations on the input value and each bit of the binary data, and sum the operation results of each bit and obtain the output value. In the case where there are multiple neural network processing units, the processor 120 can use multiple neural network processing units to perform matrix parallel operations based on the order of each bit of the binary data.

[0056] Figure 2 is a block diagram illustrating components for neural network operations including a processing unit according to an embodiment.

[0057] Figure 2 The electronic device 100 in which the block diagram is included may include an on-device artificial intelligence chip that performs neural network inference, for example, by using hardware.

[0058] The parameter values ​​of the matrix used for neural network inference by using hardware can be in a state of being quantized by using a greedy algorithm, so that important binary data can be selectively used. The importance of binary data as quantized parameter values ​​can decrease as the bit order increases.

[0059] exist Figure 2 In the electronic device 100, a scheduler 210, an adaptive controller 220, a direct memory access controller (DMAC) 230, a processing unit 240, and an accumulator 250 may be included. At least one of the scheduler 210, the adaptive controller 220, the direct memory access controller 230, the processing unit 240, or the accumulator 250 may be implemented as software and / or hardware. For example, the scheduler 210, the adaptive controller 220, the direct memory access controller 230, the processing unit 240, and the accumulator 250 may be a functional block of the processor 120. As another example, the scheduler 210, the adaptive controller 220, and the direct memory access controller 230 are functional blocks of a first processor (which is a sub-processor of the processor 120), and the processing unit 240 and the accumulator 250 may be functional blocks of a second processor (which is another sub-processor of the processor 120). The first processor is a processor responsible for the control of the second processor, and the second processor is a processor optimized for operation, for example, it may be an artificial intelligence processor or a graphics processor.

[0060] Scheduler 210 may receive resource information related to the hardware and an instruction requesting to execute a neural network model.

[0061] In response to an instruction requesting to run a neural network model, scheduler 210 may refer to lookup table 270 to determine the number of quantized bits to be used for the operation of each neural network model. The lookup table may be stored in a read-only memory (ROM) or random access memory (RAM) area of ​​processor 120, or in memory 110 external to processor 120.

[0062] The lookup table may store, for example, k number of scheduling modes. In this case, the number of quantized bits to be used for the operation of the neural network model may be predefined for each scheduling mode. For example, the number of quantized bits may be differently defined depending on the importance of the neural network model.

[0063] The scheduler 210 may first determine the hardware conditions for executing the reasoning job according to the request instruction with a degree of accuracy higher than a certain level. For example, the scheduler 210 may determine the following as hardware conditions: the total number of operations used by the processing unit 240 to execute the reasoning job; power consumption, which is the power used for the operation process of the neural network model; and delay, which is the time until the output value is obtained, which is the operation time of the neural network model. The scheduler 210 may compare the currently available hardware resource information (e.g., power consumption per time, delay) with the hardware conditions for executing the reasoning job, and determine the number of quantized bits used for each neural network model.

[0064] Will pass Figure 3 The scheduling syntax in describes in more detail the process by which scheduler 210 determines the number of quantized bits to use for each neural network model.

[0065] The adaptive controller 220 may control the operation order of the neural network models, or may perform control so that operations are performed when each neural network model has a different bit amount from each other.

[0066] For example, the adaptive controller 220 may control the processing unit 240 and the direct memory access controller 230 in consideration of the number of quantized bits used for each neural network model obtained from the scheduler 210. Alternatively, the adaptive controller 220 may obtain resource information related to hardware resources from the scheduler 210 and determine the number of quantized bits to be used for the operation of the neural network model in consideration of this.

[0067] The adaptive controller 220 may control the processing unit 240 and the direct memory access controller 230 to perform control so that quantized bits after a certain number are not used for the operation of the neural network model.

[0068] The direct memory access controller 230 may perform control such that the input value and the quantized parameter value stored in the memory 110 are provided to the processing unit 240 under the control of the adaptive controller 220. In the memory 110, the quantized parameter values may be aligned and stored according to the order of bits. For example, the quantized parameter value of the first bit may be aligned into a data format and stored, the quantized parameter value of the second bit may be aligned into a data format and stored, and subsequently, the quantized parameter value of the Nth bit may be aligned into a data format and stored. In this case, the direct memory access controller 230 may perform control such that the quantized parameter values of the first bit to the Nth bit stored in the memory 110 are sequentially provided to the processing unit 240 under the control of the adaptive controller 220. Alternatively, the direct memory access controller 230 may perform control such that the quantized parameter values of the first bit to the Kth bit (N < K) are sequentially provided to the processing unit 240.

[0069] The processing unit 240 may perform matrix operations by using the input value and the quantized parameter value received from the memory 110, and obtain the operation results for each bit order. The accumulator 250 may sum up the operation results for each bit order and obtain the output result (or output value).

[0070] As an example, in the case where the parameter value is quantized into N bits, the processing unit 240 may call the matrix multiplication operation N times and perform the operations sequentially for the first bit to the Nth bit.

[0071] Figure 3 An example of the scheduling syntax executed at the scheduler 210 according to an embodiment is shown.

[0072] In Figure 3 the definition part 310 of the scheduling syntax may be predefined: 'ExecutionModel', which has a neural network model (such as a speech recognition model, an image recognition model, etc.) as a value, and the neural network model is the object to be run; 'Constraints', which has information related to hardware resources (such as power consumption, latency) as a value;'mode', which has a scheduling mode as a value;'max_cost', which is the maximum operation cost of the hardware resources obtained from the look-up table; and 'cost', which is the operation cost obtained from the look-up table with respect to the neural network model.

[0073] The scheduling mode may be included in the look-up table, for example, and define the optimal number of bits for quantization of the operations to be used for each neural network model. As an example, in Figure 3In the , 16 scheduling modes are defined, and it can be defined that, in the case of mode 0, all neural network models run with full precision, in the case of mode 15, all neural network models are calculated by using only 1 bit, and in mode 2, for example, the speech recognition model uses 3 bits as quantized bit data and the image recognition model uses 2 bits as quantized bit data.

[0074] exist Figure 3 In the while conditional sentence 320, the hardware resources can be taken into account to compare the operation cost of the neural network model according to the current scheduling mode and the maximum operation cost.

[0075] As a result of the comparison, when determining the optimal scheduling mode for each neural network model taking into account the hardware resources, the scheduler 210 can obtain the optimal number of quantized bits to be used for each neural network model under the current hardware conditions from the lookup table as the return value 330 of the scheduling syntax, and provide it to the adaptive controller 220.

[0076] Figure 4 is a block diagram illustrating components for neural network operations including multiple processing units according to one embodiment.

[0077] exist Figure 4 In the case where there is a free space in the operation area of ​​the processor 120, a plurality of processing units 241 to 244 (e.g., a first processing unit 241, a second processing unit 242, a third processing unit 243, and a fourth processing unit 244) may be provided in the processor 120. In this case, parallel operations may be performed by using the plurality of processing units 241 to 244.

[0078] Reference above Figure 2 The scheduler 210 , the adaptive controller 220 , the direct memory access controller 230 , and the accumulator 250 are described, and thus repeated descriptions will be omitted.

[0079] exist Figure 4 In the embodiment, the plurality of processing units 241 to 244 may perform matrix parallel operations based on bit order. For example, the first processing unit 241 may perform operations with respect to the input value and the first quantized bit, the second processing unit 242 may perform operations with respect to the input value and the second quantized bit, the third processing unit 243 may perform operations with respect to the input value and the third quantized bit, and the fourth processing unit 244 may perform operations with respect to the input value and the fourth quantized bit. The adder 260 may collect the operation results of the plurality of processing units 241 to 244 and send them to the accumulator 250.

[0080] In the case of using a plurality of processing units 241 to 244, the adaptive controller 220 may control each of the plurality of processing units 241 to 244. The direct memory access controller 230 may control the memory 110 so that the quantized parameter value is input while being differentiated according to the bit order in consideration of the bit order processed by each of the plurality of processing units 241 to 244. In particular, in the case of using a plurality of processing units 241 to 244, the scheduler 210 may determine the scheduling mode in consideration of the power consumption of the processor according to the operations of the plurality of processing units 241 to 244 and the delay of the neural network operation.

[0081] In the case of performing the operation of the neural network model by using the plurality of processing units 241 to 244 , in the memory 110 , the quantized parameter values ​​may be realigned and stored for each bit sequence.

[0082] Figure 5 is a diagram illustrating a process of storing quantized parameter values ​​in the memory 110 for each bit sequence according to an embodiment.

[0083] For example, in Figure 5 In the example, the 32-bit parameter value of the real number type can exist as the parameter value of the neural network model. In this case, if the parameter value is quantized to 3 bits, the parameter value of the real number type can be represented as a coefficient factor 521 and 3 bits 522, 523, 524 of binary data ( Figure 5 The number 520 in the figure).

[0084] For efficient operation of the plurality of processing units 241 to 244, the quantized parameter values ​​may be realigned. Figure 5 As shown in the reference numeral 530 in FIG. 1 , the quantized parameter values ​​can be realigned according to the bit order. For example, with respect to Figure 4 In each of the first processing unit 241, the second processing unit 242 and the third processing unit 243, the quantized parameter value can be as follows: Figure 5 The data are realigned as shown by reference numerals 531 , 532 , and 533 in the figure and stored in the memory 110 .

[0085] Figure 6 A state in which the quantized parameter values ​​are stored in the memory 110 is shown. In an embodiment, the memory 110 may include a DRAM 600. The quantized parameter values ​​may be aligned according to a bit order and stored in the DRAM 600. In this case, 32 binary data may be included in a 32-bit word value.

[0086] As described above, the quantized parameter values ​​are stored in the memory 110 in units of words to correspond to the operation of each of the multiple processing units 241 to 244, so the quantized parameter values ​​for the neural network operation can be efficiently read from the memory 110 and sent to each of the multiple processing units 241 to 244.

[0087] Figure 7 is a flowchart illustrating a method for the electronic device 100 to perform an operation according to an embodiment.

[0088] A plurality of data having different importance levels for the operation of the neural network model may have been stored in a memory. The plurality of data having different importance levels may include parameter values ​​of a quantized matrix for the operation of the neural network model. The parameter values ​​of the quantized matrix may include binary data having different importance levels. For example, as the bit order of the binary data increases, the importance level of the binary data may decrease. The parameter values ​​of the quantized matrix may include parameter values ​​of the matrix quantized by using a greedy algorithm.

[0089] exist Figure 7 In operation 701, when the plurality of data having different degrees of importance are stored in the memory, the electronic device 100 may obtain resource information related to the hardware. The resource information related to the hardware may include, for example, at least one of the power consumption of the electronic device 100, the number, type and / or specification of the processing units that perform the operation of the neural network model, or the delay of the neural network model.

[0090] exist Figure 7 In operation 703, based on the acquired resource information, the electronic device 100 may obtain some data among the multiple data to be used for the operation of the neural network model according to the importance of each of the multiple data. For example, the electronic device 100 may refer to a lookup table in which multiple scheduling modes are defined to obtain some data for the operation of the neural network model. The electronic device 100 may obtain the number of binary data used for the neural network model among the multiple data.

[0091] exist Figure 7 In operation 705, the electronic device 100 may perform operations of the neural network model by using some of the obtained data. For example, the electronic device 100 may perform matrix operations on each bit of the input value and the binary data, and sum the operation results of each bit and obtain an output value. Alternatively, in the case where there are multiple neural network processing units, the electronic device 100 may use the multiple neural network processing units to perform matrix parallel operations.

[0092] According to an embodiment of the present disclosure, a plurality of neural network models may be provided on the electronic device 100. The plurality of neural network models may, for example, be implemented as at least one device-on-chip and provided on the electronic device 100, or may be stored as software in the memory 110 of the electronic device 100. For adaptive operation taking into account limited hardware resources, the electronic device 100 may acquire resource information related to the hardware, and determine at least one neural network model to be operated among the plurality of neural network models based on the acquired resource information. For example, the electronic device 100 may determine at least one neural network model according to priority taking into account the accuracy of reasoning or the operation speed of the neural network model.

[0093] According to an embodiment of the present disclosure, the electronic device 100 may download at least one neural network model to be operated among the multiple neural network models from an external device. For example, for adaptive operations taking into account limited hardware resources, the electronic device 100 may obtain resource information related to the hardware of the electronic device 100 and send the obtained resource information to the external device. When the external device sends at least one neural network model to the electronic device 100 based on the obtained resource information, the electronic device 100 may store the received neural network model in the memory 110 and use it when performing an inference function. In this case, a minimum neural network model for inference may be provided to the electronic device 100, thereby reducing the consumption of internal resources of the electronic device 100 or the consumption of network resources used for communication with a server, and providing fast results for inference requests.

[0094] Figure 8 is a block diagram showing a detailed configuration of the electronic device 100 according to an embodiment.

[0095] according to Figure 8 , the electronic device 100 includes a memory 110, a processor 120, a communicator 130, a user interface 140, a display 150, an audio processor 160, and a video processor 170. Figure 8 Among the components shown, Figure 1 The parts shown overlap and detailed description will be omitted.

[0096] The processor 120 controls the overall operation of the electronic device 100 by using various programs stored in the memory 110 .

[0097] Specifically, the processor 120 includes a RAM 121 , a ROM 122 , a main CPU 123 , a graphic processor 124 , first to nth interfaces 125 - 1 to 125 - n , and a bus 126 .

[0098] The RAM 121 , the ROM 122 , the main CPU 123 , the graphic processor 124 , and the first to nth interfaces 125 - 1 to 125 - n may be connected to one another through a bus 126 .

[0099] The first to nth interfaces 125-1 to 125-n are connected to the above-mentioned various components. One of the interfaces may be a network interface connected to an external device through a network.

[0100] The main CPU 123 accesses the memory 110, and performs operation by using an operating system (OS) stored in the memory 110. Then, the main CPU 123 performs various operations by using various programs stored in the memory 110, and the like.

[0101] The ROM 122 stores a set of instructions for system activation, etc. When a power-on instruction is input and power is supplied, the main CPU 123 copies the OS stored in the memory 110 to the RAM 121 according to the instructions stored in the ROM 122, and boots the system by running the OS. When the booting is completed, the main CPU 123 copies various application programs stored in the memory 110 to the RAM 121, and performs various operations by running the application programs copied to the RAM 121.

[0102] The graphic processor 124 generates a screen including various objects such as icons, images, and texts by using an operation part and a rendering part. The operation part operates attribute values ​​such as coordinate values, shape, size, and color according to the layout of the screen based on the received control instruction, and each object will be displayed by the attribute value. The rendering part generates a screen in various layouts including objects based on the attribute values ​​generated at the operation part. The screen generated in the rendering part is displayed in the display area of ​​the display 150.

[0103] The above-described operations of the processor 120 may be performed by a program stored in the memory 110 .

[0104] The memory 110 is provided separately from the processor 120 and may be implemented as a hard disk, a nonvolatile memory, a volatile memory, or the like.

[0105] The memory 110 may store a plurality of data used for the operation of the neural network model, and the plurality of data may include, for example, parameter values ​​of a quantized matrix.

[0106] According to one embodiment, the memory 110 may include at least one of an OS software module for operating the electronic device 100, an artificial intelligence model, a quantized artificial intelligence model, or a quantized module (e.g., a greedy algorithm module) for a quantized artificial intelligence model.

[0107] The communicator 130 is a component that performs communication with various types of external devices according to various types of communication methods. The communicator 130 includes a Wi-Fi chip 131, a Bluetooth chip 132, a wireless communication chip 133, a near field communication (NFC) chip 134, etc. The processor 120 performs communication with various external devices by using the communicator 130.

[0108] The Wi-Fi chip 131 and the Bluetooth chip 132 perform communication by using the Wi-Fi method and the Bluetooth method, respectively. In the case of using the Wi-Fi chip 131 or the Bluetooth chip 132, various types of connection information such as a service set identifier (SSID) or a session key are first sent and received, and by making the information perform a connection for communication, various types of information can be sent and received thereafter. The wireless communication chip 133 refers to a chip that performs communication according to various communication standards such as IEEE, ZigBee, the third generation (3G) third generation partnership project (3GPP), and long term evolution (LTE). The NFC chip 134 refers to a chip that operates in the NFC method using a 13.56MHz band among various RF-ID bands such as 135kHz, 13.56MHz, 433MHz, 860-960MHz, and 2.45GHz.

[0109] The processor 120 may receive parameter values ​​of at least one of an artificial intelligence module, a matrix included in an artificial intelligence model, or a quantized matrix from an external device through the communicator 130, and store the received data in the memory 110. Alternatively, the processor 120 may directly train the artificial intelligence model through an artificial intelligence algorithm, and store the trained artificial intelligence model in the memory 110. The artificial intelligence model may include at least one matrix.

[0110] The user interface 140 receives various user interactions. The user interface 140 can be implemented in various forms according to the implementation examples of the electronic device 100. For example, the user interface 140 can be a button provided on the electronic device 100, a microphone for receiving a user's voice, a camera for detecting a user's action, etc. When the electronic device 100 is implemented as a touch-based electronic device, the user interface 140 can be implemented as a touch screen constituting an interlayer structure with a touch pad. In this case, the user interface 140 can be used as a display 150.

[0111] The audio processor 160 is a component that performs processing of audio data. At the audio processor 160, various types of processing such as decoding or amplification of audio data, noise filtering, etc. may be performed.

[0112] The video processor 170 is a component that performs processing of video data. At the video processor 170, various types of image processing such as decoding, scaling, noise filtering, frame rate conversion, and resolution conversion of video data may be performed.

[0113] Through the method described above, the processor 120 may quantize the matrix included in the artificial intelligence model.

[0114] Embodiments of the present disclosure may be implemented as software (e.g., a program) including one or more instructions readable by a machine (e.g., a computer), the one or more instructions being stored in a machine-readable (e.g., computer-readable) storage medium (e.g., an internal memory) or an external memory. In one embodiment, a machine (e.g., a processor of an electronic device 100) may load one or more instructions stored in a storage medium and may operate according to the instructions. When the instructions are executed by the processor, the processor may perform a function corresponding to the instruction itself, or may use other components under its control. The instructions may include code generated or run by a compiler or interpreter. The machine-readable storage medium may be a non-transitory storage medium. The term "non-transitory" means that the storage medium does not include a signal and is tangible, but does not indicate whether the data is stored in the storage medium semi-permanently or temporarily.

[0115] The method according to the embodiment may be provided when stored as a computer program product. A computer program product refers to a product that can be traded between a seller and a buyer. The computer program product may be published online as a machine-readable storage medium (e.g., a compact disk ROM (CD-ROM)), or published online through an application store (e.g., play store TM). In the case of online publishing, at least a portion of the computer program product may be at least temporarily stored in a storage medium such as a memory of a manufacturer's server, an application store's server, and a relay server, or may be temporarily generated.

[0116] The above-described embodiments can be implemented in a recording medium, and the recording medium can be read by a computer or a computer-like device by using software, hardware or a combination thereof. In some cases, the above-described embodiments can be implemented as a processor itself. According to the implementation by means of software, the above-described embodiments (such as processes and functions) can be implemented as separate software modules. Each software module can perform one or more functions and operations described in this specification.

[0117] Computer instructions for performing processing operations of a machine according to an embodiment may be stored in a non-transitory computer-readable medium. When the instructions are executed by a processor of a particular machine, the computer instructions stored in such a non-transitory computer-readable medium cause the processing operations at the machine according to the embodiment to be performed by the particular machine. A non-transitory computer-readable medium refers to a medium that stores data semi-permanently and can be read by a machine, rather than a medium that temporarily stores data (such as registers, caches, and memories). As specific examples of non-transitory computer-readable media, there may be CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, ROMs, and the like.

[0118] In addition, each component (such as module or program) according to an embodiment can be composed of a single object or multiple objects. In addition, among the above-mentioned components, some components can be omitted, or other components can be further included in an embodiment. Some components (such as module or program) can be integrated into objects, and the function performed by each component before integration is performed in the same or similar manner. The operation performed by the module, program or other components according to an embodiment can be run sequentially, in parallel, repeatedly or heuristically. At least some operations can be run in different orders or omitted, or other operations can be added.

[0119] Although the embodiments of the present disclosure have been specifically shown and described with reference to the accompanying drawings, these embodiments are provided for illustrative purposes, and it will be understood by those skilled in the art that various modifications and other equivalent embodiments may be made from the present disclosure. Therefore, the true technical scope of the present disclosure is defined by the technical spirit of the appended claims.

Claims

1. A method for an electronic device to perform an operation of an artificial intelligence model, the method comprising: When a plurality of data for operation of the neural network model is stored in a memory of the electronic device, resource information for hardware of the electronic device is obtained, the plurality of data including parameter values ​​for operation of the neural network model, the parameter values ​​respectively having different importance levels from each other; quantizing the parameter value into a quantized matrix of binary data; realigning parameter values ​​of the quantized matrix based on a bit order of the binary data; selectively obtaining data for operation of the neural network model among the realigned parameter values ​​based on the acquired resource information; as well as performing operations of the neural network model by using the obtained data; Among them, the resource information of the hardware for the electronic device includes at least one of the power consumption of the electronic device, the number of processing units that execute operations of the neural network model, or a predetermined delay of the neural network model.

2. The method according to claim 1, wherein: The parameter values ​​of the quantized matrix include binary data having respective degrees of importance different from each other.

3. The method according to claim 2, wherein: In the binary data, the importance of the binary data decreases as the bit sequence of the binary data increases.

4. The method according to claim 2, wherein: The operation of executing the neural network model also includes: Perform matrix operations on each bit of input values ​​and binary data; summing the results of the operations on each bit; and Get the output value.

5. The method according to claim 2, wherein: The operation of executing the neural network model also includes using multiple neural network processing units to perform matrix parallel operations based on the order of each bit of the binary data.

6. The method according to claim 1, wherein: The parameter values ​​of the quantized matrix include parameter values ​​of the matrix quantized by using a greedy algorithm.

7. The method according to claim 1, wherein: The obtaining of data for operation of the neural network model also includes obtaining data to be used for operation of the neural network model by referring to a lookup table in which a plurality of scheduling modes are defined.

8. The method according to claim 1, wherein: The obtaining of data for operation of the neural network model also includes obtaining the number of binary data to be used for operation of the neural network model among the multiple data, and the number of binary data is less than the number of the multiple data.

9. An electronic device comprising: a memory storing a plurality of data including parameter values ​​for operation of a neural network model, wherein the parameter values ​​have different importance levels from each other; and The processor is configured as: quantizing the parameter value into a quantized matrix of binary data; realigning parameter values ​​of the quantized matrix based on a bit order of the binary data; selectively obtaining data for operation of the neural network model among the realigned parameter values ​​based on resource information for hardware of the electronic device, and performing operations of the neural network model by using the obtained data; Among them, the resource information of the hardware for the electronic device includes at least one of the power consumption of the electronic device, the number of processing units that execute operations of the neural network model, or a predetermined delay of the neural network model.

10. The electronic device according to claim 9, wherein: The parameter values ​​of the quantized matrix include binary data having respective degrees of importance different from each other.

11. The electronic device according to claim 10, wherein: In the binary data, the importance of the binary data decreases as the bit sequence of the binary data increases.

12. The electronic device according to claim 10, wherein: The processor is further configured to: Perform matrix operations on each bit of the input value and binary data, Sum the results of each bit operation, and Get the output value.

Citation Information

Patent Citations

  • Collision warning method and device using heterogeneous cameras

    KR1020190066396A

  • Neural network processor oriented automatic design method, device and optimization method

    CN107103113A

  • Deep processing unit (DPU) for implementing an artificial neural network (ANN)

    US20180046903A1