Method and device for identifying an image

By identifying and rearranging the sparseness of image data, the computational efficiency of neural network devices in image recognition tasks is optimized, and the problem of inefficient computing in the prior art is solved, and more efficient and accurate recognition results are achieved.

CN112434803BActive Publication Date: 2025-06-10SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010145984.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-26
Filing Date
2020-03-05
Publication Date
2025-06-10
Estimated Expiration
2040-03-05

AI Technical Summary

Technical Problem

When existing neural network devices handle image recognition tasks, there are a large number of unnecessary operations, resulting in inefficient computing.

Method used

By identifying the sparseness of the input image and rearranging the input data, the final feature map is generated when the convolution operation is performed on the convolution layer, thereby improving the accuracy and calculation efficiency of the recognition results.

Benefits of technology

By optimizing the sparsity processing and rearrangement of data, unnecessary convolution operations are reduced and image recognition efficiency and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112434803B_ABST
    Figure CN112434803B_ABST
Patent Text Reader

Abstract

Provided is a method and apparatus for identifying an image. The method of processing data includes: identifying sparsity of input data based on valid information included in the input data; rearranging the input data based on the form of the sparsity; and generating output data by processing the rearranged input data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of Korean Patent Application No. 10-2019-0104578, filed on Aug. 26, 2019, in the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field

[0002] The following description relates to a method and apparatus for processing data. More particularly, it relates to a method and apparatus for identifying an image. Background Art

[0003] A neural network represents a computational architecture using a biological brain as a model. According to the latest development of neural network technology, input data is analyzed by using a neural network device in various types of electronic systems, and effective information is extracted.

[0004] A neural network device performs a large number of operations on input data. Techniques capable of efficiently processing neural network operations have been studied. Neural network devices have been widely used for image recognition. Summary of the Invention

[0005] The present invention is provided to introduce a selection of concepts that are further described below in the detailed description in a simplified form. The present invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.

[0006] A method and apparatus for processing data, and a computer-readable recording medium having recorded thereon a program for executing the method on a computer are provided.

[0007] In one general aspect, a method of identifying an image includes: obtaining an input image; identifying sparsity of the input image as input data based on valid information included in the input image; rearranging the input image based on a form of the sparsity of the input image; generating a final feature map by performing a convolution operation on the rearranged input image on at least one convolutional layer; and obtaining an identification result based on the final feature map.

[0008] The step of generating the final feature map may include: identifying sparsity of input data of each convolutional layer; rearranging the input data based on a form of the sparsity; generating an output feature map by performing a convolution operation on the rearranged input data on each convolutional layer, wherein input data of a first convolutional layer among the at least one convolutional layer is the rearranged input image, and an output feature map output by a last convolutional layer among the at least one convolutional layer is the final feature map.

[0009] The step of rearranging the input data may include: rearranging the input data based on a distribution of invalid values included in the input data.

[0010] The step of rearranging the input data may include: rearranging the multiple rows included in the input data based on the number of invalid values in each of the multiple rows included in the input data.

[0011] The step of rearranging the input data may include: performing the rearrangement such that a first row of the input data among the multiple rows of the input data that includes the most invalid values is adjacent to a second row of the input data among the multiple rows of the input data that includes the least invalid values.

[0012] The step of rearranging the input data may include: shifting elements of multiple columns included in the input data according to a first rule.

[0013] The first rule may include shifting elements of the multiple columns included in the input data in the same direction by a specific size, and the first rule may be periodically applied to the multiple columns included in the input data.

[0014] The step of rearranging the input data may include: rearranging the multiple columns included in the input data to skip processing of at least one column that includes only invalid values among the multiple columns included in the input data.

[0015] The step of rearranging the input data may include: shifting a first element of a first column included in the input data to a position corresponding to a last element of a second column adjacent to the first column of the input data.

[0016] The step of generating the final feature map may include: applying one or both of a second rule and a third rule to the rearranged input data; and performing a convolution operation on the rearranged input data to which one or both of the second rule and the third rule have been applied and additional data.

[0017] In another general aspect, a non-transitory computer-readable recording medium having recorded thereon a program for performing the method on a computer.

[0018] In another general aspect, an apparatus for identifying an image includes: a memory in which at least one program is stored; and a processor configured to execute the at least one program, wherein the processor is configured to: obtain an input image; identify the sparsity of the input image as input data based on valid information included in the input image; rearrange the input image based on the form of the sparsity of the input image; generate a final feature map by performing a convolution operation on the rearranged input image on at least one convolutional layer; and obtain an identification result based on the final feature map.

[0019] The processor may also be configured to generate a final feature map by the following steps: identify the sparsity of the input data of each convolutional layer, rearrange the input data based on the form of the sparsity, and generate an output feature map by performing a convolution operation on the rearranged input data on each convolutional layer, wherein the input data of the first convolutional layer in the at least one convolutional layer is the rearranged input image, and the output feature map output by the last convolutional layer in the at least one convolutional layer is the final feature map.

[0020] In another general aspect, a device for identifying an image includes: one or more memories storing one or more programs; and one or more processors configured to execute at least one of the one or more programs to: obtain an input image, determine positions including invalid values in the input image as input data, generate rearranged data by operating on the positions including invalid values in the input data, apply a rule to the rearranged data, generate a final feature map by performing a convolution operation on the rearranged data to which the rule has been applied on at least one convolutional layer, and obtain an identification result based on the final feature map.

[0021] The one or more processors may execute at least one of the one or more programs to generate rearranged data by shifting valid values included in the input data to positions including invalid values in the input data.

[0022] The one or more processors may execute at least one of the one or more programs to generate rearranged data by shifting invalid values to other positions in the input data.

[0023] The one or more processors may execute at least one of the one or more programs to generate rearranged data by removing invalid values from the input data.

[0024] The one or more processors may execute at least one of the one or more programs to apply a rule to valid values included in a window of the rearranged data to minimize the total number of invalid values included in the input layer of the window to be input to a logic circuit.

[0025] The rule may include: shifting at least one valid value included in a layer adjacent to the input layer in a window of the rearranged data to a corresponding position including an invalid value in the input layer.

[0026] The rule may include: shifting at least one valid value included in a layer adjacent to the input layer in a window of the rearranged data to a horizontal position including an invalid value in the input layer.

[0027] In another general aspect, a method for processing data includes: identifying sparsity of input data based on valid information included in the input data; rearranging the input data based on the form of the sparsity; and generating output data by processing the rearranged input data.

[0028] In another general aspect, a device for processing data includes: a memory in which at least one program is stored; and a processor configured to execute the at least one program to: identify sparsity of input data based on valid information included in the input data; rearrange the input data based on the form of the sparsity; and generate output data by processing the rearranged input data.

[0029] In another general aspect, a device for processing data includes: one or more memories storing one or more programs; and one or more processors configured to execute at least one of the one or more programs to: determine positions including invalid values in input data, generate rearranged data by operating on the positions including invalid values in the input data, and apply a rule to the rearranged data.

[0030] Other features and aspects will be apparent from the following detailed description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a diagram showing the architecture of a neural network.

[0032] Figure 2 and Figure 3 is a diagram showing an example of a convolution operation in a neural network.

[0033] Figure 4 is a block diagram of an example of a device for processing data.

[0034] Figure 5 is a flowchart of an example of a method for processing data.

[0035] Figure 6A and Figure 6B is a diagram showing an example of a processor identifying sparsity of input data.

[0036] Figure 7 is a diagram showing an example of a processor rearranging input data.

[0037] Figure 8 is a diagram showing an example of a processor rearranging input data.

[0038] Figure 9 is a diagram showing an example of a processor rearranging input data.

[0039] Figure 10 is a diagram showing an example in which a processor rearranges input data.

[0040] Figure 11 is a flowchart showing an example in which a processor generates output data by processing the rearranged data.

[0041] Figure 12 is a diagram for describing an example in which a processor applies a second rule to the rearranged data; and

[0042] Figure 13 is a diagram showing an example in which a processor applies a third rule to the rearranged data.

[0043] Throughout the drawings and the detailed description, unless otherwise described or provided, the same reference numerals will be understood to represent the same elements, features, and structures. The drawings may not be to scale, and for clarity, illustration, and convenience, the relative sizes, proportions, and depictions of the elements in the drawings may be exaggerated. Detailed Description

[0044] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a particular order. Additionally, descriptions of features known after understanding the disclosure of this application may be omitted for increased clarity and conciseness.

[0045] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein that will be apparent after understanding the disclosure of this application.

[0046] Throughout the specification, when an element is described as being "connected to" or "coupled to" another element, the element can be directly "connected to" or "coupled to" the other element, or there can be one or more other elements intervening between them. In contrast, when an element is described as being "directly connected to" or "directly coupled to" another element, there can be no other elements intervening between them. Similarly, like expressions (e.g., "between" versus "immediately between" and "adjacent to" versus "immediately adjacent to") should be interpreted in the same manner. As used herein, the term "and / or" includes any one of the associated listed items or any combination of any two or more of the associated listed items.

[0047] Although terms such as "first", "second", and "third" may be used herein to describe various members, components, regions, layers, or sections, these members, components, regions, layers, or sections are not limited by these terms. Rather, these terms are only used to distinguish one member, component, region, layer, or section from another member, component, region, layer, or section. Thus, a first member, first component, first region, first layer, or first section as referred to in the examples described herein may also be referred to as a second member, second component, second region, second layer, or second section without departing from the teachings of the examples.

[0048] The terms used herein are for the purpose of describing various examples only and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of stated features, quantities, operations, members, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, members, elements, and / or combinations thereof.

[0049] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains based on the disclosure of this application and the present disclosure. Unless explicitly defined herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the disclosure of this application, and should not be interpreted in an idealized or overly formal manner. The use of the term "can" with respect to an example or embodiment herein (e.g., with respect to what an example or embodiment can include or achieve) means that there is at least one example or embodiment that includes or achieves such a feature, and all examples are not limited thereto.

[0050] Hereinafter, examples will be described in detail with reference to the accompanying drawings.

[0051] Figure 1 is a diagram showing the architecture of a neural network.

[0052] Reference Figure 1 , neural network 1 can be an architecture of a deep neural network (DNN) or an n-layer neural network. The DNN or the n-layer neural network can correspond to a convolutional neural network (CNN), a recurrent neural network (RNN), a deep belief network, or a restricted Boltzmann machine. For example, neural network 1 can be a CNN, but is not limited thereto. In Figure 1 , some convolutional layers of the CNN corresponding to the example of neural network 1 are shown, but in addition to the shown convolutional layers, the CNN can also include pooling layers or fully connected layers.

[0053] Neural network 1 can be implemented as an architecture having multiple layers including an input image, feature maps, and an output. In neural network 1, a convolution operation is performed on the input image using a filter called a kernel, and as a result, a feature map is output. The convolution operation is performed again on the feature map that is the output of the input feature map using the kernel, and a new feature map is output. When the convolution operation is repeatedly performed in this way, the recognition result of the features of the input image can be finally output through neural network 1.

[0054] For example, when an input image having a size of 24×24 pixels is input to Figure 1 neural network 1, the input image can be output as a feature map of four channels, each having a size of 20×20 pixels (this is simply referred to as 4@20×20) through the convolution operation with the kernel. Then, the size of the 20×20 feature map can be reduced by repeatedly performing the convolution operation with the kernel (for example, 4@20×20 can be sequentially reduced to 4@10×10 (that is, a feature map of four channels, each having a size of 10×10 pixels), 8@8×8 (that is, a feature map of eight channels, each having a size of 8×8 pixels), 8@4×4 (that is, a feature map of eight channels, each having a size of 4×4 pixels), and finally, a feature map of 20 channels, each having a size of 1×1 pixel (that is, 20@1×1) can be output. In neural network 1, the convolution operation and the downsampling (or pooling) operation can be repeatedly performed in some layers to filter and output the robust features that can represent the entire input image from the input image, and the recognition result of the input image can be obtained through the finally output features.

[0055] Figure 2 and Figure 3 are diagrams showing examples of convolution operations in a neural network.

[0056] Reference Figure 2, the input feature map 210 has a size of 6×6 pixels, the kernel 220 has a size of 3×3 pixels, and the output feature map 230 has a size of 4×4 pixels, but the sizes are not limited thereto. The neural network may include feature maps and kernels of various sizes. In addition, the values defined in the input feature map 210, the kernel 220, and the output feature map 230 are only examples and are not limited thereto.

[0057] The kernel 220 performs a convolution operation while sliding over the input feature map 210 in units of regions (or blocks) having a size of 3×3 pixels, where Figure 2 The symbol * in represents a convolution operation. The convolution operation represents an operation in which all values obtained by multiplying each pixel value in an arbitrary region of the input feature map 210 by the weights of the corresponding elements of the kernel 220 (represented by the symbol are added (represented by the symbol ) to obtain each pixel value of the output feature map 230.

[0058] The kernel 220 may first perform a convolution operation on the first region 211 of the input feature map 210. In other words, the pixel values 1, 2, 3, 4, 5, 6, 7, 8, and 9 of the first region 211 are respectively multiplied by the weights -1, -3, +4, +7, -2, -1, -5, +3, and +1 of the elements of the kernel 220. As a result, the values -1, -6, 12, 28, -10, -6, -35, 24, and 9 are obtained. Then, the values -1, -6, 12, 28, -10, -6, -35, 24, and 9 are added to obtain the value 15. Therefore, the pixel value 231 in the first row and first column of the output feature map 230 is determined to be the value 15. Here, the pixel value 231 in the first row and first column of the output feature map 230 corresponds to the first region 211.

[0059] Similarly, a convolution operation is performed between the second region 212 of the input feature map 210 and the kernel 220. Therefore, the pixel value 232 in the first row and second column of the output feature map 230 is determined to be 4. Finally, a convolution operation is performed between the sixteenth region 213 (i.e., the last window) of the input feature map 210 and the kernel 220. Therefore, the pixel value 233 in the fourth row and fourth column of the output feature map 230 is determined to be 11.

[0060] The two-dimensional (2D) convolution operation has been described with reference to Figure 2 However, optionally, the convolution operation may correspond to a three-dimensional (3D) convolution operation, where, as will be described with reference to Figure 3 There are input feature maps, kernels, and output feature maps with multiple channels.

[0061] Refer to Figure 3, the input feature map 201 may have a 3D size. There are X input channels in the input feature map 201, and the 2D input feature map of each input channel may have a size of H rows and W columns, where X, W, and H are each positive integers. The kernel 202 may have a 4D size, and there may be as many 2D kernels as there are X input channels and Y output channels (specifically, there may be as many 2D kernels as the product of X and Y). Each 2D kernel has a size of R rows and S columns, where R, S, and Y are each positive integers. In other words, the kernel 202 may have a number of channels corresponding to the number X of input channels of the input feature map 201 and the number Y of output channels of the output feature map 203. Among them, the 2D kernel of each channel may have a size of R rows and S columns. The output feature map 203 including multiple output activations may be generated through a 3D convolution operation between the 3D input feature map 201 and the 4D kernel 202, and there may be Y channels based on the result of the 3D convolution operation.

[0062] The process of generating an output feature map through the convolution operation between a 2D input feature map and a 2D kernel is as described above with reference to Figure 2 that, and the 2D convolution operation described in Figure 2 is repeatedly executed between the input feature map 201 with X input channels and the kernel 202 with Y output channels to generate the output feature map 203 with Y output channels.

[0063] Figure 4 is a block diagram of an example of a device for processing data.

[0064] Referring to Figure 4 , the device 400 for processing data may include a memory 410 and a processor 420. Although not shown in Figure 4 , the device 400 for processing data may be connected to an external memory. Figure 4 The device 400 for processing data shown in Figure 4 may include components associated with the current example. Therefore, it will be clear to those of ordinary skill in the art that the device 400 for processing data may also include other general components in addition to the

[0065] The device 400 for processing data may be the one referred to in Figures 1 to 3Devices implementing the above neural network. For example, a device 400 for processing data can be implemented using various types of devices such as a personal computer (PC), a server device, a mobile device, an embedded device, etc. Specifically, the device 400 for processing data can be included in a smart phone, a tablet device, an augmented reality (AR) device, an Internet of Things (IoT) device, an autonomous driving vehicle, a robotic device, or a medical device that performs speech recognition, image recognition, and image classification using a neural network, but is not limited thereto. The device 400 for processing data can correspond to a dedicated hardware (HW) accelerator installed on such a device and can be an HW accelerator (such as a neural processing unit (NPU), a tensor processing unit (TPU), or a neural engine) that is a dedicated module for driving a neural network.

[0066] The memory 410 stores various data processed in the device 400 for processing data. For example, the memory 410 can store data that has been processed or will be processed in the device 400 for processing data. In addition, the memory 420 can store an application or a driver to be driven by the device 400 for processing data.

[0067] For example, the memory 410 can include a random access memory (RAM) (such as a dynamic random access memory (DRAM) or a static random access memory (SRAM)), a read-only memory (RAM), an electrically erasable programmable read-only memory (EEPROM), a CD-ROM, a Blu-ray disc, an optical disc storage device, a hard disk drive (HDD), a solid state drive (SSD), or a flash memory.

[0068] The processor 420 can control the overall functions for driving the neural network in the device 400 for processing data. For example, the processor 420 can generally control the device 400 for processing data by executing a program stored in the memory 410. The processor 420 can be implemented as a central processing unit (CPU), a graphics processing unit (GPU), or an application processor (AP) included in the device 400 for processing data, but is not limited thereto.

[0069] The processor 420 can read data (such as image data, feature map data, or kernel data) from the memory 410 or write data (such as image data, feature map data, or kernel data) to the memory 410 and execute a neural network by using the read data / written data. When the neural network is executed, the processor 420 can drive a processing unit provided therein to repeatedly perform a convolution operation between an input feature map and a kernel, thereby generating data related to an output feature map. Here, the total number of operations of the convolution operation can be determined based on various factors such as the number of channels of the input feature map, the number of channels of the kernel, the size of the input feature map, the size of the kernel, and the precision of values.

[0070] For example, the processing unit may include logic circuitry for performing convolution operations. That is, the processing unit may include an arithmetic unit implemented using a combination of a multiplier, an adder, and an accumulator. The multiplier may include a combination of multiple sub-multipliers, and the adder may also include a combination of multiple sub-adders.

[0071] The processor 420 may further include an on-chip memory and a dispatcher. The on-chip memory manages the cache function for processing convolution operations, and the dispatcher allocates various operands (such as pixel values of the input feature map and weights of the kernel). For example, the dispatcher may allocate the operands (such as pixel values and weight values) required for the arithmetic operations performed by the processing unit from the data stored in the memory 410 to the on-chip memory. Then, the dispatcher may re-allocate the operands allocated to the on-chip memory to the processing unit for convolution operations.

[0072] The processor 420 performs convolution operations between the input feature map data and the kernel data. Therefore, when the data being operated on includes invalid information, the operation will be an unnecessary operation. For example, when the data being operated on is 0, the output of the convolution operation between the data is 0, so this unnecessary operation only increases the computational load of the processor 420.

[0073] Meanwhile, the input feature map data and the kernel data can be represented as matrices of M rows and N columns, where M and N are positive integers. That is, the input feature map matrix and the kernel matrix may include multiple elements, and among the multiple elements, the number of elements including 0 is proportional to the number of unnecessary operations.

[0074] The device 400 for processing data may rearrange the input data based on the valid information (such as data other than 0) included in the input data (for example, the input feature map data and the kernel data). Here, the rearrangement of the input data may represent an operation of changing the original architecture of the matrix (such as changing the positions of some elements included in the matrix or skipping some rows or columns included in the matrix). In some examples, the rearrangement of the input data may represent an operation of shifting the invalid information to another position in the input data or removing the invalid information from the input data.

[0075] Therefore, the device 400 for processing data can output valid results without performing unnecessary operations, thereby reducing the total amount of computation while outputting the desired results.

[0076] Hereinafter, reference will be made to Figures 5 to 13 Examples of the apparatus 400 for processing data rearranging the input data and processing the rearranged data to generate output data will be described.

[0077] Figure 5 is a flowchart showing an example of a method for processing data.

[0078] Referring to Figure 5 , the method for processing data may include operations performed in time series by the device 400 for processing data shown in Figure 4 . Therefore, it can be seen that the content described above regarding the device 400 for processing data shown in Figure 4 (although omitted below) also applies to the method for processing data shown in Figure 5 .

[0079] In operation 510, the processor 420 may identify the sparsity of the input data based on the valid information included in the input data.

[0080] The input data may represent an object on which the processor 420 will perform a convolution operation. For example, the input data may include image data, feature map data, or kernel data. The feature map data may be input feature map data or output feature map data. The processor 420 may perform convolution operations in multiple layers, and the output feature map data in the previous layer may be the input feature map data in the next layer. Therefore, the input data for operation 510 may be input feature map data or output feature map data. As described in detail with reference to Figure 4 , the input data may be a matrix including elements as data.

[0081] The valid information may represent data on which a meaningful convolution operation can be performed. Generally, information can be represented as a number, so the valid information may represent data as non-zero numbers. In other words, data of meaningless information may be represented as 0.

[0082] The processor 420 may identify the sparsity of the input data. Here, the sparsity may represent the presence or absence of blanks in the data, or the state of whether the data includes blanks. As described above, the valid information may be represented as data as non-zero numbers. Therefore, zero data may represent meaningless information that can be interpreted as blank data (i.e., no data). Therefore, when the processor 420 identifies the sparsity of the input data, it may represent that the processor 420 identifies the distribution of 0s in the input data.

[0083] Hereinafter, examples of the processor 420 identifying the sparsity of the input data will be described with reference to Figure 6A and Figure 6B .

[0084] Figure 6A and Figure 6B are diagrams showing examples of the processor identifying the sparsity of the input data.

[0085] Figure 6A and Figure 6B schematically shows the convolution operation performed by the processor 420. The processor 420 can generate output data by performing a convolution operation among the input data 610, 620, 630, and 640. For example, the input data 610, 620, 630, and 640 can be represented as matrices, and the processor 420 can generate output data by performing a sum-of-product calculation among the elements of the channels included in the matrices.

[0086] In Figure 6A input feature map data 610 and kernel data 620 are shown as input data, and in Figure 6B input feature map data 630 and kernel data 640 are shown. Hereinafter, for convenience, the elements included in the input feature map data 610 and 630 will be referred to as activations (a), and the elements included in the kernel data 620 and 640 will be referred to as weights (w).

[0087] When the kernel data 620 is compared with the kernel data 640, blanks are included in a part of the kernel data 640. Here, the blanks can be interpreted as weights 0. That is, the kernel data 640 can be sparser than the kernel data 620, which can indicate that more weights included in the kernel data 640 have 0 compared to the weights included in the kernel data 620.

[0088] Meanwhile, in Figure 6A and Figure 6B it is shown that 0 is included in the kernel data 640, but the disclosure is not limited thereto. In other words, 0 can be included in at least one of the input data 610, 620, 630, and 640, and the number and the form of the distribution of 0 in the input data 610, 620, 630, and 640 can vary.

[0089] The processor 420 can identify the sparsity of the input data 610, 620, 630, and 640 based on the valid information (e.g., non-zero numbers) included in the input data 610, 620, 630, and 640. In other words, the processor 420 can identify the distribution of 0 in the input data 610, 620, 630, and 640.

[0090] Returning to reference Figure 5 in operation 520, the processor 420 can rearrange the input data based on the form of the sparsity of the input data.

[0091] The processor 420 may rearrange the input data based on the distribution of 0s in the input data. For example, the processor 420 may rearrange a plurality of rows based on the number of 0s included in each of the plurality of rows included in the input data. In another example, the processor 420 may shift the elements of each of the plurality of columns of the input data according to a first rule. In another example, the processor 420 may rearrange the plurality of columns to skip processing of at least one column that includes only 0s among the plurality of columns of the input data. In another example, the processor 420 may shift the first element of the first column of the input data to the position corresponding to the last element of the second column adjacent to the first column.

[0092] Reference will be made to Figures 7 to 10 examples in which the processor 420 rearranges the input data will be described.

[0093] Figure 7 is a diagram showing an example in which a processor rearranges input data.

[0094] In Figure 7 the input feature map data 710 and the kernel data 720 are shown as the input data. The input feature map data 710 is shown as a 6-row 6-column matrix, and the kernel data 720 is shown as a 6-row 4-column matrix, but the configuration is not limited thereto.

[0095] A part of the input feature map data 710 may include blanks. Here, the blanks may be interpreted as the absence of valid information. For example, the activation corresponding to the blanks may be equivalent to 0. In Figure 7 it is shown that blanks are included in the input feature map data 710, but the configuration is not limited thereto. That is, 0s may also be included in at least one of the weights (e.g., weights a 0 ,..., a 5 , b 0 ,..., b 5 , c 0 ,..., c 5 , d 0 ,..., d 5 ) included in the kernel data 720.

[0096] The processor 420 may rearrange the input feature map data 710 based on the form of the sparsity of the input feature map data 710. For example, the processor 420 may rearrange the plurality of rows row 0 to row 5 based on the number of blanks included in each of the plurality of rows row 0 to row 5 included in the input feature map data 710.

[0097] More specifically, with reference to the input feature map data 710 and the feature map data 711, the processor 420 can perform rearrangement such that one of the rows with the most blanks among the multiple rows row 0 to row 5 (e.g., row 2 among row 2 and row 5) and row 0 with the fewest blanks are adjacent to each other. The processor 420 can also perform rearrangement such that row 4 with the second most blanks and row 3 with the second fewest blanks among the multiple rows row 0 to row 5 are adjacent to each other. In this way, the processor 420 can generate the feature map data 711 by rearranging the multiple rows row 0 to row 5 of the input feature map data 710 based on the number of blanks included.

[0098] Using the feature map data 711 generated by rearrangement, the processor 420 can minimize the execution of unnecessary operations. For example, for the convolution operation with the kernel data 720, the processor 420 can input the feature map data 711 into the logic circuit 730 for each part. The processor 420 can input the activations included in the window 740 of the input feature map data 711 into the logic circuit 730.

[0099] The processor 420 can also input the weights included in the window 750 into the logic circuit 730 by applying a window 750 having the same size as the size of the window 740 to the kernel data 720. The processor 420 can rearrange the kernel data 720 to correspond to the feature map data 711. The order of the activations input into the logic circuit 730 in the feature map data 711 and the order of the activations input into the logic circuit 730 in the input feature map data 710 are different from each other. Therefore, when the weights are input into the logic circuit 730 without rearranging the kernel data 720, inaccurate operation results may be output.

[0100] The processor 420 can rearrange the kernel data 720 such that the weights to be calculated together with the activations input into the logic circuit 730 are accurately input into the logic circuit 730. The processor 420 can input the weights into the logic circuit 730 according to the rearranged kernel data 720. Therefore, even with the feature map data 711, accurate operation results can be output from the logic circuit 730.

[0101] When the kernel data 720 is rearranged, the processor 420 can rearrange the input feature map data 710 in the same manner as described above and input the rearranged input feature map data 710 into the logic circuit 730.

[0102] The processor 420 can prevent the execution of unnecessary convolution operations by adjusting the positions of the activations included in the window 740. It will be referred to Figures 11 to 13An example is described in which the processor 420 performs a convolution operation by adjusting the positions of the activations included in the window 740.

[0103] Figure 8 FIG. is a diagram showing another example in which the processor rearranges the input data.

[0104] In Figure 8 input feature map data 810 and kernel data 820 are shown as input data. A portion of the input feature map data 810 may include blanks. In Figure 8 it is shown that blanks are included in the input feature map data 810, but the configuration is not limited thereto. That is, 0 may also be included in at least one of the weights included in the kernel data 820.

[0105] The processor 420 may rearrange the input feature map data 810 based on the form of the sparsity of the input feature map data 810. For example, the processor 420 may shift the elements of each of the multiple columns col 0 to col 5 included in the input feature map data 810 according to a first rule.

[0106] The first rule may shift the elements of each of the multiple columns col 0 to col 5 in the same direction by a specific size. Here, the specific size may be adaptively changed by the processor 420 based on the form of the sparsity of the input feature map data 810, and the sizes applied to each of the multiple columns col 0 to col 5 may be different. For example, referring to the input feature map data 810 and the feature map data 811 generated by rearrangement, the processor 420 may generate the second column col 1 of the feature map data 811 by shifting the activation included in the second column col 1 of the input feature map data 810 by one grid. The processor 420 may generate the fifth column col 4 of the feature map data 811 by shifting the activation included in the fifth column col 4 of the input feature map data 810 by two grids. According to the form of the sparsity of the input feature map data 810, the processor 420 may not shift the activations of the other columns col 0, col 2, col 3, and col 5 of the input feature map data 810.

[0107] The first rule may be periodically applied to the multiple columns col 0 to col 5. As Figure 8 shown, the processor 420 may periodically apply a shift rule of "0-1-0-0-2-0" to the next input feature map data of the input feature map data 810. For example, the period may be, but is not limited to, the same as the size of the kernel data 820 (e.g., the number of columns of the kernel data 820). Through this process, the processor 420 can prevent unnecessary convolution operations from being performed.

[0108] The processor 420 may rearrange the kernel data 820 to correspond to the feature map data 811. For example, the processor 420 may rearrange the kernel data 820 such that the weights to be calculated together with the activations input to the logic circuit 730 are accurately input to the logic circuit. The processor 420 may input the weights to the logic circuit according to the rearranged kernel data. Therefore, even with the feature map data 811, accurate operation results can be output from the logic circuit.

[0109] When the kernel data 820 is rearranged, the processor 420 may rearrange the input feature map data 810 in the same manner as described above, and input the rearranged input feature map data 811 to the logic circuit 730.

[0110] will be referred to Figures 11 to 13 to describe an example in which the processor 420 generates output data by processing the feature map data 811 and the kernel data 820.

[0111] Figure 9 is a diagram for showing another example in which a processor rearranges input data.

[0112] In Figure 9 the input feature map data 910 and the kernel data 920 are shown as input data. A part of the input feature map data 910 may include blanks. In Figure 9 it is shown that blanks are included in the input feature map data 910, but the configuration is not limited thereto. That is, 0 may also be included in at least one of the weights (e.g., weights f 0 ,..., f 5 , e 0 ,..., e 5 ) included in the kernel data 920.

[0113] The processor 420 may rearrange the input feature map data 910 based on the form of the sparsity of the input feature map data 910. For example, the processor 420 may shift the first element (activation) of column col 1 included in the input feature map data 910 to the position corresponding to the last element (activation) of column col 0 adjacent to column col 1.

[0114] More specifically, valid information is included in the first positions of column col 1 and column col 0. No valid information is included in the last position of column col 0. In this case, the processor 420 may shift the elements in the first position of column col 1 to the last position of column col 0. Through this process, the processor 420 can prevent unnecessary convolution operations from being performed. Similarly, the processor 420 may shift the elements in the second position of column col 1 to the third position of column col 0, and may shift the elements in the fifth position of column col 1 to the fifth position of column col 0.

[0115] As referred to above Figure 7 and Figure 8 stated, when the input feature map data 910 is rearranged, the kernel data 920 may also be rearranged.

[0116] Figure 10 is a diagram showing another example of the processor rearranging the input data.

[0117] In Figure 10 is shown the input feature map data 1010. A portion of the input feature map data 1010 may include blanks. Specifically, some of the columns col 1 to col 3 of the input feature map data 1010 may include only blanks.

[0118] The processor 420 may rearrange the input feature map data 1010 based on the form of the sparsity of the input feature map data 1010. For example, the processor 420 may rearrange the input feature map data 1010 to skip processing for the columns col 1 to col 3 that include only blanks among the multiple columns col 0 to col 5 included in the input feature map data 1010.

[0119] For example, the processor 420 may omit columns col 1 to col 3 from the input feature map data 1010 and generate the feature map data 1020 using only the other columns col 0, col 4, and col 5. The processor 420 may record the omission of columns col 1 to col 3 in the memory 410. Through this process, the processor 420 can prevent unnecessary convolution operations from being performed.

[0120] As referred to above Figure 7 and Figure 8 stated, when the input feature map data 1010 is rearranged, the kernel data may also be rearranged.

[0121] Returning to Figure 5 , in operation 530, the processor 420 may generate output data by processing the rearranged input data.

[0122] For example, the processor 420 can generate output data by performing a convolution operation using the rearranged input data. However, the processor 420 can alternatively apply a second rule or a third rule to the rearranged data of operation 520 to reduce unnecessary operations.

[0123] Hereinafter, reference will be made to Figures 11 to 13 describe an example in which the processor 420 generates output data.

[0124] Figure 11 is a flowchart showing an example in which a processor generates output data by processing rearranged data.

[0125] In operation 1110, the processor 420 can apply at least one of the second rule and the third rule to the rearranged data.

[0126] As referred to above Figure 7 described, the processor 420 can sequentially input the rearranged data into the logic circuit. For example, the processor 420 can apply a window of a specific size to the rearranged data and input the elements included in the window into the logic circuit. When some of the elements included in the window include invalid information (e.g., 0 or blank), the processor 420 can rearrange the elements included in the window by applying the second rule or the third rule.

[0127] In operation 1120, the processor 420 can perform a convolution operation on the rearranged data to which at least one rule has been applied and additional data. For example, the processor 420 can perform a convolution operation by inputting the rearranged activation or the rearranged weights into the logic circuit.

[0128] Hereinafter, reference will be made to Figure 12 describe an example in which the processor 420 applies the second rule to the rearranged data, and reference will be made to Figure 13 describe an example in which the processor 420 applies the third rule to the rearranged data.

[0129] Figure 12 is a diagram showing an example in which a processor applies a second rule to rearranged data.

[0130] In Figure 12 feature map data 1210 and kernel data 1220 are shown. Hereinafter, it is assumed that the feature map data 1210 is the rearranged data of operation 520.

[0131] The processor 420 may input a portion of the feature map data 1210 into the logic circuit 1230. For example, the processor 420 may input the activations included in the window 1240 of the feature map data 1210 into the logic circuit 1230. The processor 420 may input the maximum number of activations into the logic circuit 1230 by applying a second rule to the activations included in the window 1240. That is, the processor 420 may apply the second rule to the activations included in the window 1240 to minimize the blanks in the input layer 1231 of the logic circuit 1230 (e.g., the content of the input layer of the logic circuit is from the columns of the window to be input into the logic circuit). Here, the second rule may represent a rule that shifts the activations of a column (e.g., col 1) to the same position in an adjacent column (e.g., col0).

[0132] For example, the processor 420 may identify the blanks in columns col 0 and col 1 in the window 1240 and assign the activations of column col 1 to the blanks in column col 0. Refer to Figure 12 , the activations 2 and 4 of column col 1 may be shifted to the same position in column col 0.

[0133] The processor 420 may input the activations to which the second rule has been applied into the input layer 1231 of the logic circuit 1230. Compared with column col0, the number of blanks in the input layer 1231 may be less than the number of blanks in column col 0. A blank has the same effect as including data 0, so the output is 0 regardless of the value of the weight corresponding to the blank. Therefore, as the number of blanks included in the input layer 1231 increases (i.e., the number of 0s included in the input layer 1231 increases), the number of unnecessary operations increases.

[0134] As described above, the processor 420 may minimize the number of blanks included in the input layer 1231 by applying the second rule. Therefore, the processor 420 may minimize the number of times the logic circuit 1230 performs unnecessary operations.

[0135] Figure 13 is a diagram showing an example in which the processor applies a third rule to the rearranged data.

[0136] In Figure 13 the feature map data 1310 and the kernel data 1320 are shown. Hereinafter, it is assumed that the feature map data 1310 is the rearranged data of the operation 520.

[0137] The processor 420 can input the most activations into the logic circuit 1330 by applying a third rule to the activations included in the window 1340. Here, the third rule can represent such a rule: shifting the activations of a column (e.g., col 1) to the horizontal position of an adjacent column (e.g., col 0), where the horizontal position represents a position that is spaced a specific size from the same position of the adjacent column.

[0138] For example, the processor 420 can identify the blanks in columns col 0 and col 1 in the window 1340 and assign the activations of column col 1 to the blanks in column col 0. Referring to Figure 13 , the activations 0, 1, and 3 of column col 1 can be shifted to the horizontal positions of column col 0.

[0139] The processor 420 can input the activations to which the third rule has been applied into the input layer 1331 of the logic circuit 1330. Compared with the input layer 1331, there are blanks in column col 0 (more specifically, there are three blanks), but there are no blanks in the input layer 1331. Therefore, the processor 420 can minimize the number of unnecessary operations performed by the logic circuit 1230.

[0140] As described in reference to Figure 12 and Figure 13 in detail, the processor 420 can apply the second rule and the third rule separately, but the configuration is not limited to this. The processor 420 can identify the sparsity of the feature map data 1210 and 1310 and the kernel data 1220 and 1320, and adaptively apply at least one of the second rule and the third rule to the feature map data 1210 and 1310 and / or the kernel data 1220 and 1320.

[0141] As described in detail, the device 400 for processing data can rearrange the input feature map data and / or kernel data to minimize the number of blanks input to the logic circuit in which the convolution operation is performed. Therefore, the device 400 for processing data can minimize the number of unnecessary operations performed.

[0142] Meanwhile, the foregoing method can be written as a program executable on a computer and can be implemented on a general-purpose digital computer operating the program by using a computer-readable recording medium. The structure of the data used in the above method can be recorded on the computer-readable recording medium using various devices. The computer-readable recording medium can include storage media such as magnetic storage media (e.g., ROM, RAM, universal serial bus (USB), floppy disk, hard disk, etc.), optical recording media (e.g., compact disc (CD)-ROM, digital versatile disc (DVD), etc.), etc.

[0143] While the present disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail may be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein should be considered only illustrative and not for purposes of limitation. The description of a feature or aspect in each example is considered applicable to a similar feature or aspect in other examples. Appropriate results may be achieved if the described techniques are performed in a different order, and / or if the components in the described system, architecture, apparatus, or circuit are combined in a different manner, and / or are replaced or supplemented by other components or their equivalents. Accordingly, the scope of the disclosure is defined not by the specific embodiments, but by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be construed as being included in the disclosure.

Claims

1. A method for identifying an image, the method comprises: obtaining an input image; obtaining input data by inputting the input image into a neural network configured to identify the input image, wherein the input data includes at least one of image data, feature map data, and kernel data corresponding to the input image; identifying the sparsity of the input data based on the distribution of zero-valued elements included in the input data; rearranging the input data based on the form of the sparsity of the input data such that the original architecture of the matrix of the input data is changed, wherein the change in the original architecture of the matrix of the input data indicates a change in the positions of the elements included in the matrix or a skip of rows or columns included in the matrix; generating a final feature map based on the rearranged input data; and obtaining an identification result based on the final feature map.

2. The method according to claim 1, the step of generating the final feature map comprises: generating an output feature map by performing a convolution operation on the rearranged input data on each convolution layer in at least one convolution layer of the neural network, wherein the input data of the first convolution layer in the at least one convolution layer is the rearranged input image, and the output feature map output by the last convolution layer in the at least one convolution layer is the final feature map.

3. The method according to claim 1, wherein the step of rearranging the input data comprises: rearranging the plurality of rows included in the input data based on the number of zero-valued elements in each of the plurality of rows included in the input data.

4. The method according to claim 3, wherein the step of rearranging the input data comprises: performing the rearrangement such that the row of the input data among the plurality of rows of the input data that includes the most zero-valued elements is adjacent to the row of the input data among the plurality of rows of the input data that includes the fewest zero-valued elements.

5. The method according to claim 1, wherein the step of rearranging the input data comprises: shifting the elements of a plurality of columns included in the input data according to a first rule.

6. The method according to claim 5, wherein the first rule includes shifting the elements of the plurality of columns included in the input data in the same direction by a specific size, and the first rule is periodically applied to the plurality of columns included in the input data.

7. The method according to claim 1, wherein the step of rearranging the input data comprises: rearranging the plurality of columns included in the input data to skip the processing of at least one column that includes only zero-valued elements among the plurality of columns included in the input data.

8. The method according to claim 1, wherein the step of rearranging the input data comprises: shifting the non-zero values of the first column included in the input data to positions corresponding to the zero values of the second column adjacent to the first column of the input data.

9. The method according to claim 1, wherein the step of generating the final feature map comprises: applying one or both of a second rule and a third rule to the rearranged input data; and performing a convolution operation on the rearranged input data to which one or both of the second rule and the third rule are applied and additional data.

10. A non-transitory computer-readable recording medium having recorded thereon a program for executing the method according to any one of claims 1 to 9 on a computer.

11. An apparatus for identifying an image, the apparatus comprising: a memory in which at least one program is stored; and a processor configured to execute the at least one program to: obtain an input image; obtain input data by inputting the input image into a neural network configured to identify the input image, wherein the input data includes at least one of image data, feature map data, and kernel data corresponding to the input image; identify the sparsity of the input data based on the distribution of zero-valued elements included in the input data; rearrange the input data based on the form of the sparsity of the input data such that the original architecture of the matrix of the input data is changed, wherein the change in the original architecture of the matrix of the input data indicates a change in the positions of the elements included in the matrix or a skip of rows or columns included in the matrix; generate a final feature map based on the rearranged input data; and obtain an identification result based on the final feature map.

12. The apparatus according to claim 11, wherein the processor is further configured to: execute the at least one program to generate a final feature map by the following steps: generate an output feature map by performing a convolution operation on the rearranged input data on each convolution layer of at least one convolution layer of the neural network, wherein the input data of the first convolution layer among the at least one convolution layer is the rearranged input image, and the output feature map output by the last convolution layer among the at least one convolution layer is the final feature map.

13. The apparatus according to claim 11, wherein the processor is further configured to: execute the at least one program to rearrange the plurality of rows included in the input data based on the number of zero-valued elements in each row among the plurality of rows included in the input data.

14. The apparatus according to claim 13, wherein the processor is further configured to: execute the at least one program to perform the rearrangement such that the row of the input data among the plurality of rows of the input data that includes the most zero-valued elements is adjacent to the row of the input data among the plurality of rows of the input data that includes the fewest zero-valued elements.

15. The apparatus according to claim 11, wherein the processor is further configured to: execute the at least one program to shift elements of a plurality of columns included in the input data according to a first rule.

16. The apparatus according to claim 15, wherein the first rule includes shifting elements of the plurality of columns included in the input data in the same direction by a specific size, and the first rule is periodically applied to the plurality of columns included in the input data.

17. The apparatus according to claim 11, wherein the processor is further configured to: execute the at least one program to rearrange a plurality of columns included in the input data to skip processing of at least one column that includes only zero-valued elements among the plurality of columns included in the input data.

18. The device according to claim 11, wherein, the processor is further configured to: execute the at least one program to shift non - zero values in a first column included in the input data to positions corresponding to zero values in a second column adjacent to the first column included in the input data.

19. The device according to claim 11, wherein, the processor is further configured to: execute the at least one program to apply one or both of a second rule and a third rule to the rearranged input data, and perform a convolution operation on the rearranged input data to which one or both of the second rule and the third rule are applied and additional data.

20. The device according to claim 11, wherein, the device includes a neural network device.

21. A device for identifying an image, comprising: one or more memories storing one or more programs; and one or more processors configured to execute at least one of the one or more programs to: obtain an input image, obtain input data by inputting the input image into a neural network configured to identify the input image, wherein the input data includes at least one of image data, feature map data, and kernel data corresponding to the input image, determine positions in the input data that include zero values, generate rearranged data by operating on the positions in the input data that include zero values such that the original architecture of the matrix of the input data is changed, wherein the change in the original architecture of the matrix of the input data indicates a change in the positions of the elements included in the matrix or a skip of rows or columns included in the matrix, apply a rule to the rearranged data, generate a final feature map based on the rearranged data, and obtain an identification result based on the final feature map.

22. The device according to claim 21, wherein, the one or more processors are further configured to: execute at least one of the one or more programs to generate rearranged data by shifting non - zero values included in the input data to positions in the input data that include zero values.

23. The device according to claim 21, wherein, the one or more processors are further configured to: execute at least one of the one or more programs to generate rearranged data by shifting zero values to other positions in the input data.

24. The device according to claim 21, wherein, the one or more processors are further configured to: execute at least one of the one or more programs to generate rearranged data by removing zero values from the input data.

25. The device according to claim 21, wherein, the one or more processors are further configured to: execute at least one of the one or more programs to apply a rule to non - zero values included in a window of the rearranged data to minimize the total number of zero values in an input layer of a window to be input to a logic circuit.

26. The device according to claim 25, wherein, The rules include: shifting at least one non-zero value included in a layer adjacent to the input layer in a window of the rearranged data to a corresponding position in the input layer that corresponds to the at least one non-zero value and includes a zero value.

27. The apparatus according to claim 25, wherein, the rules include: shifting at least one non-zero value included in a layer adjacent to the input layer in a window of the rearranged data to a horizontal position in the input layer that includes a zero value.

Citation Information

Patent Citations

  • Choosing a Transmission Mode in a Dense Wireless Network

    KR1020190104578A