Artificial neural network model and electronic device including the artificial neural network model

By using an ANN model configured with encoding layer, decoding layer and jump connection in image processing, the efficiency and processing capability problems when converting high-resolution images from CFA format to RGB images are solved, and more efficient and high-quality image conversion is achieved.

CN112149793BActive Publication Date: 2025-06-10SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010428019.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-27
Filing Date
2020-05-19
Publication Date
2025-06-10
Estimated Expiration
2040-05-19

AI Technical Summary

Technical Problem

The prior art has high time and processing power costs when converting high-resolution images from color filter array (CFA) format to RGB images.

Method used

An artificial neural network (ANN) model configured to perform image processing operations is adopted, which includes a plurality of encoding layer units and decoding layer units, and the feature map is transmitted directly from the encoding layer to the decoding layer through jump connections, adjusting the depth of the output feature map to optimize the image processing process.

Benefits of technology

Through the use of ANN models, the efficiency and processing power of image conversion are significantly improved, time loss is reduced, and possible image quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112149793B_ABST
    Figure CN112149793B_ABST
Patent Text Reader

Abstract

Describes an electronic device that includes processing logic configured to receive input image data and use an artificial neural network model to generate output image data having a different format from the input image data. The artificial neural network model includes a plurality of encoding layer units, and the plurality of encoding layer units respectively include a plurality of layers located at multiple levels. The artificial neural network model further includes a plurality of decoding layer units, and the plurality of decoding layer units include a plurality of layers and are configured to form skip connections with a plurality of encoding layers at the same level. The first encoding layer unit at the first level receives a first input feature map and outputs a first output feature map based on the first input feature map to subsequent encoding layer units and the decoding layer unit at the first level.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of Korean Patent Application No. 10 - 2019 - 0077255, filed with the Korean Intellectual Property Office on Jun. 27, 2019, the disclosure of which is incorporated herein by reference in its entirety. Technical field

[0003] The present inventive concept relates to an artificial neural network (ANN) model, and more particularly, to an ANN model configured to perform image processing operations and an electronic device including the ANN model. Background art

[0004] An image sensor may include a color filter array (CFA) to separate color information from detected light. An imaging device may use the uncompleted output color samples from the image sensor in a process called demosaicking to reconstruct an image file. The Bayer pattern mosaic is an example of a CFA that organizes colors using a square grid of sensors. A red - green - blue (RGB) image is the final output image of the image sensor after filtering has occurred. Thus, an RGB image can be the product of a Bayer pattern mosaic after image processing.

[0005] An image sensor including a CFA may produce very large image files. Thus, the process of demosaicking an image from an image sensor having a CFA (e.g., a Bayer pattern) to an output format (e.g., an RGB image) may be expensive in terms of time and processing power. Accordingly, there is a need in the art for systems and methods for efficiently converting a high - resolution image in CFA format into a usable output image. Summary of the invention

[0006] The present inventive concept provides an artificial neural network (ANN) model having a new structure for changing the format of an image, and an electronic device including the ANN model.

[0007] According to an aspect of the inventive concept, there is provided an electronic device including processing logic configured to receive input image data and generate output image data having a format different from that of the input image data using an artificial neural network model. The artificial neural network model includes: a plurality of encoding layer units, each encoding layer unit including a plurality of layers and being located at one of a plurality of levels; and a plurality of decoding layer units, each decoding layer unit including a plurality of layers and being configured to form a skip connection with an encoding layer unit at the same level as the decoding layer unit. A first encoding layer unit at a first level receives a first input feature map and outputs a first output feature map based on the first input feature map to a subsequent encoding layer unit and a decoding layer unit at the first level. The subsequent encoding layer unit may be at the next level after the first level, and the first encoding layer unit at the first level and the decoding layer unit at the first level are connected through the skip connection.

[0008] According to another aspect of the inventive concept, there is provided an electronic device including processing logic configured to perform an operation using an artificial neural network model. The artificial neural network model includes: a plurality of encoding layer units, each encoding layer unit including a plurality of layers and being located at one of a plurality of levels; and a plurality of decoding layer units, each decoding layer unit including a plurality of layers and being located at one of the plurality of levels. A first encoding layer unit at a first level among the plurality of levels receives a first input feature map, outputs a first output feature map to an encoding layer unit at the next level after the first level and a decoding layer unit at the first level, and adjusts the depth of the first output feature map based on the first level.

[0009] According to another aspect of the inventive concept, there is provided an electronic device configured to perform an image processing operation. The electronic device includes processing logic configured to receive tetra image data from a color filter array in which four identical color filters are arranged in two rows and two columns to form a pixel unit. The processing logic uses an artificial neural network model to generate output image data having a format different from that of the tetra image data. The artificial neural network model includes: a plurality of encoding layer units, each encoding layer unit including a plurality of layers and located at one of a plurality of levels; and a plurality of decoding layer units, each decoding layer unit including a plurality of layers and configured to form a skip connection with an encoding layer unit located at the same level as the decoding layer unit. A first encoding layer unit at a first level receives a first input feature map and outputs a first output feature map based on the first input feature map to a subsequent encoding layer unit and a decoding layer unit at the first level, the subsequent encoding layer unit being at the next level after the first level, and the first encoding layer unit at the first level and the decoding layer unit at the first level being connected through the skip connection. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Embodiments of the inventive concept will be more clearly understood from the following detailed description in conjunction with the accompanying drawings, in which:

[0011] Figure 1 is a block diagram of an electronic device according to an exemplary embodiment;

[0012] Figure 2 is a diagram of an example of a neural network structure;

[0013] Figure 3 is a diagram for explaining image data according to an exemplary embodiment;

[0014] Figure 4 is a block diagram of an image processing device according to a comparative example;

[0015] Figure 5 is a block diagram for explaining a neural processing unit (NPU) according to an exemplary embodiment;

[0016] Figure 6 is a diagram for explaining an artificial neural network (ANN) model according to an exemplary embodiment;

[0017] Figure 7A and Figure 7B is a block diagram for explaining an encoding layer unit and a decoding layer unit according to an exemplary embodiment;

[0018] Figure 8 is a block diagram of a feature map block according to an exemplary embodiment;

[0019] Figure 9 is a diagram for illustrating an ANN model according to an exemplary embodiment;

[0020] Figure 10 shows a table of depth values of feature maps according to an exemplary embodiment;

[0021] Figure 11 is a diagram for illustrating a downsampling operation according to an exemplary embodiment;

[0022] Figure 12 is a block diagram of an application processor (AP) according to an exemplary embodiment; and

[0023] Figure 13 is a block diagram of an imaging device according to an exemplary embodiment. DETAILED DESCRIPTION

[0024] The present disclosure provides an artificial neural network (ANN) model for changing a format of an image, and an electronic device including the ANN model. An ANN refers to a computational architecture modeled after a biological brain. That is, an ANN may be a hardware or software component including a plurality of connected nodes (also referred to as artificial neurons), which may loosely correspond to neurons in a human brain. Each connection or edge may send a signal from one node to another node (like a physical synapse in the brain). When a node receives a signal, it may process the signal and then send the processed signal to other connected nodes. In some cases, signals between nodes include real numbers, and an output of each node may be calculated as a function of a sum of its inputs. Each node and edge may be associated with one or more node weights that determine how signals are processed and sent.

[0025] In a training process, these weights may be adjusted to improve accuracy of results (i.e., by minimizing a loss function that somehow corresponds to a difference between a current result and a target result). A weight of an edge may increase or decrease an intensity of a signal sent between nodes. In some cases, a node may have a threshold below which a signal is not sent at all. Nodes may also be aggregated into layers. Different layers may perform different transformations on their inputs. An initial layer may be referred to as an input layer, and a last layer may be referred to as an output layer. In some cases, a signal may cross a particular layer multiple times.

[0026] A deep learning or machine learning model may be implemented based on an ANN. As the number of operations to be processed using an ANN increases, performing operations using an ANN becomes more efficient compared to conventional alternatives.

[0027] According to an embodiment of the present disclosure, an ANN model may be configured to convert a high-resolution image file into another file format, e.g., an RGB image. The input image of the ANN model may be a high-resolution file from an image sensor having a color filter array (CFA). The output of the ANN model may be an available RGB image. Embodiments of the present disclosure use one or more convolutional layers in image encoding and decoding to output feature maps at respective layers or steps. Feature maps of the same level may be sent from an encoder to a decoder via skip connections.

[0028] Hereinafter, embodiments of the inventive concept will be described in detail with reference to the accompanying drawings.

[0029] Figure 1 is a block diagram of an electronic device 10 according to an exemplary embodiment.

[0030] According to an embodiment of the present disclosure, an electronic device 10 may analyze input data in real time based on an artificial neural network (ANN) model 100, extract valid information, and generate output data based on the extracted information. For example, the electronic device 10 may be applied to a smart phone, a mobile device, an image display device, an image capture device, an image processing device, a measuring device, a smart TV, a robotic device (e.g., a drone and an advanced driver assistance system (ADAS)), a medical device, and an Internet of Things (IoT) device. Additionally, the electronic device 10 may be installed on one of various kinds of electronic devices. For example, the electronic device 10 may include an application processor (AP). The AP may perform various kinds of operations, and a neural processing unit (NPU) 13 included in the AP may share operations to be performed by the ANN model 100.

[0031] Referring to Figure 1 , the electronic device 10 may include a central processing unit (CPU) 11, a random access memory (RAM) 12, a neural processing unit (NPU) 13, a storage device 14, and a sensor module 15. The electronic device 10 may further include an input / output (I / O) module, a security module, and a power control device, and may further include various kinds of operating devices. For example, some or all components of the electronic device 10 (i.e., the CPU 11, the RAM 12, the NPU 13, the storage device 14, and the sensor module 15) may be installed on one semiconductor chip. For example, the electronic device 10 may include a system on chip (SoC). Components of the electronic device 10 may communicate with each other via a bus 16.

[0032] The CPU 11 can control the overall operation of the electronic device 10. The CPU 11 can include one processor core (or a single core) or multiple processor cores (or multi-cores). The CPU 11 can process or execute programs and / or data stored in the storage device 14. For example, the CPU 11 can execute the programs stored in the storage device 14 and control the functions of the NPU 13.

[0033] The RAM 12 can temporarily store programs, data, or instructions. For example, the programs and / or data stored in the storage device 14 can be temporarily stored in the RAM 12 according to the control code or leading code of the CPU 11. The RAM 12 can include memories, for example, dynamic RAM (DRAM) or static RAM (SRAM).

[0034] The NPU 13 can receive input data, perform operations based on the ANN model 100, and provide output data based on the operation results. The NPU 13 can perform operations based on various types of networks, for example: convolutional neural network (CNN), region-based convolutional neural network (R-CNN), region proposal network (RPN), recurrent neural network (RNN), stacked deep neural network (S-DNN), state space dynamic neural network (S-SDNN), deconvolution network, deep belief network (DBN), restricted Boltzmann machine (RBM), fully convolutional network, long short-term memory (LSTM) network, and classification network. However, the inventive concept is not limited thereto, and the NPU 13 can perform various operations that simulate the human neural network.

[0035] Figure 2 is a diagram showing an example of a neural network structure.

[0036] Reference Figure 2 , the ANN model 100 can include multiple layers L1 to Ln. Each of the multiple layers L1 to Ln can be a linear layer or a non-linear layer. In some embodiments, a combination of at least one linear layer and at least one non-linear layer can be referred to as one layer. For example, the linear layer can include a convolutional layer and / or a fully connected layer, and the non-linear layer can include a sampling layer, a pooling layer, and / or an activation layer.

[0037] As an example, the first layer L1 can include a convolutional layer, and the second layer L2 can include a sampling layer. The ANN model 100 can include an activation layer and can also include layers configured to perform different types of operations.

[0038] Each of the multiple layers may receive input image data or a feature map generated in a previous layer as an input feature map, perform an operation on the input feature map, and generate an output feature map. In this case, the feature map may refer to data representing various features of the input data. The first feature map FM1, the second feature map FM2, and the third feature map FM3 may have, for example, a two-dimensional (2D) matrix form or a three-dimensional (3D) matrix form. The first feature map FM1, the second feature map FM2, and the third feature map FM3 may have a width W (or referred to as columns), a height H (or referred to as rows), and a depth D that may respectively correspond to the x-axis, y-axis, and z-axis on the coordinate. In this case, the depth D may be referred to as the number of channels.

[0039] The first layer L1 may convolve the first feature map FM1 with a weight map WM and generate the second feature map FM2. The weight map WM may filter the first feature map FM1 and may be referred to as a filter or a kernel. In some examples, the depth (i.e., the number of channels) of the weight map WM may be equal to the depth of the first feature map FM1. The channels of the weight map WM may be convolved with the corresponding channels of the first feature map FM1 respectively. The weight map WM may be shifted as a sliding window by traversing the first feature map FM1. The amount of shift may be referred to as the "stride length" or "stride". During each shift, each weight included in the weight map WM may be multiplied by the feature values in the region and added to the feature values in the region. The region may be the region where each weight value included in the weight map WM overlaps with the first feature map FM1. By convolving the first feature map FM1 with the weight map WM, one channel of the second feature map FM2 may be generated. Although Figure 2 one weight map WM is indicated, multiple weight maps may be sufficiently convolved with the first feature map FM1 to generate multiple channels of the second feature map FM2. In other words, the number of channels of the second feature map FM2 may correspond to the number of weight maps.

[0040] The second layer L2 can change the spatial dimension of the second feature map FM2 and generate a third feature map FM3. As an example, the second layer L2 can be a sampling layer. The second layer L2 can perform an upsampling operation or a downsampling operation. The second layer L2 can select a part of the data included in the second feature map FM2. For example, a 2D window WD can be shifted on the second feature map FM2 in units of the size of the window WD (e.g., a 4×4 matrix). Values at specific positions (e.g., the first row and the first column) in the region overlapping with the window WD can be selected. The second layer L2 can output the selected data as the data of the third feature map FM3. In another example, the second layer L2 can be a pooling layer. In this case, the second layer L2 can select the maximum value (or average value) of the feature values in the region where the second feature map FM2 overlaps with the window WD. The second layer L2 can output the selected data as the data of the third feature map FM3.

[0041] Therefore, a third feature map FM3 with a changed spatial dimension can be generated from the second feature map FM2. The number of channels of the third feature map FM3 can be equal to the number of channels of the second feature map FM2. At the same time, according to the exemplary embodiment, the sampling layer can have a higher operation speed than the pooling layer and improve the quality of the output image (e.g., peak signal-to-noise ratio (PSNR)). For example, since the operation caused by the pooling layer involves calculating the maximum value or the average value, the operation caused by the pooling layer may take longer operation time than the operation caused by the sampling layer.

[0042] According to some embodiments, the second layer L2 is not limited to a sampling layer or a pooling layer. For example, the second layer L2 can be a convolution similar to the first layer L1. The second layer L2 can convolve the second feature map FM2 with a weight map and generate a third feature map FM3. In this case, compared with the weight map WM on which the first layer L1 performs the convolution operation, the weight map on which the second layer L2 performs the convolution operation can be different.

[0043] The Nth layer can generate the Nth feature map through multiple layers including the first layer L1 and the second layer L2. The Nth feature map can be input to a reconstruction layer located at the backend of the ANN model 100, and the output data is output from this reconstruction layer. The reconstruction layer can generate an output image based on the Nth feature map. In addition, the reconstruction layer can receive the Nth feature map and multiple feature maps, e.g., the first feature map FM1 and the second feature map FM2, and generate an output image based on the multiple feature maps.

[0044] For example, the reconstruction layer can be a convolutional layer or a deconvolutional layer. In some embodiments, the reconstruction layer can include different types of layers capable of reconstructing an image based on the feature map.

[0045] The storage device 14, which serves as a storage place for storing data, can store, for example, an operating system (OS), various programs, and various data segments. The storage device 14 can be a DRAM, but is not limited thereto. The storage device 14 can include at least one of a volatile memory and a non-volatile memory. The non-volatile memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable and programmable ROM (EEPROM), a flash memory, a phase change random access memory (PRAM), a magnetic RAM (MRAM), a resistive RAM (RRAM), and / or a ferroelectric RAM (FRAM). The volatile memory can include a DRAM, an SRAM, a synchronous DRAM (SDRAM), a PRAM, an MRAM, an RRAM, and / or a FRAM. In an embodiment, the storage device can include at least one of a hard disk drive (HDD), a solid state drive (SSD), a compact flash (CF) memory, a secure digital (SD) memory, a micro SD memory, a mini SD memory, an extreme digital (xD) memory, and a memory stick.

[0046] The sensor module 15 can collect information about an object sensed by the electronic device 10. For example, the sensor module 15 can be an image sensor module. The sensor module 15 can sense or receive an image signal from the outside of the electronic device 10 and convert the image signal into image data (i.e., an image frame). As a result, the sensor module 15 can include at least one of the following various types of sensing devices: for example, an image capturing device, an image sensor, a light detection and ranging (LIDAR) sensor, an ultrasonic (UV) sensor, and an infrared (IR) sensor.

[0047] Figure 3 is a diagram for explaining image data according to an exemplary embodiment. Hereinafter, reference will be made to Figure 1 and Figure 2 to describe Figure 3 .

[0048] Referring to Figure 3 , the input data IDTA can be image data received by the electronic device 10. For example, the input data IDTA can be the input data IDTA received by the NPU 13 to generate an output image.

[0049] The input data IDTA can be a four-unit image. The four-unit image can be an image similar to the image of the input data IDTA shown in Figure 3 obtained by an image sensor having a color filter array. For example, the color filter array can be a color filter array in which four identical color filters are arranged in two rows and two columns. This arrangement can form a pixel unit.

[0050] The input data IDTA may include a plurality of pixels PX. The pixel PX may be an image unit representing a color and may include, for example, a red pixel PX_R, a green pixel PX_G, and a blue pixel PX_B. The pixel PX may include at least one sub-pixel SPX. The sub-pixel SPX may be a data unit obtained by an image sensor. For example, the sub-pixel SPX may be data obtained by a pixel circuit included in the image sensor and including a color filter. The sub-pixel SPX may express a color, for example, one of red, green, and blue.

[0051] The input data IDTA may include a plurality of pixel groups PG. The pixel group PX_G may include a plurality of pixels PX. For example, the pixel group PG may include multiple colors (e.g., red, green, and blue) forming the input data IDTA. For example, the pixel group PG may include a red pixel PX_R, a green pixel PX_G, and a blue pixel PX_B. Additionally, the colors are not limited to red, green, and blue as described above and may be implemented as various other colors, for example, magenta, cyan, yellow, white, etc.

[0052] The input data IDTA may be image data having various formats. For example, the input data IDTA may include the four images as described above. In another example, the input data IDTA may include various formats of images, for example, a Bayer pattern image, a red-green-blue-emerald (RGBE) pattern image, a cyan-yellow-yellow-magenta (CYYM) pattern image, a cyan-yellow-cyan-magenta (CYCM) pattern image, and a red-green-blue-white (RGBW) pattern image. For example, the inventive concept is not limited by the format of the input data IDTA.

[0053] The output data ODTA may be image data having a format different from that of the input data IDTA. For example, the output data ODTA may be a red-green-blue (RGB) image. The RGB image may be data indicating the degree to which red, green, and blue are expressed. The RGB image may be based on a color space of red, green, and blue. For example, the RGB image may be an image that does not distinguish patterns and is different from the input data IDTA. In other words, data indicating red values, green values, and blue values may be assigned to each pixel included in the RGB image. The output data ODTA may have a format different from that of the input data IDTA. The output data ODTA is not limited thereto and may be, for example, YUV data. Here, Y may be a luminance value, and U and V may be chrominance values. The format of the output data ODTA is not limited to RGB data and YUV data and may be implemented as various types of formats. In the related art, when the output data ODTA has a format different from that of the input data IDTA, a large number of processing operations may be used as described below.

[0054] Figure 4 It is a block diagram of an image processing device according to a comparative example.

[0055] According to the comparative example, multiple processing and conversion operations can be used to receive the input data IDTA and generate the output data ODTA. Refer to Figure 4 , the pre-processor 21 can receive the input data IDTA and perform pre-processing operations such as crosstalk correction operations and defect correction operations. The pre-processor 21 can perform the pre-processing operations based on the input data IDTA and output the corrected input data IDTA_C. The Bayer converter 22 can receive the corrected input data IDTA_C and convert the pattern of the image into a Bayer pattern. For example, the Bayer converter 22 can perform demosaicing operations, noise reduction operations, and sharpening operations. The RGB converter 23 can receive the Bayer image IDTA_B, convert the Bayer image IDTA_b into an RGB image, and output the output data ODTA.

[0056] According to the comparative example, since the process of converting the format of the input data IDTA and generating the output data ODTA involves a large number of operations, time loss may occur. In addition, in some cases, the embodiments of the comparative example may not substantially improve the image quality.

[0057] Figure 5 It is a block diagram for explaining the neural processing unit (NPU) 13 according to an example embodiment.

[0058] Refer to Figure 5 , the NPU 13 can include processing logic 200 configured to perform operations based on the ANN model 100. The processing logic 200 can receive the input data IDTA, perform operations (e.g., image processing operations) based on the input data IDTA, and output the output data ODTA. For example, the input data IDTA can be a quad image, and the output data ODTA can be an RGB image. The inventive concept is not limited to specific input data or output data. The processing logic 200 can generally control the ANN model 100. For example, the processing logic 200 can control various parameters, configurations, functions, operations, and connections included in the ANN model 100. More specifically, the processing logic 200 can make various modifications to the ANN model 100 according to the situation. For example, the processing logic 200 can activate or deactivate skip connections included in the ANN model 100.

[0059] According to an example embodiment, the processing logic 200 may perform an image processing operation based on the ANN model 100. The ANN model 100 may include a plurality of layer units. The layer units may include a plurality of layers, and each of the plurality of layers may be provided to perform an operation (e.g., a convolution operation). Then the operation is assigned to the layer (e.g., a convolutional layer). Hereinafter, the plurality of layers included in the encoding layer unit LUa will be referred to as a plurality of encoding layers. The plurality of layers included in the decoding layer unit LUb will be referred to as a plurality of decoding layers.

[0060] According to an example embodiment, the layer units may include an encoding layer unit LUa and a decoding layer unit LUb. The encoding layer unit LUa and the decoding layer unit LUb may be implemented symmetrically with respect to each other. For example, each of the plurality of encoding layer units LUa may have a corresponding level. Additionally, each of the plurality of decoding layer units LUb may also have a corresponding level. In other words, the plurality of encoding layer units LUa may have a plurality of levels, and the plurality of decoding layer units LUb may also have a plurality of levels. For example, the number of levels of the plurality of encoding layer units LUa may be equal to the number of levels of the plurality of decoding layer units LUb. For example, the encoding layer unit LUa and the decoding layer unit LUb at the same level may include the same type of layer. Additionally, the encoding layer unit LUa and the decoding layer unit LUb at the same level may establish a skip connection. The encoding layer unit LUa may sequentially encode the image data. Additionally, the decoding layer unit LUb may sequentially decode the data encoded by the encoding layer unit LUa and output the output data.

[0061] Reference Figure 5 The decoding layer unit LUb may receive the data encoded by the encoding layer unit LUa. For example, the encoded data may be a feature map. The decoding layer unit LUb may receive the encoded data due to the skip connection established with the encoding layer unit LUa. For example, a skip connection refers to a process of directly propagating data from the encoding layer unit LUa to the decoding layer unit LUb without propagating the data to an intermediate layer unit. The intermediate layer unit may be located between the encoding layer unit Lua and the decoding layer unit LUb. In other words, the encoding layer unit LUa may directly propagate the data to the decoding layer unit LUb at the same level as the encoding layer unit LUa. Alternatively, due to the skip connection, the encoding layer unit LUa and the decoding layer unit LUb may be selectively connected to each other. In this case, the encoding layer unit LUa and the decoding layer unit LUb are at the same level. The skip connection may be activated or deactivated according to the skip level. For example, the skip connection may be a selective connection relationship based on the skip level.

[0062] The processing logic 200 may load the ANN model 100, perform operations based on the input data IDTA, and output output data DOTA based on the operation results. The processing logic 200 may control various parameters of the ANN model 100. For example, the processing logic 200 may control at least one of the width W, height H, and depth D of the feature map output by the encoding layer unit LUa or the decoding layer unit LUb.

[0063] Figure 6 is a diagram for explaining the ANN model 100 according to an exemplary embodiment.

[0064] Reference Figure 6 , the ANN model 100 may include an input layer IL, an output layer OL, an encoding layer unit LUa, and a decoding layer unit LUb. Due to the input layer IL, the encoding layer unit Lua, the decoding layer unit LUb, and the output layer OL, the ANN model 100 may receive the input data IDTA and calculate the eigenvalue of the input data IDTA. For example, the ANN model 100 may receive four images and perform an operation for converting the four images into RGB images.

[0065] According to an exemplary embodiment, the input layer IL may output a feature map FMa0 to the first encoding layer unit LUa1. For example, the input layer IL may include a convolutional layer. In a similar manner as described in reference Figure 2 , the input layer IL may perform a convolution operation on the input data IDTA and a weight map. In this case, the weight map may perform a convolution operation with a constant stride value while traversing the input data IDTA.

[0066] The encoding layer unit LUa may receive the feature map output by the previous encoding layer unit and perform operations that can be assigned to each encoding layer unit (e.g., LUa1). For example, the first encoding layer unit LUa1 may receive the feature map FMa0 and perform operations caused by the respective layers included in the first encoding layer unit LUa1. For example, the encoding layer unit LUa may include a convolutional layer, a sampling layer, and an activation layer. The convolutional layer may perform a convolution operation. The sampling layer may perform a downsampling operation, an upsampling operation, an average pooling operation, or a max pooling operation. The activation layer may perform operations caused by a rectified linear unit (ReLU) function or a sigmoid function. The first encoding layer unit LUa1 may output a feature map FMa1 based on the operation results.

[0067] The feature map FMa1 output by the first encoding layer unit LUa1 can have a smaller width, a smaller height, and a greater depth than the input feature map FMa0. For example, the first encoding layer unit LUa1 can control the width, height, and depth of the feature map FMa1. For example, the first encoding layer unit LUa1 can control the depth of the feature map FMa1 from not increasing. The first encoding layer unit LUa1 can have a parameter for setting the depth of the feature map FMa1. At the same time, the first encoding layer unit LUa1 can include a downsampling layer DS. The downsampling layer DS can select predetermined feature values from the feature values included in the input feature map FMa0 and output the selected predetermined feature values as the feature values of the feature map FMa1. In other words, the downsampling layer DS can control the width and height of the feature map FMa1. The second encoding layer unit LUa2 and the third encoding layer unit LUa3 can also perform operations similar to those of the first encoding layer unit LUa1. For example, each of the second encoding layer unit LUa2 and the third encoding layer unit LUa3 can receive a feature map from the previous encoding layer unit, perform operations caused by a plurality of layers included in the current layer unit, and output the feature map including the operation result to the next encoding layer unit.

[0068] The encoding layer unit LUa can output the feature map to the next encoding layer unit Lua or the decoding layer unit LUb at the same level as the encoding layer unit LUa. Each encoding layer unit LUa can be fixedly connected to the next encoding layer unit Lua and connected to the decoding layer unit LUb at the same level through one or more skip connections (e.g., the first skip connection SK0 to the fourth skip connection SK3). For example, when the ordinal number of one layer unit from the input layer IL is equal to the ordinal number of another layer unit from the output layer OL, the two layer units can be said to be at the same level. The layer units at the same level can be, for example, the first encoding layer unit LUa1 and the first decoding layer unit LUb1.

[0069] Introducing the skip connection SK can improve the training of the deep neural network. In the absence of a skip connection, the introduction of additional layers can sometimes cause a decrease in the output quality (e.g., due to the vanishing learning gradient problem). Therefore, implementing one or more skip connections SK between the encoding layer unit and the decoding layer unit can improve the overall performance of the ANN by enabling more effective training of deeper layers.

[0070] According to an example embodiment, the processing logic 200, the NPU 13, or the electronic device 10 may select at least some of the plurality of skip connections (e.g., the first skip connection SK0 to the fourth skip connection SK3). For example, the processing logic 200 may receive information about the skip level. When the skip level of the ANN model 100 is set, some of the first skip connection SK0 to the fourth skip connection SK3 corresponding to the preset skip level may be activated. For example, when the skip level of the ANN model 100 is equal to 2, the first skip connection SK0 and the second skip connection SK1 may be activated. Due to the activated skip connections, the encoding layer unit LUa may output the feature map to the decoding layer unit LUb. The unactivated skip connections (e.g., SK2 and SK3) may not propagate the feature map.

[0071] According to an example embodiment, layer units at the same level (e.g., LUa1 and LUb1) may process feature maps having substantially the same size. For example, the size of the feature map FMa0 received by the first encoding layer unit LUa1 may be substantially equal to the size of the feature map FMb0 output by the first decoding layer unit LUb1. For example, the size of the feature map may include at least one of width, height, and depth. Additionally, the size of the feature map FMa1 output by the first encoding layer unit LUa1 may be substantially equal to the size of the feature map FMb1 of the first decoding layer unit LUb1.

[0072] According to an example embodiment, the encoding layer unit LUa and the decoding layer unit LUb at the same level may have substantially the same sampling size. For example, the downsampling size of the first encoding layer unit LUa1 may be equal to the upsampling size of the first decoding layer unit LUb1.

[0073] The decoding layer unit LUb may receive the feature map from the previous decoding layer unit LUb or from the encoding layer unit LUa at the same level. The decoding layer unit LUb may use the received feature map to process operations. For example, the decoding layer unit LUb may include a convolutional layer, a sampling layer, and an activation layer.

[0074] The feature map FMa1 output by the first encoding layer unit LUa1 may have a smaller width, a smaller height, and a larger depth than the input feature map FMa0. For example, the first encoding layer unit LUa1 may control the width, height, and depth of the feature map FMa1. For example, the first encoding layer unit LUa1 may control that the depth of the feature map FMa1 does not increase. The first encoding layer unit LUa1 may have a parameter for setting the depth of the feature map FMa1.

[0075] The upsampling layer US can adjust the size of the input feature map. For example, the upsampling layer US can adjust the width and height of the feature map. The upsampling layer US can perform an upsampling operation. The upsampling operation can use each eigenvalue of the input feature map and the eigenvalues adjacent to each eigenvalue. As an example, the upsampling layer US can be a layer configured to write the same eigenvalue to the output feature map using the nearest neighbor method. In another example, the upsampling layer US can be a transposed convolutional layer and upsample the image using a predetermined weight map.

[0076] The output layer OL can reconstruct the feature map FMb0 output by the first decoding layer unit LUb1 into the output data ODTA. The output layer OL can be a reconstruction layer configured to convert the feature map into image data. For example, the output layer OL can be one of a convolutional layer, a deconvolutional layer, and a transposed convolutional layer. For example, the converted image data can be RGB data.

[0077] Figure 7A and Figure 7B are block diagrams for illustrating the encoding layer unit LUa and the decoding layer unit LUb according to an exemplary embodiment.

[0078] Reference Figure 7A , the encoding layer unit LUa can include a feature map block FMB and a downsampling layer DS, and can also include a convolutional layer CONV. The feature map block FMB can include a plurality of convolutional layers, a plurality of activation layers, and a summator, which will be described in detail below with reference to Figure 8 In the following.

[0079] The input layer IL can form a skip connection with the output layer OL. For example, the feature map FMa0 output by the input layer IL can be output to the first encoding layer unit LUa1 and the output layer OL. In another example, when the skip level is equal to 0, the input layer IL may not directly output the feature map to the output layer OL. In yet another example, when the skip level is equal to 1 or greater, the input layer IL can directly output the feature map to the output layer OL.

[0080] The multiple layers included in the encoding layer unit LUa can form skip connections with the multiple layers of the decoding layer unit LUb corresponding to them respectively. For example, at least some of the multiple layers included in the encoding layer unit LUa can form skip connections with at least some of the multiple layers included in the decoding layer unit LUb at the same level as the encoding layer unit Lua.

[0081] The multiple layers included in the encoding layer unit LUa can form skip connections with the multiple layers included in the decoding layer unit LUb configured to perform symmetric operations therewith. For example, the convolutional layer CONV and the feature map block FMB can form skip connections. The convolutional layer CONV and the feature map block FMB are included in the encoding layer unit LUa and the decoding layer unit LUb and are at the same level. In addition, the downsampling layer DS and the upsampling layer US at the same level can form skip connections. Refer to Figure 7A and Figure 7B , the downsampling layer La13 and the upsampling layer Lb11 at the same level can form skip connections. The feature map block La22 and the feature map block Lb22 at the same level can form skip connections. For the sake of brevity, the above description is provided. The feature map block La12 and the feature map block Lb21 at the same level can form skip connections.

[0082] When setting the skip level, the ANN model 100 can activate some of the multiple skip connections based on a preset skip level. For example, the ANN model 100 can directly propagate data from the encoding layer unit LUa to the decoding layer unit LUb at a level based on the preset skip level. In the example, when the skip level is equal to 0, the skip connection can be deactivated. When the skip connection is deactivated, the feature map can not be propagated from the encoding layer unit Lua through the skip connection. In another example, when the skip level is equal to 1, the first skip connection SK0 can be activated. When the first skip connection SK0 is activated, the input layer IL can propagate the feature map FMa0 to the output layer OL. In yet another example, when the skip level is equal to 2, the layers included in the first encoding layer unit LUa1 and at least some of the input layer IL can propagate the feature map to the layers included in the first decoding layer unit LUb1 and at least some of the output layer OL.

[0083] Figure 8 is a block diagram of the feature map block FMB according to an example embodiment.

[0084] Refer to Figure 8 , the feature map block FMB can include multiple layers and a summator SM. The multiple layers can include multiple convolutional layers CL0, CL1, and CLx and multiple activation layers AL0, AL1, and ALn.

[0085] According to an example embodiment, the feature map block FMB can receive the feature map FMc. Based on the feature map FMc, the input activation layer AL0 can output the feature map FMd to the intermediate layer group LG and the summator SM. Additionally, the intermediate layer group MID can output the feature map FMf. The summator SM can sum the feature maps FMd and FMf received from the intermediate layer group MID and the input activation layer AL0 and output the feature map FMg.

[0086] The leading layer group FR may include multiple layers (e.g., CL0 and AL0). The leading layer group FR may be located at the front end of the feature map block FMB and receive the feature map FMc. The feature map FMc is received by the feature map block FMB. As an example, the leading layer group FR may include one convolutional layer CL0 and one activation layer AL0. In another example, the leading layer group FR may include at least one convolutional layer and at least one activation layer. The leading layer group FR may output the feature map FMd to the intermediate layer group MID and the summator SM.

[0087] The intermediate layer group MID may include multiple layers CL1, AL1, ……, and CLx. In an example, the intermediate layer group MID may include multiple convolutional layers and multiple activation layers. The positions of the multiple convolutional layers and multiple activation layers included in the intermediate layer group MID may be alternately set. In this case, the feature map FMe output by the convolutional layer CL1 may be received by the activation layer AL1. The convolutional layer CL1 may be set at the very front end of the intermediate layer group MID. Alternatively, the convolutional layer CLx may be located at the very end of the intermediate layer group MID. In other words, the feature map FMd received by the convolutional layer CL1 may be the same as the feature map FMd received by the intermediate layer group MID. The feature map FMf output by the convolutional layer CLx may be the same as the feature map FMf output by the intermediate layer group MID.

[0088] The output activation layer ALn may receive the feature map FMg output by the summator SM. Then, the output activation layer ALn may activate the characteristics of the feature map FMg and output the feature map FMh. The output activation layer ALn may be located at the very end of the feature map block FMB.

[0089] Figure 9 is a diagram for explaining the ANN model 100 according to an exemplary embodiment.

[0090] Reference Figure 9 , the ANN model 100 may include feature maps having different widths W, heights H, and depths D according to each level LV. For example, the feature map FMa0 output by the input layer IL may have the lowest depth D. As the operation performed by the encoding layer unit LUa is repeated, the depth D may increase. Additionally, when the depth D increases exponentially, the amount of operations to be processed per unit time may increase rapidly, thereby increasing the operation time. Here, the amount of operations may be expressed in units such as trillion operations per second (TOPS).

[0091] According to an example embodiment, the ANN model 100 may perform operations assigned to each layer and control the depth D of the feature map. For example, the encoding layer unit LUa and the decoding layer unit LUb may output a feature map having a depth D corresponding to each level. For example, the depth D of the output feature map may correspond to a function of the encoding layer unit LUa and the decoding layer unit LUb according to the level. In an example, the processing logic 200 may control the depth D of the feature map output by the ANN model 100. In another example, a function of the depth D may be stored in each layer. In yet another example, a function of the depth D may be stored in an external memory of the ANN model 100 and applied to each layer.

[0092] According to an example embodiment, a function of the depth D may be expressed as a function of the level LV of each layer. In an example, the function of the depth D may be linear with respect to the level LV. In this case, the function of the depth D may be a linear function with the level LV as a parameter, and the function FD of the depth D may be expressed as the equation shown: FD = a * LV + b. In another example, the function of the depth D may be an exponential function with the level LV as the base. For example, the function FD of the depth D may be expressed as the equation shown: FD = a * (LV^2) + b. Alternatively, the function FD of the depth D may be expressed as the equation shown: FD = a * (LV^c) + b. In yet another example, the function FD of the depth D may be a logarithmic function of the level LV, and the base of the logarithmic function may be arbitrarily selected. For example, the function of the depth D may be expressed as the equation shown: FD = b * log(LV - 2), where a, b, and c are constants, and LV represents the level of each layer. Alternatively, the constant a may satisfy the inequality: a ≥ b / 2.

[0093] According to an example embodiment, the function FD of the depth D may be a function in which the level LV of each layer is not exponential. For example, the function FD of the depth D may not be b * (2^LV). Alternatively, the function FD of the depth D may have a smaller depth D than a function with the level LV of each layer as the exponent. In this case, the smaller depth D is due to an increase in the operation time caused by the ANN model 100.

[0094] Figure 10 A table of depth values of the feature map according to an example embodiment is shown. Figure 10 Exemplarily shown above with reference to Figure 9 the function FD described.

[0095] According to an exemplary embodiment, function FD1 and function FD2 may be included in ANN model 100, while function FD3 may not be included in ANN model 100. Even when level LV increases, functions FD1 and FD2 of ANN model 100 may relatively monotonically increase the depth of the feature map. However, since function FD3 has level LV as an exponent, the depth of the feature map may increase sharply. Accordingly, ANN model 100 may have a function that does not have level LV as an exponent to shorten the operation time. For example, as a result of not including a function having level LV as an exponent, encoding layer unit LUa and decoding layer unit LUb may adjust the depth of the output feature map.

[0096] Figure 11 FIG. is for illustrating a downsampling operation according to an exemplary embodiment.

[0097] Reference Figure 11 , downsampling layer DS may receive feature map 31 and control the width W and height H of feature map 31. For example, downsampling layer DS may output output feature map 32 with the width W and height H controlled based on sampling information SIF.

[0098] Downsampling layer DS may perform a downsampling operation based on sampling information SIF. In other words, downsampling layer DS may select some of the feature values included in feature map 31. The selected feature values may constitute output feature map 32. For example, output feature map 32 may have a smaller size (e.g., width W or height H) and include a smaller number of feature values than feature map 31. Meanwhile, sampling information SIF may be received by processing logic 200. Sampling information SIF may be information written to downsampling layer DS.

[0099] Sampling information SIF may include sampling size information, sampling position information, and sampling window size information. Downsampling layer DS may define the size of output feature map 32 based on the sampling size information. For example, when the sampling size is equal to 2, at least one of the width W and height H of output feature map 32 may be equal to 2. For example, when the width W of output feature map 32 is equal to 2, output feature map 32 may have two columns. When the height H of output feature map 32 is equal to 3, output feature map 32 may have three rows.

[0100] Downsampling layer DS may select feature values at the same position in each feature map region FAl to FA4 based on the sampling position information. For example, when the sampling position information indicates the values in the first row and the first column, downsampling layer DS may calculate 12, 30, 34, and 37. In this case, the calculated values are the values in the first row and the first column in each feature map region FAl to FA4, and output feature map 32 is generated.

[0101] The downsampling layer DS may define the sizes of the respective feature map regions FA1 to FA4 based on the sampling window size information. For example, when the sampling window size is equal to 2, at least one of the width and height of a feature map region may be equal to 2.

[0102] According to an example embodiment, the downsampling layer DS may output an output feature map 32, which has a higher operation speed and higher image quality than a pooling layer. For example, the pooling layer may be a max pooling layer or an average pooling layer. For example, the operation time taken for the downsampling operation performed by the downsampling layer DS may be shorter than the pooling operation time taken by the pooling layer.

[0103] Figure 12 is a block diagram of the AP 400 according to an example embodiment.

[0104] Figure 12 The system shown may be the AP 400, and the AP 400 may include a system-on-chip (SoC) as a semiconductor chip.

[0105] The AP 400 may include a processor 410 and an operation memory 420. Although not shown in Figure 12 , the AP 400 may further include: at least one intellectual property (IP) module, connected to the system bus. The operation memory 420 may store software related to the operation of the system to which the AP 400 is applied, for example, various programs and instructions. As an example, the operation memory 420 may include an operating system 421 and an ANN module 422. The processor 410 may execute the ANN module 422 loaded in the operation memory 420. The processor 410 may perform operations based on the ANN model 100 including the encoding layer unit LUa and the decoding layer unit LUb according to the above embodiments.

[0106] Figure 13 is a block diagram of an imaging device 5000 according to an example embodiment.

[0107] Referring to Figure 13 , the imaging device 5000 may include an image capturing unit 5100, an image sensor 500, and a processor 5200. For example, the imaging device 5000 may be an electronic device capable of performing image processing operations. The imaging device 5000 may capture an image of an object S and obtain an input image. The processor 5200 may provide control signals and / or information for the operation of each component to the lens driver 5120 and the timing controller 520.

[0108] The image capturing unit 5100 may be a component configured to receive light, and includes a lens 5110 and a lens driver 5120, and the lens 5110 may include at least one lens. In addition, the image capturing unit 5100 may further include an iris and an iris driver.

[0109] The lens driver 5120 can send information about focus detection to the processor 5200 and receive information about focus detection from the processor 5200, and can adjust the position of the lens 5110 in response to a control signal provided by the processor 5200.

[0110] The image sensor 500 can convert incident light into image data. The image sensor 500 can include a pixel array 510, a timing controller 520, and an image signal processor 530. The optical signal transmitted through the lens 5110 can reach the light receiving surface of the pixel array 510 and form an image of the object S.

[0111] The pixel array 510 can be a complementary metal oxide semiconductor (CMOS) image sensor (CIS) configured to convert an optical signal into an electrical signal. The exposure time and sensitivity of the pixel array 510 can be adjusted by the timing controller 520. As an example, the pixel array 510 can include a color filter array for obtaining the four images described above with reference to Figure 3 the filter described.

[0112] The processor 5200 can receive the image data from the image signal processor 530 and perform various image post - processing operations on the image data. For example, the processor 5200 can convert an input image (e.g., the four - image) into an output image (e.g., an RGB image) based on the ANN model 100 according to the above - described embodiments. At the same time, the inventive concept is not limited thereto, and the image signal processor 530 can also perform operations based on the ANN model 100. Alternatively, various operation processing devices located inside or outside the imaging device 5000 can convert the format of the input image based on the ANN model 100 and generate an output image.

[0113] According to this embodiment, the imaging device 5000 can be included in various electronic devices. For example, the imaging device 5000 can be installed on an electronic device such as a camera, a smart phone, a wearable device, an IoT device, a tablet personal computer (PC), a laptop PC, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, and a display device. In addition, the imaging device 5000 can be installed on an electronic device included as a component in a vehicle, furniture, manufacturing equipment, a door, and various measurement devices.

[0114] Although the inventive concept has been specifically shown and described with reference to embodiments of the present invention, it should be understood that various changes in form and detail can be made without departing from the spirit and scope of the appended claims.

Claims

1. An electronic device, comprising: processing logic configured to receive input image data and use an artificial neural network model to generate output image data having a format different from that of the input image data, wherein the artificial neural network model comprises: a plurality of encoding layer units including a plurality of layers, wherein each layer of the plurality of layers includes an encoding level, and the encoding level includes the layer ordinal number from the input layer; and a plurality of decoding layer units including a plurality of other layers, wherein each of the plurality of other layers includes a decoding level, the decoding level includes the layer ordinal number from the output layer, and each other layer is configured to form a skip connection with a corresponding layer in the plurality of layers, the corresponding layer having the same layer ordinal number as that of this other layer, wherein a first encoding layer unit of a first level receives a first input feature map and outputs a first output feature map based on the first input feature map to a subsequent encoding layer unit and a decoding layer unit at the first level, and wherein the processing logic receives skip level information and activates or deactivates the skip connection based on the skip level indicated by the skip level information.

2. The electronic device according to claim 1, wherein the layers of the encoding layer unit and the layers of the decoding layer unit connected by the skip connection perform symmetric operations with each other.

3. The electronic device according to claim 1, wherein the first encoding layer unit adjusts the depth of the first output feature map based on the first level.

4. The electronic device according to claim 3, wherein the first encoding layer unit adjusts the depth of the output feature map based on a function having the first level as a parameter, and the function is a linear function having the first level as a parameter.

5. The electronic device according to claim 1, wherein the convolutional layers of the encoding layer unit and the decoding layer unit at the same level are connected by the skip connection, the feature map blocks of the encoding layer unit and the decoding layer unit at the same level are connected by the skip connection, or the downsampling layer of the encoding layer unit and the upsampling layer of the decoding layer unit at the same level are connected by the skip connection.

6. The electronic device according to claim 1, wherein the encoding layer unit and the decoding layer unit at each level include feature map blocks, wherein each of the feature map blocks includes: a leading layer group configured to output a first feature map; an intermediate layer group configured to receive the first feature map and output a second feature map; a summer configured to sum the first feature map and the second feature map and output a third feature map; and an output activation layer configured to output a fourth feature map based on the third feature map.

7. The electronic device according to claim 1, wherein the encoding layer unit includes: a downsampling unit configured to receive a fifth feature map, select some of the feature values included in the fifth feature map, and output the first output feature map having a size smaller than that of the fifth feature map.

8. The electronic device according to claim 1, wherein The input image data includes four images, and the output image data includes a red-green-blue image.

9. An electronic device, comprising: processing logic configured to perform operations using an artificial neural network model, wherein the artificial neural network model includes: a plurality of encoding layer units including a plurality of layers, each layer of the plurality of layers including an encoding level, and the encoding level including a layer ordinal number from the input layer; and a plurality of decoding layer units including a plurality of other layers, each other layer of the plurality of other layers including a decoding level, the decoding level including a layer ordinal number from the output layer, and each other layer being configured to form a skip connection with a corresponding layer in the plurality of layers, the corresponding layer having the same layer ordinal number as that of this other layer, wherein a first encoding layer unit at a first level among the plurality of levels receives a first input feature map, outputs a first output feature map to an encoding layer unit at a next level of the first level and a decoding layer unit at the first level, and adjusts the depth of the first output feature map based on the first level, and wherein the processing logic receives skip level information and activates or deactivates the skip connection based on the skip level indicated by the skip level information.

10. The electronic device according to claim 9, wherein, the first encoding layer unit adjusts the depth of the output feature map based on a function having the first level as a parameter, and compared with a function having the first level as an exponent, the function having the first level as a parameter adjusts the depth of the output feature map to a smaller value.

11. The electronic device according to claim 9, wherein, the layers of the encoding layer units and the layers of the decoding layer units at the same level are selectively connected to each other.

12. The electronic device according to claim 11, wherein, the layers of the encoding layer units and the layers of the decoding layer units at the same level perform operations symmetric to each other.

13. The electronic device according to claim 11, wherein, the convolutional layers of the encoding layer units and the decoding layer units at the same level are selectively connected to each other, the feature map blocks of the encoding layer units and the decoding layer units at the same level are selectively connected to each other, or the downsampling layers of the encoding layer units and the upsampling layers of the decoding layer units at the same level are connected to each other.

14. The electronic device according to claim 9, wherein, the encoding layer units and the decoding layer units at each level include feature map blocks, wherein each of the feature map blocks includes: a leading layer group configured to output a first feature map; an intermediate layer group configured to receive the first feature map and output a second feature map; a summator configured to sum the first feature map and the second feature map and output a third feature map; and an output activation layer configured to output a fourth feature map based on the third feature map.

15. The electronic device according to claim 14, wherein, wherein the leading layer group includes a convolutional layer and an activation layer, and The intermediate layer group is configured to sequentially connect a plurality of convolutional layers and a plurality of activation layers.

16. The electronic device according to claim 9, wherein, the encoding layer unit includes: a downsampling unit configured to receive a fifth feature map, select some of the feature values included in the fifth feature map, and output the first output feature map having a size smaller than that of the fifth feature map.

17. The electronic device according to claim 16, wherein, the processing logic receives sampling position information, and the downsampling unit selects the feature values located at positions based on the sampling position information from the feature values included in the fifth feature map.

18. An electronic device configured to perform an image processing operation, the electronic device comprising: processing logic configured to receive four-image data from a color filter array, in which four identical color filters are arranged in two rows and two columns to form a pixel unit, the processing logic being configured to use an artificial neural network model to generate output image data having a format different from that of the four-image data, wherein the artificial neural network model includes: a plurality of encoding layer units including a plurality of layers, where each of the plurality of layers includes an encoding level, and the encoding level includes the layer ordinal number from the input layer; and a plurality of decoding layer units including a plurality of other layers, where each of the plurality of other layers includes a decoding level, the decoding level includes the layer ordinal number from the output layer, and each of the other layers is configured to form a skip connection with a corresponding layer in the plurality of layers, the corresponding layer having the same layer ordinal number as that of this other layer, wherein a first encoding layer unit receives a first input feature map and outputs a first output feature map based on the first input feature map to a subsequent encoding layer unit and a decoding layer unit at a first level, and wherein the processing logic receives skip level information and activates or deactivates the skip connection based on the skip level indicated by the skip level information.

Citation Information

Patent Citations

  • Smart glasses

    KR1020190077255A

  • Conditional generative adversarial network-based monocular image depth estimation method

    CN108564611A

  • Method, system, and computer-readable medium for processing images using cross-stage skip connections

    CN112997479A

  • Region proposal for image regions that include objects of interest using feature maps from multiple layers of a convolutional neural network model

    US20190073553A1