Method and system for providing a rotationally invariant neural network

By designing an end-to-end symmetric neural network and utilizing symmetric kernels and weight sharing, the problem of high computational complexity of deep convolutional neural networks for input images in different directions is solved, achieving output directional invariance and improved computational efficiency.

CN111667047BActive Publication Date: 2025-10-03SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010142256.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-25
Filing Date
2020-03-04
Publication Date
2025-10-03
Estimated Expiration
2040-03-04

AI Technical Summary

Technical Problem

Deep convolutional neural networks have high computational complexity when processing input images in different orientations, which increases computational costs, and existing technologies make it difficult to ensure the directional invariance of the output.

Method used

By designing an end-to-end symmetric neural network, utilizing symmetric kernels and weight sharing, the network is trained so that its output does not depend on the input direction, and low-complexity convolution algorithms such as FFT domain point-by-point multiplication and block circulant matrix constraints are adopted to reduce computational complexity.

Benefits of technology

This achieves similarity in output when input images are in different orientations, reduces computational cost and storage requirements, and improves the network's directional invariance and computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111667047B_ABST
    Figure CN111667047B_ABST
Patent Text Reader

Abstract

A method and system for providing a rotationally invariant neural network is disclosed herein. According to one embodiment, the method for providing a rotationally invariant neural network includes receiving a first input of an image in a first orientation, and training a kernel symmetric such that an output corresponding to the first input is the same as an output corresponding to a second input of the image in a second orientation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 814,099, filed in the U.S. Patent and Trademark Office on March 5, 2019, and assigned U.S. Patent Application No. 16 / 452,005, filed in the U.S. Patent and Trademark Office on June 25, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to neural networks and more particularly to a method and system for providing a rotationally invariant neural network. Background Art

[0004] Deep convolutional neural networks have emerged as the state-of-the-art in machine learning, for example, object detection, image classification, scene segmentation, and image quality improvement such as super-resolution and disparity estimation.

[0005] There has been recent interest in developing specialized hardware accelerators for training and running deep convolutional neural networks. A convolutional neural network (CNN) consists of multiple layers of convolutional filters (also called kernels). This process is computationally expensive due to the large number of feature maps in each layer, the large dimensionality of the filters (kernels), and the increasing number of layers in a deep neural network. The computational complexity increases with larger input sizes (e.g., full high-dimensional (HD) images), which translates into larger widths and heights for the input feature maps and all intermediate feature maps. Convolution is performed by repeatedly using an array of multiply-accumulate (MAC) units. A typical MAC is a sequential circuit that computes the product of two received values ​​and accumulates the result in a register. Summary of the Invention

[0006] According to one embodiment, a method includes receiving a first input of an image at a first orientation, and training a kernel to be symmetric such that an output corresponding to the first input is the same as an output corresponding to a second input of the image at a second orientation.

[0007] According to one embodiment, a method includes receiving, by a neural network, a first input of an image at a first orientation, generating a first loss function based on a first output associated with the first input, receiving, by the neural network, a second input of an image at a second orientation, generating a second loss function based on a second output associated with the second input, and training the neural network to minimize a sum of the first loss function and the second loss function.

[0008] According to one embodiment, a system includes a memory and a processor configured to receive a first input of an image in a first orientation and train a kernel to be symmetric such that an output corresponding to the first input is the same as an output corresponding to a second input of the image in a second orientation. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other aspects, features and advantages of some embodiments of the present disclosure will become more apparent through the following detailed description taken in conjunction with the accompanying drawings, in which:

[0010] Figure 1 is a schematic diagram of a neural network according to one embodiment;

[0011] Figure 2 is a schematic diagram of a method for training a symmetric neural network according to one embodiment;

[0012] Figure 3 is a flowchart of a method for training an end-to-end symmetric neural network in which the kernel is asymmetric, according to one embodiment;

[0013] Figure 4 is a flow chart of a method for training an end-to-end neural network in which the kernel is symmetric, according to one embodiment;

[0014] Figure 5 is a schematic diagram of a block circulant matrix according to one embodiment; and

[0015] Figure 6 is a block diagram of an electronic device in a network environment according to one embodiment. DETAILED DESCRIPTION

[0016] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that the same elements will be designated by the same reference numerals even though they are shown in different drawings. In the following description, specific details such as detailed configurations and components are provided only to help fully understand the embodiments of the present disclosure. Therefore, it will be apparent to those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and structures have been omitted. The terms described below are defined in consideration of the functions in the present disclosure and may vary depending on the user, user intent or custom. Therefore, the definition of the terms should be determined based on the content throughout this specification.

[0017] The present disclosure may have various modifications and various embodiments, wherein the embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to the embodiments, but includes all modifications, equivalents and substitutes within the scope of the present disclosure.

[0018] Although the terms including ordinal numbers (such as first, second, etc.) can be used to describe various elements, structural elements are not limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of this disclosure, a first structural element can be referred to as a second structural element. Similarly, a second structural element can also be referred to as a first structural element. As used herein, the term "and / or" includes any and all combinations of one or more associated items.

[0019] The terms used herein are only used to describe various embodiments of the present disclosure and are not intended to limit the present disclosure. The singular form is intended to include the plural form, unless the context clearly indicates otherwise. In the present disclosure, it should be understood that the term "including" or "having" indicates the presence of a feature, number, step, operation, structural element, component, or a combination thereof, and does not exclude the presence of one or more other features, numbers, steps, operations, structural elements, components, or a combination thereof or the possibility of adding one or more other features, numbers, steps, operations, structural elements, components, or a combination thereof.

[0020] Unless defined differently, all terms used herein have the same meaning as understood by those skilled in the art to which the present disclosure belongs. Unless explicitly defined in the present disclosure, terms such as those defined in general dictionaries should be interpreted as having the same meaning as in the context of the relevant art and should not be interpreted as having an ideal or overly formal meaning.

[0021] The electronic device according to one embodiment may be one of various types of electronic devices. The electronic device may include, for example, a portable communication device (e.g., a smart phone), a computer, a portable multimedia device, a portable medical device, a camera, a wearable device, or a household appliance. According to one embodiment of the present disclosure, the electronic device is not limited to those described above.

[0022] The terms used in this disclosure are not intended to limit the disclosure, but are intended to include various changes, equivalents or replacements to corresponding embodiments. With respect to the description of the accompanying drawings, similar figure numerals may be used to refer to similar or related elements. The singular form of the noun corresponding to an item may include one or more items, unless the relevant context clearly indicates otherwise. As used herein, each of the phrases such as "A or B", "at least one of A and B", "at least one of A or B", "A, B or C", "at least one of A, B and C" and "at least one of A, B or C" may include all possible combinations of the items listed together in a corresponding phrase in the phrase. As used herein, terms such as "first" and "second" may be used to distinguish a corresponding component from another component, but are not intended to limit the component in other respects (e.g., importance or order). It is intended that if an element (e.g., a first element) is referred to as being “coupled,” “coupled to,” “connected to,” or “connected to” another element (e.g., a second element), with or without the term “operably” or “communicatively,” this indicates that the element may be coupled to the other element directly (e.g., by wire), wirelessly, or via a third element.

[0023] As used herein, the term "module" may include units implemented in hardware, software, or firmware, and may be used interchangeably with other terms, such as "logic," "logic block," "component," and "circuit." A module may be a single integral component, or a minimum unit or portion thereof, adapted to perform one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0024] The present disclosure provides a system and method for training convolutional neural networks so that their outputs do not depend on the orientation of the input. This may be desirable in applications such as image enhancement, where similar image enhancement performance can be achieved regardless of whether the input image is oriented horizontally or vertically, or whether the input image is upside down or flipped left to right. Another application is pattern recognition (e.g., correctly classifying objects in an image regardless of the orientation of the input image or the orientation of the objects in the image).

[0025] This disclosure provides a method and system for designing and training neural networks to improve the final results and make the network orientation invariant. In some applications, such as image denoising or image super-resolution, different results may be observed if the network is applied to different left-right or top-down mirror flips of the input image. However, in practical applications, the orientation of the input image is unknown. By designing the neural network to be end-to-end symmetric, the final results can be guaranteed to be similar for all input orientations, and the output is invariant to flips or transpositions in the input.

[0026] The present disclosure provides a method and system for designing a low-complexity implementation of convolution. Methods exist for implementing convolution with lower complexity. Such methods include point-wise multiplication in the Fast Fourier Transform (FFT) domain or the Winograd domain. By designing a neural network so that its kernel is symmetric, a lower-complexity algorithm for implementing the convolution operator in the FFT domain is provided.

[0027] In one embodiment, the system includes a neural network kernel such that the kernel is symmetric with respect to a desired degree of flipping or transposition. Thus, by making each individual kernel symmetric, the end-to-end network is also symmetric with respect to the input.

[0028] The present disclosure provides a method and system for training a network so that the output of the network is invariant to different rotations in the input. In one embodiment, the present system and method provide for training an end-to-end symmetric neural network, but the constituent filters do not have to be symmetric. In another embodiment, the present system and method provide for training a deep neural network whose constituent filters are symmetric and their responses are invariant to input rotations, and therefore, if the execution is forced between all filters in the network, the resulting network is also end-to-end symmetric.

[0029] Figure 11 is a schematic diagram of a neural network 100 according to one embodiment. An input 102 is sent to the neural network 100, which includes five layers. The first layer 104 is a convolutional layer (CONV) with a rectified linear unit (ReLU) activation function and a 5×5×32 kernel. The second layer 106 is a convolutional layer with a ReLU activation function and a 1×1×5 kernel. The third layer 108 is a convolutional layer with a ReLU activation function and a 3×3×5 kernel. The fourth layer 110 is a convolutional layer with a ReLU activation function and a 1×1×32 kernel. The fifth layer 112 is a deconvolutional (DECONV) layer with a 9×9 kernel with a stride of 3. The weights of the first layer 104 are shared with the weights of the fifth layer 112. The network produces an output 114, wherein a loss function 116 is generated based on the output 114.

[0030] Figure 2 is a diagram 200 of a method for training a symmetric neural network according to one embodiment. Using a neural network such as Figure 1 The depicted neural network shows a first training iteration 202, and subsequent training iterations 204, 206, and 208 are shown using copies of the neural network from the first training iteration 202. In each training iteration 202-208, an image 210 is input at a different orientation. The input is processed through the various layers of the network, an output 212 is generated, and a loss function 214 is calculated, which can be minimized. A loss function 214 is generated for each output 212, and each loss function 214 corresponds to an orientation of the image 210.

[0031] The system forces the network as a whole to learn symmetric transformations. For example, consider a network that only learns that the output is invariant to left-right mirror flips. The system can feed an image to the network and feed the left-right flipped version of the image to another copy of the network (via weight sharing). The system can modify the loss function to have an additional term corresponding to the added loss from the flipped image (for example, using mean square error (MSE)) to indicate to the network and let the network know that the two outputs should be the same or similar, this additional term forces the network output of the normal image and the reverse flipped output of the copy network. This can be extended to 9 networks to know all flips. However, in inference, only one network is used (for example, a single network). This network has more degrees of freedom than a layer-by-layer symmetric network, so it can have better performance.

[0032] This system can be applied to super-resolution. It derives convolutional neural networks (CNNs) whose outputs are independent of the input's orientation. Instead of training a CNN using only the original data, the system can augment data with different orientations. The system can train a unified network using the augmented dataset, but calculate separate loss functions for data with different orientations. These loss functions are minimized simultaneously while assigning the same weights. This is similar to training multiple CNNs using data with different orientations, while sharing the same architecture and weights.

[0033] In addition, the method can also be applied to a cascade trained-super resolution convolutional neural network (CT-SRCNN)-5 layer as described in U.S. patent application Ser. Nos. 15 / 655,557 and 16 / 138,279, both titled “Systems and methods for designing efficient super-resolution deep convolutional neural networks through cascade network training, cascade network pruning, and dilated convolutions” (these patent applications are incorporated herein by reference).

[0034] The main difference of this system compared to typical methods is that the convolutional neural network is only applied once during testing. Because the loss function for each direction is given the same weight for all training samples, the network will implicitly learn approximately direction-invariant.

[0035] As shown in Table 1, the above can be tested using images from Set 14 with different orientations. Figure 1 CT-SRCNN-5 network.

[0036] Table 1

[0037] enter Set14PSNR Set14SSIM original 29.23 0.8193 Flip upside down 29.23 0.8190 Flip left and right 29.23 0.8191 Up and down + left and right flip 29.22 0.8192 average 29.23 0.8192

[0038] It is observed that the peak signal to noise ratio (PSNR) and structural similarity index (SSIM) are almost the same in different directions. Therefore, it can be confirmed that Figure 1 The data augmentation architecture in

[15] can achieve approximate direction invariance.

[0039] Figure 3300 is a flow chart of a method for training an end-to-end symmetric neural network in which the kernel is asymmetric, according to one embodiment. At 302, the system receives a first input of an image in a first orientation. The first input can be received by a neural network. The image can have various orientations, such as the original orientation, a flipped left-to-right (flipLR) orientation, a flipped up-to-down (flipUD) orientation, etc., such as Figure 2 The direction of input 202-208.

[0040] At 304, the system receives a second input of an image in a second orientation. For example, if the system receives an image in a second orientation Figure 2 The image of the direction shown in 202, the second direction can be Figure 2 The second input may be received by a copy of the neural network that received the first input.

[0041] At 306, the system modifies the loss function of the neural network. The loss function is modified to include an additional term that forces the network output for the first input and the second input so that the network knows that the first output and the second output should both be the same or similar.

[0042] At 308, the system trains the neural network based on minimizing the loss function modified at 306. Thus, the system learns a symmetric convolutional neural network, in which each of the kernels in the filters of the convolutional neural network is symmetric. One advantage is that this ensures that the output is exactly the same for different rotations, and the resulting network is very efficient. Less memory space is required to store the network weights, and computational costs (e.g., real FFT) can also be reduced.

[0043] Figure 4 4 is a flow chart of a method for training an end-to-end neural network in which the kernel is symmetric, according to one embodiment. At 402, the system receives an image through the neural network. The image can be in different orientations.

[0044] Considering a linear interpolation convolution filter, for ×4 upsampling, the filter is sampled at intervals of 4 in the +ve and -ve directions, giving the following symmetric filter coefficients as shown in equation (1):

[0045] conv(xx, y)=flipLR(conv(xx, flipLR(y))) (1)

[0046] At 404, the system performs learning via a neural network. The neural network can be trained to have symmetric outputs and save 8 computations in inference in several ways. First, a fully convolutional CT-SRCNN-like architecture can be utilized, where before the first layer, the input is upsampled to the desired resolution by a symmetric traditional interpolation filter (e.g., bicubic interpolation). Using a greedy algorithm, if the result after each layer is invariant to flipping in any direction, then the result of the entire network will be invariant to flipping. A sufficient condition is that the two-dimensional (2d) plane of all kernels is symmetric. For example, consider the 2d 3×3 kernel below as in equation (2).

[0047]

[0048] These 9 parameters are constrained to have 4 degrees of freedom instead of 9, which makes the kernel symmetric in both x and y directions. Nonlinear operations following the symmetric kernel can be exploited because they are applied element-wise.

[0049] Second, a fast super resolution CNN (FSRCNN) architecture with deconvolution in the last layer can be used with upsampled feature maps of low resolution (LR) input and high resolution (HR) output. The sufficient condition for symmetry is that the deconvolution filter should be symmetric and the input should be upsampled by appropriate zero insertion. Therefore, all 2D planes of the convolution kernel and the deconvolution kernel are symmetric, which is sufficient for the entire network to be symmetric.

[0050] Third, an efficient sub-pixel CNN (ESPCNN) architecture can be utilized. A sufficient condition is that all filters before the final sub-pixel rearrangement layer are symmetric. The individual filters in the sub-pixel rearrangement layer are not symmetric, but they are polyphase implementations of symmetric filters and are sampled from the same symmetric deconvolution filters described above for FSRCNN.

[0051] In order to train the network to have symmetric filters in the 2D plane (a 3D kernel is a group of 2D filters), the system can adjust the filter implementation by using weight tying or weight sharing. For all filter coefficients belonging to the same value, for example, at all w1 positions above, the system can initialize them with similar values ​​and update them with the average gradient of the filter coefficients (the average gradient is the average of all weight gradients at the position of w1), so that after each update step, they are forced to be the same layer by layer. This should also be equivalent to averaging the filter coefficients at all positions of w1 after each update step.

[0052] The condition of a symmetric filter at each filter is sufficient, but not necessary. For example, the convolution of an asymmetric filter and its flipping in a symmetric filter results in the following equation (3).

[0053] conv([1 2],[2 1])=[2 5 2] (3)

[0054] This indicates that the network can be trained end-to-end such that its end-to-end transfer function is symmetric (i.e., the network function is invariant to flipping).

[0055] Several examples for training a neural network such that each kernel is symmetric are disclosed below with reference to W shown in equation (4).

[0056]

[0057] At 404A, the system learns by forcing weights to be tied within each kernel. The system can enforce weight tying by averaging the gradients of coordinates belonging to the same value. This training method can be evaluated using the current CT-SRCNN model for symmetric kernels for image super-resolution. Table 2 summarizes the PSNR and SSIM results for the Set14 image dataset with and without the symmetric kernel constraint.

[0058] Table 2

[0059]

[0060]

[0061] It is observed that the loss due to the symmetric kernel constraint is marginal.

[0062] At 404B, the system adds additional regularization on the weights in the loss function. The system can add regularization terms to force the condition for each kernel so that the kernel is symmetric. Note that this may be too prohibitive for deep networks. For a kernel, the system sums the regularization over these constraints to provide a symmetric filter as shown below in equations (5) and (6).

[0063] w3=w1=w7=w9; w2=w8; w4=w6 (5)

[0064]

[0065] At 404C, the system performs learning by providing weight sharing. A 2d symmetric filter can be represented as a sum of 4 filters as in Equation (7):

[0066] W+flipLR(W)+flipUD(W)+flipUD(flipLR(W)) (7)

[0067] Where, as shown in equation (8):

[0068]

[0069] The sum is as in equation (9).

[0070]

[0071] Each convolutional layer W is replaced by W+Y+Z+V. Training is done by weight sharing between the 4 filters, which is determined by the positions in the filter with the same weight value.

[0072] Alternatively, if weight sharing with permutation cannot be used, the system can force weight sharing with permutation by finding the average gradient matrix G of the gradient matrices of the rotated filters. Then, each of the rotated weight matrices is updated using the corresponding rotation of the average gradient matrix G, as shown in equation (10).

[0073]

[0074] Each layer is then invariant to rotation because its output is the sum of the responses from the different rotated filters (i.e., the response due to W+Y+Z+V (which is a symmetric filter)). As shown in equations (11), (12), (13), and (14), a symmetric 2D matrix V4 can be obtained from a random matrix.

[0075] V=rand(3) (11)

[0076] V4=V+flipLR(V)+flipUD(V)+flipLR(flipUD(V)) (12)

[0077]

[0078]

[0079] If the filter output also needs to be symmetric in the transpose direction, the system can add 4 more constraints, or equivalently, the filter can be expressed as a sum of 8 filters, as shown in equations (15) and (16).

[0080]

[0081]

[0082] In one embodiment, the system can utilize symmetric convolution to reduce the implementation complexity of the neural network. One advantage of a symmetric filter is that its FFT is real. The system can use a theorem that states that the Fourier transform of a real even function is real even, as shown in equation (17).

[0083]

[0084] Since x(n) and cos(ωn) are both real and even, their product is real and even, and the sum of their samples is also real and even, whereas x(n) is real and even and sin(ωn) is real and odd, and the sum of its samples is zero.

[0085] In one dimension (1D), 1D symmetry guarantees that its FFT coefficients are real. To have real FFT coefficients in 2D, at least V4 type symmetry (e.g., symmetry over all LR, UD, LR(UD)) is required (e.g., ignoring imaginary coefficients smaller than 1e-14 due to numerical issues).

[0086] Using the above example for V and V4, the FFT of V is complex, as shown in equation (18).

[0087]

[0088] The FFT of V4 is real, as shown in equation (19).

[0089]

[0090] The FFT of a real signal has Hermitian symmetry, which reduces computational complexity. Observe that the FFT of V4 is also 2D symmetric. Because these values ​​are also real numbers, the missing values ​​are the same, and there is no need to perform complex conjugation operations.

[0091] The FFT of V5 is real and, like V5, is bisymmetric, as shown in equation (20).

[0092]

[0093] The FFT of V5 is real and symmetric. For a 3×3 kernel, 2D symmetry requires only 4 real FFT computations, and 2D symmetry and 2D transpose symmetry require only 3 real FFT computations. However, if V is doubly symmetric, the computational gain of computing the 2D FFT increases with the kernel size. For a kernel N×N with double 2D symmetry, the number of different parameters in the spatial domain (or FFT computations in the FFT domain for each kernel) is given by Equation (21).

[0094]

[0095] For a single 2D symmetry, the number of different parameters (or required FFT calculations) is given by Equation (22).

[0096]

[0097] The convolution can be implemented in the Fourier domain A as in equation (23).

[0098] conv B=IFFT(FFT(A) ⊙FFT(B)) (23)

[0099] Therefore, by forcing the kernel to be symmetric, the complexity can be further reduced, so that only the real coefficients of the FFT need to be calculated, and due to the 2D symmetry of the FFT, only calculations need to be performed at specific locations. By reducing the implementation complexity of each convolutional layer, the implementation complexity of the deep convolutional neural network can be reduced proportionally.

[0100] In one embodiment, the system can further reduce the FFT computational complexity by constraining the circulant matrix to be symmetric. Such a block can be referred to as a block circulant and symmetric matrix. The weight matrix of the fully connected layer is constrained to be a block circulant matrix so that matrix multiplication can be performed with low complexity using fast FFT-based multiplication.

[0101] Figure 5is a schematic diagram of a block circulant matrix according to one embodiment. The block circulant matrix 500 is an unstructured weight matrix including 18 parameters. After processing 502 by a reduction ratio, the block circulant matrix 504 is constrained and includes only 6 parameters, thereby reducing computational complexity.

[0102] When computing WX (where W is a weight matrix of size N×M and X is an input feature map of size M×1), it is assumed that W consists of a small circulant matrix W of size n×n ij (where 1≤i≤l=N / n and 1≤j≤k=M / n). For simplicity, assume that N and M are integer multiples of n, where any weight matrix can be given by zero padding if necessary. X can be partitioned similarly and is easily shown as shown in equation (24):

[0103]

[0104] Each W ij are all circulant matrices, and the FFT and inverse FFT (IFFT) can then be used to calculate W ij X i , as shown in equation (25):

[0105] W ij X i =IFFT(FFT(w ij )⊙FFT(X1)), (25)

[0106] where w ij is the matrix W ij is the first column vector of , and ⊙ represents the element-wise product.

[0107] Given a vector x=[x1, x2…, x N ]. If for all 1≤i≤N, x N-i =x i , then for all 1≤i≤N, where X = [X1, X2, ...X N ] is the FFT output of x. That is, the system only needs to calculate the first half of the FFT and take the conjugate to get the other half. In addition, as shown above, because x represents the filter weight and is a real number, the DTFT is both real and symmetric. Therefore, there is no need to calculate the complex coefficients. If each circulant matrix W ij is constrained to be symmetric (i.e., so that w ij Symmetric), then the computational complexity of FFT can be reduced.

[0108] Figure 6is a block diagram of an electronic device 601 in a network environment 600 according to one embodiment. Figure 6 , the electronic device 601 in the network environment 600 can communicate with the electronic device 602 via the first network 698 (e.g., a short-range wireless communication network), or communicate with the electronic device 604 or the server 608 via the second network 699 (e.g., a long-range wireless communication network). The electronic device 601 can communicate with the electronic device 604 via the server 608. The electronic device 601 may include a processor 620, a memory 630, an input device 650, a sound output device 656, a display device 660, an audio module 670, a sensor module 676, an interface 677, a haptic module 679, a camera module 680, a power management module 688, a battery 689, a communication module 690, a subscriber identification module (SIM) 696, or an antenna module 697. In one embodiment, at least one of these components (e.g., the display device 660 or the camera module 680) may be omitted from the electronic device 601, or one or more other components may be added to the electronic device 601. In one embodiment, some components may be implemented as a single integrated circuit (IC). For example, the sensor module 676 (eg, a fingerprint sensor, an iris sensor, or an illumination sensor) may be embedded in the display device 660 (eg, a display).

[0109] The processor 620 may execute, for example, software (e.g., program 640) to control at least one other component (e.g., hardware or software component) of the electronic device 601 coupled to the processor 620, and may perform various data processing or calculations. As at least part of the data processing or calculation, the processor 620 may load commands or data received from another component (e.g., sensor module 676 or communication module 690) into the volatile memory 632, process the commands or data stored in the volatile memory 632, and store the resulting data in the non-volatile memory 634. The processor 620 may include a main processor 621 (e.g., a central processing unit (CPU) or an application processor (AP)) and an auxiliary processor 622 (e.g., a graphics processing unit (GPU), an image signal processor (ISP), a sensor hub processor, or a communication processor (CP)), the auxiliary processor 622 operating independently of the main processor 621 or in conjunction with the main processor 621. Additionally or alternatively, the auxiliary processor 622 may be adapted to consume less power than the main processor 621, or to perform specific functions. The auxiliary processor 622 may be implemented separate from the main processor 621 or as part of the main processor 621.

[0110] The auxiliary processor 622 may control at least some of the functions or states related to at least one component (e.g., the display device 660, the sensor module 676, or the communication module 690) among the components of the electronic device 601, instead of the main processor 621 while the main processor 621 is in an inactive (e.g., sleep) state, or control together with the main processor 621 while the main processor 621 is in an active state (e.g., executing an application). According to one embodiment, the auxiliary processor 622 (e.g., an image signal processor or a communication processor) may be implemented as part of another component (e.g., the camera module 680 or the communication module 690) that is functionally related to the auxiliary processor 622.

[0111] The memory 630 may store various data used by at least one component of the electronic device 601 (e.g., the processor 620 or the sensor module 676). The various data may include, for example, input data or output data of software (e.g., the program 640) and commands related thereto. The memory 630 may include a volatile memory 632 or a non-volatile memory 634.

[0112] The program 640 may be stored as software in the memory 630 and may include, for example, an operating system (OS) 642 , middleware 644 , or applications 646 .

[0113] The input device 650 may receive commands or data from outside the electronic device 601 (eg, a user) to be used by other components of the electronic device 601 (eg, the processor 620). The input device 650 may include, for example, a microphone, a mouse, or a keyboard.

[0114] The sound output device 656 can output sound signals to the outside of the electronic device 601. The sound output device 656 may include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as playing multimedia or recording, while the receiver can be used for receiving incoming calls. According to one embodiment, the receiver can be implemented as being separated from the speaker or being a part of the speaker.

[0115] The display device 660 can visually provide information to the outside of the electronic device 601 (e.g., a user). The display device 660 may include, for example, a display, a holographic device, or a projector, and a control circuit that controls a corresponding one of the display, the holographic device, and the projector. According to one embodiment, the display device 660 may include a touch circuit suitable for detecting a touch, or a sensor circuit suitable for measuring the strength of a force caused by a touch (e.g., a pressure sensor).

[0116] The audio module 670 can convert sound into an electrical signal, and vice versa. According to one embodiment, the audio module 670 can obtain sound via the input device 650, or output sound via the sound output device 656 or via headphones of an external electronic device 602 directly (e.g., wired) or wirelessly coupled to the electronic device 601.

[0117] The sensor module 676 can detect the operating state (e.g., power or temperature) of the electronic device 601 or the environmental state (e.g., the state of the user) outside the electronic device 601, and then generate an electrical signal or data value corresponding to the detected state. The sensor module 676 can include, for example, a posture sensor, a gyroscope sensor, an atmospheric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illumination sensor.

[0118] The interface 677 may support one or more designated protocols for directly (e.g., wiredly) or wirelessly coupling the electronic device 601 with the external electronic device 602. According to one embodiment, the interface 677 may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB), a secure digital (SD) card interface, or an audio interface.

[0119] The connection terminal 678 may include a connector through which the electronic device 601 can be physically connected to the external electronic device 602. According to one embodiment, the connection terminal 678 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0120] The haptic module 679 may convert electrical signals into mechanical stimulation (eg, vibration or motion) or electrical stimulation, which the user may recognize through tactile or kinesthetic sense. According to one embodiment, the haptic module 679 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.

[0121] The camera module 680 may capture still images or moving images. According to one embodiment, the camera module 680 may include one or more lenses, image sensors, image signal processors, or flashes.

[0122] The power management module 688 may manage power supplied to the electronic device 601. The power management module 688 may be implemented as, for example, at least a portion of a power management integrated circuit (PMIC).

[0123] The battery 689 may supply power to at least one component of the electronic device 601. According to one embodiment, the battery 689 may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0124] The communication module 690 can support establishing a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device 601 and an external electronic device (e.g., electronic device 602, electronic device 604, or server 608), and perform communication via the established communication channel. The communication module 690 may include one or more communication processors that operate independently of the processor 620 (e.g., AP) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module 690 may include a wireless communication module 692 (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module 694 (e.g., a local area network (LAN) communication module or a power line communication (PLC) module). Corresponding ones of these communication modules can communicate with external electronic devices via a first network 698 (e.g., a short-range communication network such as Bluetooth, wireless-fidelity (Wi-Fi) direct, or an Infrared Data Association (IrDA) standard) or a second network 699 (e.g., a long-range communication network such as a cellular network, the Internet, or a computer network (e.g., a LAN or a wide area network (WAN))). These different types of communication modules can be implemented as a single component (e.g., a single IC), or can be implemented as multiple components separated from each other (e.g., multiple ICs). The wireless communication module 692 can use user information (e.g., international mobile subscriber identity (IMSI)) stored in the user identification module 696 to identify and authenticate the electronic device 601 in a communication network such as the first network 698 or the second network 699.

[0125] The antenna module 697 can transmit or receive signals or power to or from the outside of the electronic device 601 (e.g., an external electronic device). According to one embodiment, the antenna module 697 may include one or more antennas, and at least one antenna suitable for the communication scheme used in the communication network (such as the first network 698 or the second network 699) can be selected from among them, for example, by the communication module 690 (e.g., the wireless communication module 692). Signals or power can then be transmitted or received between the communication module 690 and the external electronic device via the selected at least one antenna.

[0126] At least some of the above components can be coupled to each other via an inter-peripheral communication scheme (e.g., a bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industry processor interface (MIPI)) and communicate signals (e.g., commands or data) therebetween.

[0127] According to one embodiment, commands or data can be sent or received between electronic device 601 and external electronic device 604 via server 608 coupled to second network 699. Each of electronic devices 602 and 604 can be of the same or different type as electronic device 601. All or some operations to be performed at electronic device 601 can be performed at one or more of external electronic devices 602, 604, or 608. For example, if electronic device 601 is to automatically perform a function or service, or in response to a request from a user or another device, electronic device 601 can request one or more external electronic devices to perform at least a portion of the function or service instead of, or in addition to, performing the function or service. The one or more external electronic devices that receive the request can perform at least a portion of the requested function or service, or additional functions or services related to the request, and pass the results of the execution to electronic device 601. Electronic device 601 can provide the results (with or without further processing) as at least part of a response to the request. To this end, for example, cloud computing, distributed computing, or client-server computing technologies can be used.

[0128] One embodiment may be implemented as software (e.g., program 640) comprising one or more instructions stored in a storage medium (e.g., internal memory 636 or external memory 638) readable by a machine (e.g., electronic device 601). For example, a processor of electronic device 601 may call at least one of the one or more instructions stored in the storage medium and execute it with or without one or more other components under the control of the processor. Thus, the machine may be operated to perform at least one function in accordance with the at least one instruction called. The one or more instructions may include compiler-generated code or interpreter-executable code. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but the term does not distinguish between cases where data is semi-permanently stored in the storage medium and cases where data is temporarily stored in the storage medium.

[0129] According to one embodiment, the method of the present disclosure may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CD-ROM)) or via an app store (e.g., an Android store). TM (Play Store TM )) online distribution (e.g., download or upload), or directly between two user devices (e.g., smartphones). If distributed online, at least a portion of the computer program product may be temporarily generated or at least temporarily stored in a machine-readable storage medium, such as a memory of a manufacturer's server, a server of an application store, or a relay server.

[0130] According to one embodiment, each component of the above-mentioned components (e.g., a module or a program) may include a single entity or multiple entities. One or more of the above-mentioned components may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., modules or programs) may be integrated into a single component. In this case, the integrated component may still perform one or more functions of each of the multiple components in the same or similar manner as they were performed by a corresponding one of the multiple components before integration. The operations performed by a module, a program or another component may be performed sequentially, in parallel, repeatedly or heuristically, or one or more operations may be performed in a different order or omitted, or one or more other operations may be added.

[0131] Although certain embodiments of the present disclosure have been described in the detailed description of the present disclosure, the present disclosure can be modified in various forms without departing from the scope of the present disclosure. Therefore, the scope of the present disclosure should not be determined based solely on the described embodiments, but rather on the appended claims and their equivalents.

Claims

1. A method for providing a rotationally invariant neural network, comprising: receiving a first input of an image in a first orientation; training the kernel to be symmetric so that an output corresponding to the first input is the same as an output corresponding to a second input of the image in a second orientation; and By training multiple kernels in different layers to be symmetrical, the convolutional neural network is trained to be symmetrical. The second orientation is a flipped version of the image relative to the first orientation of the image.

2. The method according to claim 1, wherein Training the kernel to be symmetric also includes forcing weight tying within the kernel.

3. The method according to claim 2, wherein: Enforcing weight tying also involves averaging the gradients of coordinates belonging to the same value.

4. The method according to claim 1, wherein Training the kernel to be symmetric also includes adding regularization with respect to the weights in the loss function associated with the output.

5. The method according to claim 1, wherein Training the kernel to be symmetric also includes providing weight sharing among multiple filters.

6. The method according to claim 5, wherein: Sharing weights among a plurality of filters also includes expressing the kernel as a sum of the plurality of filters.

7. The method according to claim 5, wherein: Weights are shared among the plurality of filters based on an average gradient matrix of the gradient matrices of the plurality of filters.

8. The method of claim 1, further comprising applying the trained symmetric kernel to a block-circular weight matrix.

9. A system for providing a rotationally invariant neural network, comprising: Memory; and The processor is configured to: receiving a first input of an image in a first orientation; training the kernel to be symmetric so that an output corresponding to the first input is the same as an output corresponding to a second input of the image in a second orientation; and By training multiple kernels in different layers to be symmetrical, the convolutional neural network is trained to be symmetrical. The second orientation is a flipped version of the image relative to the first orientation of the image.

10. The system according to claim 9, wherein: The processor is further configured to train the kernel to be symmetric by causing weight tying within the kernel.

11. The system according to claim 10, wherein: The processor is further configured to facilitate weight binding by averaging gradients belonging to coordinates of the same value.

12. The system according to claim 9, wherein: The processor is further configured to train the kernel to be symmetric by adding regularization with respect to the weights in a loss function associated with the output.

13. The system according to claim 9, wherein: The processor is further configured to train the kernel to be symmetric by providing weight sharing among a plurality of filters.

14. The system according to claim 13, wherein: The weight sharing among the plurality of filters further comprises expressing the kernel as a sum of the plurality of filters.

15. The system according to claim 13, wherein: Weights are shared among the plurality of filters based on an average gradient matrix of the gradient matrices of the plurality of filters.

16. The system according to claim 9, wherein: The processor is further configured to apply the trained symmetric kernel to the block-circular weight matrix.

17. A method for providing a rotationally invariant neural network, comprising: Receiving, by a neural network, a first input of an image at a first orientation; generating a first loss function based on a first output associated with the first input; receiving, by the neural network, a second input of the image in a second orientation; generating a second loss function based on a second output associated with the second input; training the neural network to minimize the sum of the first loss function and the second loss function; and By training multiple kernels in different layers to be symmetrical, the convolutional neural network is trained to be symmetrical. The second orientation is a flipped version of the image relative to the first orientation of the image.

18. The method of claim 17, further comprising modifying the second loss function to have an additional term corresponding to an added loss of the second input from the image.

Citation Information

Patent Citations

  • System and method for designing efficient super resolution deep convolutional neural networks by cascade network training, cascade network trimming, and dilated convolutions

    US10803378B2

  • System and method for designing efficient super resolution deep convolutional neural networks by cascade network training, cascade network trimming, and dilated convolutions

    US11354577B2

  • Training convolutional neural networks on graphics processing units

    US20070047802A1

  • Method and apparatus for image processing

    US9799098B2