Neural network device for retrieving images and method of operation thereof
Patent Information
- Application Number
- CN202110347114.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-04
- Filing Date
- 2021-03-31
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2041-03-31
AI Technical Summary
[0005]然而,当数据库中的标记图像不足以检索图像时,基于卷积神经网络的散列函数可能无法容易地获得期望的性能,因此,期望一种将未标记图像与标记图像一起使用的方法
Smart Images

Figure CN113496277B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to Korean Patent Application No. 10-2020-0041076, filed on April 3, 2020, with the Korean Intellectual Property Office, and Korean Patent Application No. 10-2021-0016274, filed on February 4, 2021, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The embodiments of the present invention relate to a neural network device and its operating method, and more specifically, to a neural network device and its operating method for retrieving images. Background Technology
[0004] The use of binary hash codes obtained from hash functions has demonstrated remarkable performance in retrieving images with minimal storage space and high retrieval rates. With the development of deep learning methods, hash functions based on convolutional neural networks have been proposed.
[0005] However, when the number of labeled images in the database is insufficient to retrieve images, hash functions based on convolutional neural networks may not easily achieve the desired performance. Therefore, a method that can be used together with labeled images is desired. Summary of the Invention
[0006] An embodiment of the present invention provides a neural network and a method for operating the same, which can robustly improve image retrieval performance by using unlabeled images in addition to labeled images.
[0007] According to an embodiment of the present invention, a neural network apparatus is provided, comprising: a processor that performs operations to train a neural network; a feature extraction module that extracts unlabeled feature vectors corresponding to unlabeled images and labeled feature vectors corresponding to labeled images; and a classifier that classifies a category of a query image, wherein the processor performs a first learning operation with respect to a plurality of codebooks by using labeled feature vectors and performs a second learning operation with respect to the plurality of codebooks by optimizing entropy based on both labeled and unlabeled feature vectors.
[0008] According to other embodiments of the present invention, a method for operating a neural network is provided, the method comprising the steps of: extracting unlabeled feature vectors corresponding to unlabeled images and labeled feature vectors corresponding to labeled images; performing a first learning about a plurality of codebooks by using the labeled feature vectors; and performing a second learning about the plurality of codebooks by optimizing the entropy based on the labeled feature vectors and the unlabeled feature vectors.
[0009] According to other embodiments of the present invention, a method for operating a neural network device is provided, the method comprising the steps of: receiving a plurality of biometric images including at least one labeled image and at least one unlabeled image; extracting feature vectors corresponding to the plurality of biometric images; performing learning about a plurality of codebooks using the feature vectors; receiving a query image and calculating a distance value based on the plurality of biometric images; estimating an expected value of a category for classifying the plurality of biometric images based on the calculated distance value; and performing biometric authentication when the maximum value among the expected values exceeds a preset threshold. Attached Figure Description
[0010] Figure 1 This is a block diagram of a neural network device according to an embodiment of the present invention.
[0011] Figure 2 The processing of the feature extraction module according to an embodiment is illustrated.
[0012] Figure 3A A neural network device according to an embodiment of the present invention is shown.
[0013] Figure 3B An example is shown illustrating the result of performing product quantization in the feature space according to an embodiment of the present invention.
[0014] Figure 3C The relationship between subvectors, codebooks, and quantized subvectors is illustrated in an embodiment of the present invention.
[0015] Figure 3D An example of a learning process using a classifier, according to an embodiment of the present invention, is shown.
[0016] Figure 4 This is a flowchart illustrating the operation of a neural network device according to an embodiment of the present invention.
[0017] Figure 5 This is a flowchart illustrating the learning process of a neural network according to an embodiment of the present invention, which shows in detail... Figure 4 Operation S110 in the middle.
[0018] Figure 6 This is a flowchart illustrating the entropy optimization of a subspace according to an embodiment of the present invention, which is shown in detail in the figure. Figure 5 Operation S230 in the middle.
[0019] Figure 7 An improved image retrieval result is shown according to an embodiment of the concept of the present invention.
[0020] Figure 8 This is a flowchart of a neural network operation according to an embodiment of the present invention. Detailed Implementation
[0021] In the following, embodiments of the inventive concept will be described in detail with reference to the accompanying drawings.
[0022] Figure 1 This is a block diagram of a neural network device according to an embodiment of the present invention.
[0023] Reference Figure 1 The neural network device 10 according to the embodiment includes a processor 100, a random access memory (RAM) 200, a storage device 300, and a camera 400.
[0024] The neural network device 10 according to the embodiment can analyze input data and extract useful information, and can generate output data based on the extracted information. The input data may be image data including various objects, and the output data may be image retrieval results, such as the classification result of the query image, or the image most similar to the query image.
[0025] In this embodiment, the image data may include labeled images and unlabeled images. Labeled images include information about the classification of objects shown in the image. Unlabeled images do not include information about the classification of objects shown in the image.
[0026] For example, when an image includes the shape of an animal such as a cat, the labeled image includes information about categories used to classify objects that have the shape shown in the image and are mapped to that image, such as animals with higher category information and cats with lower category information. For example, when an image does not include information about categories used to classify objects that have the shape shown in the image, the image is an unlabeled image.
[0027] According to various embodiments, the neural network device 10 can be implemented as a personal computer (PC), an Internet of Things (IoT) device, or a portable electronic device. Portable electronic devices can be included in various devices such as laptops, mobile phones, smartphones, tablet PCs, personal digital assistants (PDAs), enterprise digital assistants (EDAs), digital still cameras, digital video cameras, audio devices, portable multimedia players (PMPs), personal navigation devices (PNDs), MP3 players, handheld game consoles, e-books, or wearable devices.
[0028] According to various embodiments, processor 100 controls neural network device 10. For example, processor 100 trains the neural network based on learning images and performs image retrieval about a query image using the trained neural network.
[0029] According to various embodiments, processor 100 includes a central processing unit (CPU) 110 and a neural processing unit (NPU) 120. CPU 110 controls all operations of the neural network device 100. CPU 110 may be a single-core processor or a multi-core processor. CPU 110 can process or execute programs or data stored in storage device 300. For example, CPU 110 controls the functionality of NPU 120 by executing programs or modules stored in storage device 300.
[0030] According to various embodiments, the NPU 120 generates a neural network, trains or learns the neural network, performs operations based on the training data, and generates information signals based on the execution results, or retrains the neural network.
[0031] According to various embodiments, there are various types of neural network models, such as convolutional neural networks (CNNs) (such as GoogleNet, AlexNet, or VGG NewWork), region representation networks (R-CNNs) with convolutional neural networks, region extraction networks (RPNs), recurrent neural networks (RNNs), stack-based deep neural networks (S-DNNs), state-space dynamic neural networks (S-SDNs), deconvolutional networks, deep belief networks (DBNs), restricted Boltzmann machines (RBMs), fully convolutional networks, long short-term memory (LSTM) networks, or classification networks, etc., and the embodiments are not limited to the models mentioned above.
[0032] According to various embodiments, the NPU 120 also includes additional memory for storing programs corresponding to the neural network model. The NPU 120 also includes additional coprocessor blocks for processing operations required to drive the neural network. For example, the additional coprocessor blocks may include a graphics processing unit (GPU) or an accelerator, such as a floating-point accelerator, for rapidly performing specific operations.
[0033] According to various embodiments, RAM 200 temporarily stores programs, data, or instructions. For example, RAM 200 may temporarily load programs or data stored in storage device 300 to control or boot CPU 110. For example, RAM 200 may be dynamic RAM (DRAM), static RAM (SRAM), or synchronous DRAM (SDRAM).
[0034] According to various embodiments, the storage device 300 can store an operating system (OS), various programs, and various types of data. For example, the storage device 300 corresponds to non-volatile memory. For example, the storage device 300 can be one or more of read-only memory (ROM), flash memory, phase-change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), and ferroelectric RAM (FRAM). According to embodiments, the storage device 300 can be implemented as a hard disk drive (HDD) or a solid-state drive (SSD), etc.
[0035] According to various embodiments, storage device 300 stores feature extraction module 302, hash table 304, classifier 306, and lookup table (LUT) 308. Feature extraction module 302 can generate feature vectors from the image. Classifier 306 can classify feature vectors using labels and includes information about the prototype vector representing each label. Hash table 304 maps the feature vectors obtained from feature extraction module 302 to binary hash codes. Mapping to binary hash codes includes: dividing the feature vectors into at least one feature sub-vector based on product quantization, and replacing each of the at least one feature sub-vector with a binary index value that refers to the codeword with the shortest distance in the codewords of hash table 304. LUT 308 can perform distance calculations on the query image using the learned codebook and store the resulting values. LUT 308 is used for classification of the query image.
[0036] Figure 2 The processing operations of the feature extraction module are shown.
[0037] Reference Figure 2 According to various embodiments, a neural network NN includes multiple layers L1 to Ln. Each of the multiple layers L1 to Ln can be a linear layer or a non-linear layer, and according to embodiments, at least one linear layer and at least one non-linear layer can be combined with each other and referred to as a layer. For example, a linear layer can be a convolutional layer or a fully connected layer, and a non-linear layer can be one of a sampling layer, a pooling layer, and an activation layer.
[0038] According to an embodiment, the first layer L1 is a convolutional layer, and the second layer L2 is a sampling layer. The neural network NN also includes activation layers, and may further include at least one layer that performs another operation.
[0039] According to an embodiment, each of the multiple layers receives input image data or an input feature map generated from the previous layer, and generates an output feature map by performing operations on the input feature map. In this case, the feature map is data representing various properties of the input data.
[0040] According to an embodiment, feature maps FM1, FM2, and FM3 have, for example, the form of a two-dimensional matrix or a three-dimensional matrix. Each of feature maps FM1, FM2, and FM3 has a width W (or column), a height H (or row), and a depth D, which correspond to the x-axis, y-axis, and z-axis in a coordinate system, respectively. In this case, the depth can be referred to as the number of channels.
[0041] According to an embodiment, the first layer L1 generates the second feature map FM2 by convolving a first feature map FM1 with a weight map WM. The weight map WM filters the first feature map FM1 and can be referred to as a filter or kernel. The depth of the weight map WM (i.e., the number of channels in the weight map WM) is the same as the depth of the feature map FM1 (i.e., the number of channels in the first feature map FM1), and the same channels in the weight map WM and the first feature map FM1 are convolved with each other. The weight map WM is shifted by traversing the first feature map FM1 as a sliding window. The shift amount can be referred to as the "stride length" or "step size". In each shift, each weight in the weight map WM is multiplied by all the feature values in the region overlapping with the first feature map FM1 and summed. The channels of the second feature map FM2 are generated when the first feature map FM1 and the weight map WM are convolved.
[0042] although Figure 2 Only one weight map (weight map WM) is shown, but in reality, multiple weight maps can be convolved with the first feature map FM1, and multiple channels of the second feature map FM2 can be generated. In other words, the number of channels in the second feature map FM2 corresponds to the number of weight maps.
[0043] According to an embodiment, the second layer L2 generates the third feature map FM3 by changing the spatial size of the second feature map FM2. For example, the second layer L2 is a sampling layer. The second layer L2 performs upsampling or downsampling and selects some data from the second feature map FM2. For example, a two-dimensional window WD is shifted on the second feature map FM2 in units of the size of the window WD (such as a 4×4 matrix), and values at specific positions (such as the first row or the first column) are selected in the region overlapping with the window WD. The second layer L2 outputs the selected data as the data for the third feature map FM3. In another example, the second layer L2 is a pooling layer. In this case, the second layer L2 selects the maximum pooling value or the average pooling value of each feature value in the region of the second feature map FM2 that overlaps with the window WD. The second layer L2 outputs the selected data as the data for the third feature map FM3.
[0044] By doing so, according to an embodiment, a third feature map FM3 is generated having a spatial size different from that of the second feature map FM2. The third feature map FM3 has the same number of channels as the second feature map FM2. According to an embodiment of the invention, the sampling layer operates at a higher rate than the pooling layer, and the sampling layer can improve the quality of the output image, such as peak signal-to-noise ratio (PSNR). For example, since the pooling layer calculates a maximum or average value, the operations performed by the pooling layer take longer than those performed by the sampling layer.
[0045] According to the embodiment, the second layer L2 is not limited to a sampling layer or a pooling layer. That is, the second layer L2 can be a convolutional layer similar to the first layer L1. The second layer L2 can generate a third feature map FM3 by convolving the second feature map FM2 with the weight map. In this case, the weight map used by the second layer L2 is different from the weight map WM used by the first layer L1.
[0046] According to an embodiment, after multiple layers including a first layer L1 and a second layer L2, an Nth feature map is generated in the Nth layer. The Nth feature map is input to a reconstruction layer at the back end of the neural network NN, from which output data is output. The reconstruction layer generates an output image based on the Nth feature map. Additionally, the reconstruction layer receives multiple feature maps such as a first feature map FM1, a second feature map FM2, and the Nth feature map, and generates an output image based on these multiple feature maps.
[0047] According to an embodiment, the third layer L3 can generate a feature vector (FV). The generated FV can be used to determine the category (CL) of the input data by combining the features of the third feature map FM3. Additionally, the third layer L3 generates a recognition signal (REC) corresponding to the category. For example, the input data can be still image data or video frames. In this case, the third layer L3 extracts the category corresponding to the object in the image represented by the input data based on the third feature map FM3, thereby identifying the object and generating a recognition signal (REC) corresponding to the identified object.
[0048] Figure 3A A neural network device according to an embodiment of the present invention is shown.
[0049] Reference Figure 3A According to an embodiment, the feature extraction module 302 extracts compact feature vectors using images as input data. These feature vectors are used to train a neural network. The neural network learns optimal codewords for classification based on the feature vectors of labeled and unlabeled images, and can retrieve query images using a codebook consisting of groups of learned codewords. Feature vectors extracted from similar images have short distances to each other.
[0050] According to various embodiments, Figure 3AThe objective function used in the convolutional neural network shown as follows:
[0051] Equation 1:
[0052]
[0053] here, It is an objective function of N pairs of product quantizations, and it corresponds to an objective function that is transformed from a metric learning-based objective function into an objective function suitable for image retrieval. The objective function corresponding to the labeled image, and The objective function corresponds to the unlabeled image. B corresponds to the number of learning images, and λ1 and λ2 refer to the learning weights for labeled and unlabeled images, respectively. The relative values of the learning weights indicate which image, labeled or unlabeled, is favored more during the learning process. For example, even when the neural network learns using both unlabeled and labeled images, the neural network device 10 can configure the learning to be more influenced by labeled images by setting λ1 to be greater than λ2. For example, for λ2 greater than λ1, the neural network device 10 performs a small amount of optimal codeword learning for the prototype vector using labeled images and a large amount of optimal codeword learning using unlabeled images.
[0054] According to various embodiments, the neural network device 10 performs neural network learning to generate a database for retrieving images. The neural network receives labeled and unlabeled images as input images. The feature extraction module 302 extracts feature vectors from each of the labeled and unlabeled images. In the following text, for ease of explanation, the feature vector of the labeled image will be referred to as the labeled feature vector, and the feature vector of the unlabeled image will be referred to as the unlabeled feature vector.
[0055] According to various embodiments, the neural network performs inner normalization for each of the labeled and unlabeled feature vectors. The inner normalization divides each of the labeled and unlabeled feature vectors into multiple sub-vectors.
[0056] Figure 3B An example of the result of performing product quantization in the feature space according to an embodiment of the present invention is shown; Figure 3C The relationship between subvectors, codebooks, and quantized subvectors is illustrated in an embodiment of the present invention.
[0057] Reference Figure 3B According to various embodiments, circles formed by solid lines correspond to feature vectors, and circles formed by dashed lines correspond to quantized vectors. According to various embodiments, Figure 1 The hash table in the middle receives f Land f U As input, and for f L and f U Classify the images to easily distinguish their categories. L and f U These represent feature vectors based on labeled images and feature vectors based on unlabeled images, respectively.
[0058] According to various embodiments, feature vectors are not arranged in the feature space before product quantization is performed on hash table 304, so that feature vectors of the same class are close to each other. Figure 3B The upper part of the image represents the feature space before product quantization is performed. x1 + x4 + q4 + and q1 + Constitutes the first category, x2 - q2 - x3 - and q3 - This constitutes the second category. (See reference...) Figure 3B By performing product quantization using hash table 304, the quantized feature vectors in the first and second categories can be arranged close to each other.
[0059] Reference Figure 3A and Figure 3B Hash table 304 performs learning for product quantization based on the following equation.
[0060] Equation 2:
[0061]
[0062] here, The objective function is the cross-entropy function, S. b Y is the cosine similarity between the b-th eigenvector and all quantized vectors. b It is the similarity between the b-th marker in a batch and all other markers.
[0063] Reference Figure 3A and Figure 3C According to various embodiments, the first to fourth sub-vectors are shown on the left. The sub-vectors are obtained by dividing the k-dimensional feature vectors received by the neural network device 10 and extracted by the feature extraction module 302 at regular intervals.
[0064] According to various embodiments, the first to fourth quantized subvectors q are calculated according to the following equations. m .
[0065] Equation 3:
[0066]
[0067] Here, x m Corresponding to the m-th sub-vector, z mk This corresponds to the k-th codeword in the m-th subspace of the codebook. Equation 3 shows that the quantized subvector is calculated based on the cosine similarity between the subvector and each codeword in the corresponding subspace. In other words, the quantized subvector is quantized in the following way: it reflects a large number of codewords with similar characteristics among multiple codewords in the corresponding subspace, and a small number of codewords with different data characteristics. For example, refer to... Figure 3C By calculating the first subvector #1 and multiple codewords C in the first subspace corresponding to the first subvector #1 11 C 12 C 13 and C 14 The cosine similarity between them is used to obtain the first quantized subvector #1.
[0068] According to various embodiments, hash table 304 obtains the binary hash code of all feature vectors by replacing or representing subvectors with binary index values that represent the codewords with the shortest distance in the subspace corresponding to each divided subvector. For example, refer to Figure 3C The first to fourth subvectors are respectively associated with the C in the codewords of their respective subspaces. 11 C 24 C 32 and C 43 It has high cosine similarity. In this case, hash table 304 stores the binary values representing the codewords with the highest cosine similarity in each subspace, instead of storing the real values of the first to fourth quantized feature vectors. For example, hash table 304 stores the feature vectors as (00, 11, 01, 10).
[0069] Figure 3D An example of a learning process using a classifier, according to an embodiment of the present invention, is shown.
[0070] Reference Figure 3A and Figure 3D According to various embodiments, classifier 306 performs neural network learning about each of the labeled and unlabeled images.
[0071] According to various embodiments, classifier 306 performs learning about the neural network by classifying labels about labeled images based on the following objective function.
[0072] Equation 4:
[0073]
[0074] Here, β is the scaling factor, y is the label that includes the feature vector of the labeled image, and This represents the probability that a labeled image belongs to a specific category. Classifier 306 obtains sub-prototype vectors by learning codewords from labeled images corresponding to each category. (See reference...) Figure 3D In (a) and (b), classifier 306 uses Equation 4 to learn codewords that minimize the cross-entropy of the labeled images. Thus, in (a), the subvectors of the labeled images are arranged to have high entropy on a two-dimensional hyperspherical manifold, but in (b), the subvectors of the labeled images are rearranged based on the subprototype vectors such that the first class is close to each other and the second class is close to each other.
[0075] Reference Figure 3D In (c) and (d) of various embodiments, the classifier 306 performs learning about the neural network based on the feature vectors of the unlabeled image based on the following objective function.
[0076] Equation 5:
[0077]
[0078] Here, P U It represents the probability that an image belongs to a specific category, and therefore it is derived over all M subvectors and over N. c The value is calculated per marker. Additionally, β is a scaling factor.
[0079] That is, according to various embodiments, classifier 306 adjusts the sub-prototype vectors such that for sub-vectors of unlabeled images distributed on the manifold, the entropy of the sub-prototype vectors is maximized, and codewords that minimize the entropy of unlabeled images distributed in the subspace are learned based on the rearranged sub-prototype vectors.
[0080] Figure 4 This is a flowchart illustrating the operation of a neural network device 10 according to an embodiment of the present invention.
[0081] Reference Figure 4 According to various embodiments, in operation S110, the neural network device 10 trains the neural network using labeled and unlabeled images. Specifically, the neural network device 10 constructs an optimal codebook for image retrieval through the following steps: learning codewords that minimize the cross-entropy of the feature vectors of the labeled image and determine the sub-prototype vector; modifying the determined sub-prototype vector by learning codewords that maximize the entropy of the sub-prototype vector; and learning codewords that minimize the entropy of the unlabeled feature vector based on the modified sub-prototype vector. Reference will be made below. Figure 5 Describe the details of operation S110.
[0082] In operation S120, according to various embodiments, the neural network device 10 obtains a binary hash code by performing binary hashing on multiple codebooks. The binary hash code is obtained by replacing the feature vector with a binary value representing the codeword in the subspace that has the highest cosine similarity to each subvector.
[0083] In operation S130, according to various embodiments, the neural network device 10 performs a distance comparison between the feature vector of the query image and its binary hash code. The neural network device 10 extracts the feature vector for image retrieval. The neural network device 10 extracts the feature vector of the received query image using the feature extraction module 302. The neural network device 10 generates lookup tables that store the distance differences between multiple codebooks generated during the training phase and the extracted feature vector of the query image. The size of the lookup tables is the same as the size of the codebooks. Subsequently, the neural network device 10 sums the values represented by each binary hash code in each subspace of each generated lookup table and calculates the distance between the query image and each training image. For example, when training the neural network device 10 with five hundred images, five hundred codebooks are generated and stored. The neural network device 10 generates lookup tables including the distance differences between the feature vector and these five hundred codebooks, and extracts only the values represented by each binary hash code corresponding to each of the codebooks and sums them to obtain the distance between the query image and these five hundred training images.
[0084] In operation S140, according to various embodiments, the neural network device 10 determines the category of the query image as the category of the image with the shortest distance to the result value calculated in operation S130. The training images with the shortest distance include the characteristics most similar to the characteristics of the query image. Therefore, the neural network device 10 retrieves and classifies the query image by determining the category corresponding to the training image with the shortest distance to the category of the query image.
[0085] Figure 5 This is a flowchart illustrating the learning process of a neural network according to an embodiment of the present invention, which shows in detail... Figure 4 Operation S110 in the middle.
[0086] Reference Figure 5 According to various embodiments, in operation S210, the neural network device 10 receives labeled and unlabeled images and extracts feature vectors. Both labeled and unlabeled images are included in the training images. The feature extraction module 302 of the neural network device 10 obtains labeled feature vectors and unlabeled feature vectors from the received labeled and unlabeled images.
[0087] In operation S220, according to various embodiments, the neural network device 10 performs learning about multiple codebooks by using labeled feature vectors. That is, the neural network device 10 learns codebooks using feature vectors of labeled images. To obtain codewords that minimize the cross-entropy between labeled feature vectors, the neural network device 10 learns codewords through backpropagation. When the learned codebook is used to quantize the feature vectors of labeled images, the quantized labeled feature vectors can be tightly arranged on a two-dimensional manifold.
[0088] In operation S230, according to various embodiments, the neural network device 10 performs entropy optimization on the subspace by using labeled feature vectors and unlabeled feature vectors to learn multiple codebooks. Entropy optimization of the subspace is achieved by performing both entropy maximization of the sub-prototype vectors and entropy minimization of the unlabeled feature vectors. That is, the neural network device 10 performs entropy maximization of the sub-prototype vectors by changing the sub-prototype vector values obtained in operation S220 based on the unlabeled feature vectors, and learns multiple codebooks to minimize the entropy of the unlabeled feature vectors on the manifold of the subspace. (Refer to later...) Figure 6 Describe its details.
[0089] Figure 6 This is a flowchart illustrating the entropy optimization of a subspace according to an embodiment of the present invention, which is shown in detail in the figure. Figure 5 Operation S230 in the middle.
[0090] Reference Figure 6 According to various embodiments, in operation S310, the neural network device 10 maximizes the entropy value of the sub-prototype vector by changing the value of the sub-prototype vector based on the unlabeled feature vector. For example... Figure 5 As shown, the neural network device 10 determines the sub-prototype vector for each subspace based on the values of the labeled feature vectors. The determined sub-prototype vectors correspond to... Figure 3D The sub-prototype vector is shown in (b). In one embodiment, the sub-prototype vector is determined based on one of the mean, median, and mode of the labeled feature vectors. The neural network device 10 maximizes the entropy of the sub-prototype vector by changing the value of the sub-prototype vector using unlabeled feature vectors distributed in a subspace. That is, the sub-prototype vector is shifted to a region on the manifold in which multiple unlabeled feature vectors are distributed.
[0091] In operation S320, according to various embodiments, the neural network device 10 learns multiple codebooks based on the modified sub-prototype vector values to minimize the entropy of the unlabeled feature vectors. In other words, the neural network device 10 learns multiple codebooks to be adjacent to the adjusted sub-prototype vectors on the manifold. In other words, when the feature vectors of the unlabeled image are quantized using the codebooks before learning, these feature vectors are separated from the labeled feature vectors in similar categories on the manifold. However, when the feature vectors of the unlabeled image are quantized using the learned codebooks, these feature vectors are shifted to be closer on the manifold to the labeled feature vectors in similar categories.
[0092] Figure 7 An improved image retrieval result is shown according to an embodiment of the concept of the present invention.
[0093] Figure 7 The comparison results of the image retrieval performance of the neural network of the present invention and the semi-supervised deep hash (SSDH) neural network are shown. CIFAR-10 and NUS-WIDE correspond to the datasets of standard benchmark programs for image retrieval. In both the CIFAR-10 and NUS-WIDE datasets, it was found that the mean average precision (mAP) of the neural network based on the embodiment of the present invention is greater than that of the SSDH-based neural network.
[0094] According to various embodiments, when using the neural network conceived according to the present invention, it has been found that the overall image retrieval performance is improved by using unlabeled images. Furthermore, in neural networks not based on product quantization, improvements in image retrieval performance are expected simply by adding classifier 306. In neural networks that do not retrieve images but perform image classification and recognition, the objective function in Equation 5 can be additionally used to improve the network's robustness to unlabeled images.
[0095] Figure 8 This is a flowchart illustrating the operation of a neural network device 10 according to an embodiment of the present invention.
[0096] Reference Figure 8According to various embodiments, in operation S410, the neural network device 10 trains the neural network and obtains a hash table using multiple biometric images. Here, the multiple biometric images are used to identify or verify an individual based on human physical characteristics. For example, the multiple biometric images may include fingerprint images, iris images, vein images, facial images, etc. In the following description, for ease of description, it will be assumed that the multiple biometric images are facial images; however, the embodiments are not limited thereto. For example, the neural network device 10 trains the neural network using five facial images of each of ten people who are biometrically secure. In an embodiment, the neural network device 10 maps information about which of the ten people are biometrically secure to each of the multiple biometric images and learns the mapped information. That is, the multiple biometric images all correspond to labeled images. In another embodiment, to prevent biased labeled images and to maintain robustness in more cases, the neural network device 10 selects at least one of five hundred images as an unlabeled image without mapping information.
[0097] According to various embodiments, by using multiple biometric images, the neural network device 10 learns five hundred codebooks such that the feature vectors of similar images are arranged to be adjacent to each other in each of multiple categories. For example, the multiple categories could be ten categories corresponding to ten people who have been biometrically authenticated.
[0098] In operation S420, according to various embodiments, the neural network device 10 receives the query image. The neural network device 10 uses... Figure 1 The camera 400 in the system acquires a facial image of a person requiring biometric authentication. The neural network device 10 receives the acquired facial image as a query image.
[0099] In operation S430, according to various embodiments, the neural network device 10 calculates the expected value of the query image for each category. The neural network device 10 uses the feature extraction module 302 to extract feature vectors from the query image and determines the similarity between the extracted feature vectors and multiple biometric images. For example, the neural network device 10 generates multiple lookup tables by subtracting feature vectors from multiple codebooks, loading binary hash codes corresponding to the multiple codebooks respectively, and then extracting and summing the values corresponding to the binary hash codes for all subspaces in the codebooks in the lookup tables. This allows it to calculate the distance between the query image and a specific image.
[0100] In this embodiment, the neural network device 10 increases the class expectation value of the specific image only when the distance between the calculated query image and the specific image is greater than a threshold distance. That is, the neural network device 10 only increases the class expectation value when it determines that the query image and the specific image have a similarity equal to or greater than a certain magnitude. Therefore, in a neural network with poor image retrieval and classification performance due to a small number of training images, a query image can increase each of at least two or more class expectations values.
[0101] In operation S440, according to various embodiments, the neural network device 10 identifies the category with the maximum value and determines whether the expected value of the identified category exceeds a threshold. An example of the expected value measured by comparing the distance between the query image and multiple biometric images may be as follows:
[0102] [Table 1]
[0103] 1 Kim xx 1 2 Lee xx 8 ... ... 10 Shin xx 2
[0104] Referring to Table 1, each category corresponds to a specific person according to various embodiments. As described above, there are ten categories when it is assumed that ten people can be safely identified through biometrics. The expected value refers to the number of images that are significantly different from the query image and are identified as similar. That is, referring to Table 1, as a result of comparing the distance between a query image and each training image (i.e., five hundred biometric images), it is found that one image in the "Kim xx" category is identified as having a similarity to the query image equal to or greater than a threshold distance, eight images in the "Lee xx" category are identified as having a similarity to the query image equal to or greater than a threshold distance, and two images in the "Shin xx" category are identified as having a similarity to the query image equal to or greater than a threshold distance. The neural network device 10 identifies the category "Lee xx" as having the maximum expected value and determines whether the expected value of this category exceeds the threshold.
[0105] In operation S450, according to various embodiments, if it is determined that the expected value identified in operation S440 does not exceed a threshold, the neural network device 10 processes the query image as a biometric authentication failure. According to embodiments, when the biometric security level is very high, the threshold is also very high. For example, when the threshold is 9, the expected value for the category "Lee xx" is less than the threshold; therefore, biometric authentication is processed as a failure and the process terminates.
[0106] In operation S460, according to various embodiments, if it is determined that the expected value identified in operation S440 exceeds a threshold, the neural network device 10 processes the query image as successfully biometrically authenticated. That is, the neural network device 10 determines that the person currently requiring biometric authentication is a person in the "Lee xx" category. Furthermore, the neural network device 10 retrains the neural network using the successfully biometrically authenticated query image. The query image received during biometric authentication is an unlabeled image without information about its category. Therefore, each time a new unlabeled image is received, the neural network device 10 uses the new unlabeled image to learn the optimal codeword for classification, thereby improving the image retrieval and classification performance of the neural network.
[0107] Although embodiments of the inventive concept have been specifically shown and described with reference to the disclosed embodiments thereof, it will be understood that various changes in form and detail may be made therein without departing from the spirit and scope of the appended claims.
Claims
1. A neural network device, comprising: The processor, which performs the operation of training the neural network; and The feature extraction module extracts the unlabeled feature vector corresponding to the unlabeled image and the labeled feature vector corresponding to the labeled image. The processor performs a first learning operation on multiple codebooks by using the labeled feature vectors to minimize the cross-entropy of the labeled feature vectors and determine a prototype vector. It then performs a second learning operation on the multiple codebooks by using all of the labeled feature vectors and the unlabeled feature vectors to maximize the entropy of the prototype vectors and minimize the entropy of the unlabeled feature vectors.
2. The neural network device of claim 1, wherein, In the first learning process, the processor calculates a representative value of the labeled feature vector, determines the calculated vector as the prototype vector, and learns the plurality of codebooks corresponding to the labeled feature vector so that the labeled feature vector is close to the prototype vector.
3. The neural network device of claim 2, wherein, The representative value corresponds to one of the mean, median, and mode of the labeled feature vector.
4. The neural network device according to claim 2, wherein, In the second learning process, the processor maximizes the entropy of the prototype vector by adjusting the value of the prototype vector to be close to the unlabeled feature vector.
5. The neural network device according to claim 4, wherein, In the second learning, the processor learns the plurality of codebooks corresponding to the unlabeled feature vectors, such that the unlabeled feature vectors are adjacent to the adjusted prototype vectors.
6. The neural network device according to claim 5, wherein, The processor generates a hash table comprising binary hash codes for each of the plurality of codebooks, wherein the binary hash codes represent codewords similar to feature vectors.
7. The neural network device according to claim 6, wherein, The processor receives a query image as an unlabeled image, extracts the feature vector of the received query image, and stores the difference between the extracted feature vector and the multiple codebooks in a lookup table.
8. The neural network device according to claim 7, wherein, The processor calculates the distance value between the query image, the unlabeled image, and the labeled image based on the lookup table and the hash table.
9. The neural network apparatus of claim 8, wherein the processor identifies an image corresponding to the minimum of the calculated distance values, and the neural network apparatus further includes a classifier configured to classify the category of the identified image as the category of the query image.
10. A method for operating a neural network, the method comprising the steps of: Extract the unlabeled feature vector corresponding to the unlabeled image and the labeled feature vector corresponding to the labeled image; A first learning process is performed on multiple codebooks using the labeled feature vectors, minimizing the cross-entropy of the labeled feature vectors and determining the prototype vector; and A second learning process is performed on the plurality of codebooks by using all of the labeled feature vectors and the unlabeled feature vectors, such that the entropy of the prototype vectors is maximized and the entropy of the unlabeled feature vectors is minimized.
11. The method according to claim 10, wherein, The steps for performing the first learning include: Calculate the representative value of the labeled feature vector; The calculated representative value is determined as the value of the prototype vector; and Learn multiple codebooks corresponding to the labeled feature vector, such that the labeled feature vector is close to the prototype vector.
12. The method according to claim 11, wherein, The second learning step further includes optimizing the entropy of the prototype vector by adjusting the value of the prototype vector to make the prototype vector closer to the unlabeled feature vector.
13. The method according to claim 12, wherein, The second learning step further includes learning multiple codebooks corresponding to the unlabeled feature vectors, such that the unlabeled feature vectors correspond to the adjusted prototype vectors.
14. The method of claim 13, further comprising the step of: generating a hash table comprising binary hash codes for each of the plurality of codebooks, wherein the binary hash codes represent codewords similar to feature vectors.
15. The method of claim 14, further comprising the step of: Receive the query image; Extract the feature vector of the received query image; and The difference between the extracted feature vector and the multiple codebooks is stored in the lookup table.
16. The method of claim 15, further comprising the step of: calculating a distance value between the query image, the unlabeled image, and the labeled image based on the lookup table and the hash table.
17. The method of claim 16, further comprising the step of: Identify the image corresponding to the minimum value among the calculated distance values; and The identified images are classified into categories belonging to the query images.
18. A method for operating a neural network, the method comprising the steps of: Receive multiple biometric images, including at least one labeled image and at least one unlabeled image; Extract feature vectors corresponding to the multiple biometric images; Learning about multiple codebooks is performed by using the feature vectors; Receive a query image and calculate the distance between the query image and the plurality of biometric images; The expected value of the category used to classify the plurality of biometric images is estimated based on the calculated distance value; as well as When the maximum value in the expected value exceeds a preset threshold, biometric authentication is performed.
19. The method of claim 18, further comprising the step of: retraining the plurality of codebooks by using the query image in response to biometric authentication.
20. The method according to claim 19, wherein, The step of estimating the expected value of the category for classifying the plurality of biometric images based on the calculated distance value further includes: counting the number of images whose calculated distance value exceeds a threshold.
Citation Information
Patent Citations
Eddy current heating system reducing energy consumption by breaking torque
KR1020200041076A
Patterning process of a semiconductor structure with enhanced adhesion
KR1020210016274A
Learning method and learning device for adjusting parameters of CNN by using multi-scale feature maps and testing method and testing device using same
CN109670512A
Apparatus and method for training a classification network for character recognition
JP2019528520A