Electronic device for identifying user's identity on basis of hand image and system including same
Patent Information
- Application Number
- PCT/KR2025/009590
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-07-04
- Publication Date
- 2026-10-01
Smart Images

Figure KR2025009590_01102026_PF_FP_ABST
Abstract
Description
Electronic device for identifying user identity based on hand image and system including the same
[0001] The present disclosure relates to an electronic device for verifying and identifying a user's identity based on a user's hand image and a system including the same. More specifically, the present disclosure relates to an electronic device capable of verifying and identifying a user's identity based on a user's hand image using a classification model pre-trained to classify a plurality of attributes (e.g., gender, skin color, presence or absence of accessories, presence or absence of nail polish, hand direction, identity, etc.) and a system including the same.
[0002] With the advancement of communication and network technologies, various systems capable of managing personal and user information on servers have been developed. To ensure the security of such personal and user information, biometric technologies are being developed to verify and identify user identities.
[0003] Biometric technology is a technology that verifies identity by utilizing an individual's unique physiological or behavioral characteristics as measurable units. Facial recognition utilizes various identifiable features such as eyes, nose, and mouth, but it has the disadvantage of being vulnerable to external factors like makeup or masks. Fingerprint recognition provides high reliability by using unique fingerprint patterns, but it has limitations such as the inconvenience of the contact method, the possibility of forgery, and a decrease in recognition rates due to physical damage.
[0004] To overcome these limitations, biometric technology utilizing objects that can be easily captured non-contactually and possess various characteristics is gaining attention.
[0005] Meanwhile, an artificial intelligence system is a computer system that implements human-level intelligence, in which the machine learns and makes judgments autonomously, and its recognition rate improves with use. Artificial intelligence technology consists of machine learning (deep learning) technology, which utilizes algorithms to classify and learn the characteristics of input data autonomously, and component technologies that mimic functions such as cognition and judgment of the human brain by utilizing machine learning algorithms. The component technologies may include, for example, at least one of linguistic understanding technology that recognizes human language / characters; visual understanding technology that perceives objects like human vision; reasoning / prediction technology that judges information to logically infer and predict; knowledge representation technology that processes human experience information into knowledge data; and motion control technology that controls autonomous driving of vehicles and the movement of robots.
[0006] Recently, there is a sufficient need for technology that performs biometric recognition using such artificial intelligence technology.
[0007] The object of the present disclosure is to provide an electronic device that verifies and identifies the identity of a user based on an image of the user's hand, and a system including the same.
[0008] More specifically, the object of the present disclosure is to provide an electronic device and a system including the same that can verify and identify a user's identity based on an image of a user's hand using a classification model pre-trained to classify a plurality of attributes (e.g., gender, skin color, presence or absence of accessories, presence or absence of nail polish, hand direction, identity, etc.).
[0009] The purposes of the present disclosure are not limited to those mentioned above, and other purposes and advantages of the present disclosure not mentioned may be understood from the following description and will be more clearly understood from the embodiments of the present disclosure. Furthermore, it will be readily apparent that the purposes and advantages of the present disclosure can be realized by the means and combinations thereof set forth in the claims.
[0010] An electronic device according to the present disclosure comprises a memory for storing at least one instruction and a processor for executing said at least one instruction, wherein the processor receives a target image of a user’s hand, generates a target embedding for said received target image using a classification model pre-trained based on a neural network, and can identify said user by comparing said target embedding with a reference embedding for at least one pre-defined reference image.
[0011] In addition, the reference image is stored in the memory, and the processor can generate a reference embedding for the reference image using the classification model and verify the identity of the user by comparing the target embedding with the reference embedding.
[0012] In addition, the above classification model may include a ResNet50 model.
[0013] In addition, the classification model includes a GAP layer (Global Average Pooling Layer) that receives a feature map as input and outputs an embedding vector, and an FC layer (Fully-Connected Layer) that receives the embedding vector as input and outputs a classification result, and the FC layer may include a first layer to a sixth layer defined to perform a first to sixth task predefined in relation to the classification task.
[0014] In addition, the processor can train the classification model during the training process.
[0015] In addition, the processor can control each of the first to sixth layers to perform each of the first to sixth tasks on a learning embedding vector corresponding to a predefined learning image during the learning step, and train the classification model by calculating a loss function for the output value of each of the first to sixth tasks based on the learning embedding vector.
[0016] In addition, the processor can generate the target embedding and the reference embedding through the GAP layer included in the classification model during the inference process.
[0017] Additionally, the processor can identify the identity of the user by calculating the cosine similarity between the target embedding and the reference embedding in the inference step.
[0018] A non-transient computer-readable recording medium storing at least one instruction that is executed by a processor of an electronic device according to the present disclosure to cause the electronic device to perform the operation method of claims 1 through 8 may include the steps of receiving a target image of a user's hand, generating a target embedding for the received target image using a classification model that is pre-trained based on a neural network, and verifying the identity of the user by comparing the target embedding with a reference embedding for at least one pre-defined reference image.
[0019] A system comprising a user terminal for a user according to the present disclosure and an electronic device that communicates with the user terminal to verify the identity of the user may include a user terminal that generates a target image of the user's hand and transmits the generated target image to the electronic device, and an electronic device that receives the target image from the user terminal, generates a target embedding for the received target image using a classification model pre-trained based on a neural network, and verifies the identity of the user by comparing the target embedding with a reference embedding for at least one pre-defined reference image.
[0020] The present disclosure can overcome the side effects and disadvantages inherent in existing biometric technologies, such as facial recognition and fingerprint recognition, by verifying and identifying the user's identity based on an image of the user's hand. More specifically, the present disclosure can overcome limitations such as the possibility of forgery and reduced recognition rates due to physical damage by performing biometric recognition using an image of the user's hand that can be easily captured non-contactually and contains various features, while ensuring high reliability and user convenience.
[0021] In addition, since the user's hand image is less affected by changes in appearance, the present disclosure allows for easy acquisition even in a controlled environment compared to conventional methods such as facial recognition and fingerprint recognition, and accordingly, the present disclosure can further improve user convenience related to biometric recognition.
[0022] In addition, the present disclosure can further improve biometric accuracy by verifying and identifying the user's identity based on the user's hand image using a classification model that has been pre-trained to classify a plurality of attributes (e.g., gender, skin color, presence or absence of accessories, presence or absence of nail polish, hand direction, identity, etc.).
[0023] Aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description with reference to the accompanying drawings.
[0024] FIG. 1 is a drawing for explaining a biometric recognition system including a user terminal and an electronic device according to one embodiment of the present disclosure.
[0025] FIG. 2 is a block diagram for explaining the configuration of an electronic device according to one embodiment of the present disclosure.
[0026] FIG. 3 is a block diagram illustrating a classification model stored in the memory of an electronic device according to one embodiment of the present disclosure.
[0027] FIGS. 4 and 5 are drawings for illustrating a learning step for a classification model performed by a processor according to one embodiment of the present disclosure.
[0028] FIGS. 6 and 7 are drawings for illustrating an inference step using a classification model performed by a processor according to one embodiment of the present disclosure.
[0029] FIGS. 8 and 9 are experimental data for explaining the effects of an electronic device according to one embodiment of the present disclosure.
[0030] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0031] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.
[0032] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.
[0033] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0034] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0035] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.
[0036] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0037] Where it is stated that a component (e.g., Component 1) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., Component 2), it should be understood that the component may be directly connected to the other component or connected through the other component (e.g., Component 3).
[0038] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between the certain component and the other component.
[0039] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.
[0040] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.
[0041] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for the 'module' or 'part' that needs to be implemented in specific hardware.
[0042] Meanwhile, various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0043] Hereinafter, embodiments according to the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them.
[0044]
[0045] FIG. 1 is a drawing for explaining a biometric recognition system including a user terminal and an electronic device according to one embodiment of the present disclosure.
[0046] Referring to FIG. 1, a biometric recognition system (1) according to one embodiment of the present disclosure may include a user terminal (100) and an electronic device (200).
[0047] The user terminal (100) is a device that the user carries and stores.
[0048] For example, a user terminal (100) can generate a target image by capturing an object and transmit the generated target image to an electronic device (200).
[0049] At this time, the object may include a part of the user's body. In other words, the user terminal (100) may capture a part of the user's body to generate a target image and then transmit it to the electronic device (200). For example, the object may include the user's hand, and the target image may include an image of the user's hand, but the embodiments of the present invention are not limited thereto.
[0050] Meanwhile, the user terminal (100) may include at least one of a smartphone, a tablet PC, a laptop PC, a netbook computer, a mobile device, and a wearable device, but is not limited thereto.
[0051] The electronic device (200) may be a device that performs biometric recognition for the user through communication with the user terminal (100).
[0052] For example, an electronic device (200) can identify the identity of a user based on visual understanding of a target image received from a user terminal (100). In this case, visual understanding is a technology that recognizes and processes objects like human vision, and includes object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, image enhancement, etc.
[0053] For example, an electronic device (200) may receive a target image of a user's hand from a user terminal (100) and then perform identity verification for the target image using a classification model that has been pre-trained based on a neural network. For example, the electronic device (200) may generate embeddings by inputting the target image and a pre-stored reference image into the classification model, and perform identity verification for the user by comparing each generated embedding. A detailed explanation of this will be provided later through FIGS. 2 to 7.
[0054] Meanwhile, the electronic device (200) may be, for example, a server which is a computer that provides services to a user terminal (100) via a network. The server may be an FTP server, a web server, a database server, or a cloud-type server, and the server may be built with an operating system such as Linux. In this case, the server may include a learning function for the aforementioned classification model, an identity verification function using the classification model, etc., and may not necessarily be a single device but may be distributed across multiple devices to implement each function. However, the electronic device (200) according to one embodiment of the present disclosure is not limited to the devices described above, and the electronic device (200) may be implemented as an electronic device having two or more functions of the devices described above.
[0055] Hereinafter, the operation of an electronic device (200) according to one embodiment of the present disclosure will be described in more detail with reference to FIG. 2.
[0056]
[0057] FIG. 2 is a block diagram for explaining the configuration of an electronic device according to one embodiment of the present disclosure. FIG. 3 is a block diagram for explaining a classification model stored in the memory of an electronic device according to one embodiment of the present disclosure.
[0058] Referring to FIGS. 1 to 3, the electronic device (200) may include a communication interface (210), a memory (220), and a processor (230).
[0059] The communication interface (210) may be configured to perform communication with the user terminal (100). For example, the communication interface (210) may receive a target image of the user's hand from the user terminal (100). For another example, when identity verification is completed based on the target image, the communication interface (210) may output the identity verification result to the user terminal (100) through a display, speaker, etc. However, the embodiments of the present invention are not limited thereto, and the operation and role of the aforementioned communication interface (210) may be integrated into the processor (230) described later.
[0060] Meanwhile, the communication interface (210) may include a wireless communication interface, a wired communication interface, an input interface, etc.
[0061] A wireless communication interface can perform communication with various external devices using wireless communication technology or mobile communication technology. Such wireless communication technologies may include, for example, Bluetooth, Bluetooth Low Energy, CAN communication, Wi-Fi, Wi-Fi Direct, ultrawide band (UWB), Zigbee, infrared data association (IrDA), or near field communication (NFC), and mobile communication technologies may include 3GPP, Wi-Max, LTE (Long Term Evolution), 5G, etc. The wireless communication interface may be implemented using an antenna, a communication chip, a substrate, etc., capable of transmitting electromagnetic waves to the outside or receiving electromagnetic waves transmitted from the outside.
[0062] A wired communication interface can communicate with various devices based on a wired communication network. Here, the wired communication network can be implemented using physical cables, such as, for example, pair cables, coaxial cables, fiber optic cables, or Ethernet cables.
[0063] Depending on the embodiment, either the wireless communication interface or the wired communication interface may be omitted. Accordingly, the electronic device (200) may include only a wireless communication interface or only a wired communication interface. Furthermore, the electronic device (200) may be equipped with an integrated communication interface that supports both wireless connection via the wireless communication interface and wired connection via the wired communication interface. Additionally, the electronic device (200) is not limited to including a single communication interface that performs a communication connection in one manner, but may include multiple communication interfaces that perform communication connections in multiple manners.
[0064] The memory (220) stores various programs or data temporarily or non-temporarily and transmits the stored information to the processor (230) upon the call of the processor (230). Additionally, the memory (220) can store various information required for the operation, processing, or control operation of the processor (230) in an electronic format.
[0065] The memory (220) may include, for example, at least one of a main memory and an auxiliary memory. The main memory may be implemented using a semiconductor storage medium such as ROM and / or RAM. The ROM may include, for example, a conventional ROM, EPROM, EEPROM and / or MASK-ROM. The RAM may include, for example, a DRAM and / or SRAM. The auxiliary memory may be implemented using at least one storage medium capable of storing data permanently or semi-permanently, such as a flash memory device, an SD (Secure Digital) card, a solid state drive (SSD), a hard disk drive (HDD), an optical recording medium such as a magnetic drum, a compact disc (CD), a DVD, or a laser disc, a magnetic tape, a magneto-optical disc and / or a floppy disk.
[0066] Memory (220) may store commands, information, and / or data associated with the operation of each component included in the electronic device (200). For example, memory (220) may store instructions that enable the processor (230) to perform various operations described in this document during execution. For another example, memory (220) may store various algorithms or models that can be used when the processor (230) performs biometric recognition or identity verification, such as a classification model (221) described below. For yet another example, memory (220) may store a reference image used when the processor (230) performs biometric recognition or identity verification.
[0067] The processor (230) controls the overall operation of the electronic device (200). Specifically, the processor (230) is connected to the configuration of the electronic device (200) including the memory (220) as described above, and can control the overall operation of the electronic device (200) by executing at least one instruction stored in the memory (220) as described above.
[0068] At this time, the processor (230) may be implemented in various ways. For example, the processor (230) may be implemented as a single processor as well as as a plurality of processors. At this time, one or more processors may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. One or more processors may control one or any combination of other components of the electronic device (200) and may perform operations or data processing related to communication. One or more processors may execute one or more programs or instructions stored in memory (220). For example, one or more processors may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory (220).
[0069] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an artificial intelligence dedicated processor).
[0070] One or more processors may be implemented as a single-core processor comprising one core, or as one or more multicore processors comprising multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.
[0071] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one of the plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in a multi-core processor, or the first operation and the second operation may be performed by a first core included in a multi-core processor and the third operation may be performed by a second core included in a multi-core processor.
[0072] In the embodiments of the present disclosure, the processor (230) may mean a system-on-chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, GPU, APU, MIC, DSP, NPU, hardware accelerator, or machine learning accelerator, but the embodiments of the present disclosure are not limited thereto.
[0073] Meanwhile, the processor (230) can perform identity verification for the user based on a target image of the user's hand transmitted from the user terminal (100).
[0074] For example, the processor (230) can train a classification model (221) as shown in FIG. 3 and then use the trained classification model (221) to verify the identity of the user. At this time, the classification model (221) may be stored in memory (220), but the embodiments of the present invention are not limited thereto, and the processor (230) may load and use a classification model (221) existing outside the electronic device (200).
[0075] The classification model (221) may be an artificial intelligence-based neural network model.
[0076] To explain in more detail, deep learning, a type of machine learning, involves learning by descending to deep levels in multiple stages based on data. In other words, deep learning represents a set of machine learning algorithms that extract key data from multiple datasets by progressively increasing the levels.
[0077] For example, neural networks can utilize various known deep learning structures. For instance, neural networks can utilize structures such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), DBN (Deep Belief Network), GNN (Graph Neural Network), GAN (Generative Adversarial Network), Transformer, and Autoencoder.
[0078] Specifically, a Convolutional Neural Network (CNN) is a model that mimics the function of the human brain, based on the assumption that when humans recognize an object, they extract basic features of the object, perform complex calculations within the brain, and then recognize the object based on the results. CNNs may include, but are not limited to, well-known structures such as LeNet, AlexNet, VGGNet, GoogleNet, and ResNet (e.g., ResNet50).
[0079] Recurrent Neural Networks (RNNs) are widely used in natural language processing and are an effective structure for processing time-series data that changes over time; they can be constructed by stacking layers at every moment.
[0080] A Deep Belief Network (DBN) is a deep learning structure constructed by stacking Restricted Boltzmann Machines (RBMs), a deep learning technique, in multiple layers. When the Restricted Boltzmann Machine (RBM) training is repeated until a certain number of layers are reached, a Deep Belief Network (DBN) with that number of layers can be constructed.
[0081] A Graphic Neural Network (GNN) represents an artificial neural network structure implemented by deriving similarities and feature points between modeling data using modeling data modeled based on data mapped between specific parameters.
[0082] A Generative Adversarial Network (GAN) represents an artificial neural network structure that uses a generative neural network and a discriminative neural network to generate new data in a form similar to input data. GANs may include known DCGAN (Deep Convolutional GAN), CGAN (Conditional GAN), WGAN (Wasserstein GAN), StyleGAN (Style-Based GAN), CycleGAN, etc., but embodiments of the present invention are not limited thereto.
[0083] The Transformer is an artificial neural network with an encoder-decoder structure utilizing attention, capable of grasping the overall meaning between input and output sequences. By employing an attention mechanism, the Transformer ensures that every element of the input sequence influences the output sequence, allowing both the encoder and decoder to consider the entire sequence. The Transformer can use natural language and time-series data, as well as patched images, as input.
[0084] An autoencoder is a deep learning architecture that performs the role of extracting and reconstructing data features. Typically, an autoencoder includes an encoder that compresses input values and a decoder that restores the compressed data. The encoder transforms input values into low-dimensional latent representations, while the decoder restores the latent representations to the same dimension as the input values. In this process, both the encoder and decoder can be composed of Multilayer Perceptrons (MLPs). When training an autoencoder, input data is used, and weights and biases are trained to minimize the difference between the output and input values. An autoencoder trained in this way can effectively extract features from input data and restore noisy input data. Autoencoders are primarily utilized in fields such as data compression, dimensionality reduction, noise removal, and data generation; they can also be applied in areas such as image recognition, natural language processing, and speech recognition.
[0085] Meanwhile, artificial neural network training can be achieved by adjusting the weights of the connections between nodes (and adjusting bias values if necessary) to produce a desired output for a given input. Additionally, artificial neural networks can continuously update weight values through learning. Furthermore, methods such as backpropagation can be used for the training of artificial neural networks.
[0086] In this case, machine learning methods for artificial neural networks, such as unsupervised learning, semi-supervised learning, and supervised learning, can be used. Additionally, depending on the settings, the neural network can be controlled to automatically update the artificial neural network structure to output analysis data after training.
[0087] For some examples, the classification model (221) in the present disclosure may include a CNN-based ResNet50 model, but embodiments of the present invention are not limited thereto.
[0088] Hereinafter, with reference to FIGS. 4 and 5, a learning step in which a processor (230) according to one embodiment of the present disclosure learns a classification model (221) will be described, and with reference to FIGS. 6 and 7, a performing step in which a processor (230) according to one embodiment of the present disclosure uses the learned classification model (221) to verify the identity of a user will be described.
[0089] FIGS. 4 and 5 are drawings for illustrating a learning step for a classification model performed by a processor according to one embodiment of the present disclosure.
[0090] Referring to FIGS. 2 to 5, a processor (230) according to one embodiment of the present disclosure can train a classification model (221). In other words, the processor (230) can perform a training process to train a classification model (221).
[0091] To this end, first, the processor (230) can receive a training image (301) (S100). At this time, the training image (301) may be data that is pre-stored in memory (220) as training data for training a classification model (221). For example, the training image (301) may include image data of multiple users' hands.
[0092] At this time, the processor (230) can preprocess the received training image (301). For example, the processor (230) may perform processes such as image resizing, data augmentation (e.g., rotation, flipping, etc.), tensor transformation, and normalization on the training image (301), but embodiments of the present invention are not limited thereto.
[0093] Next, the processor (230) can control the classification model (221) to perform the first to sixth tasks by applying the training image (301) to the classification model (221) (S200).
[0094] At this time, the classification model (221) may include a feature extraction layer (221_1) that extracts a feature map for input data, a GAP (Global Average Pooling Layer) layer that receives the feature map and outputs an embedding vector, and an FC layer (Fully-Connected Layer, 221_2) that receives the embedding vector and outputs a classification result.
[0095] To explain in more detail, first, the processor (230) can input a training image (301) into a feature extraction layer (221_1) to extract a feature map (303). At this time, the feature extraction layer (221_1) can generate a feature map (303) through a convolution operation on the input training image (301). For example, since the classification model (221) may include a ResNet50 model as described above, the feature extraction layer (221_1) may include four convolutional block groups, and each convolutional block may include at least one residual block that performs a skip connection. Accordingly, the classification model (221) in the present invention can mitigate problems such as gradient vanishing or explosion that may occur as the network depth increases.
[0096] Next, the processor (230) can control the GAP layer to generate a learning embedding vector by applying the extracted feature map (303) to the GAP layer. In other words, by applying the extracted feature map (303) to the GAP layer, the processor (230) can control the GAP layer to extract a learning embedding vector, which is a learning embedding vector, based on the feature map (303).
[0097] Next, the processor (230) can normalize the learning embedding vector. For example, the classification model (221) may further include a dropout layer, and the processor (230) can perform the normalization process by applying the learning embedding vector output from the GAP layer to the dropout layer.
[0098] At this time, the dropout layer can reduce the risk of overfitting by randomly deactivating some of the neurons during training and prevent the classification model (221) from relying excessively on specific features while maintaining robustness in all tasks. However, in the present invention, normalization of the learning embedding vector may be omitted.
[0099] Next, the processor (230) can control the FC layer (221_2) to output a classification result by applying the extracted learning embedding vector or the normalized learning embedding vector.
[0100] At this time, the FC layer (221_2) may include a plurality of layers that perform a task of classifying hand images according to predefined attributes (304), and the processor (230) may control each layer to perform its own task by applying a learning embedding vector to each of the plurality of layers. At this time, each layer may operate independently.
[0101] For example, the FC layer (221_2) may include a first layer that performs a first task of classifying hand images according to a first attribute (304_1) related to gender, a second layer that performs a second task of classifying hand images according to a second attribute (304_2) related to skin color, a third layer that performs a third task of classifying hand images according to a third attribute (304_3) related to accessories, a fourth layer that performs a fourth task of classifying hand images according to a fourth attribute (304_4) related to nail polish, a fifth layer that performs a fifth task of classifying hand images according to a fifth attribute (304_5) related to aspect of hand, and a sixth layer that performs a sixth task of classifying hand images according to a sixth attribute (304_6) related to identity.
[0102] At this time, the class of the classification result output by the first layer may be 2 (male, female), the class of the classification result output by the second layer may be 4 (dark, medium, pair, very pair), the class of the classification result output by the third layer may be 2 (present, non-present), the class of the classification result output by the fourth layer may be 2 (present, non-present), the class of the classification result output by the fifth layer may be 4 (right palm, back of right hand, left palm, back of left hand), and the class of the classification result output by the sixth layer may be 170 (0 to 169), but the embodiments of the present invention are not limited thereto.
[0103] Next, the processor (230) can train a classification model (221) by calculating a loss function for the output values of the first to sixth tasks (S300).
[0104] For example, the processor (230) can train a classification model (221) through a cross-entropy loss function such as Equation 1 below.
[0105]
[0106] In the above mathematical formula 1 represents the actual class label, and represents the predicted probability for the i-th class through the softmax function, and represents the loss function for each classification task t. Here, the loss function ( The processor (230) calculates the error by summing the discrepancies between the predicted probability and the actual value across all classes. The final loss value is the sum of the cross-entropy losses calculated for the six classification labels (304_1 to 304_6) associated with the input training image (301). During this training process, the processor (230) can train the classification model (221) in a direction that minimizes this final loss value.
[0107] When the model trained in the manner described above is ResNet150, as shown in Figure 9 to be described later, the mAP value is 75.48 and the rank-1 value is 100, and experimental results showed that it is the best for both items, and through this, a significant effect compared to other models or other training methods can be recognized.
[0108] FIGS. 6 and 7 are drawings for illustrating an inference step using a classification model performed by a processor according to one embodiment of the present disclosure.
[0109] Referring to FIGS. 6 and FIGS. 7, the processor (230) can verify the identity of a user using a classification model (221) learned through the process described above. In other words, the processor (230) can perform an inference process, an execution step, to verify the identity of a user using the learned classification model (221).
[0110] To this end, first, the processor (230) can receive a target image (401) (S400). At this time, the target image (401) may be data received from a user terminal (100) as input data input to the classification model (221). For example, the target image (301) may include image data of a specific user's hand that is the subject of identity verification.
[0111] Next, the processor (230) can generate a target embedding (403) based on the target image (401) (S500).
[0112] For example, the processor (230) can output a target embedding (403) from the classification model (221) by inputting the target image (401) into the classification model (221) learned through the process of FIGS. 4 and FIGS. 5. More specifically, the processor (230) can generate a target embedding (403) by inputting the target image (401) into a feature extraction layer (221_1 in FIG. 5) to extract a feature map, and inputting the extracted feature map into a GAP layer. At this time, the processor (230) may also perform normalization on the generated target embedding (403).
[0113] Next, the processor (230) can generate a reference embedding (404) based on the reference image (402) (S600).
[0114] For example, the processor (230) can output a reference embedding (404) from the classification model (221) by loading a plurality of reference images (402) from memory (220) and then inputting the reference images (402) into the classification model (221) learned through the process of FIGS. 4 and FIGS. 5. More specifically, the processor (230) can generate a reference embedding (404) by inputting the reference images (402) into a feature extraction layer (221_1 in FIG. 5) to extract a feature map, and inputting the extracted feature map into a GAP layer. At this time, the processor (230) may perform normalization on the generated reference embedding (404). Meanwhile, the number of reference images (402) may be multiple, and accordingly, the number of reference embeddings (404) may also be multiple.
[0115] Next, the processor (230) can identify the identity of the user based on the target embedding and the reference embedding (S700).
[0116] For example, the processor (230) can determine whether the user corresponding to the target embedding (403) corresponds to one of the multiple users corresponding to the multiple reference embeddings (404) by comparing the target embedding (403) with the multiple reference embeddings (404). In other words, the processor (230) can perform an identity verification process to determine whether the user corresponding to the target image (401) corresponds to one of the multiple users corresponding to the multiple reference images (402) based on the result of the comparison between the target embedding (403) and the reference embeddings (404).
[0117] For example, the processor (230) can perform identity verification based on the similarity between the target embedding (403) and a plurality of reference embeddings (404). For instance, the processor (230) can calculate the similarity between the target embedding (403) and each reference embedding (404) through a cosine similarity calculation such as Equation 2 below.
[0118]
[0119] In the above mathematical formula 2, A represents the target embedding (403), and B represents any one of the multiple reference embeddings (404). represents the dot product of the target embedding (403) and the reference embedding (404), and is the size of the target embedding (403), represents the size of any one of the multiple reference embeddings (404).
[0120] Next, the processor (230) can generate a sorted list (405) by sorting the reference embeddings (404) according to the magnitude of the calculated cosine similarity. That is, the processor (230) can generate a sorted list (405) by sorting the reference embeddings (404) in order of highest cosine similarity with the target embedding (403).
[0121] Next, the processor (230) can recognize and identify the user corresponding to the reference embedding (404) at the top of the sorting list (405), that is, the reference embedding (404) having the maximum similarity to the target embedding (403), as the user for the target image (403).
[0122] However, unlike the above, the processor (230) may not perform steps (S500 to S700) and may perform a classification operation for the target image (401) after step (S400). That is, the processor (230) may not determine whether the target image (401) is a hand image of a user by performing steps (S500 to S700), but may identify the attributes of the target image (401) by performing a classification operation for the target image (401) after step (S400). For example, the processor (230) may input the target image (401) into a feature extraction layer (221_1 in FIG. 5) to extract a feature map, input the extracted feature map into a GAP layer to generate a target embedding (403), and input the generated target embedding (403) into multiple layers (layers 1 through 6) of an FC layer (221_2 in FIG. 5) to generate multiple classification results (results of the execution of the first to sixth tasks) for the target image (401).
[0123] FIGS. 8 and 9 are experimental data for explaining the effects of an electronic device according to one embodiment of the present disclosure.
[0124] Referring to FIGS. 2, FIG. 3, and FIG. 8, FIG. 8 illustrates the results of measuring top-1 accuracy according to the attribute (304 in FIG. 5) defined in the present invention. Here, top-1 accuracy refers to the ratio of cases where the class with the highest probability predicted by the model is the actual correct class. As shown in FIG. 8, the classification model (221) of the present invention shows 100% accuracy in predicting gender, skin color, and nail polish, and also shows very high accuracy approaching 100% in predicting other attributes, namely accessories, aspect of hand, and identity.
[0125] Referring to FIGS. 2, FIGS. 3, FIGS. 7, and FIG. 9, FIG. 9 illustrates the results of comparing the accuracy of user identity verification by CNN-based backbone network. FIG. 9 illustrates the results of measuring mAP (Mean Average Precision) and rank-1 as examples of accuracy measurement criteria. mAP is an indicator that measures the overall matching accuracy between a target image (401) and a reference image (402), and rank-1 accuracy refers to the ratio of reference images (402) with the highest similarity to the target image (401) having the same ID as the target image (401). As shown in FIG. 9, it can be seen that the accuracy is highest when the backbone network of the classification model (221) is set to ResNet50.
[0126] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0127] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. In an electronic device, Memory that stores at least one instruction; and It includes a processor that executes at least one of the above instructions, The above processor is, Receives a target image of the user's hand, and A target embedding for the received target image is generated using a classification model pre-trained based on a neural network, and An electronic device that identifies the identity of the user by comparing a reference embedding for at least one predefined reference image with the target embedding.
2. In Paragraph 1, The above reference image is stored in the above memory, and The processor generates the reference embedding for the reference image using the classification model, and An electronic device that identifies the identity of the user by comparing the target embedding and the reference embedding.
3. In Paragraph 2, The above classification model is an electronic device including a ResNet50 model.
4. In Paragraph 3, The above classification model is, A GAP layer (Global Average Pooling Layer) that takes a feature map as input and outputs an embedding vector, and It includes a Fully-Connected Layer (FC layer) that receives the above embedding vector as input and outputs a classification result, and The above FC layer is an electronic device comprising a first layer to a sixth layer defined to perform a first to sixth task predefined in relation to a classification task.
5. In Paragraph 4, The above processor is an electronic device that trains the classification model in the training phase.
6. In Paragraph 5, The above processor, in the learning step, Controls each of the first to sixth layers to perform each of the first to sixth tasks on a learning embedding vector corresponding to a predefined learning image, and An electronic device for training the classification model by calculating a loss function for each of the output values of the first to sixth tasks based on the learning embedding vectors.
7. In Paragraph 4, The above processor is an electronic device that generates the target embedding and the reference embedding through the GAP layer included in the classification model during the inference process.
8. In Paragraph 7, The above processor is an electronic device that identifies the identity of the user by calculating the cosine similarity between the target embedding and the reference embedding in the inference step.
9. A non-transient computer-readable recording medium storing at least one instruction that is executed by a processor of an electronic device to cause said electronic device to perform the method of operation of claims 1 through 8, Step of receiving a target image of the user's hand; A step of generating a target embedding for the received target image using a classification model pre-trained based on a neural network; and A computer-readable recording medium comprising the step of verifying the identity of the user by comparing a reference embedding for at least one predefined reference image with the target embedding.
10. A system comprising a user terminal for a user and an electronic device that communicates with said user terminal to verify the identity of said user, A user terminal that generates a target image of the user's hand and transmits the generated target image to the electronic device; and A system comprising an electronic device that receives a target image from a user terminal, generates a target embedding for the received target image using a classification model pre-trained based on a neural network, and verifies the identity of the user by comparing the target embedding with a reference embedding for at least one pre-defined reference image.