Electronic apparatus and method for controlling thereof

KR102999940B1Active Publication Date: 2026-08-05SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2020-03-25
Publication Date
2026-08-05

Smart Images

  • Figure 112020031361268-PAT00001_ABST
    Figure 112020031361268-PAT00001_ABST
Patent Text Reader

Abstract

A method for controlling an electronic device is disclosed. The control method according to the present disclosure comprises: acquiring a neural network model trained to detect an object corresponding to at least one class; acquiring a user command for detecting a first object corresponding to a first class; and, if the first object does not correspond to at least one class, acquiring a new neural network model based on information regarding the neural network model and the first object.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to an electronic device for customizing a neural network model and a method for controlling the same, and more specifically, to an electronic device for acquiring a new neural network model by adding a new class according to a user command and a method for controlling the same. Background Technology

[0002] Conventional artificial intelligence (AI)-based object recognition models are trained offline based on data regarding predetermined categories or classes, and once training is complete, they are applied to devices such as smartphones, robots / robotic devices, or other image and / or voice recognition systems.

[0003] Meanwhile, once a trained object recognition model is applied to a device, modifying the model—such as adding new recognizable classes (or categories)—is difficult. This is because adding new classes requires a large number of samples for those classes and cloud computing resources for retraining the model, which is not feasible in terms of time and cost.

[0004] Recently, as consumer demand to customize object recognition models applied to smartphones and other devices increases, there is a growing need for technology to customize object recognition models that have completed training and are applied to products. The problem to be solved

[0005] The technical problem that the present invention aims to solve is to provide an electronic device capable of obtaining a new neural network model by adding a new class corresponding to a user command to a neural network model, and recognizing an object corresponding to the new class based on the obtained neural network model.

[0006] The technical problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem

[0007] According to an exemplary embodiment of the present disclosure for solving the above-described technical problem, a control method for an electronic device may be provided, comprising: a step of acquiring a neural network model trained to detect an object corresponding to at least one class; a step of acquiring a user command to detect a first object corresponding to a first class; and a step of acquiring a new neural network model based on information regarding the neural network model and the first object if the first object does not correspond to the at least one class.

[0008] According to another exemplary embodiment of the present disclosure for solving the technical problem described above, an electronic device may be provided comprising: a memory including at least one instruction; and a processor; wherein the processor acquires a neural network model trained to detect an object corresponding to at least one class, acquires a user command to detect a first object corresponding to a first class, and if the first object does not correspond to the at least one class, acquires a new neural network model based on information regarding the neural network model and the first object.

[0009] The means for solving the problem of the present disclosure are not limited to the means for solving the problem described above, and means for solving the problem not mentioned will be clearly understood by those skilled in the art to which the present disclosure belongs from the present specification and the attached drawings. Effects of the invention

[0010] According to various embodiments of the present disclosure as described above, an electronic device can recognize an object corresponding to a new class added to a neural network model according to a user command.

[0011] Furthermore, other effects that can be obtained or predicted by the embodiments of the present disclosure will be disclosed directly or implicitly in the detailed description of the embodiments of the present disclosure. For example, various effects predicted according to the embodiments of the present disclosure will be disclosed in the detailed description to be set forth below. Brief explanation of the drawing

[0012] FIG. 1 is a drawing for explaining an electronic device according to one embodiment of the present disclosure. FIG. 2 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure. Figure 3a is a diagram illustrating a conventional object recognition model using infinitesimal learning. Figure 3b is a diagram illustrating a training method for a conventional object recognition model. FIG. 4a is a diagram illustrating a method for obtaining a new neural network model according to one embodiment of the present disclosure. FIG. 4b is a drawing for explaining a method of learning a neural network model according to one embodiment of the present disclosure. FIG. 5a is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure. FIG. 5b is a drawing for explaining a method to obtain a new neural network model according to one embodiment of the present disclosure. FIG. 6a is a diagram illustrating a method for customizing a neural network model using image samples according to one embodiment of the present disclosure. FIG. 6b is a diagram illustrating the customization of a neural network model using video frame samples according to one embodiment of the present disclosure. FIG. 7 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure. Figure 8 is a flowchart of an example of checking whether a user request class already exists in a neural network model. Specific details for implementing the invention

[0013] The terms used in this specification will be briefly explained, and the present disclosure will be described in detail.

[0014] The terms used in the embodiments of this disclosure have been selected to be as widely used as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant explanatory section of this disclosure. Therefore, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.

[0015] The embodiments of the present disclosure are subject to various modifications and may have various embodiments; therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the scope of the disclosed spirit and technology. In describing the embodiments, if it is determined that a detailed description of related prior art may obscure the essence, such detailed description is omitted.

[0016] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. Terms are used solely for the purpose of distinguishing one component from another.

[0017] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "consisting of" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0018] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0020] FIG. 1 is a drawing for explaining an electronic device according to an embodiment of the present disclosure. The electronic device (100) can capture the surrounding environment to acquire an image (30). For example, the electronic device (100) may be a robot device. The electronic device (100) may recognize an object included in the acquired image (30). Specifically, the electronic device (100) may recognize an object included in the image (30) by inputting the acquired image (30) into a neural network model trained to recognize objects included in the image. For example, the electronic device (100) may recognize a cup included in the image (30).

[0021] Meanwhile, the electronic device (100) may receive a command from the user (10) to find the user's (10) cup (20). At this time, for the electronic device (100) to recognize the user's (10) cup (20), the electronic device (100) must recognize the user's (10) cup (20) using a neural network model trained to recognize the user's (10) cup (20). However, the electronic device (100) generally distributed to the user (10) uses a neural network model trained to be suitable for the most common and general purposes (e.g., identifying whether an object is a cup or a book). Therefore, even if the electronic device (100) can recognize multiple cups included in the image (30) using a conventional object recognition model (or neural network model), it cannot recognize the user's (10) cup (20) among the multiple cups.

[0022] Meanwhile, when the electronic device (100) according to the present disclosure obtains a command (e.g., a command to find the cup (20)) related to the user (10) which is an object that cannot be identified by a previously stored neural network model from the user (10), it can extract features for the cup (20) and obtain a new neural network model by changing the weight vector values ​​of the previously stored neural network model based on the extracted features. Then, the electronic device (100) can perform a function corresponding to the user command (e.g., a function to provide the location of the cup (20) to the user) using the obtained new neural network model.

[0023] Below, a method for controlling an electronic device (100) to perform such a function will be described.

[0024] FIG. 2 is a flowchart illustrating a control method of an electronic device (100) according to one embodiment of the present disclosure.

[0025] An electronic device (100) may acquire a neural network model trained to detect objects corresponding to at least one class (S210). At this time, the electronic device (100) may acquire a neural network model trained from an external server or an external device. Then, the electronic device (100) may use the acquired neural network model to detect objects corresponding to at least one class. For example, the electronic device (100) may acquire an image of the surroundings and input the acquired image into a neural network model to acquire type information and location information regarding objects included in the image. Meanwhile, in the present disclosure, "class" is a term for classifying multiple objects and may be used interchangeably with "classification" and "category."

[0026] Meanwhile, the neural network model may include a feature extraction module (or feature extractor) that extracts feature values ​​(or features) for an object, and a classification value acquisition module (or classifier) ​​that acquires a classification value for an object based on the feature values ​​acquired from the feature extraction module. Additionally, the classification value acquisition module may include a weight vector comprising at least one column vector.

[0027] Neural network models can be trained using various training methods. In particular, neural network models can be trained based on machine learning methods. In this case, the neural network model can be trained based on a loss function that includes a predefined regularization function to prevent overfitting. Specifically, neural network models can be trained based on orthogonal constraints.

[0028] The electronic device (100) can obtain a user command for detecting a first object corresponding to a first class (S220). At this time, the user command may be in various forms. For example, the user command may include a command to store the first class as a new class in the electronic device (100) or a command to detect a first object corresponding to the first class. Additionally, the electronic device (100) can obtain an image of the first object.

[0029] When a user command is obtained, the electronic device (100) can determine whether the first object corresponds to at least one previously learned class that can be identified by a neural network model. Specifically, the electronic device (100) can obtain an image of the first object and input the image of the first object into a neural network model to obtain a first feature value for the first object. The electronic device (100) can determine whether it corresponds to at least one previously learned class by comparing the first feature value with the weight vector of the learned neural network model.

[0030] If the first object does not correspond to at least one class, the electronic device (100) acquires a new neural network model based on information about the neural network model and the first object (S230), and the electronic device (100) can detect the first object using the acquired new neural network model. At this time, the electronic device (100) can acquire a new neural network model based on few-shot learning. Few-shot learning refers to learning using a small amount of learning data compared to general learning. Meanwhile, unless otherwise specified, the neural network model below refers to the neural network model acquired in step S210, and the neural network model acquired through step S230 is referred to as the new neural network model.

[0031] The electronic device (100) can obtain a first feature value for the first object by inputting an image of the first object into a feature extraction module of a neural network model. Then, the electronic device (100) can obtain a new classification value acquisition module based on the first feature value and the classification value acquisition module of the neural network model. Specifically, the electronic device (100) can obtain a new classification value acquisition module by generating a first column vector based on the average value of the first feature value and adding the first column vector as a new column vector of the weight vector. In addition, the electronic device (100) can normalize the new classification value acquisition module based on a predefined normalization function.

[0032] Below, we will explain in more detail the training method for neural network models and the method for acquiring new neural network models based on the trained neural network models.

[0033] FIG. 3a is a diagram illustrating a conventional object recognition model using infinitesimal learning. When a new class is added to a previously trained object recognition model, the object recognition model can detect only objects corresponding to the new class, and cannot detect objects corresponding to classes that were previously trained before the new class was added. That is, the classifier (302) of the conventional object recognition model includes only weight vectors for the new class and does not include weight vectors for the existing class. In addition, in order to recognize objects corresponding to the new class, a process of training the classifier (302) based on training data corresponding to the new class was required.

[0034] FIG. 3b is a diagram illustrating a training method for a conventional object recognition model. Referring to FIG. 3b, the object recognition model includes a feature extractor (300) and a classifier (302). The feature extractor (300) extracts features for an input training sample, and the classifier (302) outputs a classification value for the extracted features. The classifier (302) includes a weight vector (W). The object recognition model is trained by backpropagation, which minimizes an error calculated based on a loss function defined based on the output classification value and labeling data corresponding to the training sample.

[0035] FIG. 4a is a diagram illustrating a method for obtaining a new neural network model according to an embodiment of the present disclosure. An electronic device (100) can train a neural network model based on a very small number of learnings. The neural network model may include a feature extraction module (410) and a classification module (or classification value acquisition module) (420). The classification module (420) may include a weight vector composed of a plurality of column vectors or row vectors. Each column vector of the weight vector may correspond to a respective class. The classification module (420) may include a base portion (421) and a novel portion (422). The base portion (421) includes a vector corresponding to a previously learned class, and the novel portion (422) may include a vector corresponding to a novel class based on user input. The novel portion (422) may also be used interchangeably with the local portion.

[0036] The electronic device (100) can obtain feature values ​​for an object (41) by inputting the object (41) corresponding to a previously learned class (i.e., included in the base region) into a feature extraction module (410). The electronic device (100) can obtain classification values ​​for an object (41) by inputting the obtained feature values ​​into a classification module (420). At this time, the electronic device (100) can perform an inner product operation between the feature values ​​for the object (41) and the weight vector included in the classification module (420). Additionally, the neural network model can output classification values ​​based on the cosine distance between the vector for the feature values ​​and the weight vector of the classification module (420).

[0037] Meanwhile, the electronic device (100) may receive a request from a user for a first object (42) that does not correspond to a previously learned class (i.e., is not included in the base area (421)). At this time, the electronic device (100) may input the first object (42) into a feature extraction module (410) to obtain a first feature value (43). Based on the first feature value (43), the electronic device (100) may obtain a new weight vector (or column vector) to be assigned or stored in a new area (422). For example, the electronic device (100) may average the first feature value (43) and store the averaged first feature value (43) in the new area (422). Accordingly, the new area (422) may include a weight vector corresponding to a new class corresponding to the first object (42). In this way, the electronic device (100) can obtain a new neural network model by updating a new region (422) of the classification module (420) based on the first feature value (43).

[0038] Meanwhile, referring again to FIG. 3a, when a conventional neural network model (or an electronic device to which a neural network model is applied) is trained to recognize an object corresponding to a new class, the classifier (302) is updated to include a weight vector corresponding to the new class, and thus does not include a weight vector corresponding to a previously trained class. Accordingly, the conventional neural network model could no longer recognize an object corresponding to a previously trained class. On the other hand, the new neural network model according to the present disclosure includes both a base region (421) and a new region (422), even when trained to recognize an object corresponding to a new class. Therefore, the electronic device (100) can recognize not only the first object (42) but also the previously trained object (41) using the new neural network model.

[0039] FIG. 4b is a drawing for explaining a method of learning a neural network model according to one embodiment of the present disclosure.

[0040] The neural network model may include a feature extraction module (410) and a classification value acquisition module (420). The feature extraction module (410) is trained to extract features (or feature values) for an input training sample. The classification value acquisition module (420) is trained to output a classification value for a training sample based on the feature values ​​output from the feature extraction module (410). The neural network model may be trained based on an orthogonality score calculated based on a predefined function, a classification value, and a loss function defined based on labeling data corresponding to the training sample. The neural network model may be trained according to an error inversion method that minimizes the error calculated based on the loss function. Here, the orthogonality score may be a normalized value of the weight vector of the classification value acquisition module (420).

[0041] Meanwhile, the conventional object recognition model according to FIG. 3b is not trained based on orthogonality scores, so normalization of the classifier (302) is not performed. Consequently, the conventional object recognition model is overfitted to a previously trained class and cannot recognize objects of a new class. In contrast, the neural network model according to the present disclosure is normalized based on orthogonality scores, so it is not overfitted to a specific class. In addition, compatibility between weight vectors included in the base region (421) and weight vectors added to the new region (422) can be improved. That is, if the neural network model is trained based on orthogonality constraints, the user can customize the neural network model more easily. Furthermore, the electronic device (100) can perform normalization based on orthogonality constraints on the weight vectors added to the new region (422) as well.

[0042] FIG. 5a is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.

[0043] The electronic device (100) may include a camera (110), a microphone (120), a communication interface (130), a memory (140), a display (150), a speaker (160), and a processor (170). For example, the electronic device (100) may be any user device such as a smartphone, tablet, laptop, computer or computing device, virtual assistant device, robot and robotic device, consumer goods / home appliance device (e.g., smart refrigerator), Internet of Things device, or video recording system / device, but is not limited thereto. Each component is described below.

[0044] The camera (110) can capture images around the electronic device (100) to acquire images. Additionally, the camera (110) can acquire user commands. For example, it can acquire images of objects provided by the user or video footage capturing the user's gestures. The camera (110) can be implemented as various types of cameras. For example, the camera (110) can be either a 2D-based RGB camera or an IR camera. Alternatively, the camera (110) can be either a 3D-based Time of Flight (ToF) camera or a stereo camera.

[0045] A microphone (120) is configured to receive a user's voice and may be provided within an electronic device (100), but this is merely one embodiment, and may be connected to the electronic device (100) externally via a wired or wireless connection. In particular, the microphone (120) may receive a user's voice for searching for a specific object.

[0046] The communication interface (130) includes at least one circuit and can perform communication with various types of external devices. For example, the communication interface (130) can perform communication with an external server or a user terminal. Additionally, the communication interface (130) can perform communication with external devices according to various types of communication methods. The communication interface (130) can perform data communication wirelessly or wired. When performing communication with an external device using a wireless communication method, the communication interface (130) may include at least one of a Wi-Fi communication module, a cellular communication module, a 3G (3rd generation) mobile communication module, a 4G (4th generation) mobile communication module, a 4th generation LTE (Long Term Evolution) communication module, and a 5G (5th generation) mobile communication module. Meanwhile, according to one embodiment of the present disclosure, the communication interface (130) may be implemented as a wireless communication module, but this is merely one embodiment, and it may be implemented as a wired communication module (e.g., LAN, etc.).

[0047] The memory (140) can store commands or data related to an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and the components of the electronic device (100). To this end, the memory (140) can be implemented as non-volatile memory (e.g., hard disk, SSD (Solid state drive), flash memory), volatile memory, etc.

[0048] For example, memory (140) may store instructions that, when executed, cause the processor (150) to acquire type information and location information for objects included in the image when an image is acquired from the camera (110). Additionally, memory (140) may store a neural network model for recognizing objects. In particular, the neural network model may be executed by a conventional general-purpose processor (e.g., CPU) or a separate AI-dedicated processor (e.g., GPU, NPU, etc.). Furthermore, memory (140) may store data for an application that allows a user to request the addition of a new class to the neural network model.

[0049] The display (150) can display various screens. For example, the display (150) can display an application execution screen, allowing the user to input a request for a new class using the application. Additionally, the display (150) can display an object requested by the user, or display prompts or notifications generated by the electronic device (100). Meanwhile, the display (150) can be implemented as a touch screen. In this case, the processor (170) can obtain the user's touch input through the display (150).

[0050] The speaker (160) may be a component that outputs various audio data received externally, as well as various notification sounds or voice messages. At this time, the electronic device (100) may include an audio output device such as the speaker (160), but may also include an output device such as an audio output terminal. In particular, the speaker (160) may provide response results and operation results for user voice in the form of voice.

[0051] The processor (170) can control the overall operation of the electronic device (100). For example, the processor (170) may acquire a neural network model trained to detect objects corresponding to at least one pre-set class. The acquired neural network model may be a neural network model trained to acquire type information about objects included in an image. The neural network model may include a feature extraction module that extracts feature values ​​for objects and a classification value acquisition module that acquires classification values ​​for objects based on the feature values ​​acquired from the feature extraction module.

[0052] The processor (170) can obtain a user command to detect a first object corresponding to a first class. The processor (170) inputs an image of the first object into a learned neural network model to obtain a first feature value for the first object, and compares the first feature value with the weight vector of the learned neural network model to determine whether the first object belongs to at least one class.

[0053] If the first object does not correspond to at least one class, the processor (170) can obtain a new neural network model based on the neural network model and information about the first object. Specifically, the processor (170) can obtain an image of the first object and input the obtained image of the first object into a feature extraction module to obtain a first feature value for the first object. Then, the processor (170) can obtain a new classification value acquisition module based on the first feature value and the classification value acquisition module. At this time, the classification value acquisition module may include a weight vector containing a plurality of column vectors. The processor (170) can obtain a new classification value acquisition module by generating a first column vector based on the average value of the first feature value and adding the first column vector as a new column vector of the weight vector. The processor (170) can normalize the obtained new classification value acquisition module based on a predefined normalization function.

[0054] The processor (170) can customize the neural network model. The processor (170) receives a user request for a new class and can determine whether the new class is new and should be added to the neural network model. If the new class is determined to be new, the processor (170) can obtain at least one sample representing the new class. At this time, the processor (170) can obtain at least one of an image, an audio file, an audio clip, a video, and a frame of a video as a sample.

[0055] The processor (170) can obtain at least one feature extracted from at least one sample from a neural network model including a feature extraction module and a basic area of ​​a classification value acquisition module. At this time, the processor (170) can transmit a user request and at least one sample to an external server including the neural network model through a communication interface (130). The processor (170) can receive a feature for at least one sample from the external server. The processor (170) can store the extracted at least one feature as a representative of a new class. At this time, the processor (170) can store a weight vector of the classification value acquisition module corresponding to the new class in memory (140).

[0056] The processor (170) can obtain at least one keyword related to a new class. The processor (170) can determine whether at least one keyword matches one of a plurality of predefined keywords in the base area of ​​the classification value acquisition module of the neural network model. If at least one keyword matches one of the plurality of predefined keywords, the processor (170) can identify a class corresponding to the matched predefined keyword. The processor (170) can control a display (150) or a speaker (160) to output an example sample corresponding to the identified class and a suggestion to assign at least one keyword to the identified class. At this time, if user confirmation is obtained that at least one keyword should be assigned to the identified class, the processor (170) can assign at least one keyword to the identified class. Conversely, if user input is obtained that disallows the assignment of at least one keyword to the identified class, the processor (170) can add a new class to the neural network model.

[0057] Meanwhile, the electronic device (100) according to the present disclosure may be effective in terms of protecting the user's privacy. This is because the new class is stored in the electronic device (100) rather than being stored in a cloud accessible to other users or added to the base part of the classifier. However, the user may wish to share the new class defined by the user with other devices of the user (e.g., from a smartphone to a laptop, virtual assistant, robot butler, smart refrigerator, etc.). Accordingly, the processor (170) may share the new class stored in the local area of ​​the classification value acquisition module with external devices. The processor (170) may share the new class stored in the local area of ​​the classification value acquisition module with an external server including the base area of ​​the classification value acquisition module. This may be executed automatically. For example, if a neural network model is used as part of a camera application, when the neural network model is updated on the user's smartphone, the neural network model may be automatically shared with other devices of the user where the same camera application is running. Thus, sharing may be part of software application synchronization between many devices.

[0058] In particular, the artificial intelligence-related functions according to the present disclosure are operated through a processor (170) and a memory (140). The processor (170) may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as a CPU, AP, DSP (Digital Signal Processor), graphics-dedicated processors such as a GPU, VPU (Vision Processing Unit), or artificial intelligence-dedicated processors such as an NPU. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in the memory (140). Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0059] The predefined behavioral rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined behavioral rules or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, Generative Adversarial Networks, or reinforcement learning, but are not limited to the examples described above.

[0060] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.

[0061] Meanwhile, the electronic device (100) according to the present disclosure enables the neural network model to be customized efficiently in terms of time, resources, and cost while ensuring the accuracy of the neural network model is maintained. This is achieved by locally extending the classification value acquisition module of the neural network model in the electronic device (100). In other words, the electronic device (100) does not make changes to the entire previously learned classification value acquisition module. This means that the neural network model does not need to be retrained from scratch, and the neural network model can be updated quickly. Furthermore, this means that there is no need to use expensive cloud computing to update / customize the neural network model.

[0062] Meanwhile, the electronic device (100) can obtain a neural network model from an external server.

[0063] FIG. 5b is a diagram illustrating a method for obtaining a new neural network model according to an embodiment of the present disclosure. An electronic device (100) may obtain a neural network model (180) including a feature extraction module (181) and a classification value acquisition module (182) from an external server (500). The classification value acquisition module (182) may include a base area (182-1). The electronic device (100) may obtain a new classification value acquisition module from the classification value acquisition module (182) obtained from the external server (500). Specifically, the electronic device (100) may average the feature values ​​for an object corresponding to a new class and store them in a local area (182-2). Accordingly, the electronic device (100) may detect an object corresponding to a new class. Meanwhile, the operation of obtaining feature values ​​for an object corresponding to a new class may be performed by the external server (500). In this case, the electronic device (100) can obtain information about an object corresponding to a new class from the user and transmit it to an external server (500), and receive information about an object corresponding to a new class extracted by the external server (500).

[0064] FIG. 6a is a diagram illustrating a method for customizing a neural network model using image samples according to an embodiment of the present disclosure. In this embodiment, a user may try to find a picture of the user's dog on an electronic device (100). The electronic device (100) may be a smartphone. However, the electronic device (100) may store images of numerous other dogs in addition to the user's dog. An image gallery application of the electronic device (100) may search for the location of all photos containing a dog using a text-based keyword search, but it may not be able to search for the location of an image containing the user's dog because there is no keyword or class corresponding to the user's dog.

[0065] In step (S600), the electronic device (100) can obtain a user command to select an image gallery application installed on the electronic device (100). Then, the user can enter the settings section of the image gallery application installed on the electronic device (100). The electronic device (100) can display a screen for "adding a new search category." The user can add a "new search category" on the screen displayed on the electronic device (100). Accordingly, the electronic device (100) can perform an action to customize the neural network model. The electronic device (100) can induce the user to input a request for a new class. The electronic device (100) can induce the user to input a new category keyword. In this case, in step (S602), the user can input keywords related to the new category, such as "German Shepherd," "my dog," and "Laika" (dog name).

[0066] In step (S604), the user can take a video of the dog using a camera or add a photo of the user's dog from an image gallery to the application. Accordingly, the electronic device (100) can store the new category entered by the user. Then, when the user enters the keyword "my dog" into the search function of the image gallery application, the electronic device (100) can display an image of the user's dog and provide it to the user (S606).

[0067] Additionally, users can use the settings section of the image gallery application to delete categories that are no longer requested by removing keywords associated with those categories. This allows the classifier weight vector associated with those keywords to be removed from the local portion of the classifier.

[0068] FIG. 6b illustrates the customization of a neural network model using a video frame sample according to an embodiment of the present disclosure. In this embodiment, the user may not like the default palm gesture (as in S610) used on the user's smartphone to give a command to the smartphone camera when taking a selfie. The user may want to register a new gesture for "taking a selfie." In step (S612), the user enters the camera settings section and selects "Set new selfie taking gesture." The user starts recording a custom gesture of moving the head from left to right (S614). In step (S616), the user can confirm the user's selection. Thus, the user can use the user's new gesture to activate the selfie taking action (S618).

[0069] As described above, it may be desirable to register a new class in the local part of a machine learning model's classifier only when the base part of the classifier (or the local part if it already exists) does not already contain a similar or substantially identical class. If a similar or substantially identical class already exists in the classifier, the user may notify the existence of the class and suggest connecting the corresponding keyword to the existing class. The user may accept or reject this suggestion; if rejected, the process of adding the class provided by the user to the model continues.

[0070] FIG. 7 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.

[0071] The electronic device (100) can obtain a user command for the first class (S710). The user can make this request in an appropriate way, such as by installing an application that allows the user to interact with the neural network model. The application may be, for example, an application related to the camera of the electronic device (100) or an application used to compare images and videos captured by the camera.

[0072] The electronic device (100) can determine whether the first class is a new class and whether it should be added to a neural network model already stored in the electronic device (100) (S720). This determination can be made to avoid class duplication in the model, which may result in inefficient model operation. An example of a method for determining whether the first class requested by the user is a new class will be described below with reference to FIG. 8.

[0073] If the first class is determined to be a new class in step (S720), the electronic device (100) may acquire at least one sample representing the first class (S730). The at least one sample may be one or more of an image, an audio file, an audio clip, a video, and a frame of a video. Generally, the at least one sample may be a set of images representing the same object (or feature) to be used to define the first class. For example, if a user wants a neural network model to identify the user's dog in images and videos, the user may provide one or more photos of the user's dog as input samples representing the first class. If multiple samples are acquired, the samples may all be of the same type / file format (e.g., images) or different types (e.g., images and videos). In other words, the user may provide both photos and videos of the user's dog as input samples.

[0074] One sample representing the first class may be sufficient to customize the neural network model. However, as with other machine learning techniques, usually more samples can be used to improve or obtain better results. If the quality of the acquired samples is poor or insufficient to define the first class and add it to the neural network model, the electronic device (100) may output a message requesting the user to input more samples.

[0075] In some cases, the user request at step (S710) may include a sample representing a new class (i.e., a first class). In this case, at step (S730), the electronic device (100) may simply use a sample that has already been received. In some cases, the user request at step (S710) may not include any sample. In this case, at step (S730), the electronic device (100) may include a guide message that induces the user to provide / input a sample. Alternatively, the sample may be received at step (S720), and thus at step (S730), the electronic device (100) may use a sample obtained at step (S720).

[0076] The electronic device (100) can obtain feature values ​​for the acquired sample (S740). At this time, the electronic device (100) can obtain feature values ​​for the sample by inputting the acquired sample into a feature extraction module included in the neural network model. All or part of the neural network model may be implemented in the electronic device (100) or a remote server / cloud server.

[0077] The electronic device (100) can store the acquired feature values ​​in a local area of ​​the classification value acquisition module (S750). By doing so, the electronic device (100) can acquire a neural network model capable of recognizing objects corresponding to the first class.

[0078] Figure 8 is a flowchart of an example of checking whether a user request class already exists in a neural network model.

[0079] The electronic device (100) receives a user request for a new class (S800) and can receive at least one keyword associated with the new class (S802). The electronic device (100) can determine whether the received at least one keyword matches a predefined keyword (S804). As a result of the determination, if the local area of ​​the classifier already exists, the electronic device (100) can match keywords associated with the base area and the local area of ​​the classifier.

[0080] If the received keyword matches a predefined keyword, the electronic device (100) can identify a class corresponding to the matched keyword (S806). Then, the electronic device (100) can output a proposal to assign the received keyword to the identified existing class (S808). At this time, the electronic device (100) can also output an example sample corresponding to the identified class to explain why the new class is similar / identical to the identified existing class.

[0081] The electronic device (100) can determine whether the user has accepted the proposal (S810). At this time, the electronic device (100) can determine whether the user has accepted the proposal based on the user's response. If it is determined that the user has accepted the proposal, the electronic device (100) can assign the keyword to the identified class (S812). On the other hand, if it is determined that the user has not accepted the proposal, the electronic device (100) can perform an operation to add a new class to the neural network model, and this operation can lead to step S730 of FIG. 7.

[0082] If the keyword received in step (S804) does not match any of the predefined keywords, the electronic device (100) may receive at least one sample representing the received class (S814). Then, the electronic device (100) may obtain a feature value of the received sample and determine whether the obtained feature value matches an existing class (S818). The electronic device (100) may determine this by calculating the inner product of the feature vector extracted from the received sample and each classifier weight vector of the classifier.

[0083] If it is determined that the feature value of the sample matches an existing class, the electronic device (100) can output a proposal to assign the received keyword to the identified class (S808). On the other hand, if it is determined that the feature value of the sample does not match an existing class, the electronic device (100) can perform an operation to add a new class to the neural network model, and this operation can lead to step S730 of FIG. 7.

[0084] Meanwhile, the various embodiments described above may be implemented in a recording medium readable by a computer or a similar device using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented as the processor itself. According to software implementation, embodiments such as the procedures and functions described herein may be implemented as separate software modules. Each of the software modules may perform one or more functions and operations described herein.

[0085] Meanwhile, the various embodiments described above may be implemented in a recording medium readable by a computer or a similar device using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented as the processor itself. According to software implementation, embodiments such as the procedures and functions described herein may be implemented as separate software modules. Each of the software modules may perform one or more functions and operations described herein.

[0086] Meanwhile, computer instructions for performing processing operations according to the various embodiments of the present disclosure described above may be stored in a non-transitory computer-readable medium. When computer instructions stored in such a non-transitory computer-readable medium are executed by a processor, they may cause a specific device to perform processing operations according to the various embodiments described above.

[0087] A non-transient computer-readable medium refers to a medium that stores data semi-permanently and can be read by a device, unlike media that store data for a short period of time such as registers, caches, and memory. Specific examples of non-transient computer-readable media include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.

[0088] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory storage medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.

[0089] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0090] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure. Explanation of the symbols

[65535] 100 : Electronic device

Claims

Claim 1 A control method for an electronic device comprises: a step of obtaining from an external server a neural network model trained to detect an object corresponding to at least one class; a step of obtaining a user command for detecting a first object corresponding to the first class, the command including at least one keyword associated with the first class; and, if the first object does not correspond to the at least one class, a step of obtaining a new neural network model based on information regarding the neural network model obtained from the external server and the first object; wherein the new neural network model is stored in the electronic device and is obtained by updating the neural network model obtained from the external server in the electronic device to detect not only the object corresponding to the at least one class but also the first object corresponding to the first class. Claim 2 A control method according to claim 1, wherein the learned neural network model comprises a feature extraction module for extracting feature values ​​for an object and a classification value acquisition module for acquiring a classification value for the object based on the feature values ​​obtained from the feature extraction module. Claim 3 A control method according to claim 2, wherein the step of acquiring the user command includes the step of acquiring an image of the first object, and the step of acquiring the new neural network model includes the step of inputting the acquired image of the first object into the feature extraction module to acquire a first feature value for the first object, and the step of acquiring a new classification value acquisition module based on the first feature value and the classification value acquisition module. Claim 4 In claim 3, the classification value acquisition module includes a weighting vector comprising a plurality of column vectors, and the step of acquiring the new classification value acquisition module comprises generating a first column vector based on the average value of the first feature value and adding the first column vector as a new column vector of the weighting vector to acquire the new classification value acquisition module. Claim 5 A control method further comprising, in claim 4, a step of normalizing the acquired new classification value acquisition module based on a previously defined normalization function. Claim 6 A control method according to claim 1, wherein the learned neural network model is learned based on a loss function including a predefined regularization function to prevent overfitting. Claim 7 A control method according to claim 1, further comprising the step of determining whether the first object corresponds to the at least one class when the user command is obtained; wherein the determining step comprises obtaining an image of the first object, inputting the image of the first object into the learned neural network model to obtain a first feature value for the first object, and comparing the first feature value with the weight vector of the learned neural network model to determine whether the first object corresponds to the at least one class. Claim 8 An electronic device comprising: a memory including at least one instruction; and a processor; wherein the processor obtains from an external server a neural network model trained to detect an object corresponding to at least one class, and obtains a user command for detecting a first object corresponding to the first class, which includes at least one keyword associated with the first class, and if the first object does not correspond to the at least one class, obtains a new neural network model based on information regarding the neural network model obtained from the external server and the first object, wherein the new neural network model is stored in the electronic device and is obtained by updating the neural network model obtained from the external server in the electronic device to detect not only the object corresponding to the at least one class but also the first object corresponding to the first class. Claim 9 In claim 8, the learned neural network model comprises an electronic device including a feature extraction module for extracting feature values ​​for an object and a classification value acquisition module for acquiring a classification value for the object based on the feature values ​​obtained from the feature extraction module. Claim 10 An electronic device according to claim 9, wherein the processor acquires an image of the first object, inputs the acquired image of the first object into the feature extraction module to acquire a first feature value for the first object, and acquires a new classification value acquisition module based on the first feature value and the classification value acquisition module. Claim 11 An electronic device according to claim 10, wherein the classification value acquisition module comprises a weight vector including a plurality of column vectors, and the processor generates a first column vector based on the average value of the first feature value and acquires the new classification value acquisition module by adding the first column vector as a new column vector of the weight vector. Claim 12 In claim 11, the processor is an electronic device that normalizes the acquired new classification value acquisition module based on a predefined normalization function. Claim 13 In claim 8, the electronic device wherein the learned neural network model is learned based on a loss function including a predefined regularization function to prevent overfitting. Claim 14 In claim 8, the processor acquires an image of the first object, inputs the image of the first object into the learned neural network model to acquire a first feature value for the first object, and compares the first feature value with the weight vector of the learned neural network model to determine whether the first object belongs to at least one class. Claim 15 A computer-readable recording medium having a program recorded thereon for performing a method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • User-definded machine learning apparatus for smart phone user and method thereof

    KR1020190023787A

  • Machine learning device and classification device for accurately classifying into category to which content belongs

    US20160125273A1

  • Neural network learning method and device for recognizing class

    WO2019050247A2