Electronic device capable of recognizing user, and control method therefor

The electronic device effectively addresses the challenge of identifying the correct user among multiple users by using gesture recognition and account information management to perform control operations based on user inputs.

WO2025121693A1PCT designated stage expired Publication Date: 2025-06-12SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017415
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing electronic devices struggle to accurately identify the user with control authority when multiple users share the same device, using only trigger gestures or voices.

Method used

An electronic device equipped with an interface, memory, and processor that identifies a user's gesture for account registration, provides candidate words based on gesture characteristics, and stores account information including selected words and gesture data. When a voice and gesture are identified, the device performs control operations based on corresponding account information.

Benefits of technology

The solution enables accurate user recognition in a multi-modal manner, allowing the device to perform operations corresponding to the user's gesture or voice input, thereby addressing the challenge of identifying the correct user among multiple users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017415_12062025_PF_FP_ABST
    Figure KR2024017415_12062025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device comprises an interface, a memory and one or more processors, wherein the one or more processors can: provide one or more candidate words on the basis of feature information of a gesture when the gesture for account registration is identified on the basis of sensing data acquired through the interface; store, in the memory, account information including the selected candidate word and the feature information of the gesture when one or more is selected from among the one or more one candidate words; and perform a control operation on the basis of the account information corresponding to the voice and gesture of the user from among a plurality of pieces of account information stored in the memory when the voice and gesture of the user are identified.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device capable of recognizing a user and method for controlling the same

[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device that recognizes a user through a detected gesture of the user and a method for controlling the same.

[0002] Recently, various electronic devices such as mobile phones, tablets, and TVs are increasingly recognizing users' gestures or voices and performing actions corresponding to those gestures or voices.

[0003] In order to use an electronic device in gesture recognition mode or voice recognition mode, you may need to input a specific trigger gesture or trigger voice.

[0004] For example, when the voice "Hi XXX" is input, the electronic device can initiate voice recognition mode, recognize the user's voice input thereafter, and perform corresponding actions. Specifically, if the user utters "Tell me today's weather," the electronic device can provide a response corresponding to the utterance. If the voice "Hi XXX" is not input, the electronic device may not respond in any way.

[0005] However, when multiple users use one electronic device, there is a problem in that it is difficult for the electronic device to identify the user with control authority using only a trigger gesture or trigger voice.

[0006] Therefore, there is a growing need for a technology that identifies and responds to a recognized gesture or voice input by which user among multiple users.

[0007] According to one aspect of the present disclosure, an electronic device includes an interface, a memory, and at least one processor, wherein the at least one processor, when a gesture for account registration is identified based on sensing data acquired through the interface, provides at least one candidate word based on characteristic information of the gesture, and when at least one of the at least one candidate word is selected, stores account information including the selected candidate word and characteristic information of the gesture in the memory, and when a user's voice and gesture are identified, performs a control operation based on account information corresponding to the user's voice and gesture among a plurality of pieces of account information stored in the memory.

[0008] According to one aspect of the present disclosure, a control method of an electronic device may include, when a gesture for account registration is identified, providing at least one candidate word based on characteristic information of the gesture; when at least one of the at least one candidate word is selected, storing account information including the selected candidate word and characteristic information of the gesture; and when a voice and gesture of a user are identified, performing an operation based on account information corresponding to the voice and gesture of the user among a plurality of stored account information.

[0009] According to one aspect of the present disclosure, a computer-readable recording medium including a program for executing a control method of an electronic device may include, when a gesture for account registration is identified, providing at least one candidate word based on characteristic information of the gesture; when at least one of the at least one candidate word is selected, storing account information including the selected candidate word and characteristic information of the gesture; and when a voice and gesture of a user are identified, performing an operation based on account information corresponding to the voice and gesture of the user among a plurality of stored account information.

[0010] FIG. 1 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.

[0011] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.

[0012] FIG. 3 is a diagram illustrating a word providing method of an electronic device according to one or more embodiments of the present disclosure.

[0013] FIG. 4 is a diagram illustrating a method for providing a live view of an electronic device according to one or more embodiments of the present disclosure.

[0014] FIG. 5 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.

[0015] FIG. 6 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.

[0016] FIG. 7 is a diagram for explaining a gesture recognition mode operation method of an electronic device according to one or more embodiments of the present disclosure.

[0017] FIG. 8 is a diagram illustrating a method for switching user accounts of an electronic device according to one or more embodiments of the present disclosure.

[0018] FIG. 9 is a diagram illustrating a method for providing a candidate word of an electronic device according to one or more embodiments of the present disclosure.

[0019] FIG. 10 is a block diagram illustrating a detailed configuration of an electronic device according to one or more embodiments of the present disclosure.

[0020] FIGS. 11 and 12 are flowcharts illustrating a method for storing account information in an electronic device according to one or more embodiments of the present disclosure.

[0021] FIG. 13 is a flowchart illustrating a method for identifying a user account of an electronic device according to one or more embodiments of the present disclosure.

[0022] Hereinafter, terms used in this specification will be briefly described, and the present disclosure will be described in detail. In this disclosure, the expression “at least one of a, b, or c” can refer to “a,” “b,” “c,” “a and b,” “a and c,” “b and c,” “all of a, b, and c,” or variations thereof.

[0023] The terms used in this disclosure are selected from widely used, common terms, taking into account the functions of the disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names, but rather based on the meanings of the terms and the overall content of the disclosure.

[0024] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein. Furthermore, terms containing ordinal numbers, such as "first" or "second," used herein may be used to describe various components, but such components should not be limited by such terms. Such terms are used solely to distinguish one component from another.

[0025] When a part of the specification is said to "include" a component, unless otherwise specifically stated, this does not exclude other components but rather implies the inclusion of other components. Furthermore, terms such as "part" and "module" used in the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.

[0026] The present disclosure relates to an electronic device and method for more accurately recognizing a user in a multi-modal manner and efficiently performing an action corresponding to the user's gesture or voice. An electronic device according to an embodiment of the present disclosure can acquire a user's gesture and voice and identify which user account among multiple user accounts stored in the electronic device corresponds to the user's gesture and voice.

[0027] In the present disclosure, "account information" may be identification information registered for each user in an electronic device. For example, the account information may include a word corresponding to a gesture, a word corresponding to an object, a word corresponding to a distance, a word corresponding to a section, and a wake-up word. Here, a gesture may refer to various actions taken by a user while holding his or her body or other items. An object may refer to an item held by the user or designated by the user. For example, if a user holds and shakes a cup or a mobile phone, a word corresponding to the cup, mobile phone, etc. may be registered in the account information. The distance may refer to the distance between the user and the electronic device, and the section may refer to information indicating the direction in which the user is located relative to the center of the electronic device.

[0028] In the present disclosure, the "wake-up word" is a word registered in user account information, and may be a word that induces an electronic device to initiate an operation to identify a user account, or a word that commands the electronic device to initiate a specific mode. Upon input of the "wake-up word," the electronic device initiates an operation to compare a user gesture input to the electronic device with characteristic information of the gesture already stored in the electronic device. Furthermore, the "wake-up word" may be replaced by various words that can initiate an operation of the electronic device, such as a "trigger word," an "activation word," or a "start word."

[0029] "Account information" for multiple users may be stored on an electronic device and used to identify each user. Furthermore, "account information" may be replaced by other terms indicating information that can identify each user, such as "user identification information" and "user personal information."

[0030] In the present disclosure, "gesture feature information" may be information indicating a feature that differentiates a detected gesture from other gestures. Gesture feature information may include information about a word corresponding to the gesture, a word corresponding to the area where the gesture was detected, a word corresponding to the distance where the gesture was detected, and the like.

[0031] In the present disclosure, "object feature information" may be information indicating features that differentiate a detected object from other objects. Object feature information may include information about a word corresponding to the object, a word corresponding to the area where the object was detected, a word corresponding to the distance at which the object was detected, and the like.

[0032] In the present disclosure, the "feature map of a gesture" and the "feature map of an object" may include data in the form of a two-dimensional array obtained through a trained neural network model. Specifically, the "feature map" is data that can be obtained by inputting a photographed image into the trained neural network model, and is distinct from the "feature information of a gesture" and the "feature information of a gesture" described above.

[0033] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts irrelevant to the description are omitted for clarity of description of the present disclosure, and similar parts are designated with similar reference numerals throughout the specification.

[0034] The present disclosure will be described below with reference to the attached drawings.

[0035] FIG. 1 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.

[0036] In one embodiment, the electronic device (100) can identify a user's gesture (200), recognize the identified gesture, and provide a candidate word to the user. While FIG. 1 illustrates a case where the user makes a gesture using his or her hand, the user can also make a gesture using any object.

[0037] The electronic device (100) may provide multiple candidate words based on the user's gesture (200), and the user may select one candidate word and store it as a wake-up word in the user account information corresponding to the user. The wake-up word may be information that causes the electronic device to perform an operation to identify the user account or an operation to switch to a specific mode, as described above. The candidate word may be a word likely to be used as a wake-up word.

[0038] For example, if a user makes a gesture (200) of spreading out two fingers to form a shape similar to the letter "V," the electronic device (100) detects the gesture (200) and inputs the detected data into a learned neural network model. Based on the output of the neural network model, the electronic device (100) can provide various candidate words, such as "Victory," "V," "Two," "Second," etc.

[0039] The user can select at least one of the candidate words as the wake-up word. For example, if the user selects the provided candidate word "V" as the wake-up word, the electronic device (100) stores the candidate word in the user account information. In this state, if the electronic device (100) recognizes the word "V" through the user's voice, it can identify the user's account information among the multiple account information stored in the electronic device (100).

[0040] The electronic device (100) can store not only the wake-up word in the user account information, but also the characteristic information of the gesture (200) in the user account information. Since the characteristic information of the gesture (200) is also stored in the user account information, the electronic device (100) can accurately identify the account information of the corresponding user by comparing the characteristic information of the user's gesture (200) with the gesture characteristic information of a plurality of previously stored account information.

[0041] The characteristic information of the gesture (200) may include information about a word corresponding to the user's gesture (200) recognized by the electronic device (100), a word corresponding to the distance between the detected user's gesture (200) and the electronic device (100), a word corresponding to the section in which the user's gesture (200) is detected by the electronic device (100), and may also include information about the time the user maintained the gesture (200).

[0042] Specifically, if there are two users who have registered the wake-up word “V,” the electronic device (100) can identify the two user account information that have registered the wake-up word when it is determined that the wake-up word “V” has been uttered.

[0043] Alternatively, even if the gesture "V" (200) is registered in both user account information, the distance and the detected area of ​​the gesture "V" (200) may be different. Therefore, the electronic device (100) can finally identify one of the two user account information by comparing the characteristic information of the gesture (200). That is, assuming that user A registers the V gesture while standing on the left side with respect to the center of the electronic device (100) at a distance of 2 meters from the electronic device (100), and user B registers the V gesture while standing in front of the electronic device (100) at a distance of 3 meters from the electronic device (100), the electronic device (100) can identify that user A is connected when the user makes the "V" gesture on the left side of the electronic device (100) at a distance of 2 meters and utters "V".

[0044] Accordingly, the electronic device (100) according to the present disclosure recognizes the user in a 'multi-modal' manner that detects both voice and gesture corresponding to the wake-up word, thereby being able to recognize the user more accurately than the conventional 'single modal' manner.

[0045] In one embodiment, the electronic device (100) may be various types of devices capable of detecting a user's voice, gestures, distance from the user, etc. For example, the electronic device (100) may be various types of electronic devices capable of detecting a user's gestures through sensors, such as a TV, a smart phone, a tablet PC, a laptop PC, a desktop PC, a set-top box, a cleaning robot, a speaker, etc., but is not limited thereto. The electronic device (100) may directly have a built-in camera or microphone for receiving input of the user's gestures and voice, but is not necessarily limited thereto, and may be connected to a camera, a microphone, or an external device having these built-in, and may receive a captured image of the user or a voice signal of the user from these devices, or may receive a recognition result for a gesture or a recognition result for a voice signal.

[0046] Additionally, the user's gesture (200) may include a variety of motions, including not only hand motions such as a "V", but also a raising of the hand, a forward movement of the hand, a waving of the hand, a heart shape with the fingers, a scissors or a fist movement, etc.

[0047] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.

[0048] According to FIG. 2, the electronic device (100) includes an interface (110), a memory (120), and a processor (130).

[0049] The interface (110) may include a communication interface, a manipulation interface, an input / output interface, and the like. For example, the communication interface is a configuration for performing communication with at least one external device. The communication interface may include at least one wireless communication module, at least one wired communication module, and the like. Each communication module may be implemented in the form of at least one hardware chip. The wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, or other communication modules. In addition, the communication interface may include at least one communication chip that performs communication according to various wireless communication standards such as Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), and the like. The wired communication module may include, for example, at least one among a LAN (Local Area Network) module, an Ethernet module, a paired cable, a coaxial cable, a fiber optic cable, or a UWB (Ultra Wide-Band) module. The communication interface is implemented in various forms like this, and by performing communication with an external device, various data can be received from the external device.

[0050] The operation interface is a configuration for receiving user operation input. The operation interface may include various buttons, a touch screen, etc. provided on the main body of the electronic device (100).

[0051] The input / output interface is a configuration for inputting and outputting various external signals. The input / output interface can be connected to various external memories or external sources (e.g., web servers, user terminal devices, etc.) and can input various data. The input / output interface can be implemented as at least one interface among HDMI (High Definition Multimedia Interface), MHL (Mobile High-Definition Link), USB (Universal Serial Bus), USB C-type, DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (Dsubminiature), and DVI (Digital Visual Interface). At least some of the input / output interfaces may be connected to communication interfaces. For example, the input / output interface can transmit information received from an external device to the communication interface or transmit information received through the communication interface to the external device.

[0052] The interface (110) can be connected to a sensor or an external device that incorporates a sensor. The sensor is configured to detect a user's gesture, an object of the user, the distance between the user and the electronic device, etc. The sensor may include, for example, a camera, a LiDAR sensor, a TOF sensor, etc.

[0053] For example, if a camera is connected via a USB port in the interface (110), the electronic device (100) can acquire an image of the user captured by the camera. The processor (130) can detect the user's gesture or object through the acquired image.

[0054] Specifically, the camera can capture an image incident at the time of photographing the user using an image sensor and store the captured image data. The electronic device (100) can obtain an image corresponding to the user through the captured image data and detect the user's gesture or the user's object based on the image corresponding to the user.

[0055] As another example, when the interface (110) is connected to a lidar sensor or an external device including a lidar sensor, the lidar sensor can emit laser beams in various directions and collect the emitted laser beams when they are reflected from surrounding objects. The lidar sensor can obtain the round trip time of the laser beam. The processor (130) can perform sensing of the position, distance, and height of surrounding objects based on the information about the round trip time obtained from the lidar sensor.

[0056] In FIG. 2, a case in which sensors, etc. are connected through an interface (110) is illustrated and described, but depending on the embodiment, the electronic device (100) may also have a built-in camera, lidar sensor, etc.

[0057] When the electronic device (100) includes both a camera and a lidar sensor, the processor (130) can obtain visual information through images captured by the camera, and can obtain spatial information about the location, distance, etc. of objects based on the sensing values ​​of the lidar sensor. The processor (130) can obtain diverse and accurate sensing data by using multiple sensors.

[0058] The above-described camera, lidar sensor, etc. are only examples of sensors, and the electronic device (100) may be connected to or may have built-in various types of sensors that detect user gestures, user objects, the distance between the user and the electronic device, etc.

[0059] The memory (120) can store various programs, data, commands, etc. used in the electronic device (100). The memory (120) can also store information on multiple user accounts. In addition, the memory (120) can store various data and learning models according to various embodiments of the present disclosure, such as sensing data acquired through sensors and learned neural network models.

[0060] A memory (120) according to an example of the present disclosure may be implemented as an internal memory such as a ROM (e.g., an electrically erasable programmable read-only memory (EEPROM)) or a RAM included in one or more processors (130), or may be implemented as a separate memory from one or more processors (130). In this case, the memory (120) may be implemented as a memory embedded in the electronic device (100) or as a memory detachable from the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for an extended function of the electronic device (100) may be stored in a memory detachable from the electronic device (100).

[0061] Meanwhile, in the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)), and in the case of memory that can be detachably attached to the electronic device (100), it may be implemented as a memory card (e.g., compact flash (CF), secure digital (SD), micro secure digital (Micro-SD), mini secure digital (Mini-SD), extreme digital (xD), multi-media card (MMC), etc.), external memory that can be connected to a USB port (e.g., USB memory), etc. there is.

[0062] The processor (130) controls the overall operation of the electronic device (100). Specifically, the processor (130) is connected to the interface (110) and the memory (120), and can control the overall operation of the electronic device (100) by executing one or more commands or programs stored in the memory (120).

[0063] The processor (130) can obtain sensing data regarding the user's gesture through a sensor connected via the interface (110) or a built-in sensor. As described above, the sensor may correspond to a camera or a lidar sensor, and the processor (130) can obtain an image of the user's gesture through the sensor or obtain a 3D cloud point regarding the user's gesture.

[0064] Additionally, the processor (130) can obtain various types of sensing data through various types of sensors, such as being able to obtain distance information about a user's gesture when the sensor is a ToF sensor.

[0065] Additionally, the processor (130) can identify the user's gesture based on the acquired sensing data. Here, "identifying the user's gesture" may mean acquiring characteristic information of the user's gesture based on the sensing data.

[0066] Gesture identification can be implemented in various ways. For example, the processor (130) can analyze the pixel values ​​of each pixel within a plurality of image frames sequentially captured by the camera to identify a user, and track the identified user to identify the gesture.

[0067] Specifically, the processor (130) divides all pixels included in each of a plurality of consecutive image frames into a plurality of block units each consisting of n*m pixels. The processor (130) can detect a representative value representing the characteristics of the pixels in each block. The representative value may be, but is not limited to, an average pixel value of the pixels in each block, and may also be a maximum pixel value, a minimum pixel value, or an RMS (Root Means Square) value.

[0068] The processor (130) connects blocks having representative values ​​of a similar range among a plurality of blocks and arranged in consecutive positions to form a closed loop, and can identify the closed loop as an edge of an object included in a photographed image. The processor (130) can identify the size of the object based on the number of blocks included in the edge. In addition, the processor (130) can identify the shape of the object based on the shape of the edge, and thereby identify the type of the object. In addition, the processor (130) can determine in which area the user makes a gesture based on the position of the block corresponding to the edge.

[0069] When a user raises his or her hand and makes a gesture, the processor (130) can detect an edge connecting a plurality of blocks arranged similarly to the shape of a human hand. The processor (130) can track the change in position of blocks corresponding to the same edge in a plurality of consecutive image frames to identify whether the user is making a gesture of raising his or her hand in a V shape.

[0070] Alternatively, the processor (130) may identify the user's gesture using a learned neural network model.

[0071] For example, if a user makes a gesture called "V" at a distance of 2 m from an electronic device and in the upper right area relative to the center point of the electronic device, the processor (130) can obtain information that the gesture called "V" was detected at a distance of 2 m from the electronic device and in the upper right area of ​​the electronic device based on the sensing data. In addition, the processor (130) can input an image of the gesture called "V" into the learned neural network model stored in the memory (120) and output as a result that the gesture in the captured image corresponds to the word "V". As a result, the processor (130) can obtain characteristic information of the gesture that the user's gesture corresponds to the word "V", the distance sensed from the electronic device is "2 m", and the area sensed by the electronic device is "upper right" based on the sensing data. The method for obtaining a word corresponding to a gesture using the learned neural network model described above will be described in detail in the description of FIG. 3 described below.

[0072] Additionally, the processor (130) may provide at least one candidate word based on characteristic information of the gesture. In one embodiment, the processor (130) may select and provide a word corresponding to the gesture, a word corresponding to the distance at which the gesture was detected, and a word corresponding to the area at which the gesture was detected, respectively, as candidate words. For example, the processor (130) may provide the user with various candidate words, such as "V," "Victory," "Two," "second," "2m," or "RIGHT TOP."

[0073] In another embodiment, the processor (130) may provide candidate words by combining characteristic information of the gesture. For example, the processor (130) may provide the user with the candidate word "V 2m" by combining a word corresponding to the gesture and a word corresponding to the distance at which the gesture was detected, or may provide the user with the candidate word "2m RIGHT TOP" by combining a word corresponding to the distance at which the gesture was detected and a word corresponding to the area.

[0074] Additionally, the processor (130) may provide only one candidate word to the user, or alternatively, may provide multiple candidate words at once. For example, the processor (130) may provide multiple candidate words to the user, such as "V", "2m", "RIGHT TOP", "V 2m", "2m RIGHT TOP", etc.

[0075] The processor (130) can provide candidate words to the user in various ways. For example, if the electronic device (100) is a device equipped with a display, such as a TV, smart monitor, laptop PC, mobile phone, tablet PC, or kiosk, the processor (130) can control the display to display a UI screen providing candidate words. In another embodiment, the electronic device (100) may be connected to an external display device through the interface (110). For example, the electronic device (100) may be a desktop PC, a set-top box, a content integration source, a server device, or the like. In this case, the processor (130) can provide data regarding the UI screen providing candidate words and a control signal for displaying the UI screen to the external display device. Based on the control signal and data, the external display device can display a UI screen capable of providing candidate words.

[0076] As another embodiment, if the electronic device (100) has a speaker capable of outputting sound, or is connected to an external electronic device having a speaker and can provide a control signal to output sound to the external electronic device, the candidate word can be provided to the user by outputting the candidate word as a voice signal.

[0077] The embodiments of providing candidate words described above are merely examples, and it is to be understood that at least one candidate word may be provided to the user through various methods that may provide a display or voice signal to the user.

[0078] Additionally, the processor (130) may receive a candidate word selection from the user. In one embodiment, the user may select the candidate word by directly touching the UI screen provided by the processor (130) or through various input means provided or connected to the electronic device (100). For example, if the electronic device (100) is a TV, the user may select the candidate word using a remote control.

[0079] The user is not required to select a single candidate word, and in some embodiments, multiple candidate words may be selected. In another embodiment, the processor (130) may receive a voice signal from the user, such as "Select V," to select a candidate word. The processor (130) may receive a voice signal from the user through a microphone provided in the electronic device (100), as described above, and may also obtain a voice signal from an external electronic device equipped with a microphone, such as a remote control or a smartphone.

[0080] Methods for recognizing speech from a voice signal can be implemented in various ways. For example, the processor (130) removes noise from an input speech signal and then extracts a feature vector. The processor (130) can generate a phoneme sequence from the feature vector using an acoustic model. A phoneme sequence can be a continuous arrangement of phonemes, which are the smallest units of phonology that can distinguish meaning. The processor (130) can generate a word sequence by decoding a phoneme sequence into word units, or can generate a syllable sequence by decoding a phoneme sequence into syllable units. A word sequence can be a continuous arrangement of words separated by spaces in a text that is a result of speech recognition. A syllable sequence can be a continuous arrangement of syllables, which are units of speech sounds having a single phonetic value. The processor (130) can generate a plurality of words or a plurality of syllables from a phoneme sequence based on dictionary data stored in the memory (120).

[0081] The processor (130) can determine text as a result of recognition of a speech signal based on at least one of a word sequence and a syllable sequence. The processor (130) can combine the word sequence and the syllable sequence by replacing at least one word element among the word sequence with a syllable sequence, and determine the combined result as text as a result of speech recognition. The processor (130) can recognize the meaning of the text based on previously stored dictionary data. The processor (130) can identify which candidate word the user has selected based on the recognized meaning. In the above-described section, speech recognition is described as being performed directly by the processor (130). However, depending on the embodiment, the task of converting the speech signal into text, the task of recognizing the meaning of the text, etc. may be performed by at least one external server device. When a candidate word is selected by a user, the processor (130) can store characteristic information of the selected candidate word and gesture in the memory (120) as account information corresponding to the user. Account information corresponding to a user may mean information that is used to identify whether the input voice signal or gesture was input by the user when a voice signal or gesture is input to the electronic device (100), and may mean information that is compared with the input voice signal and gesture.

[0082] For example, if a candidate word called "V" is selected by the user, the processor (130) can store the candidate word called "V" as a wake-up word in the memory (120). While storing the wake-up word, the processor (130) can also store gesture characteristic information, such as a word corresponding to the gesture, a word corresponding to the distance at which the gesture was detected, and a word corresponding to the area at which the gesture was detected, in the memory (120) as account information corresponding to the user.

[0083] At this time, the processor (130) may store both the word corresponding to the distance at which the gesture was detected and the word corresponding to the area at which the gesture was detected in the memory (120) as account information corresponding to the user, or may store only one word as account information in the memory (120), or may not store the word corresponding to the distance and area in the account information.

[0084] An example of user account information stored in memory (120) is shown in Table 1 below.

[0085] User 1 account information User 2 account information Gesture VV Distance 2m 1m Area Top right Wake-up word "V" "V" , "V 1m"

[0086] If user 1 selects the candidate word "V", the processor (130) can register "V" as a wake-up word and also register the word corresponding to the distance (2m) and the detected area (upper right) at which the gesture V was detected in the account information of user 1. In the case of user 2, since he selected not only the candidate word "V" but also the candidate word "V 1m", the processor (130) can register both "V" and "V 1m" as wake-up words. In addition, user 2 can register only the word (1m) corresponding to the distance at which the gesture was detected among the characteristic information of the gesture in the account information. The processor (130) can store account information corresponding to each of a plurality of users in the memory (120) according to the above-described method. In addition, the user can modify or delete the account information as needed even after storing the account information in the memory (120). For example, the word corresponding to a gesture can be changed from "V" to "HI", and the word corresponding to a detected distance can be changed from "2m" to "3m".

[0087] In addition, the processor (130) can obtain the user's voice. In one embodiment, if the electronic device (100) is equipped with a microphone, the processor (130) can obtain the user's voice through the microphone equipped in the electronic device (100). In another embodiment, the electronic device (100) can also obtain the user's voice through a remote control. Specifically, the electronic device (100) can communicate with the remote control through an IR (Infrared Red) method, and the remote control can digitize the user's voice signal input through a microphone built into the remote control and provide the digitized signal to the electronic device (100) as an IR signal. Alternatively, the electronic device (100) can communicate with various external electronic devices capable of receiving the user's voice signal, such as a smartphone or an AI speaker, through a Bluetooth or Wi-Fi method, and the external electronic device can digitize the input user's voice signal and provide the digitized signal to the electronic device (100).

[0088] The processor (130) can acquire a word corresponding to the acquired voice signal based on voice recognition technology. At this time, the processor (130) can transmit the acquired voice signal to an external server, and the external server can convert the transmitted voice signal into text and transmit it again to the electronic device (100). Here, the external server may include, but is not limited to, an STT (Speech To Text) server, and may include various server devices capable of performing the function of an STT server. Alternatively, the processor (130) may independently acquire a word corresponding to the acquired voice signal through a pre-trained model built into the electronic device (100). Since a specific voice recognition method has been described in detail in the above-mentioned section, a redundant description thereof will be omitted.

[0089] The processor (130) can identify at least one piece of account information among the plurality of user account information stored in the memory (120) based on the acquired word. The processor (130) can identify only account information in which the acquired word matches the wake-up word registered in the account information of each of the plurality of users stored in the memory (120). In one embodiment, if the processor (130) identifies the word corresponding to the user's voice as "V," the processor (130) can identify the user account that has registered the word matching "V" as the wake-up word. If the user account information as shown in Table 1 described above is stored in the memory (120), the processor (130) can identify the account information of User 1 and User 2 that have registered "V" as the wake-up word.

[0090] In another embodiment, if the processor (130) identifies the word corresponding to the user's voice as "V 1m", the processor (130) can identify only the account information of user 2 who registered "V 1m" as the wake-up word. If only one piece of account information is identified, the account information identification task based on the characteristic information of the gesture described below is not performed, and the processor (130) can immediately perform an operation based on the identified account information.

[0091] In addition, the processor (130) can identify at least one piece of account information based on the input user's voice, then receive a user's gesture to obtain characteristic information of the gesture, and identify the user's account information among the at least one piece of account information identified based on the obtained characteristic information of the gesture.

[0092] The processor (130) can obtain characteristic information of a gesture, such as a word corresponding to the input user gesture as described above, a word corresponding to the distance at which the gesture was detected, a word corresponding to the area at which the gesture was detected, and the like, and can compare the obtained characteristic information with characteristic information of the gesture registered in the user account information stored in the memory (120). Through the comparison, the processor (130) can identify a user account in which characteristic information matching the obtained characteristic information of the gesture is registered.

[0093] For example, assume that user account information such as Table 1 described above is stored in the memory (120). User 1 can make a gesture of "V" while looking at the electronic device (100) and utter the word "V" in the upper right area of ​​the electronic device (100) from a distance of 2 m from the electronic device (100). The processor (130) can identify the account information of User 1 and User 2 who have registered "V" as a wake-up word through the input of the word "V" uttered by the user. In addition, the processor (130) can detect the gesture of "V" made by the user and obtain information about the word ("V") corresponding to the gesture of "V", the distance (2 m) at which the gesture was detected, and the area (upper right) at which the gesture was detected, and can identify that the characteristic information of the obtained gesture matches the information registered in the account information of User 1. As a result, the processor (130) can identify the words and gestures entered into the electronic device (100) as having been entered by user 1 by identifying the account information of user 1.

[0094] Additionally, when one account information is identified by the above-described method, the processor (130) can perform an operation based on the identified account information.

[0095] When powering on the electronic device (100) or executing a specific application on the electronic device (100), the user can log in through his / her user account, and the electronic device (100) can operate while the user account is logged in and can operate to perform only the control commands of the logged-in user. For example, if User 1 logs in to the electronic device (100) with User 1's account, the control authority (Right for Control) of the electronic device (100) can be given only to User 1. Accordingly, when User 1 inputs a command to the electronic device (100) through a gesture or voice, the electronic device (100) can perform a corresponding operation, but even if User 2 inputs a command to the electronic device (100), the electronic device (100) may not perform a corresponding operation.

[0096] Therefore, the processor (130) can identify whether the user who is currently logged in and has control authority over the electronic device (100) matches the user who input the voice and gesture in the above-described manner based on the account information stored in the memory (120), and if they match, can perform an action corresponding to the subsequent user's voice or gesture.

[0097] For example, if the user currently having control authority over the electronic device (100) is User 1, and the processor (130) has identified that the user who inputted voice and gestures into the electronic device (100) is also User 1, the processor (130) can receive the user's command without performing an operation to switch user accounts. In this case, the user's command can be received via voice or through the user's gesture.

[0098] An operation mode or operation state that receives user commands via voice may be referred to as a "voice recognition mode" or a "voice recognition state." Similarly, an operation mode or operation state that receives user commands via gestures may be referred to as a "gesture recognition mode" or a "gesture recognition state." The voice recognition mode or voice recognition state does not necessarily recognize only the user's voice, but may also be implemented as a mode that recognizes and operates on various other audio signals such as the sound of applause or the sound of a musical instrument. The gesture recognition mode or gesture recognition state may also be implemented so that it recognizes a user's actions, such as changing a facial expression, rather than making a specific gesture, and performs a corresponding action. The voice recognition mode may also be referred to as an audio recognition mode, and the gesture recognition mode may also be referred to as a motion recognition mode or a movement recognition mode.

[0099] Specifically, when the processor (130) operates in voice recognition mode, the user can input a voice command such as "Turn up the speaker volume," and the processor (130) can recognize the user's voice and convert it into text. In addition, the processor (130) can perform natural language processing and pattern recognition on the converted text using a pre-trained model, thereby identifying the user's intention of "Turn up the speaker volume." The processor (130) can control the configuration of the electronic device (100) or an external device connected to the electronic device (100) according to the identified user's intention.

[0100] In addition, the processor (130) may operate in gesture recognition mode. The gesture recognition mode will be described in detail in the description of FIG. 6 below. Although the above description exemplifies that the processor (130) can operate in voice recognition mode or gesture recognition mode, the processor (130) may also operate in various operation modes, such as a user-customized operation mode, a peripheral device recognition mode, and a surrounding environment recognition mode, in which the processor can receive a user's command or recognize the user's surrounding environment and perform corresponding operations. For example, when the processor (130) operates in a user-customized operation mode, the processor (130) may analyze the user's viewing history or search history to provide the user with customized content.

[0101] Additionally, if the processor (130) identifies that the user who is currently logged in and has control authority over the electronic device (100) does not match the user who input voice and gestures in the manner described above, it may perform a user account switch or register a new user account.

[0102] In one embodiment, if the processor (130) identifies that the user currently having control authority over the electronic device (100) is User 1 and the user who inputted the voice and gesture is User 2, the processor (130) may perform a user account switch so that User 2 has control authority over the electronic device (100). When the user account switch is performed, the processor (130) may immediately receive the voice or gesture of User 2 and perform corresponding operations. Alternatively, the processor (130) may not immediately perform the user account switch, but may perform a process to determine whether to perform the user switch for User 2. The question "Do you want to switch users?" may be provided to User 2 through a UI screen or through voice.

[0103] In another embodiment, the processor (130) may perform new user account registration. Specifically, if only account information for User 1 and User 2 were stored in the electronic device (100), and User 3 input a wake-up word that is not registered in the User 1 account information and the User 2 account information into the electronic device (100), the processor (130) may proceed with a process for registering a user account for User 3. For the user account registration, the processor (130) may receive various information from User 3, such as User 3's name, User 3's gesture characteristic information, and User 3's wake-up word, and may register the information in User 3's user account information and store the information in the memory (120). The process of receiving a new user's gesture and selecting a wake-up word has been described in detail in the above description, and thus a description thereof will be omitted.

[0104] As described above, if the processor (130) identifies that the currently logged-in user account and the user who input the voice or gesture do not match, the processor (130) may perform a user account switch or register a new user account, but is not limited thereto, and the processor (130) may also perform various operations corresponding to the input user's voice or gesture.

[0105] The specific operations of the processor (130) detecting a user's gesture, acquiring characteristic information of the gesture, and storing account information corresponding to the user or identifying whether a user account needs to be switched based on the acquired characteristic information will be described in more detail through the drawings and descriptions thereof described below.

[0106] Meanwhile, the above has described an embodiment in which changes such as granting control authority to a user when a user's account is identified, but the processor (130) may also perform various other operations. For example, if a user account stores various pieces of information such as a background screen, background music, preferred channel information, preferred content information, preferred volume information, and display option setting information, when a user with a registered user account is recognized, the processor (130) may change the background screen or background music based on the information stored in the user account of the user, and may immediately play and output preferred content of the preferred channel. In addition, the speaker volume may be adjusted according to the preferred volume information, and various display options such as contrast, screen ratio, resolution, and color properties may also be changed based on the setting information.

[0107] FIG. 3 is a diagram illustrating a word providing method of an electronic device according to one or more embodiments of the present disclosure.

[0108] According to FIG. 3, the processor (130) can input a user's gesture (200) and output a word "V" (220) according to a learned neural network model stored in the memory (120). Here, the learned neural network model is an artificial intelligence model that is trained to input a user's gesture, extract a feature map (210) of the gesture, and provide a word (220) based on the extracted feature map (210), and may include a feature extraction unit (310) and a word provision unit (320).

[0109] The feature extraction unit (310) can receive an image of a user's gesture (200) and generate a feature map (210) of the user's gesture (200). The feature extraction unit (310) can be composed of multiple neural network layers that extract features of the image.

[0110] For example, the feature extraction unit (310) may be composed of a convolution layer and a pooling layer. When an image capturing a user's gesture (200) is input to the convolution layer of the feature extraction unit (310), the captured image is divided into multiple regions and a filter capable of performing convolution on each region is applied. As a result of the convolution, feature values ​​for multiple regions can be obtained, and a feature map can be generated based on the obtained feature values. The generated feature map can be input to the pooling layer of the feature extraction unit (310), and the feature map input from the pooling layer undergoes operations such as max pooling and average pooling. Through these operations, the amount of information in the feature map input to the pooling layer can be reduced (downsampled). As described above, the processor (130) can obtain a feature map (210) corresponding to the user's gesture, which has an amount of information reduced to an appropriate size through the feature extraction unit (310).

[0111] The word providing unit (320) can receive the feature map (210) output from the feature extraction unit (310), classify the image into predefined words, and provide the image to the user. Here, the predefined words may refer to words already stored in the word providing unit (320). For example, if the word providing unit (320) can classify the user's gestures into only the words "V", "Hello", "Promise", "Gun", and "Good luck", the predefined words may include "V", "Hello", "Promise", "Gun", and "Good luck".

[0112] In one embodiment, the word providing unit (320) may be configured as a fully connected layer. In the word providing unit (320), the feature map (210) may be input to the fully connected layer so that the feature map (210) in the form of a two-dimensional array may be flattened into the form of a one-dimensional array. Thereafter, the flattened one-dimensional array may pass through various activation functions, such as Softmax, to output the probability of which word among the predefined words the input user's gesture (200) corresponds to. If the user's gesture (200) has the highest probability of corresponding to the word "V" (220), the word providing unit (320) may provide the word "V" (220) to the user. If the user's gesture was a gesture of showing a palm with five fingers raised, the word providing unit (320) may determine that the user's gesture has the highest probability of corresponding to the word "Hello (HI)" and may provide the word "Hello (HI)" to the user.

[0113] The feature extraction unit (310) or the word provision unit (320) may be composed of a plurality of neural network layers. At least one layer has at least one weight value and performs the operation of the layer through the operation result of the previous layer and at least one defined operation. Examples of the neural network include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), and deep Q-networks, and a transformer, and the neural network in the present disclosure is not limited to the examples described above.

[0114] The processor (130) can obtain a word corresponding to the user's gesture through the learned neural network model stored in the memory (120) as described above, and can also receive the word from an external server. For example, the processor (130) can transmit sensing data regarding the user's gesture (200) obtained through a sensor connected to the interface (110) or a sensor built into the electronic device (100) to an external server, and the external server can input the received sensing data into the learned neural network model to output a word corresponding to the gesture (200). The external server can transmit the outputted word to the electronic device (100), and the processor (130) can provide the word received from the external server to the user.

[0115] In addition, the processor (130) may only perform the role of generating the feature map (210) of the gesture (200), and the external server may only perform the role of outputting a word corresponding to the feature map (210) based on the feature map (210). In order to output a word corresponding to the feature map (210), all pattern data for pre-defined words (e.g., “V”, “hello”) must be stored. Therefore, as the number of pre-defined words increases, the size of the data to be stored in the memory (120) may become very large. Accordingly, the memory (120) may only store the neural network layer corresponding to the feature extraction unit (310), and the processor (130) may transmit the feature map (210) corresponding to the gesture (200) to the external server so that the external server may perform the role of the word provision unit (320), and receive the word (220) from the external server.

[0116] As described above, the processor (130) can detect a user's gesture and generate a word corresponding to the gesture using a trained neural network model. The specific operations by which the processor (130) acquires information about the distance and area where the gesture was detected will be described in detail in FIG. 4 and its description below.

[0117] FIG. 4 is a diagram illustrating a method for providing a live view of an electronic device according to one or more embodiments of the present disclosure.

[0118] According to FIG. 4, the electronic device (100) may include a display (140), and the processor (130) may divide the display (140) into a plurality of areas to display a live view screen.

[0119] Here, "live view" can refer to a real-time screen that displays the image captured through the camera lens. When a shooting command is input while live view is displayed, the camera captures the image captured at that time using the image sensor and stores the captured image data.

[0120] The processor (130) can control the display (140) to display a live view screen provided from a camera included in the electronic device (100) or an external camera or a built-in camera device connected to the electronic device (100). The processor (130) can display a graphic object for area division on the live view screen, or can express multiple areas without separate borders. FIG. 4 illustrates a state in which a first area is displayed in a circle in the middle, and second to fifth areas are displayed around it. The user can immediately know in which area among the multiple areas his or her gesture is displayed through the live view screen. Although not illustrated in FIG. 4, the processor (130) can also display a distance on the live view screen so that the user can check how far away his or her gesture is detected from the electronic device (100).

[0121] For example, as illustrated in FIG. 4, the processor (130) can display a live view screen divided into a center (1), a left upper portion (LEFT TOP) (2), a left lower portion (LEFT BOTTOM) (3), a right lower portion (RIGHT BOTTOM) (4), and a right upper portion (RIGHT TOP) (5). Through this live view screen, in the process of storing gesture characteristic information in account information corresponding to the user or in the process of executing various operation modes of the electronic device (100) by inputting the user's gesture, the user can check the area where his / her gesture is detected.

[0122] In one embodiment, if a user wishes to register a gesture called "V" and an area called "top right" in his / her account information, the user can confirm through the live view screen that his / her gesture is accurately detected in the top right area. In another embodiment, if a word corresponding to the gesture ("V") and a word corresponding to the area where the gesture is detected ("top right") are registered in his / her account information, the user must input a gesture so that the gesture "V" is detected in the "top right" area in order to perform account switching to his / her account, and the user can confirm through the live view screen whether his / her gesture is accurately detected in the top right area.

[0123] In addition, although not illustrated in detail in FIG. 4, the processor (130) may also display information about the distance at which a gesture is detected through the live view screen. For example, the processor (130) may measure the distance between the user and the electronic device (100) through various sensors capable of measuring the distance to the user, such as a lidar sensor or a TOF sensor, which are connected to the interface (110) or built into the electronic device (100), and may display the measured distance as “3 m,” “2.5 m,” “1.3 m,” etc. on the live view screen to provide information about the distance between the user and the electronic device (100). The distance information may be displayed for each area. The user may check the distance at which his or her gesture is detected through the provided distance information.

[0124] Additionally, the processor (130) may also provide candidate words based on a word corresponding to an area where the user's gesture is identified and a word corresponding to a distance where the user's gesture is identified. In this case, "first information" is defined as information about an area where the user's gesture is identified among multiple areas of the live view screen, and "second information" is defined as information about the distance between the user and the electronic device (100) provided through the live view screen.

[0125] For example, if the user's gesture is displayed at the upper right of the live view screen and the distance from the user displayed on the live view screen is 3 m, the first information may be "upper right" and the second information may be "3 m." Based on the first information and the second information, the processor (130) may provide the user with candidate words such as "upper right," "3 m," or "upper right 3 m."

[0126] When a user selects a specific candidate word from among candidate words provided based on the first information and the second information, the processor (130) can store the selected candidate word as account information corresponding to the user in the memory (120).

[0127] For example, if a user selects a candidate word such as "top right," the word "top right" can be registered as the user's wake-up word in the user account information corresponding to the user. Accordingly, if the user inputs the voice "top right" into the electronic device (100), the processor (130) can identify the account information of the user who registered the word "top right" as the wake-up word. The operation of registering the wake-up word in the user account information or identifying the user account information based on the wake-up word has been described in detail in the description of FIG. 2, and therefore, a detailed description thereof will be omitted.

[0128] The processor (130) can obtain the first information by analyzing the live view screen that captures the user. Here, analyzing the live view screen may mean identifying which area of ​​the live view screen is where the user's gesture is detected. For example, the processor (130) can continuously monitor the live view screen to detect the user's gesture that is distinct from the surrounding environment fixed in the captured image, and can identify that the detected gesture was performed in the area corresponding to the upper right of the live view screen.

[0129] Additionally, the processor (130) may obtain the first information by receiving a user's selection. For example, the processor (130) may receive a user's selection of the upper right area among multiple areas of the live view screen, in which case the "upper right" area may become the first information.

[0130] The processor (130) may perform the specific operations described above through the external display device by providing a control signal to the external display device. For example, if the electronic device (100) is connected to a TV through an output port of the electronic device (100), the processor (130) may transmit a control signal to the TV through the output port to display a live view screen including an image captured by the user and divided into a plurality of regions. In addition, the TV may receive an input from the user to select one region among the plurality of regions of the live view screen and transmit first information about the input region to the electronic device (100), and the processor (130) may provide at least one candidate word to the user based on the first information received from the TV.

[0131] In the above description, only a TV is exemplified as an external display device, but various external display devices capable of displaying images, such as a desktop PC, laptop PC, tablet PC, or smartphone, can be connected to the electronic device (100) and transmit and receive various signals.

[0132] In the above description, only the case of connecting to an external display device through the output port of the electronic device (100) is exemplified, but this is only one example, and it is of course possible for the processor (130) to transmit and receive various signals with the external display device through a wireless communication method such as Bluetooth or Wi-FI.

[0133] FIG. 5 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.

[0134] According to FIG. 5, when a user inputs a gesture (200) and the word “V” (220) into the electronic device (100), the processor (130) can identify account information corresponding to the user.

[0135] For example, assuming that User 1 has saved “V” as a wake-up word, “V” as a word corresponding to the gesture, “2m” as a word corresponding to the distance at which the gesture was detected, and “top right” as a word corresponding to the area at which the gesture was detected, as shown in Table 1 above, User 1 can confirm that his / her gesture is detected at a distance of 2m from the electronic device (100) through the text “2m” displayed on the live view screen, and that his / her gesture is detected in the top right among the areas through the live view screen divided into multiple areas.

[0136] After confirming that his / her gesture (200) is detected at a distance of 2 m from the electronic device (100) and in the upper right area, the user 1 can utter the word "V" (220). The processor (130) can receive the word uttered by the user, input the user's gesture, and identify the account information of the user 1 corresponding to the user who input the word and gesture. Based on the identification of the account information of the user 1, the processor (130) can identify whether a user account switch is required, and then perform various operations based on the identification result.

[0137] Through the live view screen described above, the user can accurately check at what distance from the electronic device (100) his / her gesture is detected and in what area it is detected by the electronic device (100), and the user can input his / her gesture into the electronic device (100) more accurately.

[0138] In the description of the above-described drawings 4 and 5, only the division of the live view screen into a total of five areas was described, but this is only one example, and it is of course possible to divide the live view screen into a variety of areas by various criteria, such as dividing the areas into only three areas of “left,” “center,” and “right.”

[0139] The electronic device (100) of the present disclosure can operate based on a user's gesture, as described with respect to FIGS. 1 to 5 , and can also operate based on a user's object. Specific operations of the electronic device (100) based on a user's object will be described in detail in the description of FIG. 6 , which will be described later.

[0140] FIG. 6 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.

[0141] According to FIG. 6, the electronic device (100) can detect a user's object (e.g., a cup) (400) and store account information corresponding to the user in the memory (120), and can also identify whether a user account switch is required.

[0142] Here, the term "user's object" may refer to an object detected within an image acquired by the processor (130). Specifically, the term "user's object" may include an object held by the user. It may refer to an object held by the user in the hand, and may also include any object in contact with the user. For example, if the user is holding a mobile phone in the hand, the object held by the user may include the user's mobile phone. Additionally, if the user's foot is in contact with a ball, the object held by the user may include the ball in contact with the user's foot. As shown in FIG. 5, if the user is holding a cup (400) in the hand, the user's object may be the cup (400). Additionally, the term "user's object" may include objects other than the object in contact with the user. For example, even if the user is not present in the image acquired by the processor (130) and only the cup (400) is present, the processor (130) may identify user account information based on the cup present in the image.

[0143] First, we will explain the operation of detecting the user's object and obtaining the object's feature information.

[0144] The processor (130) can obtain sensing data about the cup (400) through a sensor connected to the interface (110) or a sensor built into the electronic device (100). The processor (130) can identify that the user's object is the cup (400) by using a learned neural network model based on the sensing data obtained through the sensor. Returning to FIG. 3, the processor (130) can classify the gesture (200) into a predefined word by using the learned neural network model including the feature extraction unit (310) and the word provision unit (320), as well as classify the user's object into a predefined word.

[0145] Specifically, when an image of a user's object is input into the feature extraction unit (310), a feature map is generated through a plurality of layers included in the feature extraction unit (310), and when the generated feature map is input into the word provision unit (320), the input feature map can be classified into a predefined word. For example, when an image of a cup (400) is input into the feature extraction unit (310), a feature map (210) corresponding to the image of the cup (400) is generated, and when the feature map (210) is input into the word provision unit (320), the word "cup" can be provided to the user based on the result that the probability that the feature map (210) corresponds to the word "cup" among the predefined plurality of words is the highest. A detailed description of a method for providing words through a learned neural network model and detailed configurations of the learned neural network model have been described in detail in the description of FIG. 3, and therefore will be omitted below.

[0146] In addition, the processor (130) can obtain information on the area and distance at which the cup (400) held by the user is detected by the sensor in the same manner as the method of obtaining information on the area and distance at which the user's gesture is detected. The processor (130) can provide candidate words to the user based on the words corresponding to the object provided through the learned neural network model, the words corresponding to the area at which the object is detected, and the words corresponding to the distance at which the object is detected.

[0147] For example, if a cup (400) held by a user is detected at a distance of 2 m from the electronic device (100) and in the upper right area based on the center point of the electronic device (100), the processor (130) may provide the user with candidate words such as “cup,” “2 m,” “upper right,” “cup 2 m,” and “cup upper right” based on characteristic information of the object such as “cup,” “2 m,” and “upper right.”

[0148] In addition, when a user selects a candidate word from among at least one candidate word provided, the processor (130) may store the candidate word selected by the user and the characteristic information of the object as account information corresponding to the user in the memory (120). Since the method for the user to select a candidate word through a UI screen or voice recognition, etc. has been described above, a detailed description thereof will be omitted.

[0149] An example of a state in which the characteristic information of candidate words and objects is stored in memory (120) as account information corresponding to the user is as shown in Table 2.

[0150] User 3 account information User 4 account information Object Cup Cup Distance 2m 1m Area Top right Wakeup word "Cup", "Cup top right" "Cup", "Cup 1m"

[0151] According to Table 2, the user 3 can be detected by the sensor while holding a cup (400) at a distance of 2 m from the electronic device (100) and in the upper right area. Based on the state of the user 3, the processor (130) can provide various candidate words such as “cup”, “cup 2 m”, “cup upper right”, “2 m”, “2 m upper right” to the user 3, and can select the candidate words “cup” and “cup upper right” from the user 3 and store “cup” and “cup upper right” as wake-up words in the memory (120). In addition, the processor (130) can obtain the user’s voice and obtain a word corresponding to the user’s voice through voice recognition, and can obtain feature information of the user’s object by obtaining an image of the user’s object. Since the detailed method of obtaining the word corresponding to the user’s voice and the detailed method of obtaining feature information through the photographed image have been described above, a detailed description thereof will be omitted.

[0152] The processor (130) can identify at least one account information among the plurality of user account information stored in the memory (120) based on a word corresponding to the acquired user's voice. The processor (130) can identify only the account information in which the acquired word matches the wake-up word registered in the account information by comparing the word corresponding to the user's voice with the wake-up word registered in the account information of each of the plurality of users stored in the memory (120). For example, if the processor (130) identifies the word corresponding to the user's voice as "cup," the processor (130) can identify the user account that has registered the word matching "cup" as the wake-up word. If the user account information as shown in Table 2 described above is stored in the memory (120), the processor (130) can identify the account information of User 3 and User 4 that have registered "cup" as the wake-up word.

[0153] In addition, the processor (130) can obtain the characteristic information of the user's object obtained through the captured image, and identify the user's account information among at least one piece of account information identified based on the characteristic information of the obtained object.

[0154] As described above, the processor (130) can obtain characteristic information of an object, such as a word corresponding to the user's object, a word corresponding to the distance at which the object was detected, and a word corresponding to the detected area, and can compare the obtained information with characteristic information of the object registered in the user account information stored in the memory (120). Through the comparison, the processor (130) can identify a user account in which characteristic information matching the obtained characteristic information of the object is registered.

[0155] For example, assume that user account information such as Table 2 described above is stored in the memory (120). User 3 can utter the word "cup" while holding a cup in the upper right area of ​​the electronic device (100) from a distance of 2 m away from the electronic device (100). The processor (130) can identify the account information of User 3 and User 4 who registered "cup" as a wake-up word through the input of the word "cup" uttered by the user. Thereafter, the processor (130) can detect the cup, which is the user's object, and obtain object characteristic information regarding the word corresponding to the object ("cup"), the word corresponding to the distance at which the object was detected (2 m), and the word corresponding to the area at which the object was detected (upper right), and can identify that the obtained object characteristic information matches the object characteristic information registered in the account information of User 3. As a result, the processor (130) can identify that the words and objects entered into the electronic device (100) were entered by user 3 by identifying the account information of user 3.

[0156] Additionally, when a piece of account information is identified by the above-described method, the processor (130) can perform an operation based on the identified account information. The various operations performed by the processor (130) based on the identified account information have been described in detail in the description of FIG. 2, and therefore, a detailed description thereof will be omitted.

[0157] When the processor (130) identifies a piece of account information, if the identified piece of account information matches the account information of the user currently having control authority, the processor may operate in gesture recognition mode without performing a user account switch. The gesture recognition mode may refer to a mode in which an action corresponding to the user's gesture is performed, and a detailed description thereof will be provided in the description section regarding FIG. 7.

[0158] FIG. 7 is a diagram for explaining a gesture recognition mode operation method of an electronic device according to one or more embodiments of the present disclosure.

[0159] According to FIG. 7, a user can input a gesture (500) into an electronic device (100), and the processor (130) can perform an operation corresponding to the input user gesture (500). In the description of FIG. 7, the user's gesture (500) is different from the gesture (200) registered in the user account information used to identify the user account, and the processor (130) can recognize the user's gesture (500) as a user command.

[0160] The memory (120) can store a user's gesture (500) and an action corresponding to the gesture. For example, the processor (130) can store a gesture (500) of moving a user's finger counterclockwise in the memory (120) by corresponding it to an action called "return to the previous screen." When the processor (130) operates in gesture recognition mode, the processor (130) detects the user's gesture through a sensor connected to the interface (110) or a sensor built into the electronic device (100), and if the detected user's gesture is identified as matching the "gesture of moving a finger counterclockwise" stored in the memory (120), the processor can perform an action corresponding to the gesture called "return to the previous screen."

[0161] In addition, when the processor (130) operates in gesture recognition mode, if the processor (130) determines that the user's gesture detected through the sensor does not match a plurality of gestures stored in the memory (120), the processor (130) may newly register the gesture detected through the sensor.

[0162] For example, if a user makes a gesture of raising a finger upward by spreading out one finger, and a gesture matching the upward finger gesture is not stored in the memory (120), the processor (130) may provide the user with a UI screen asking whether to newly register a gesture of raising a finger upward. If the user accepts the new registration of the gesture, the processor (130) may provide the user with a UI screen asking the user to select an action corresponding to the newly registered gesture. The user may select an action called “increase speaker volume” from among several action items provided on the UI screen, and the “increase speaker volume” action may be mapped to the user’s “gesture of raising a finger upward” and stored in the memory (120).

[0163] The above-described UI screen can be displayed by the processor (130) controlling the display (140) included in the electronic device (100), and can also be displayed by providing a control signal to an external display device connected to the electronic device (100).

[0164] In addition, the processor (130) may, of course, store the word for the user's gesture and the characteristic information of the gesture together in the memory (120) in order to more accurately recognize the user's gesture even in gesture recognition mode. Since performing a corresponding action every time the user turns his / her finger counterclockwise may cause the user's gesture (500) to be recognized even when the user does not want it, which may result in an undesirable result, the processor (130) may store the word for the gesture and the information on the distance and area where the gesture is to be detected together in the memory (120). For example, the processor (130) may store the gesture (500) of "turning the finger counterclockwise", the word "previous screen", the distance "1m", and the area information "lower right" together in the memory (120), and the processor (130) may perform the action corresponding to the user's gesture only when the gesture, word, distance, and area information are all satisfied.

[0165] In the above description, it was explained that the processor (130) can detect only a simple movement using the user's hand and perform a corresponding movement, but this is only one example, and it is of course possible to detect various movements using the user's body.

[0166] As described above, the processor (130) can operate in gesture recognition mode when a user account switch is not required, whereas when a user account switch is recognized as necessary, the processor can switch the user account and perform an operation based on the configuration information corresponding to the switched user account information. A description of the operation based on the configuration information corresponding to the switched user account information will be described in detail in the description of FIG. 8 described below.

[0167] FIG. 8 is a diagram illustrating a method for switching user accounts of an electronic device according to one or more embodiments of the present disclosure.

[0168] According to FIG. 8, when an account is switched from an account of user 1 to an account of user 2, the processor (130) can change the default settings of the electronic device (100) based on the setting information corresponding to the account information of user 2. That is, when account information corresponding to a voice, gesture, or object of a new user is identified while another user account is logged in, the processor (130) can switch to the user account of the new user based on the identified account information, and change the default settings of the electronic device (100) based on the setting information corresponding to the user account of the new user.

[0169] Here, the configuration information corresponding to user account information may refer to the default settings of an electronic device set by each user. For example, the configuration information corresponding to user account information may include display resolution, multi-display settings, display refresh rate, display layout, connection settings with external devices, audio settings, Bluetooth connection settings, and Internet connection settings.

[0170] When a word and gesture corresponding to the account information of User 2 are input while the account of User 1 is logged in, the processor (130) can identify that it is necessary to switch the account to the account of User 2, and based on the identification result, the processor (130) can perform the account switch to the account of User 2. The processor (130) can perform the account switch to the account of User 2, check the setting information for the electronic device (100) preset by User 2, and change the basic settings of the electronic device (100) based on the setting information.

[0171] For example, User 1 may store a single-view state (601) that allows the display (140) to display only one content in the memory (120) as setting information corresponding to User 1 account information, and User 2 may store a multi-view state (602) that allows the display to display three contents at once in the memory (120) as setting information corresponding to User 2 account information. Based on the above-described setting information, the processor (130) may change the default setting of the electronic device (100) from the display setting of the single-view state (601) to the display setting of the multi-view state (602) while performing an account switching operation from the account of User 1 to the account of User 2.

[0172] The above description only describes an operation in which the display settings are changed, such as changing the default settings of the electronic device (100) from a single-view state (601) to a multi-view state (602), but this is only one example, and it is of course possible for the default settings of the electronic device (100) to be changed according to various setting information corresponding to user account information, such as changing the number and types of external devices connected to the electronic device (100) or changing the speaker volume of the electronic device (100).

[0173] FIG. 9 is a diagram illustrating a method for providing a candidate word of an electronic device according to one or more embodiments of the present disclosure.

[0174] According to FIG. 9, the processor (130) can provide a control signal to display a UI screen (700) that can provide a plurality of candidate words to the user, and the user can select at least one candidate word from among the plurality of candidate words provided as a wake-up word.

[0175] For example, the processor (130) may provide a combination of a word corresponding to a detected user gesture (“V”), a word corresponding to a detected object (“cup”), a word corresponding to a distance at which the user gesture or object is detected (“2m”), and a word corresponding to a detected area (“top right”) as candidate words, or may provide each of them individually as candidate words. Specifically, “V,” “cup,” “2m,” and “top right” may be provided to the user as candidate words, and combinations of the above-described words, such as “V cup,” “V 2m,” and “V top right,” may be provided to the user as candidate words. The user may select only “V” as a wake-up word among a plurality of candidate words provided through the UI screen (700), or may select a plurality of candidate words of “V,” “cup,” and “V 2m” as wake-up words.

[0176] Additionally, in addition to allowing the user to select a few candidate words from among those provided to the user, the processor (130) may also receive any word from the user and register it as a wake-up word for the corresponding account information. For example, a user may wish to enter the word "bottom left" and register it as a wake-up word for their account information.

[0177] At this time, since the word "top right" is already registered as a wake-up word in the user's account information, a conflict may occur between the already registered wake-up word "top right" and the word "bottom left" that the user wishes to register in his or her account information. In order to prevent a situation where a conflict occurs between words, when the user attempts to register a word that conflicts with the already registered wake-up word as a wake-up word, the processor (130) may provide a UI screen that provides a warning message that a conflict may occur and an alternative word to the user.

[0178] For example, the processor (130) may provide a UI screen to the user that includes a message such as "The word 'bottom left' cannot be registered because it conflicts with an already registered wake-up word" or "Would you like to register the word '2m' as the wake-up word instead of the word 'bottom left'?"

[0179] The above description only describes the collisions that may occur between words corresponding to the area where the gesture and object are detected, but this is only one example, and even in cases where various collisions that may occur between words occur, such as when the wake-up word already registered in the user's account information is "2m" but an attempt is made to register "4m" as a new wake-up word, the processor (130) may provide an alternative word.

[0180] The UI screen illustrated in FIG. 9 can be displayed on the display (140) of the electronic device (100), and the external display device can also display the UI screen by providing a control signal to the processor (130) to display the UI screen that provides a plurality of candidate words to the user on the external display device connected to the electronic device (100).

[0181] Although only a UI screen displaying multiple candidate words is shown in FIG. 9, this is only one example, and the processor (130) may provide multiple candidate words to the user through various methods, such as outputting voice to provide multiple candidate words to the user.

[0182] FIG. 10 is a block diagram illustrating a detailed configuration of an electronic device according to one or more embodiments of the present disclosure.

[0183] According to FIG. 10, an electronic device (100) according to an embodiment of the present disclosure may further include an interface (110), a memory (120), a processor (130), a display (140), a sensor (150), a microphone (160), etc. However, the configurations as shown in FIGS. 2 and 10 are merely exemplary, and it goes without saying that new configurations may be added or some configurations may be omitted in addition to the configurations as shown in FIGS. 2 and 10 when implementing the present disclosure. In the description of FIG. 10, any description overlapping with the description of FIG. 2 will be omitted.

[0184] The interface (110) can be divided into an input / output interface (111) and a communication interface (112).

[0185] The input / output interface (111) is a configuration for inputting and outputting various external signals. The input / output interface can be connected to various external memories or external sources (e.g., web servers, user terminal devices, etc.) and can input various data. The input / output interface can be implemented as at least one interface among HDMI (High Definition Multimedia Interface), MHL (Mobile High-Definition Link), USB (Universal Serial Bus), USB C-type, DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (Dsubminiature), and DVI (Digital Visual Interface). At least some of the input / output interfaces may be connected to communication interfaces. For example, the input / output interface can transmit information received from an external device to the communication interface or transmit information received through the communication interface to the external device.

[0186] The communication interface (112) is a configuration for performing communication with at least one external device. The communication interface may include at least one wireless communication module, at least one wired communication module, etc. Each communication module may be implemented in the form of at least one hardware chip. The wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, or other communication modules. In addition, the communication interface may include at least one communication chip that performs communication according to various wireless communication standards such as Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc. The wired communication module may include, for example, at least one among a LAN (Local Area Network) module, an Ethernet module, a pair cable, a coaxial cable, a fiber optic cable, or a UWB (Ultra Wide-Band) module. The communication interface is implemented in various forms like this, and by performing communication with an external device, various data can be received from the external device.

[0187] The memory (120) may include a general user database (DB) (121) and a special user database (DB) (122). Here, the general user DB (121) and the special user DB (122) may refer to a database that systematically organizes general user account information and a database that systematically organizes special user account information, respectively. That is, the memory (120) may store general user account information and special user account information separately.

[0188] For example, a special user's account may include an administrator account of the electronic device (100). The memory (120) can further enhance security functions by separately managing the general user's account and the administrator account of the electronic device (100).

[0189] In one embodiment, the account information of the special user in the special user DB (122) stored in the memory (120) may include information about the time the special user maintains a gesture and / or information about consecutive gestures of the special user. Specifically, the account information of the special user may include information that the special user must maintain the "five-finger spread" gesture for "more than 20 seconds," or the account information of the special user may include information that the special user must continuously make an "OK" gesture with their fingers after making the "five-finger spread" gesture. That is, by storing the general user DB (121) and the special user DB (122) separately, the processor (130) may not consider changes in gestures over time for the general user's gestures, but may consider changes in gestures over time for the special user's gestures.

[0190] In addition, the memory (120) may include a first neural network model (123) and a second neural network model (124). Here, the first neural network model (123) and the second neural network model may refer to trained neural network models that may include a plurality of neural network layers. Specifically, the first neural network model (123) may include a trained neural network model that can input an image of a user's gesture and output a word corresponding to the user's gesture. In addition, the second neural network model (124) may include a trained neural network model that can input an image of a user's object and output a word corresponding to the user's object.

[0191] As described above, the first neural network model (123) and the second neural network model (124) may not be stored in the memory (120). If the memory (120) does not store the first neural network model (123) and the second neural network model (124), the processor (130) may transmit an image of a gesture or object to an external server, and the external server may output a word corresponding to the captured gesture or object and transmit the output word to the electronic device (100).

[0192] The display (140) can perform an operation of displaying a live view screen or providing a UI screen to the user under the control of the processor (130).

[0193] The display (140) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, it may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (140) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc.

[0194] The sensor (150) is configured to detect a user's gesture or an object of the user. The sensor (150) may include various types of sensors capable of detecting the user's movements, the user's object, the distance from the user, etc., such as a camera, a lidar sensor, or a TOF sensor.

[0195] The microphone (160) can acquire a signal for a sound or voice generated from outside the electronic device (100). Specifically, the microphone (160) can acquire vibrations according to a sound or voice generated from outside the electronic device (100) and convert the acquired vibrations into electrical signals.

[0196] In particular, the microphone (160) according to the present disclosure can acquire a voice signal for a user's voice generated by the user's speech. In addition, the acquired voice signal can be converted into a digital signal and stored in a memory (120).

[0197] Although FIG. 10 illustrates and describes the configuration of an electronic device having both a display and a microphone, as described above, at least some of these configurations may be implemented to be included in an external device.

[0198] FIG. 11 is a flowchart illustrating a method for storing account information in an electronic device according to one or more embodiments of the present disclosure.

[0199] According to FIG. 11, the electronic device can detect a user's gesture and obtain a word corresponding to the user's gesture (S1110).

[0200] For example, an electronic device may obtain sensing data regarding a user's gesture through a sensor, and input the sensing data into a trained neural network model to obtain a word corresponding to the gesture. Here, the trained neural network model may include an artificial intelligence model trained to receive sensing data regarding the gesture, such as a photographed image of the gesture, and classify it into one of predefined words. According to another embodiment, instead of using the trained neural network model, a word corresponding to the gesture may be obtained by performing direct image analysis.

[0201] Next, the electronic device may provide the user with at least one candidate word based on a word corresponding to the user's gesture, a word corresponding to the area where the gesture is detected, and a word corresponding to the distance where the gesture is detected (S1120).

[0202] For example, the electronic device may individually provide the user with a word corresponding to the user's gesture ("V"), a word corresponding to the area where the gesture was detected ("top right"), and a word corresponding to the distance where the gesture was detected ("2m") as candidate words. Alternatively, the electronic device may provide the user with a combination of the word corresponding to the user's gesture ("V"), the word corresponding to the area where the gesture was detected ("top right"), and the word corresponding to the distance where the gesture was detected ("2m") as candidate words.

[0203] Next, when a candidate word is selected by a user, the electronic device can store the characteristic information of the selected candidate word and gesture as account information corresponding to the user (S1130).

[0204] For example, if a candidate word "V" is selected by a user, the electronic device can store the word "V" as the wake-up word for that user, and store the characteristic information of the gesture together with the account information corresponding to that user.

[0205] Here, the feature information of the gesture may include a word corresponding to the user's gesture, a word corresponding to the area where the gesture is detected, and a word corresponding to the distance where the gesture is detected.

[0206] FIG. 12 is a flowchart illustrating a method for storing account information in an electronic device according to one or more embodiments of the present disclosure.

[0207] According to FIG. 12, the electronic device can detect the user's object and obtain a word corresponding to the object (S1210).

[0208] For example, an electronic device can acquire sensing data about a user's object through a sensor, and input the sensing data into a trained neural network model to obtain a word corresponding to the object. Here, the trained neural network model may include an artificial intelligence model trained to receive sensing data about the object, such as a photographed image of the object, and classify it into one of a set of predefined words. This does not necessarily require the use of a trained neural network model; an object can also be identified through image analysis and a word corresponding to the object can be obtained.

[0209] Next, the electronic device may provide the user with at least one candidate word based on a word corresponding to the user's object, a word corresponding to the area where the object is detected, and a word corresponding to the distance where the object is detected (S1220).

[0210] For example, the electronic device may individually provide the user with a word corresponding to the user's object ("cup"), a word corresponding to the area where the object was detected ("top right"), and a word corresponding to the distance where the object was detected ("2m") as candidate words. Alternatively, the electronic device may provide the user with a combination of the word corresponding to the user's object ("cup"), the word corresponding to the area where the object was detected ("top right"), and the word corresponding to the distance where the object was detected ("2m") as candidate words.

[0211] Next, when a candidate word is selected by a user, the electronic device can store the selected candidate word and characteristic information of the object as account information corresponding to the user (S1230).

[0212] For example, if a candidate word "cup" is selected by a user, the electronic device can store the word "cup" as the wake-up word for the user, and store the characteristic information of the object together with the account information corresponding to the user.

[0213] Here, the object feature information may include a word corresponding to the user's object, a word corresponding to the area where the object is detected, and a word corresponding to the distance where the object is detected.

[0214] FIG. 13 is a flowchart illustrating a method for identifying a user account of an electronic device according to one or more embodiments of the present disclosure.

[0215] According to FIG. 13, the electronic device can receive words and gestures from a user (S1310).

[0216] Next, the electronic device identifies whether there is user account information that has registered a word identical to the word input by the user as a wake-up word (S1320).

[0217] Next, if it is identified that account information exists (S1330:Y), it is identified whether user account information exists that has registered characteristic information of a gesture identical to the characteristic information of a gesture input from the user (S1340).

[0218] Next, if it is identified that account information exists (S1350:Y), it is identified whether user account switching is required (S1360).

[0219] Next, if it is determined that user account switching is not required (S1360:N), the electronic device may operate in gesture recognition mode (S1380). On the other hand, if it is determined that user account switching is required (S1360:Y), the electronic device may perform user account switching (S1370).

[0220] The various methods described in FIGS. 11 to 13 can be performed by an electronic device having the configuration shown in FIG. 2 or FIG. 10, but are not necessarily limited thereto, and can be performed by electronic devices having various configurations.

[0221] While various embodiments have been described individually or in combination above, each embodiment is not necessarily implemented independently. That is, the various embodiments described above may be implemented together in whole or in part with at least one other embodiment in a single product.

[0222] Meanwhile, the methods according to the various embodiments of the present disclosure described above may be implemented in the form of an application that can be installed on an existing electronic device.

[0223] Additionally, the methods according to the various embodiments of the present disclosure described above can be implemented only with a software upgrade or a hardware upgrade for an existing electronic device.

[0224] Additionally, the various embodiments of the present disclosure described above may also be performed through an embedded server provided in an electronic device, or at least one external server.

[0225] Meanwhile, according to a temporary example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored in the storage medium and operate according to the called instructions, and may include an electronic device according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or under the control of the processor by using other components. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between whether data is stored semi-permanently or temporarily in the storage medium.

[0226] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0227] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0228] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, interface; memory; and comprising at least one processor; At least one processor of the above, When a gesture for account registration is identified based on sensing data acquired through the above interface, at least one candidate word is provided based on characteristic information of the gesture, When at least one of the at least one candidate word is selected, account information including the selected candidate word and characteristic information of the gesture is stored in the memory, An electronic device that, when a user's voice and gesture are identified, performs a control operation based on account information corresponding to the user's voice and gesture among a plurality of account information stored in the memory.

2. In paragraph 1, At least one processor of the above, An electronic device that identifies a user's voice and gesture while another user account is logged in, and identifies account information corresponding to the user's voice and gesture, and switches to the user account of the user based on the identified account information.

3. In paragraph 1, At least one processor of the above, An electronic device that operates in a gesture recognition state in which a function is performed according to the user's gesture or in a voice recognition state in which a function is performed according to the user's voice when the user's voice and gesture are identified in a state in which there is no logged-in user account and account information corresponding to the user's voice and gesture is identified.

4. In paragraph 2, At least one processor of the above, An electronic device that changes the default settings of the electronic device based on the setting information included in the identified account information when the user account is switched based on the identified account information.

5. In paragraph 1, display; including more, At least one processor of the above, Control the display to display a live view screen obtained through a sensor connected to the above interface by dividing it into multiple areas, When a user's gesture is identified within one of the plurality of areas, at least one candidate word is provided based on the area in which the user's gesture is identified and the distance between the user and the electronic device. An electronic device, wherein when at least one of the at least one candidate word provided is selected, the selected candidate word is included in account information corresponding to the user and stored in the memory.

6. In paragraph 5, At least one processor of the above, An electronic device that identifies an area selected by the user among the plurality of areas included in the live view screen as an area where the user's gesture is located.

7. In paragraph 1, At least one processor of the above, When an image corresponding to a user is acquired, the image corresponding to the user is input into a learned neural network model to acquire a feature map of the user's gesture. An electronic device that identifies and provides at least one candidate word corresponding to the feature map of the acquired gesture.

8. In paragraph 1, The at least one processor, when an object for account registration is identified based on sensing data acquired through the interface, provides at least one candidate word based on characteristic information of the object, An electronic device that stores account information including the selected candidate word and characteristic information of the object in the memory when at least one of the at least one candidate word is selected.

9. In paragraph 8, At least one processor of the above, An electronic device, which, when a user's voice and an object are identified, performs an action based on account information corresponding to the user's voice and object among a plurality of account information stored in the memory.

10. In paragraph 9, At least one processor of the above, An electronic device that, when a voice and an object of a user are identified while another user account is logged in, and account information corresponding to the voice and the object of the user are identified, switches to the user account of the user based on the identified account information.

11. In a method for controlling an electronic device, When a gesture for account registration is identified, a step of providing at least one candidate word based on characteristic information of the gesture; When at least one of the at least one candidate word is selected, a step of storing account information including the selected candidate word and characteristic information of the gesture; and A method for controlling an electronic device, comprising: when a user's voice and gesture are identified, performing an operation based on account information corresponding to the user's voice and gesture among a plurality of stored account information; 12. In paragraph 11, The step of performing an action based on the above identified account information is: A method for controlling an electronic device, comprising: a step of identifying a user's voice and gesture while another user account is logged in, and identifying account information corresponding to the user's voice and gesture; and switching to the user account of the user based on the identified account information.

13. In paragraph 11, The steps to perform different actions depending on whether the above user account switching is required are: A method for controlling an electronic device, comprising: a step of operating in a gesture recognition state in which a function is performed according to the user's gesture or in a voice recognition state in which a function is performed according to the user's voice when a user's voice and gesture are identified in a state in which there is no logged-in user account; 14. In paragraph 11, A step of displaying the live view screen by dividing it into multiple areas; When a gesture of the user is identified within one of the plurality of areas, providing at least one candidate word based on the area in which the gesture of the user is identified and the distance between the user and the electronic device; and A method for controlling an electronic device, comprising: a step of storing, when at least one of the provided candidate words is selected, the selected candidate word by including it in account information corresponding to the user; 15. A non-transitory computer-readable recording medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the operation comprising: When a gesture for account registration is identified, a step of providing at least one candidate word based on characteristic information of the gesture; When at least one of the at least one candidate word is selected, a step of storing account information including the selected candidate word and characteristic information of the gesture; and A computer-readable recording medium, comprising: a step of performing an action based on account information corresponding to the user's voice and gesture among a plurality of stored account information when the user's voice and gesture are identified;

Citation Information

Patent Citations

  • Apparatus for detecting hand motion and method thereof

    KR101553484B1

  • Reduced air resistance Float for Seaweed Farming

    KR1020240123170A

  • Refrigerator

    KR1020250046938A

  • Medicament reaction analyzing method using medicament reaction container

    KR102563773B1

  • Inverter cooling system and method for vehicle

    KR102729871B1