Electronic device and method of controlling same

Through the microphone and voice recognition model combined with UI screen metadata, the target object is identified and the display is controlled to display the corresponding UI screen, which solves the voice control problem without voice control API applications and achieves the accurate execution of user intentions.

CN120380452APending Publication Date: 2025-07-25SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089653.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-12-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, there is a lack of effective methods to implement voice control for applications that do not provide voice control-related APIs, making it difficult to achieve operations that match user intentions through user voice.

Method used

Receive user voice through a microphone, obtain text information using the voice recognition model, and combine the metadata of the UI screen to identify the target object, and control the display to display the corresponding UI screen to execute user commands.

Benefits of technology

Voice control for applications that do not provide voice control APIs, can accurately identify user intentions and perform corresponding operations, and improve the convenience of user interface control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380452A_ABST
    Figure CN120380452A_ABST
Patent Text Reader

Abstract

An electronic device and a control method of the electronic device are disclosed. Specifically, the electronic device according to the present disclosure comprises: a microphone; a display; a memory; and a processor: if a user speech is received through the microphone while a first user interface (UI) screen is displayed on the display, obtaining text information corresponding to the user speech by inputting the user speech to the speech recognition model; obtaining, based on the text information, first information including information on a command included in the text information and information on an execution target of the command; obtaining, based on the metadata on the first UI screen, second information including information on functions corresponding to the plurality of objects and information on text included in the plurality of objects; identifying whether a target object corresponding to the user voice exists in a plurality of objects based on a comparison result between the first information and the second information; and if the target object is recognized, controlling the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an electronic device and a method of controlling the electronic device. More specifically, the present disclosure relates to an electronic device capable of performing voice control matching a user's intention and a control method thereof. Background Art

[0002] Recently, with the development of technologies associated with artificial intelligence, a technology of controlling an electronic device based on a user's voice has been drawing attention. More specifically, the development of a technology of controlling an operation associated with a user interface (UI) object by inputting a user's voice instead of touch interaction while a UI screen provided by an application in a display is in a display state has been accelerating more.

[0003] However, in the case of an application providing an application programming interface (API) associated with voice control, it may be easy to perform an operation corresponding to a user's voice or display a UI screen corresponding to a user's voice based on the provided API. However, for an application not providing an API associated with voice control, there may be a problem that voice control is not easy.

[0004] Therefore, there is an increasing need for a technology capable of performing voice control matching a user's intention by effectively constructing a UI graph capable of showing attributes of each UI screen.

[0005] The above information is presented only as background information to help understand the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above is applicable as prior art with respect to the present disclosure. Summary of the Invention

[0006] Technical Solution Aspects of the present disclosure are to at least solve the above problems and / or disadvantages and at least provide the following advantages. Accordingly, an aspect of the present disclosure is to provide an electronic device capable of performing voice control matching a user's intention using a user interface (UI) graph and a control method thereof.

[0007] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments.

[0008] According to one aspect of the present disclosure, an electronic device is provided. The electronic device includes: a microphone; a display; a memory; and a processor configured to: obtain text information corresponding to the user voice by inputting the user voice into a speech recognition model based on the user voice received through the microphone while a first UI screen including a plurality of objects is being displayed on the display, obtain first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command, obtain second information based on metadata about the first UI screen, the second information including information about functions corresponding to the plurality of objects and information about text included in the plurality of objects, identify whether a target object corresponding to the user voice exists in the plurality of objects based on a comparison result between the first information and the second information, and control the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command based on the target object being identified.

[0009] The processor is configured to: identify an object including the first text among the plurality of objects as the target object based on the information about the execution target corresponding to the first text included in the plurality of objects.

[0010] The processor is configured to: identify an object corresponding to the one function among the plurality of objects as the target object based on the information about the execution target corresponding to a second text included in the plurality of objects and the command corresponding to one of the plurality of functions corresponding to the plurality of objects.

[0011] The processor is configured to: control the display to maintain the display of the first UI screen based on the target object not being identified.

[0012] The processor is configured to: control the display to display a third UI screen for performing an operation corresponding to the command based on the target object not being identified.

[0013] The processor is configured to: identify a variable area and a non-variable area different from the variable area based on the metadata, in the variable area, information displayed within the first UI screen changes; and obtain second information from at least one object included in the non-variable area.

[0014] In the memory, a UI graph is stored for each application. The UI graph includes a plurality of nodes corresponding to types of a plurality of UI screens and a plurality of edges showing connection relationships between the plurality of nodes according to operations performed through transitions between the plurality of UI screens. The processor is configured to: based on the target object being recognized, obtain a first embedding vector corresponding to second information based on metadata of a first UI screen, based on a comparison result between the first information and the second information and a comparison result between the first embedding vector corresponding to the second information and second embedding vectors respectively corresponding to the plurality of nodes, recognize a first node corresponding to the first UI screen among the plurality of nodes, based on the information about the command and the information about an execution target of the command, recognize a second node for performing an operation corresponding to the command from at least one node connected to the first node, and control the display to display a second UI screen corresponding to the second node.

[0015] According to another aspect of the present disclosure, a method of controlling an electronic device is provided. The method includes: based on receiving user speech while a first UI screen including a plurality of objects is displayed on a display of the electronic device, obtaining text information corresponding to the user speech by inputting the user speech into a speech recognition model, obtaining first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command, obtaining second information based on metadata of the first UI screen, the second information including information about functions corresponding to the plurality of objects and information about text included in the plurality of objects, based on a comparison result between the first information and the second information, recognizing whether a target object corresponding to the user speech exists among the plurality of objects, and based on the target object being recognized, controlling the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.

[0016] The step of recognizing whether the target object exists includes: based on the information about the execution target corresponding to a first text included in the plurality of objects, recognizing an object including the first text among the plurality of objects as the target object.

[0017] The step of recognizing whether the target object exists includes: based on the information about the execution target corresponding to a second text included in the plurality of objects and the command corresponding to one of a plurality of functions corresponding to the plurality of objects, recognizing an object corresponding to the one function among the plurality of objects as the target object.

[0018] The method of controlling the electronic device includes: based on the target object not being recognized, controlling the display to maintain the display of the first UI screen.

[0019] The method for controlling an electronic device includes: based on the target object not being recognized, controlling the display to display a third UI screen for performing an operation corresponding to the command.

[0020] The step of obtaining second information includes: based on the metadata, identifying a variable region and a non-variable region different from the variable region, in which information within the first UI screen changes in the variable region; and obtaining second information from at least one object included in the non-variable region.

[0021] The method for controlling an electronic device includes: obtaining a UI graph, the UI graph including a plurality of nodes corresponding to types of a plurality of UI screens and a plurality of edges showing connection relationships between the plurality of nodes according to operations performed through transitions between the plurality of UI screens; based on the target object being recognized, obtaining a first embedding vector corresponding to second information based on the metadata regarding the first UI screen; based on a comparison result between the first information and the second information and a comparison result between the first embedding vector corresponding to the second information and second embedding vectors respectively corresponding to the plurality of nodes, identifying a first node corresponding to the first UI screen among the plurality of nodes; based on the information regarding the command and the information regarding the execution target of the command, identifying a second node for performing an operation corresponding to the command from among at least one node connected to the first node, and controlling the display to display a second UI screen corresponding to the second node.

[0022] According to another aspect of the present disclosure, there is provided one or more non-transitory computer-readable storage media storing computer-executable instructions, which, when executed by at least one processor of an electronic device, configure the electronic device to perform a method. The method includes: based on receiving user speech while a first UI screen including a plurality of objects is displayed on a display of the electronic device, obtaining text information corresponding to the user speech by inputting the user speech into a speech recognition model, obtaining first information based on the text information, the first information including information regarding a command included in the text information and information regarding an execution target of the command, obtaining second information based on metadata regarding the first UI screen, the second information including information regarding functions corresponding to the plurality of objects and information regarding text included in the plurality of objects, based on a comparison result between the first information and the second information, identifying whether a target object corresponding to the user speech exists among the plurality of objects, and based on the target object being recognized, controlling the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.

[0023] The detailed description of various embodiments of the present disclosure is disclosed in conjunction with the following drawings, and other aspects, advantages, and significant features of the present disclosure will become apparent to those skilled in the art. Description of the Drawings

[0024] Through the following description in conjunction with the drawings, the above and other aspects, features, and advantages of certain embodiments of the present disclosure will become more apparent, wherein: Figure 1 is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure; Figure 2 is a block diagram showing a plurality of modules according to an embodiment of the present disclosure; Figure 3 , Figure 4 and Figure 5 are diagrams showing the process of identifying a target object from a user interface (UI) screen based on a comparison result between first information and second information according to various embodiments of the present disclosure; Figure 6 is a block diagram showing a plurality of modules used in the process of generating a UI diagram according to an embodiment of the present disclosure; Figure 7 is a flowchart for showing the process of generating a UI diagram according to an embodiment of the present disclosure; Figure 8 is a flowchart showing the configuration of an electronic device according to an embodiment of the present disclosure; and Figure 9 is a flowchart showing a method of controlling an electronic device according to an embodiment of the present disclosure.

[0025] Throughout the drawings, it should be noted that the same reference numerals are used to describe the same or similar elements, features, and structures. Detailed Description of Specific Embodiments

[0026] The following description with reference to the drawings is provided to assist in a comprehensive understanding of various embodiments of the present disclosure defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details should be considered merely exemplary. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure.

[0027] In addition, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0028] The terms and words used in the following description and claims are not limited to their written meanings, but are used only by the inventors to enable a clear and consistent understanding of the present disclosure. Thus, it will be apparent to those skilled in the art that the following description of the various embodiments of the present disclosure is for illustrative purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0029] It will be understood that, unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" include plural referents. Thus, for example, a reference to "a component surface" includes a reference to one or more such surfaces.

[0030] In the present disclosure, expressions such as "has," "may have," "includes," "may include," etc. are used to specify the existence of corresponding characteristics (e.g., elements such as numerical values, functions, operations, or components), without excluding the existence or possibility of additional characteristics.

[0031] In the present disclosure, expressions such as "A or B," "at least one of A and / or B," or "one or more of A and / or B" may include all possible combinations of the items listed together. For example, "A or B," "at least one of A and B," or "at least one of A or B" may refer to all cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.

[0032] Expressions such as "first," "second," "1st," "2nd," etc. used herein may be used to refer to various elements regardless of order and / or importance. In addition, it should be noted that these expressions are only used to distinguish one element from another element and do not limit the relevant elements.

[0033] When an element (e.g., a first element) is indicated as "coupled / couples to" or "connected to" another element (e.g., a second element) "(operatively or communicatively)", it may be understood that the element is directly coupled / couples to the other element or is understood to be coupled through other elements (e.g., a third element).

[0034] On the other hand, when an element (e.g., a first element) is indicated as "directly coupled to" or "directly connected to" another element (e.g., a second element), it may be understood that no other elements (e.g., a third element) exist between the element and the other element.

[0035] The expression "configured to (or set to)" used in the present disclosure may be interchangeably used with, for example, "suitable for", "capable of", "designed to", "adapted to", "manufactured to", or "able to" depending on the context. The term "configured to (or set to)" does not necessarily mean "specially designed to" in terms of hardware.

[0036] On the contrary, in some cases, the expression "a device configured to" may mean that the device "can perform certain operations" together with another device or component. For example, the phrase "a sub - processor configured to (or set to) execute A, B, or C" may represent a dedicated processor (e.g., an embedded processor) for performing the corresponding operations, or a general - purpose processor (e.g., a central processing unit (CPU) or an application processor) capable of performing the corresponding operations by executing one or more software programs stored in a memory device.

[0037] The term "module" or "component" used in the embodiments herein performs at least one function or operation and can be implemented in hardware or software, or in a combination of hardware and software. In addition, except for "modules" or "components" that need to be implemented as specific hardware, multiple "modules" or multiple "components" can be integrated into at least one module and implemented in at least one processor.

[0038] Meanwhile, various elements and regions of the drawings have been schematically shown. Therefore, the technical concept of the present disclosure is not limited by the relative sizes and distances shown in the drawings.

[0039] Hereinafter, one or more embodiments according to the present disclosure will be described with reference to the drawings to assist those skilled in the art in understanding.

[0040] Figure 1 is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure.

[0041] Figure 2 is a block diagram showing a plurality of modules according to an embodiment of the present disclosure.

[0042] Referring to Figure 1 , an electronic device 100 according to one or more embodiments of the present disclosure may include a microphone 110, a display 120, a memory 130, and a processor 140.

[0043] The microphone 110 may obtain a sound or a signal for a voice generated outside the electronic device 100. Specifically, the microphone 110 may obtain a sound or a vibration according to a voice generated outside the electronic device 100 and convert the obtained vibration into an electrical signal.

[0044] Specifically, the microphone 110 according to the present disclosure can receive user speech. For example, the microphone 110 can obtain a speech signal for the user speech generated by the user's words, and the obtained signal can be converted into a signal in digital form and stored in the memory 130. The microphone 110 can include an analog-to-digital converter (A / D converter), and can operate in association with an A / D converter located outside the microphone 110.

[0045] The display 120 can output image data under the control of the processor 140. Specifically, the display 120 can output an image pre-stored in the memory 130 under the control of the processor 140. The display 120 can be implemented as a liquid crystal display (LCD) panel, an organic light-emitting diode (OLED), etc., and the display 120 can also be implemented as a flexible display 120, a transparent display 120, etc. according to circumstances. However, the display 120 according to the present disclosure is not limited to a specific type.

[0046] Specifically, the display 120 according to the present disclosure can display a user interface (UI) screen for each of the applications stored in the memory 130. The display 120 can identify a target object corresponding to the user speech, and display a UI screen corresponding to the identified object for performing an operation according to the user speech.

[0047] In the memory 130, at least one instruction associated with the electronic device 100 can be stored. In addition, an operating system (O / S) for driving the electronic device 100 can be stored in the memory 130. Additionally, various software programs or applications for operating the electronic device 100 according to various embodiments of the present disclosure can be stored in the memory 130. Furthermore, the memory 130 can include a semiconductor memory such as a flash memory, a magnetic storage medium such as a hard disk, etc.

[0048] Specifically, various software modules for operating the electronic device 100 according to various embodiments of the present disclosure can be stored in the memory 130, and the processor 140 can control the operation of the electronic device 100 by executing various software modules stored in the memory 130. For example, the memory 130 can be accessed by the processor 140, and reading / writing / modifying / deleting / updating, etc. of data can be executed by the processor 140.

[0049] In addition, in the present disclosure, the term "memory 130" can be used to mean including the memory 130, the read-only memory (ROM) in the processor 140, the random access memory (RAM), or a memory card (e.g., a micro secure digital (SD) card, a memory stick) installed in the electronic device 100.

[0050] Specifically, according to various embodiments of the present disclosure, at least one instruction associated with the electronic device 100 may be stored in the memory 130. Specifically, the UI according to the present disclosure may be stored in the memory 130 for each application. Specifically, the UI diagram may include a plurality of nodes corresponding to types of a plurality of UI screens and a plurality of edges showing connection relationships between the plurality of nodes according to operations performed by transitions between the plurality of UI screens.

[0051] For example, the UI diagram may be a diagram that shows "information about UI screens" sequentially provided from the current UI screen until a UI screen corresponding to the purpose of a user command is provided as a plurality of nodes, and shows "information about operations" sequentially performed from the current UI screen until a UI screen corresponding to the purpose of the user command is provided as a plurality of edges. For example, operations shown by the plurality of edges may include an operation of clicking an object, an operation of activating a text input field, an operation of displaying information about an object, and the like.

[0052] Specifically, the plurality of nodes according to the present disclosure may be divided based on a similarity between an embedding vector corresponding to each UI screen and symbolic information represented by a plurality of objects included in each UI screen.

[0053] Here, the embedding vector may refer to a vector obtained based on metadata about the UI screen through a neural network model (e.g., a neural network encoder described below), and the symbolic information may be a concept opposite to the embedding vector and include information about text of a plurality of objects included in the UI screen and information about functions of the plurality of objects. For example, the symbolic information may include first information, second information, etc. described below. In addition, the metadata may be provided by an application developer and include UI screens for each application and information showing attributes of objects for each UI screen in the UI screens.

[0054] For example, when dividing the plurality of nodes by only considering the embedding vector corresponding to each UI screen in the UI screen, the number of nodes included in the UI diagram may increase infinitely, and when dividing the plurality of nodes by only considering the symbolic information represented by the plurality of objects, it may be difficult to clearly identify nodes matching the intention of the user command. Therefore, the symbolic information included in the UI screen and the embedding vector corresponding to the UI screen may be considered together to generate the UI diagram according to the present disclosure. A process of generating the UI diagram according to the present disclosure will be described with reference to Figure 6 Describe the process of generating the UI diagram according to the present disclosure.

[0055] The processor 140 may control the overall operation of the electronic device 100. Specifically, the processor 140 may be connected to the configuration of the electronic device 100 including the microphone 110, the display 120, and the memory 130, and control the overall operation of the electronic device 100 by executing at least one instruction stored in the memory 130 as described above.

[0056] The processor 140 may be implemented in various ways. For example, the processor 140 may be implemented as at least one of an application specific integrated circuit (ASIC), an embedded processor, a microprocessor, hardware control logic, a hardware finite state machine (FSM), and a digital signal processor (DSP). In addition, the term "processor 140" in the present disclosure may be used to mean including a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), etc.

[0057] Specifically, according to various embodiments of the present disclosure, the processor 140 may use multiple modules to control the UI screen according to the present disclosure. Refer to Figure 2 , the multiple modules may include a voice recognition module 210, a first information acquisition module 220, a second information acquisition module 230, a target object recognition module 240, and a UI screen control module 250. Examples of implementing the processor 140 according to various embodiments of the present disclosure using multiple modules will be described below.

[0058] The processor 140 may receive a user voice through the microphone 110 while a first UI screen including multiple objects is being displayed on the display 120. In addition, the processor 140 may obtain text information corresponding to the user voice by inputting the user voice into the voice recognition module 210.

[0059] Specifically, the voice recognition module 210 may refer to a module capable of outputting text information corresponding to the input user voice, and the voice recognition module 210 may include a neural network model called automatic speech recognition (ASR).

[0060] The processor 140 may obtain first information based on the text information, where the first information includes information about a command included in the text information and information about an execution target of the command. The "first information" may refer to symbol information included in the text information corresponding to the user voice, and specifically, includes information about a command and information about an execution target of the command. The processor 140 may obtain the first information by using the first information acquisition module 220.

[0061] Refer to Figure 2, when receiving text information from the speech recognition module 210, the first information acquisition module 220 can obtain information indicating which of a predefined plurality of commands is included in the text information and information indicating what the execution target of the recognized command is, and send the information to the target object recognition module.

[0062] For example, if the first UI screen is a UI screen provided by a music application and the text information corresponding to the user's voice is "Play LOVE DIVING", the processor 140 can obtain first information including information that the command corresponding to the text "Play" is "Play" and the execution target of the command "Play" is "LOVE DIVING". Here, the command "Play" can be a command indicating music playback, and "LOVE DIVING" as the execution target of the command can be the title of the song as the target of music playback.

[0063] The process of identifying the command included in the text information from the text information corresponding to the user's voice can be performed by using a neural network model (also referred to as natural language understanding (NLU)), but is not limited thereto, and can be performed by a matching process of a predefined plurality of commands with the text information.

[0064] In addition, the processor 140 can obtain second information based on the metadata about the first UI screen, and the second information includes information about functions corresponding to a plurality of objects and information about text included in the plurality of objects. The "second information" can refer to symbol information included in the first UI screen, and includes information about functions corresponding to a plurality of objects and information about text included in the plurality of objects. The processor 140 can obtain the second information by using the second information acquisition module 230.

[0065] Refer to Figure 2 , when inputting the metadata about the first UI screen, the second information acquisition module 230 can obtain the second information by analyzing the attributes included in the first UI screen and extracting information about a plurality of functions corresponding to a plurality of objects and information about a plurality of texts included in the plurality of objects.

[0066] According to one or more embodiments of the present disclosure, when the first object among the plurality of objects includes text, the second information acquisition module 230 can identify the text included in the first object. Additionally, when the second object among the plurality of objects includes an icon, the second information acquisition module 230 can identify the text representing the function of the icon. Furthermore, the processor 140 can obtain second information including the text included in the first object and the text corresponding to the icon of the second object.

[0067] For example, the first object among multiple objects can be the text "LOVE DIVING", the second object can be an icon with text attributes (such as a "playback button"), and in addition to the above, the second information acquisition module 230 can also obtain text corresponding to the attributes of the objects included in each UI screen.

[0068] According to one or more embodiments of the present disclosure, the processor 140 can identify variable regions and invariant regions of the first UI screen based on metadata. In addition, the processor 140 can obtain second information from at least one object included in the invariant region. For example, the processor 140 can obtain second information based only on at least one object included in the invariant region without considering the objects included in the variable region of the first UI screen.

[0069] Here, the variable region can refer to a region where the information displayed within the UI screen can change, and the invariant region can refer to a region where the information displayed within the UI screen does not change. For example, if the first UI screen is a UI screen provided by a music application, the region including a list showing the titles of the searched music can be the variable region, and the region showing the title of the currently playing music can be the invariant region.

[0070] The processor 140 can identify whether a target object corresponding to the user voice exists among the multiple objects based on the comparison result between the first information and the second information. Here, the "target object" can refer to an object corresponding to the intention of the user who emits the user voice among the multiple objects included in the first UI screen. The processor 140 can use the target object recognition module 240 to identify the target object.

[0071] Refer to Figure 2 , when receiving the first information from the first information acquisition module 220 and the second information from the second information acquisition module 230, the target object recognition module 240 can identify the target object by comparing the information respectively included in the first information and the second information, and can obtain information about the identified target object and send it to the UI screen control module 250.

[0072] Specifically, the target object recognition module 240 may identify a target object based on whether the information about the command and the information about the execution target of the command included in the first information "match" the information about the function and the information about the text included in the second information. The target object recognition module 240 may identify a target object based on whether the information about the command and the information about the execution target of the command included in the first information are "included" in the information about the function and the information about the text included in the second information. In addition, the target object recognition module 240 may identify a target object based on whether the information about the command and the information about the execution target of the command included in the first information are "similar" to the information about the function and the information about the text included in the second information. Here, predefined rules or a trained neural network model may be used to identify similarity.

[0073] In addition, as described above, the processor 140 may identify a target object based on a combination of whether the information matches, whether the information is included, and whether the information is similar, and in addition to the above, various rules for identifying a target object may also be applied.

[0074] According to one or more embodiments of the present disclosure, if the information about the execution target of the command corresponds to the first text included in a plurality of objects, the processor 140 may identify the first text as the target object from the plurality of objects. In other words, the processor 140 may compare only the execution target of the command included in the first information with the text included in the second information and identify the target object.

[0075] For example, if the command included in the text information is "play", and the execution target of the command "play" is "LOVE DIVING", then in the case where there is an object including the text "LOVE DIVING" among the plurality of objects, the processor 140 may identify the object including the text "LOVE DIVING" as the target object without considering the command "play".

[0076] According to one or more embodiments of the present disclosure, if the information about the execution target of the command corresponds to the second text included in a plurality of objects, and the command corresponds to one of the plurality of functions corresponding to the plurality of objects, the processor 140 may identify the object corresponding to the one function as the target object. In other words, the processor 140 may identify a target object by comparing the information about the execution target of the command included in the first information with the text included in the second information, and by comparing the information about the command included in the first information with the information about the plurality of functions included in the second information.

[0077] For example, if the text information corresponding to the user's voice is "Open LOVE DIVING", and the command included in the text information is "Play" and the execution target of the command "Play" is "LOVE DIVING", then when there is an object including the text "LOVE DIVING" among the multiple objects on the first UI screen, and when there is an object corresponding to the command "Play", the processor 140 may identify the object corresponding to the command "Play" as the target object.

[0078] Optionally, if the information about the execution target does not correspond to the second text among the multiple texts included in the multiple objects, or if the command does not correspond to the multiple functions corresponding to the multiple objects, then the processor 140 may identify that the target object does not exist among the multiple objects.

[0079] For example, if the text information corresponding to the user's voice is "Open LOVE DIVING", and the command included in the text information is "Play" and the execution target of the command "Play" is "LOVE DIVING", then when there is an object including the text "LOVE DIVING", but there is no object corresponding to the command "Play" among the multiple objects on the first UI screen, the processor 140 may identify that the target object does not exist among the multiple objects.

[0080] If the target object is identified, the processor 140 may control the display 120 to display a second UI screen, which corresponds to the target object and is used to perform an operation corresponding to the command. Specifically, the processor 140 may control the UI screen being displayed on the display 120 by using the UI screen control module 250.

[0081] Refer to Figure 2 , when receiving information about the target object from the target object recognition module 240, the UI screen control module 250 may identify the second UI screen to be displayed after the first UI screen, obtain information about the identified second UI screen, and control the display of the second UI screen.

[0082] The process of identifying the second UI screen to be displayed after the first UI screen may be performed based on the UI diagram described below. As described above, in the memory 130, a UI diagram may be stored for each application, and the UI diagram includes multiple nodes corresponding to the types of multiple UI screens and multiple edges showing the connection relationships between the multiple nodes according to the operations performed by the transitions between the multiple UI screens.

[0083] If the target object is recognized, the processor 140 may obtain a first embedding vector corresponding to the second information based on the metadata regarding the first UI screen. Specifically, the first embedding vector may be obtained by inputting the metadata regarding the first UI screen through a trained neural network model (e.g., the first embedding vector acquisition module to be described below) to obtain the first embedding vector corresponding to the second information. In addition, the processor 140 may obtain a second embedding vector showing the attributes of a plurality of nodes included in the UI graph. Here, the second embedding vector may be pre-stored in the memory 130 like the information regarding the UI graph.

[0084] The processor 140 may identify a first node corresponding to the first UI screen from among the plurality of nodes based on the comparison result between the first information and the second information and the comparison result between the first embedding vector and the second embedding vector.

[0085] According to one or more embodiments of the present disclosure, the processor 140 may identify the recognized node as the first node corresponding to the first UI screen based on that the target object corresponding to the user voice is recognized from among the plurality of objects based on the comparison between the first information corresponding to the user voice and the second information and that the node having the second embedding vector with a similarity to the first embedding vector greater than or equal to a preset threshold is recognized from among the plurality of nodes based on the comparison between the first embedding vector and the second embedding vector. In other words, the processor 140 may not only compare the embedding vectors but also compare the symbolic information therewith, and identify the first node corresponding to the first UI screen currently displayed on the display 120 from among the plurality of nodes included in the UI graph.

[0086] After that, the processor 140 may identify a second node for performing an operation corresponding to the command from among at least one node connected to the first node based on the information regarding the command and the information regarding the execution target of the command, and control the display 120 to display a second UI screen corresponding to the second node.

[0087] Specifically, the processor 140 may identify at least one sequence from the first node until the target node corresponding to the target of the user voice, and identify the second node from among at least one node connected to the first node by identifying an edge corresponding to the operation corresponding to the command according to the user voice based on the information regarding the command and the information regarding the execution target of the command.

[0088] For example, if there is an object including the text "LOVE DIVING" which is the execution target of the command, and if there is an object corresponding to the command "Play" among multiple objects in the first UI screen, the processor 140 can execute the function of the object corresponding to the command "Play" and control the display 120 to change the first UI screen to a second UI screen including information showing that music is being played back.

[0089] In addition, if the target object is not recognized, the processor 140 may control the display 120 to keep displaying the first UI screen. For example, if there is an object including the text "LOVE DIVING" as the execution target of the command, and if there is no object corresponding to the command "play" among the multiple objects of the first UI screen, the above situation may be a situation where the playback of the music "LOVE DIVING" has been performed in the first UI screen, that is, when the target of the command has been achieved. Therefore, in this case, the processor 140 may control the display 120 to maintain the display of the first UI screen.

[0090] In addition, if the target object is not recognized, the processor 140 can control the display 120 to display a third UI screen for performing an operation corresponding to the command. For example, if there is an object including the text "LOVEDIVING" as the execution target of the command, and there is no object corresponding to the command "play" among the multiple objects of the first UI screen, the processor 140 can control the display 120 to display a third UI screen including an object corresponding to the command "play", and then achieve the target of the command by performing an operation of clicking the object corresponding to the command "play".

[0091] According to one or more of the above-mentioned embodiments, the electronic device 100 can clearly identify the target object corresponding to the user voice in the UI screen currently being displayed by comparing the first information obtained from the user voice with the second information obtained from the UI screen currently being displayed, and perform voice control that matches the user's intention.

[0092] Specifically, according to the present disclosure, even for an application that does not provide an application programming interface (API) related to voice control, voice control of a UI screen currently being displayed can be achieved by obtaining embedding vectors and symbol information based on metadata about the above application and by comparing the above embedding vectors and symbol information with the embedding vectors and symbol information corresponding to the user's voice.

[0093] Figure 3 , Figure 4 and Figure 5It is a diagram showing a process of identifying a target object from a UI screen based on a comparison result between first information and second information according to various embodiments of the present disclosure.

[0094] Referring to Figure 3 , according to one or more embodiments of the present disclosure, a first UI screen may be a UI screen provided by a chat application and may be a UI screen showing information about a plurality of chat rooms.

[0095] As described above, if a user voice is received while the first UI screen is being displayed on the display 120, the processor 140 may obtain text information corresponding to the received user voice by inputting the received user voice into the voice recognition module 210. In addition, the processor 140 may obtain first information based on the text information, and the first information includes information about a command included in the text information and information about an execution target of the command.

[0096] In Figure 3 's example, the text information corresponding to the user voice may be "Send a message asking SAM, SIHO to call". When obtaining the text information as described above, the processor 140 may obtain first information showing that the command corresponding to the text "Send a message" is "Send a message", and the execution target of the command "Send a message" (i.e., the recipient of the message) is the person called "SAM, SIHO".

[0097] The processor 140 may obtain second information based on metadata about the first UI screen, and the second information includes information about functions corresponding to a plurality of objects and information about text included in the plurality of objects.

[0098] In Figure 3 's example, the processor 140 may obtain second information including information about an object (icon) showing the "Add chat room" function, information about text (such as "CHIL, SIHO", "Hi, there", "Hello", and "SAM,SIHO" 310), etc.

[0099] The processor 140 may identify whether a target object corresponding to the user voice exists in the plurality of objects based on a comparison result between the first information and the second information. Specifically, the process of identifying a target object based on a comparison result between the first information and the second information may be performed according to various embodiments described below.

[0100] According to one or more embodiments of the present disclosure, a target object may be identified based on a match between information about a command and information about an execution target of the command included in the first information and information about a function and information about text included in the second information.

[0101] In Figure 3 the example of, the first information may include information indicating that the execution target of the command is a person named "SAM, SIHO", and the second information may include the text "SAM, SIHO". For example, since "SAM, SIHO" in the first information matches "SAM, SIHO" in the second information, the processor 140 may identify the object corresponding to "SAM, SIHO" 310 included in the first UI screen as the target object.

[0102] According to one or more embodiments of the present disclosure, the target object may be identified based on whether the information about the function included in the second information and the information about the text include the information about the command and the information about the execution target of the command included in the first information.

[0103] In Figure 3 the example of, since "SAM, SIHO" in the first information is included in "SAM, SIHO" in the second information, the processor 140 may identify that the target object corresponds to "SAM, SIHO" 310 included in the first UI screen. In addition, different from the example of Figure 3 even when the second information includes "brother SAM, SIHO", since "SAM, SIHO" in the first information is included in "brother SAM, SIHO" in the second information, the processor 140 may also identify the object corresponding to "brother SAM, SIHO" included in the first UI screen as the target object.

[0104] Referring to Figure 4 according to one or more embodiments of the present disclosure, the first UI screen may be a UI screen provided by a chat application and may be a UI screen showing information about a specific chat room.

[0105] In Figure 4 the example of, the processor 140 may obtain the second information including information about texts (such as "SAM, SIHO" 410 and "ARE YOU IN CONTACT WITH CHIL, SIHO?" 420) as information about functions corresponding to multiple objects.

[0106] In Figure 4In the example, if the text information corresponding to the user's voice is "Send a message requesting a call from SAM, SIHO", the processor 140 may obtain first information indicating that the execution target of the command is the person called "SAM, SIHO". And since the above situation is one where "SAM, SIHO" in the first information matches "SAM, SIHO" in the second information, the processor 140 may identify "SAM, SIHO" 410 included in the first UI screen as the target object. In this case, since the first UI screen is the UI screen of the chat room showing the person "SAM, SIHO", the processor 140 may input a message indicating "Call me" into the text input field included in the first UI screen without switching the UI screen and execute the function corresponding to the message send button.

[0107] In addition, according to Figure 4 the embodiment, if the text information corresponding to the user's voice is "Send a message requesting a call from CHIL, SIHO", the processor 140 may obtain first information indicating that the execution target of the command is the person called "CHIL, SIHO". And since the above situation is one where "CHIL, SIHO" in the first information matches "CHIL, SIHO" in the second information, the processor 140 may identify "CHIL, SIHO" 420 included in the first UI screen as the target object. In this case, although the first UI screen is the UI screen of the chat room showing the person called "SAM, SIHO", based on the presence of the target object in the first UI screen, if a message of "Call me" is input into the text input field included in the first UI screen without switching the UI screen and the function corresponding to the message send button is executed, an operation that does not match the user's speech intention will be performed. Therefore, the processor 140 may identify the target object by determining the information to be compared from the first information and the second information.

[0108] Specifically, the processor 140 may identify a variable area where the information displayed within the first UI screen can be changed and an invariant area different from the variable area based on the metadata regarding the first UI screen, and obtain the second information from at least one object included in the invariant area. For example, the processor 140 may obtain the second information based only on at least one object included in the invariant area without considering the objects included in the variable area of the first UI screen.

[0109] For example, the processor 140 may not only divide the variable area and the invariant area based on the metadata regarding the first UI screen, but also use a predefined algorithm or a neural network model trained to divide the variable area and the invariant area included in the UI screen to divide the variable area and the invariant area.

[0110] Reference Figure 4 As shown in Figure 4 , the text "SAM, SIHO" 410 may be included in the invariant region, and the text "CHIL, SIHO" 420 may be included in the variable region. Therefore, when obtaining the second information based on the metadata regarding the first UI screen, the processor 140 may obtain the second information without considering the text "CHIL, SIHO" 420 included in the variable region but considering the text "SAM, SIHO" 410 included in the invariant region.

[0111] Accordingly, if the text information corresponding to the user's voice is "Send a message asking CHIL, SIHO to call", since "CHIL, SIHO" of the first information does not match "SAM, SIHO" of the second information and "CHIL, SIHO" of the first information is not included in "SAM, SIHO" of the second information, the processor 140 may recognize that the target object corresponding to the user's voice does not exist among the multiple objects included in the first UI screen. In this case, the processor 140 may perform an operation matching the user's speech intention by controlling the display 120 to display a third UI screen showing the chat room of the person called "CHIL, SIHO", inputting the message "Call me" in the text input field included in the third UI screen, and performing the function corresponding to the message send button.

[0112] Reference Figure 5 As shown in Figure 5 , the first UI screen according to one or more embodiments of the present disclosure may be a UI screen provided by a music application and may include information about multiple music contents, information about the currently selected music content, and the like.

[0113] In Figure 5 the example of Figure 5 , the text information corresponding to the user's voice may be "Open LOVE DIVING". When obtaining the described text information, the processor 140 may obtain the first information, which includes information indicating that the command corresponding to the text "Open" is "Play" and the execution target of the command "Play" is "LOVE DIVING".

[0114] In addition, Figure 5 in the example of Figure 5 , the processor 140 may obtain the second information, which includes information about the text 510 "LOVE DIVING" in the variable region showing information about multiple music contents, information about the text 520 "LOVE DIVING" in the invariant region showing information about the currently selected music content, information about the object 530 showing the function "Play the selected song", and the like.

[0115] The process of identifying the target object can be performed by considering whether the information about the command and the information about the execution target of the command included in the first information match the information about the function and the information about the text included in the second information, and whether the information about the command and the information about the execution target of the command included in the first information are included in the information about the function and the information about the text included in the second information.

[0116] Specifically, the processor 140 can identify the target object based on whether the information about the execution target of the command included in the first information matches the information about the text included in the second information and whether the information about the command included in the first information is included in the information about the function included in the second information.

[0117] Referring to Figure 5 , since "LOVE DIVING", which is the execution target of the command included in the first information, matches "LOVE DIVING", which is the text included in the second information, and the command "play" included in the first information is included in "play the selected song", which is the text of the function showing the object included in the second information, the processor 140 can identify the object 530 showing the function of "play the selected song" as the target object from among the multiple objects included in the first UI screen. In this case, the processor 140 can play back the music content called "LOVE DIVING" by executing the function "play the selected song" and control the display 120 to display the second UI screen, which shows that the music content called "LOVE DIVING" is being played back.

[0118] In addition, different from the example of Figure 5 , if "LOVEDIVING", which is the execution target of the command included in the first information, matches "LOVE DIVING", which is the text included in the second information, and the command "play" included in the first information is not included in the text representing the function of the object included in the second information, the processor 140 can identify that the target object does not exist in the first UI screen and control the display 120 to display the third UI screen for playing back the music content called "LOVEDIVING".

[0119] In the present disclosure, the second UI screen can refer to a UI screen for performing an operation corresponding to a command when there is a target object corresponding to the user's voice in the first UI screen, and the third UI screen can refer to a UI screen for performing an operation corresponding to a command when there is no target object corresponding to the user's voice in the first UI screen.

[0120] In addition, different from Figure 5Unlike the example, even though "LOVEDIVING" which is the execution target of the command included in the first information matches "LOVE DIVING" which is the text included in the second information, if the music content called "LOVEDIVING" has been played back in the instance (for example, an object showing the function of "pausing the selected song" is included in the first UI screen instead of the object showing the function of "playing the selected song"), the processor 140 may keep the display of the first UI screen.

[0121] In addition, different from Figure 5 the example, if "LOVEDIVING" which is the execution target of the command included in the first information does not match "LOVE DIVING" which is the text included in the second information, the processor 140 may recognize that the target object does not exist in the first UI screen without considering whether the command "play" included in the first information is included in the second information, and the processor 140 may control the display 120 to display a third UI screen for searching for the music content called "LOVEDIVING", or display a third UI screen for determining the music content to be played back.

[0122] Figure 6 is a block diagram showing a plurality of modules used in the process of generating a UI diagram according to an embodiment of the present disclosure.

[0123] Figure 7 is a flowchart for showing the process of generating a UI diagram according to an embodiment of the present disclosure.

[0124] Referring to Figure 6 , a plurality of modules according to an embodiment of the present disclosure may include a second information acquisition module 610, a third information acquisition module 620, a first embedding vector acquisition module 630, a second embedding vector acquisition module 640, a comparison module 650, and a UI diagram update module 660, and the comparison module 650 may specifically include a symbol information comparison module 651 and an embedding vector comparison module 652. The process of generating a UI diagram according to the present disclosure will be described below with reference to Figure 6 and Figure 7 the process of generating a UI diagram according to the present disclosure.

[0125] Referring to Figure 7, when a UI screen is selected in operation S710, in operation S720, the processor 140 may obtain second information based on the metadata about the UI screen. The second information includes information about functions corresponding to multiple objects and information about text included in the multiple objects. Specifically, when the metadata about the selected UI screen is input, the second information acquisition module 610 may obtain the second information as symbol information included in the UI screen and send the second information to the symbol information comparison module 651. As described above, the second information may include information about functions corresponding to multiple objects and information about text included in the multiple objects included in the UI screen.

[0126] In operation S730, the processor 140 may obtain third information based on the UI diagram. The third information includes information about functions corresponding to multiple nodes respectively and information about text corresponding to multiple nodes respectively. Specifically, when the information about the UI diagram is input, the third information acquisition module 620 may obtain the third information as symbol information included in the UI diagram and send the third information to the symbol information comparison module 651. Here, the third information refers to information about the UI screen corresponding to each node included in the UI diagram, and specifically, may include information about functions corresponding to multiple objects and information about text included in the multiple objects included in the UI screen.

[0127] In operation S740, the processor 140 may compare the second information with the third information, and in operation S750, may identify whether a node corresponding to the UI screen exists in the multiple nodes. Specifically, when receiving the second information from the second information acquisition module 610 and receiving the third information from the third information acquisition module 620, the symbol information comparison module 651 may identify whether a node corresponding to the UI screen exists in the multiple nodes based on at least one of the following: whether the second information matches the third information, whether the second information is similar to the third information, and whether the second information is included in the third information.

[0128] If - yes in operation S750, and a node corresponding to the UI screen is identified as existing in the multiple nodes, then in operation S760, the processor 140 may obtain a first embedding vector corresponding to the second information based on the metadata about the UI screen. Specifically, when the metadata about the UI screen is input, the first embedding vector acquisition module 630 may obtain the first embedding vector showing the attributes of the multiple objects included in the UI screen by encoding the metadata about the UI screen and send the first embedding vector to the embedding vector comparison module 652.

[0129] In operation S770, the processor 140 may obtain second embedding vectors corresponding to a plurality of nodes based on the UI graph. Specifically, when information about the UI graph is input, the first embedding vector acquisition module 630 may obtain second embedding vectors showing the attributes of the plurality of nodes included in the UI graph by encoding the information about the UI graph, and send the second embedding vectors to the embedding vector comparison module 652.

[0130] In operation S780, the processor 140 may compare the first embedding vector with the second embedding vector, and in operation S785, may identify whether a node having a similarity to the UI screen greater than or equal to a threshold exists among the plurality of nodes. Specifically, when receiving the first embedding vector from the first embedding vector acquisition module 630 and receiving the second embedding vector from the second embedding vector acquisition module 640, the embedding vector comparison module 652 may calculate the similarity between the first embedding vector and the second embedding vector, and identify whether a node having a calculated similarity greater than or equal to the threshold exists.

[0131] For example, the similarity between the first embedding vector and the second embedding vector may be calculated based on various methods, such as cosine similarity or Euclidean distance.

[0132] If, in operation S785 - YES, a node having a similarity to the UI screen greater than or equal to the threshold is identified as existing among the plurality of nodes, then in operation S790, the processor 140 may add information about the selected UI screen to the corresponding node. For example, when it is identified based on comparing the symbol information of the second information and the third information that the first node corresponding to the first UI screen exists among the plurality of nodes, and it is identified based on comparing the first embedding vector and the second embedding vector that the similarity between the first node and the first UI screen is greater than or equal to the threshold, the UI graph update module 660 may add the information about the first UI screen as the information about the first node. Thus, the processor 140 may identify the first UI screen as corresponding to the first node when subsequently controlling the display of the UI screen based on user voice.

[0133] If, in operation S750 - no, the node corresponding to the UI screen is recognized as not present among the multiple nodes, or if, in operation S785 - no, the node having a similarity to the UI screen greater than or equal to the threshold is recognized as not present among the multiple nodes, then in operation S795, the processor 140 may add a new node to the UI graph. For example, if the node corresponding to the first UI screen is recognized as not present among the multiple nodes, the UI graph update module 660 may update the UI graph such that the first node corresponding to the first UI screen is included in the UI graph. Thus, the processor 140 may recognize the first UI screen as corresponding to the first node when subsequently displaying the UI screen based on user voice control.

[0134] In addition, in Figure 6 and Figure 7 the following embodiments have been described, where first the second information as symbol information is compared with the third information, and the first embedding vector is compared with the second embedding vector only when the node corresponding to the UI screen is recognized as present among the multiple nodes based on the comparison of the second information with the third information, but the present disclosure is not limited thereto.

[0135] For example, the electronic device 100 may first compare the first embedding vector with the second embedding vector, and may also compare the second information with the third information only when the node having a similarity to the UI screen greater than or equal to the threshold is recognized as present among the multiple nodes based on the comparison of the first embedding vector with the second embedding vector. In addition to the above, the electronic device 100 may quantify the comparison result of the second information and the third information, and identify the node corresponding to the UI screen from among the multiple nodes by combining the quantified comparison result with the value representing the similarity between the first embedding vector and the second embedding vector.

[0136] As described above, according to the embodiments described above with reference to Figure 6 and Figure 7 the electronic device 100 according to the present disclosure may, as described above, while selecting various UI screens, construct a UI graph based on the similarity of the embedding vectors corresponding to each UI screen in the UI screen and whether the symbol information shown by the multiple objects included in each UI screen in the UI screen corresponds, and thus, may subsequently effectively perform control of the UI screen based on user voice.

[0137] Specifically, according to the present disclosure, since nodes can be generated by considering not only the similarity of the embedding vectors corresponding to each UI screen in the UI screen but also whether symbol information such as text information corresponds, the UI graph can be constructed more effectively by partitioning the nodes of the UI graph not only based on the similarity of the UI screens but also based on the functions of the objects included in the UI screens.

[0138] Figure 8 is a flowchart showing the configuration of an electronic device according to an embodiment of the present disclosure.

[0139] Referring to Figure 8 , the electronic device 100 according to an embodiment of the present disclosure may include not only a microphone 110, a display 120, a memory 130, and a processor 140, but may also include a communicator 150, an input device 160, and an output device 170. However, as Figure 1 and Figure 8 shown in the configuration is only an example, and when implementing the present disclosure, in addition to Figure 1 and Figure 8 the configuration shown in, new configurations may be added or some configurations may be omitted.

[0140] The communicator 150 may include a circuit and perform communication with an external device. Specifically, the processor 140 may receive various data or information from an external device connected through the communicator 150 and send various data or information to the external device.

[0141] The communicator 150 may include at least one of a Wi-Fi module, a Bluetooth module, a wireless communication module, an NFC module, and an ultra-wideband (UWB) module. Specifically, the Wi-Fi module and the Bluetooth module may perform communication in Wi-Fi and Bluetooth methods, respectively. When using the Wi-Fi module or the Bluetooth module, various connection information such as a service set identifier (SSID) is first sent and received, and various information may be sent and received after communicatively connecting by using the various connection information.

[0142] In addition, the wireless communication module may perform communication according to various communication standards (such as, for example but not limited to, the Institute of Electrical and Electronics Engineers (IEEE), ZigBee, third generation (3G), 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), fifth generation (5G), etc.). In addition, the NFC module may use the 13.56 MHz band in various radio frequency identification (RFID) bands (such as, for example but not limited to 135 kilohertz (kHz), 13.56 megahertz (MHz), 433 MHz, 860 - 960 MHz, 2.45 gigahertz (GHz), etc.) to perform communication in a near field communication (NFC) method. In addition, the UWB module may accurately measure the time of arrival (ToA) and the angle of arrival (AoA) through communication between UWB antennas. The time of arrival (ToA) is the time when a pulse arrives at a target object, and the angle of arrival (AoA) is the angle at which a pulse arrives from a transmitting device. Therefore, accurate distance and position identification can be performed within an error range of several tens of centimeters (cm) indoors.

[0143] Specifically, according to various embodiments of the present disclosure, the processor 140 may receive information about the UI diagram, information about the application, metadata about the UI screen, etc. from an external device through the communicator 150. In addition, the processor 140 may control the communicator 150 to send information about the user voice to a server including a speech recognition model, and receive text information corresponding to the user voice from an external device through the communicator 150.

[0144] The input device 160 may include a circuit, and the processor 140 may receive a user command for controlling the operation of the electronic device 100 through the input device 160. Specifically, the input device 160 may be implemented in the form of, for example, a camera (not shown) and a remote control signal receiver (not shown). In addition, the input device 160 may be implemented in the form of a touch screen included in the display 120.

[0145] Specifically, according to various embodiments of the present disclosure, the processor 140 may receive a user input for activating the microphone 110 through the input device 160. In addition, the processor 140 may receive a user input for selecting a UI screen through the input device 160, and receive a user input for generating a UI diagram corresponding to the selected UI screen.

[0146] The output device 170 may include a circuit, and the processor 140 may output various functions that can be executed by the electronic device 100 through the output device 170. In addition, the output device 170 may include at least one of a speaker and an indicator. The speaker may output audio data under the control of the processor 140, and the indicator may be lit under the control of the processor 140.

[0147] Specifically, according to various embodiments of the present disclosure, the processor 140 may output a message for guiding that an operation corresponding to the user voice has been executed through the output device 170. In addition, the processor 140 may output a message for guiding what object the recognized target object is among a plurality of objects through the output device 170.

[0148] Above, the microphone 110 has been described as a configuration separated from the input device 160, and the display 120 has been described as a configuration separated from the output device 170, but the microphone 110 and the display 120 may be a configuration of the input device 160 and the output device 170, respectively.

[0149] Figure 9 is a flowchart showing a method of controlling an electronic device according to an embodiment of the present disclosure.

[0150] Refer to Figure 9, at operation S910, the electronic device 100 may receive user speech while a first user interface (UI) screen including a plurality of objects is being displayed on the display 120. Then, at operation S920, the electronic device 100 may obtain text information corresponding to the user speech by inputting the user speech into a speech recognition model. Here, the speech recognition model may be stored not only in the electronic device 100 but also in an external server.

[0151] At operation S930, the electronic device 100 may obtain first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command. Specifically, the "first information" may refer to symbol information included in the text information corresponding to the user speech, and specifically, includes information about a command and information about an execution target of the command.

[0152] At operation S940, the electronic device 100 may obtain second information based on metadata about the first UI screen, the second information including information about functions corresponding to the plurality of objects and information about text included in the plurality of objects. For example, if the first object among the plurality of objects includes text, the second information obtaining module 230 may identify the text included in the first object. Additionally, if the second object among the plurality of objects includes an icon, the second information obtaining module 230 may identify text indicating the function of the icon. Further, the electronic device 100 may obtain second information including the text included in the first object and text corresponding to the icon of the second object.

[0153] At operation S950, the electronic device 100 may identify whether a target object corresponding to the user speech exists among the plurality of objects based on a comparison result between the first information and the second information.

[0154] Specifically, the target object recognition module may identify the target object based on whether the information about the command and the information about the execution target of the command included in the first information "match" the information about the function and the information about the text included in the second information. The target object recognition module may identify the target object based on whether the information about the command and the information about the execution target of the command included in the first information are "included" in the information about the function and the information about the text included in the second information. Additionally, the target object recognition module may identify the target object based on whether the information about the command and the information about the execution target of the command included in the first information are "similar" to the information about the function and the information about the text included in the second information. Here, predefined rules or a trained neural network model may be used to identify similarity.

[0155] In addition, the electronic device 100 may identify a target object based on a combination of whether it matches, whether it includes, and whether it is similar as described above, and in addition to the above, various rules for identifying the target object may be applied.

[0156] When the target object among the multiple objects is identified in operation S950 - Yes, in operation S960, the electronic device 100 may display a second UI screen for performing an operation corresponding to the command.

[0157] Specifically, when the target object is identified, the electronic device 100 may obtain a first embedding vector corresponding to the second information based on the metadata regarding the first UI screen. The electronic device 100 may obtain a second embedding vector showing the attributes of multiple nodes included in the UI diagram. The electronic device 100 may identify a first node corresponding to the first UI screen from among the multiple nodes based on the comparison result between the first information and the second information and the comparison result between the first embedding vector and the second embedding vector. Then, the electronic device 100 may identify a second node for performing an operation corresponding to the command from at least one node connected to the first node based on the information of the command and the information regarding the execution target of the command, and control the display 120 to display a second UI screen corresponding to the second node.

[0158] If the target object is not identified from among the multiple objects in operation S950 - No, in operation S970, the electronic device 100 may maintain the display of the first UI screen or display a third UI screen for performing an operation corresponding to the command.

[0159] In addition, the method for controlling the electronic device 100 according to the above - described embodiment may be implemented as a program and provided in the electronic device 100. Specifically, a program including the method for controlling the electronic device 100 may be stored and provided in a non - transitory computer - readable medium.

[0160] Specifically, with respect to a non-transitory computer-readable storage medium storing a program for a method of controlling an electronic device 100, the method of controlling the electronic device 100 includes: while displaying a first user interface (UI) screen including a plurality of objects on a display 120 of the electronic device 100 and receiving a user voice, obtaining text information corresponding to the user voice by inputting the user voice into a speech recognition model, obtaining first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command, obtaining second information based on metadata about the first UI screen, the second information including information about functions corresponding to the plurality of objects and information about text included in the plurality of objects, based on a comparison result between the first information and the second information, identifying whether a target object corresponding to the user voice exists among the plurality of objects, and based on the identified target object, controlling the display 120 to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.

[0161] In the foregoing, a method of controlling an electronic device 100 and a computer-readable recording medium storing a program for the method of controlling the electronic device 100 have been briefly described, but this is only to omit redundant descriptions, and various embodiments of the electronic device 100 can be applied to the method of controlling the electronic device 100, and even to a computer-readable storage medium storing a program for the method of controlling the electronic device 100.

[0162] According to various embodiments of the present disclosure as described above, the electronic device 100 can clearly identify a target object corresponding to a user voice from a currently displayed UI screen by comparing first information obtained from the user voice with second information obtained from the currently displayed UI screen, and perform voice control matching the user's intention.

[0163] Specifically, according to the present disclosure, even for an application that does not provide an application programming interface (API) for voice control, an embedding vector and symbol information can be obtained based on metadata about the application, and voice control can be performed on the currently displayed UI screen by comparing the above with the embedding vector and symbol information corresponding to the user voice.

[0164] Functions associated with artificial intelligence according to the present disclosure can be operated by a processor 140 and a memory 130 of the electronic device 100. The processor 140 may be formed of one or more processors 140. At this time, the one or more processors 140 may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural processing unit (NPU), but is not limited to the examples of the processor 140 described above.

[0165] The CPU can be used as a general-purpose processor 140 that can not only perform typical calculations but also perform artificial intelligence calculations, and can effectively execute complex programs through a multi-level cache structure. The CPU may be advantageous in a serial calculation method in which the previous calculation result and the subsequent calculation result can be organically connected through sequential calculation. Except for the case specified for the above CPU, the general-purpose processor 140 is not limited to the above example.

[0166] The GPU can be a processor 140 for large-scale calculations such as floating-point calculations used in graphics processing, and can execute large-scale calculations in parallel through a large number of integrated cores. Specifically, compared with the CPU, the GPU may be advantageous in parallel calculation methods such as convolution calculation. In addition, the GPU can be used as a coprocessor 140 to supplement the functions of the CPU. Except for the case specified as the above GPU, the processor 140 for large-scale calculations is not limited to the above example.

[0167] The NPU can be a processor 140 dedicated to artificial intelligence calculations using artificial neural networks, and can be implemented as hardware (e.g., silicon) for each layer forming the artificial neural network. At this time, since the NPU is specifically designed according to the specifications required by the company, it has less freedom compared with the CPU or GPU, but it can effectively process the artificial intelligence calculations required by the company. At the same time, as a processor 140 dedicated to artificial intelligence calculations, the NPU can be implemented in various forms (such as, for example but not limited to, tensor processing unit (TPU), intelligent processing unit (IPU), vision processing unit (VPU), etc.). Except for the case specified as the above NPU, the artificial intelligence processor 140 is not limited to the above example.

[0168] In addition, one or more processors 140 can be implemented as a system-on-chip (SoC). At this time, in the SoC, in addition to one or more processors 140, a memory 130 and a network interface such as a bus for data communication between the processor 140 and the memory 130 can also be included.

[0169] When multiple processors 140 are included in the SoC included in the electronic device 100, the electronic device 100 may perform computations associated with artificial intelligence (e.g., learning of an artificial intelligence model or computations associated with inference) by using some of the multiple processors 140. For example, the electronic device 100 may perform computations associated with artificial intelligence by using at least one of a GPU, an NPU, a VPU, a TPU, and a hardware accelerator dedicated to artificial intelligence computations (such as convolutional computations and matrix multiplication computations) among the multiple processors 140. However, the above is only one embodiment of the present disclosure, and a general-purpose processor 140 (such as a CPU) may be used to process computations associated with artificial intelligence.

[0170] In addition, the electronic device 100 may perform computations on functions associated with artificial intelligence by using multiple cores (e.g., dual-core, quad-core, etc.) included in one processor 140. Specifically, the electronic device 100 may use the multiple cores included in the processor 140 to perform artificial intelligence computations, such as convolutional computations and matrix multiplication computations, in parallel.

[0171] One or more processors 140 may control the processing of input data according to predefined operation rules or an artificial intelligence model stored in the memory 130. The predefined operation rules or the characteristics of the artificial intelligence model may be created by learning.

[0172] Creation by learning may refer to forming a predefined operation rule or an artificial intelligence model with desired characteristics by applying a learning algorithm to multiple learning data. The learning may be performed in the device itself that executes artificial intelligence according to the present disclosure, or by a separate server / system.

[0173] The artificial intelligence model may be formed by multiple neural network layers. At least one layer may have at least one weight value and perform layer computations through the computation results of the previous layer and at least one defined computation. Examples of neural networks may include convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), deep Q networks, and transformers, and the neural networks of the present disclosure are not limited to the above examples unless otherwise specified.

[0174] The learning algorithm may be a method for training a predetermined target machine (e.g., a robot) to make decisions or predictions on its own using multiple learning data. Examples of learning algorithms may include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithms of the present disclosure are not limited to the above examples unless the learning algorithms are specified in the present disclosure.

[0175] A machine-readable storage medium may be provided in the form of a non-transitory storage medium. In this document, "non-transitory storage medium" only means that it is a tangible device and does not include signals (e.g., electromagnetic waves), and this term does not distinguish whether data is stored semi-permanently or temporarily in the storage medium. For example, a "non-transitory storage medium" may include a buffer for temporarily storing data.

[0176] According to one or more embodiments of the present disclosure, a method according to various embodiments described in the present disclosure may be provided to be included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., PLAYSTORE™), or may be directly distributed (e.g., downloaded or uploaded) between two user devices (e.g., a smart phone). In the case of online distribution, at least a part of the computer program product (e.g., a downloadable application) may be at least temporarily stored in a storage medium readable by a device (such as the memory 130 of a manufacturer's server, an application store's server, or a relay server), or may be temporarily generated.

[0177] Each of the elements according to various embodiments of the present disclosure as described above (e.g., a module or a program) may be formed as a single entity or multiple entities, and some of the above-described sub-elements may be omitted, or other sub-elements may be further included in various embodiments. Optionally or additionally, some elements (e.g., a module or a program) may be integrated into one entity to perform the same or similar functions as those performed by the corresponding elements before integration.

[0178] Operations performed by a module, a program, or another element according to various embodiments of the present disclosure may be performed sequentially, in parallel, repeatedly, or in a heuristic manner, or at least some operations may be performed in a different order, omitted, or different operations may be added.

[0179] In addition, the term "component" or "module" used in the present disclosure may include a unit formed of hardware, software, or firmware, and may be used interchangeably with terms such as, for example but not limited to, logic, logic block, component, circuit, etc. A "component" or "module" may be an integrally formed component or the smallest unit or a part of a component that performs one or more functions. For example, a module may be formed as an application specific integrated circuit (ASIC).

[0180] Various embodiments of the present disclosure can be implemented using software including instructions stored in a machine-readable storage medium (e.g., a computer). The machine can call the instructions stored in the storage medium, and as a device operable according to the called instructions, can include an electronic device (e.g., electronic device 100) according to the above embodiments.

[0181] Based on the instructions executed by the processor, the processor can directly or using other elements perform functions corresponding to the instructions under the control of the processor. The instructions can include code generated by a compiler or executed by an interpreter.

[0182] Although the present disclosure has been shown and described with reference to various embodiments of the present disclosure, those skilled in the art will understand that various changes in form and detail can be made therein without departing from the spirit and scope of the present disclosure defined by the appended claims and their equivalents.

Claims

1. An electronic device, comprising: a microphone; a display; a memory; and a processor configured to: Based on user speech received through the microphone while a first user interface (UI) screen including a plurality of objects is being displayed on the display, obtain text information corresponding to the user speech by inputting the user speech into a speech recognition model. Obtain first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command. Obtain second information based on metadata about the first UI screen, the second information including information about functions corresponding to the plurality of objects and information about text included in the plurality of objects. Based on a comparison result between the first information and the second information, identify whether a target object corresponding to the user speech exists among the plurality of objects, and Based on the target object being identified, control the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.

2. The electronic device according to claim 1, wherein, The processor is further configured to: Based on the information about the execution target corresponding to a first text included in the plurality of objects, identify the object including the first text among the plurality of objects as the target object.

3. The electronic device according to claim 1, wherein, The processor is further configured to: Based on the information about the execution target corresponding to a second text included in the plurality of objects and the command corresponding to one of a plurality of functions corresponding to the plurality of objects, identify the object corresponding to the one function among the plurality of objects as the target object.

4. The electronic device according to claim 1, wherein, The processor is further configured to: Based on the target object not being identified, control the display to maintain the display of the first UI screen.

5. The electronic device according to claim 1, wherein The processor is further configured to: Based on the target object not being identified, control the display to display a third UI screen for performing an operation corresponding to the command.

6. The electronic device according to claim 1, wherein, The processor is further configured to: Based on the metadata, identify a variable area and a constant area different from the variable area, in which the information displayed within the first UI screen changes; and Obtain second information from at least one object included in the constant area.

7. The electronic device according to claim 1, Among them, The memory stores UI diagrams for each application, the UI diagrams including a plurality of nodes corresponding to types of a plurality of UI screens and a plurality of edges showing connection relationships between the plurality of nodes according to operations performed through transitions between the plurality of UI screens. Wherein, the processor is further configured to: Based on the target object being identified, obtain a first embedding vector corresponding to the second information based on the metadata about the first UI screen. Based on a comparison result between the first information and the second information and a comparison result between the first embedding vector corresponding to the second information and second embedding vectors corresponding to the plurality of nodes respectively, identify a first node corresponding to the first UI screen among the plurality of nodes. Based on the information about the command and the information about the execution target of the command, identify, from at least one node connected to the first node, a second node for performing an operation corresponding to the command, and control the display to display a second UI screen corresponding to the second node.

8. A method for controlling an electronic device, the method comprising: Based on user speech received while a first user interface (UI) screen including a plurality of objects is displayed on a display of the electronic device, obtain text information corresponding to the user speech by inputting the user speech into a speech recognition model; Obtain first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command; Obtain second information based on metadata about the first UI screen, the second information including information about functions corresponding to the plurality of objects and information about text included in the plurality of objects; Based on a comparison result between the first information and the second information, identify whether a target object corresponding to the user speech exists in the plurality of objects; and Based on the target object being identified, control the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.

9. The method according to claim 8, wherein, The step of identifying whether the target object exists includes: based on the information about the execution target corresponding to a first text included in the plurality of objects, identifying an object including the first text in the plurality of objects as the target object.

10. The method according to claim 8, wherein, The step of identifying whether the target object exists includes: based on the information about the execution target corresponding to a second text included in the plurality of objects and the command corresponding to one of a plurality of functions corresponding to the plurality of objects, identifying an object corresponding to the one function in the plurality of objects as the target object.

11. The method according to claim 8, further comprising: Based on the target object not being identified, control the display to maintain the display of the first UI screen.

12. The method according to claim 8, further comprising: Based on the target object not being identified, control the display to display a third UI screen for performing an operation corresponding to the command.

13. The method according to claim 8, wherein, The step of obtaining the second information includes: Based on the metadata, identify a variable area and a non-variable area different from the variable area, in the variable area, information displayed within the first UI screen changes; and Obtain the second information from at least one object included in the non-variable area.

14. The method according to claim 8, further comprising: Obtain a UI graph, the UI graph including a plurality of nodes corresponding to types of a plurality of UI screens and a plurality of edges showing connection relationships between the plurality of nodes according to operations performed through transitions between the plurality of UI screens; Based on the target object being identified, obtain a first embedding vector corresponding to the second information based on the metadata about the first UI screen; Identify a first node corresponding to a first UI screen among the multiple nodes based on a comparison result between a first piece of information and a second piece of information, and a comparison result between a first embedding vector corresponding to the second piece of information and second embedding vectors respectively corresponding to the multiple nodes; Based on the information about the command and the information about the execution target of the command, identify a second node for performing an operation corresponding to the command from at least one node connected to the first node; And Control the display to display a second UI screen corresponding to the second node.

15. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by at least one processor of an electronic device, configure the electronic device to perform operations including the following: Based on user speech received while a first user interface (UI) screen including multiple objects is displayed on a display of the electronic device, obtain text information corresponding to the user speech by inputting the user speech into a speech recognition model; Obtain first information based on the text information, the first information including information about a command included in the text information and information about an execution target of the command; Obtain second information based on metadata about the first UI screen, the second information including information about functions corresponding to the multiple objects and information about text included in the multiple objects; Based on a comparison result between the first information and the second information, identify whether a target object corresponding to the user speech exists among the multiple objects; And Based on the target object being identified, control the display to display a second UI screen corresponding to the target object for performing an operation corresponding to the command.