Electronic device for controlling object contained in screen on basis of user voice, and control method therefor

The electronic device enhances voice-based control by pre-mapping high-resolution images and texts to improve accuracy and speed in controlling screen objects, addressing challenges in diverse form factors and streaming services.

WO2026029320A1PCT designated stage Publication Date: 2026-02-05SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/005913
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-29
Filing Date
2025-04-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in accurately and efficiently controlling applications, particularly streaming services, due to manufacturer and service provider differences, leading to delays and reduced recognition performance in voice-based control methods, especially in diverse form factors like HMD devices.

Method used

An electronic device equipped with a microphone, memory, and processor that captures user voice, identifies text and images on the screen, and controls objects based on voice commands, utilizing pre-mapped information from high-resolution images to enhance accuracy and reduce delays.

Benefits of technology

Improves processing speed and accuracy in controlling screen objects by pre-mapping high-resolution images and texts, reducing errors and delays in voice-based control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025005913_05022026_PF_FP_ABST
    Figure KR2025005913_05022026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device comprises: a microphone; a memory storing one or more instructions; and one or more processors executing the one or more instructions individually or collectively. When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device may: acquire text corresponding to a user voice when the user voice is received through the microphone while a screen containing a plurality of first images is being outputted; identify a second image corresponding to the acquired text from among information related to one or more text and images corresponding to the one or more text, wherein the information related to the one or more text and the images corresponding to the one or more text is stored in the memory; and control an object in an area of the screen corresponding to the identified second image on the basis of the user voice and a captured image corresponding to the screen.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for controlling objects included in a screen based on user voice and method for controlling the same

[0001] The present disclosure relates to an electronic device and a control method thereof, and more particularly, to an electronic device and a control method thereof for controlling an object included on a screen based on a user's voice.

[0002] Advances in multimedia technology have provided consumers with a variety of streaming services. For example, a variety of streaming services that provide content via the Internet have recently become available.

[0003] However, streaming services are often installed as applications on electronic devices. This means that the manufacturers of electronic devices and the streaming service providers are different, limiting users to controlling the application using only the input methods supported by the application.

[0004] In particular, with the recent development of electronic devices of various form factors, such as HMD (Head-mounted Display) devices, problems are arising in controlling applications related to streaming services.

[0005] The embodiments of the present disclosure are susceptible to various modifications. Accordingly, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, it should be understood that the present disclosure is not limited to the specific embodiments, but rather encompasses all modifications, equivalents, and alternatives that do not depart from the spirit and scope of the present disclosure. Furthermore, detailed descriptions of well-known functions or configurations that may unnecessarily obscure the gist of the present disclosure are omitted.

[0006] According to one embodiment of the present disclosure for achieving the above object, an electronic device includes a microphone, a memory storing one or more commands, and one or more processors for individually or collectively executing the one or more commands, and when the one or more commands are individually or collectively executed by the one or more processors, when a user voice is received through the microphone while a screen including a plurality of first images is output, the electronic device obtains a text corresponding to the user voice, identifies a second image corresponding to the obtained text among information about at least one text and an image corresponding to the at least one text, and the information about the at least one text and the image corresponding to the at least one text is stored in the memory, and controls an object in an area of ​​the screen corresponding to the identified second image based on the user voice and a captured image corresponding to the screen.

[0007] In addition, the electronic device further includes a communication interface and a display, and when the one or more commands are individually or collectively executed by the one or more processors, the electronic device controls the display to output the screen based on a plurality of second images received from a server through the communication interface, and can obtain information about each of a plurality of texts obtained from the plurality of second images and information about a second image corresponding to each of the plurality of texts among the plurality of second images.

[0008] And, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can obtain a version of the second image corresponding to each of the plurality of texts having a changed resolution as information about the second image corresponding to each of the plurality of texts.

[0009] Additionally, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can obtain a compressed version of each of the second images, each of which has its resolution adjusted, as information about the second image corresponding to each of the plurality of texts.

[0010] And, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can identify one or more texts corresponding to each of the plurality of second images among the plurality of texts, identify at least one candidate text among the one or more texts corresponding to each of the plurality of second images, and obtain information about the one or more texts corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts.

[0011] In addition, the electronic device further includes a communication interface and a display, and when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can receive the screen from the server through the communication interface, control the display to output the screen, and obtain information about at least one text obtained from the captured image and at least one first image corresponding to each of the at least one text.

[0012] And, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can control the object based on a command input method supported by an application corresponding to the screen.

[0013] Additionally, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can control the object based on a command corresponding to a touch at a point in an area corresponding to the identified second image, if the command input method includes a touch input method.

[0014] And, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can control the object based on at least one first command for moving the focus included in the captured image to an area corresponding to the identified second image and a second command for executing the object after the at least one first command, if the command input method does not include a touch input method.

[0015] Additionally, when the one or more commands are individually or collectively executed by the one or more processors, the electronic device can move the focus and identify the current position of the focus by comparing the captured image and another captured image corresponding to the screen after the focus has been moved.

[0016] Meanwhile, according to one embodiment of the present disclosure, a method for controlling an electronic device may include, when a user voice is received through a microphone of the electronic device while a screen including a plurality of first images is output, a step of obtaining a text corresponding to the user voice, a step of identifying a second image corresponding to the text corresponding to the user voice among information about at least one text stored in the electronic device and an image corresponding to the at least one text, and a step of controlling an object in an area of ​​the screen corresponding to the identified second image based on the user voice and a captured image corresponding to the screen.

[0017] In addition, the method may further include a step of outputting the screen based on a plurality of second images received from the server, and a step of obtaining information about each of the plurality of texts obtained from the plurality of second images and information about a second image corresponding to each of the plurality of texts among the plurality of second images.

[0018] And, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts among the plurality of second images may obtain a version of the second image corresponding to each of the plurality of texts having a changed resolution as information about the second image corresponding to each of the plurality of texts.

[0019] In addition, the step of obtaining information about each of the plurality of texts and a second image corresponding to each of the plurality of texts among the plurality of second images may obtain a compressed version of each of the second images with adjusted resolution as information about the second image corresponding to each of the plurality of texts.

[0020] And, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts among the plurality of second images may include identifying one or more texts corresponding to each of the plurality of second images among the plurality of texts, identifying at least one candidate text among the one or more texts corresponding to each of the plurality of second images, and obtaining information about the one or more texts corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts.

[0021] ,,,,,,,Meanwhile, according to one embodiment of the present disclosure, a non-transitory computer-readable medium has stored thereon instructions that, when executed by at least one processor, cause the at least one processor to execute a method for controlling an electronic device, the method may include: when a user voice is received through a microphone of the electronic device while a screen including a plurality of first images is output, obtaining a text corresponding to the user voice; identifying a second image corresponding to the text corresponding to the user voice among information about at least one text stored in the electronic device and an image corresponding to the at least one text; and controlling an object in an area of ​​the screen corresponding to the identified second image based on the user voice and a captured image corresponding to the screen.

[0022] In addition, with respect to a non-transitory computer-readable medium, the method may further include a step of outputting the screen based on a plurality of second images received from a server, and a step of obtaining information about each of a plurality of texts obtained from the plurality of second images and information about a second image corresponding to each of the plurality of texts among the plurality of second images.

[0023] And, with respect to the non-transitory computer-readable medium, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts among the plurality of second images may obtain a version of the second image corresponding to each of the plurality of texts having a changed resolution as information about the second image corresponding to each of the plurality of texts.

[0024] In addition, with respect to a non-transitory computer-readable medium, the step of obtaining information about each of the plurality of texts and a second image corresponding to each of the plurality of texts among the plurality of second images may obtain a compressed version of each of the second images with adjusted resolution as information about the second image corresponding to each of the plurality of texts.

[0025] And, with respect to the non-transitory computer-readable medium, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts among the plurality of second images may include identifying one or more texts corresponding to each of the plurality of second images among the plurality of texts, identifying at least one candidate text among the one or more texts corresponding to each of the plurality of second images, and obtaining information about the one or more texts corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts.

[0026] The above and other aspects and features of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0027] FIGS. 1A, 1B, and 1C are drawings to illustrate the difficulties of screen control to help understand the present disclosure.

[0028] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.

[0029] FIG. 3 is a block diagram showing a detailed configuration of an electronic device according to an embodiment of the present disclosure.

[0030] FIG. 4 is a block diagram showing the configuration of an electronic system according to an embodiment of the present disclosure.

[0031] FIG. 5 is a flowchart illustrating an operation of storing mapping information according to an embodiment of the present disclosure.

[0032] FIG. 6 is a flowchart illustrating an operation of processing an object corresponding to a user's voice based on mapping information according to an embodiment of the present disclosure.

[0033] FIG. 7 is a drawing for explaining an operation of identifying a position of focus according to one embodiment of the present disclosure.

[0034] FIG. 8 is a diagram illustrating an operation of pre-identifying information about a poster before a user's voice is received according to an embodiment of the present disclosure.

[0035] FIG. 9 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.

[0036] The purpose of the present disclosure is to provide an electronic device and a control method thereof for increasing the speed and accuracy of controlling an object included on a screen through a user's voice.

[0037] It should be understood that the various embodiments and terms used in this document are not intended to limit the technical features described in this document to specific embodiments, but rather to include various modifications, equivalents, or substitutes of the embodiments.

[0038] In connection with the description of the drawings, similar reference numerals may be used for similar or related components.

[0039] The singular form of a noun corresponding to an item may include one or more items, unless the context clearly indicates otherwise.

[0040] In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in that phrase, or all possible combinations thereof.

[0041] Terms such as "first," "second," or "first" or "second" may be used simply to distinguish one component from another and do not qualify the components in any other respect (e.g., importance or order).

[0042] When a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component), with or without the terms “functionally” or “communicatively,” it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0043] The terms “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in this document, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0044] When a component is said to be “connected,” “coupled,” “supported,” or “in contact with” another component, this includes not only cases where the components are directly connected, coupled, supported, or in contact, but also cases where the components are indirectly connected, coupled, supported, or in contact through a third component.

[0045] When we say that a component is “on” another component, this includes not only cases where the component is in contact with the other component, but also cases where there is another component between the two components.

[0046] The term “and / or” includes any combination of a plurality of related described elements or any one of a plurality of related described elements.

[0047] The operating principle and embodiments of the present invention will be described with reference to the attached drawings below.

[0048] In connection with any method or process described herein, identification codes may be used for convenience of explanation, but are not intended to describe the order of each step or operation. Each step or operation may be implemented in a different order than the illustrated order, unless the context clearly dictates otherwise. Unless the context clearly dictates otherwise, one or more steps or operations may be omitted.

[0049] The various operations, actions, blocks, steps, etc. of the flowchart may be performed in the order presented, in a different order, or simultaneously. Furthermore, in one or more embodiments, some of the operations, actions, blocks, steps, etc. may be omitted, added, modified, skipped, etc. without departing from the scope of the present disclosure.

[0050] FIGS. 1A, 1B, and 1C are drawings to illustrate the difficulties of screen control to help understand the present disclosure.

[0051] With the recent development of electronic devices in various form factors, control methods are being provided using gesture recognition, eye tracking, or AI assistants based on voice recognition. In other words, control methods that do not utilize conventional remote controls are being developed. Regarding voice recognition, linguistic understanding refers to the technology for recognizing, applying, and processing human language and text, encompassing natural language processing, machine translation, dialogue systems, question-answering, and voice recognition / synthesis.

[0052] However, as illustrated in Figures 1a and 1b, it is challenging for electronic devices to accurately recognize user gestures from a distance of 3 to 4 meters and control detailed UIs based on these gestures. For example, various issues can arise, such as reduced recognition performance at long distances or in low light, lack of comfortable control movements, and gesture recognition for multiple people. Furthermore, gaze tracking poses challenges even greater than gesture recognition, and thus, methods utilizing voice recognition are actively being researched.

[0053] Voice has the advantage of being relatively unrestricted in environments such as long-distance and low-light environments, boasting high recognition accuracy and no linguistic limitations on the information it can express. However, its ability to express spatial (visual) information is severely limited to directional information (e.g., up / down / left / right).

[0054] Most user interfaces on electronic devices are optimized for four-way remote controls, necessitating voice optimization. However, most streaming services are provided by third parties, making UI modifications difficult. Furthermore, existing voice-based methods suffer from delays and low accuracy.

[0055] For example, if the manufacturer of the electronic device and the company providing the streaming service are different, screen capture and text recognition are performed in the process of voice controlling the application providing the streaming service, as shown in Fig. 1c, and this process may result in reduced accuracy and delay.

[0056] For example, when a user's voice is received, an electronic device can capture an image of the screen, perform OCR (Optical Character Reader) on the captured image to identify text included in the screen, identify the location of the text corresponding to the user's voice, and control an object corresponding to the text corresponding to the user's voice. However, since this process is performed after the user's voice is received, a delay may occur. In addition, if a screen such as Fig. 1c has a resolution of, for example, 1280×720, each of the multiple images included in the screen has a resolution of only approximately 180×60, and errors may occur in the process of identifying text from each of the multiple images. Therefore, there is a need to reduce delay time and errors.

[0057] FIG. 2 is a block diagram showing the configuration of an electronic device (100) according to one embodiment of the present disclosure.

[0058] The electronic device (100) may be a device that controls the electronic device (100) based on a user's voice. In particular, the electronic device (100) may be a device that controls an object displayed on a screen based on a user's voice, and may be a device equipped with a display, such as a TV, a desktop PC, a laptop, a smartphone, a tablet PC, smart glasses, a smart watch, an HMD, etc., and that controls an object displayed on the screen based on the user's voice. Here, the screen may be a screen of a 3rd party application.

[0059] However, the present invention is not limited thereto, and the electronic device (100) may be a device that provides screen information to an external display device and controls objects included in a screen displayed on the external display device based on a user's voice. For example, the electronic device (100) may be a computer body, a set-top box (STB), or the like.

[0060] According to FIG. 2, the electronic device (100) includes a microphone (110), a memory (120), and a processor (130).

[0061] The microphone (110) is configured to receive sound and convert it into an audio signal. The microphone (110) is electrically connected to the processor (130) and can receive sound under the control of the processor (130).

[0062] For example, a microphone (110) can receive a user's voice in analog form, digitize it, and provide the digitized signal to a processor (130).

[0063] The microphone (110) may be formed as an integrated unit integrated into the upper side, front side, side direction, etc. of the electronic device (100). Alternatively, the microphone (110) may be provided in a remote control, etc., separate from the electronic device (100). In this case, the remote control may receive sound through the microphone (110) and provide the received sound to the electronic device (100).

[0064] The microphone (110) may include various configurations such as a microphone that collects analog sound, an amplifier circuit that amplifies the collected sound, an A / D conversion circuit that samples the amplified sound and converts it into a digital signal, and a filter circuit that removes noise components from the converted digital signal.

[0065] The microphone (110) may be implemented in the form of a sound sensor, and any method that can collect sound may be used.

[0066] Memory (120) may refer to hardware that stores information such as data in an electrical or magnetic form so that a processor (130) or the like can access it. To this end, memory (120) may be implemented as at least one piece of hardware from among non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), solid state drive (SSD), RAM, ROM, etc.

[0067] The memory (120) may store at least one instruction required for the operation of the electronic device (100) or the processor (130). Here, the instruction is a unit of code that instructs the operation of the electronic device (100) or the processor (130), and may be written in machine language, which is a language that a computer can understand. Alternatively, the memory (120) may store a plurality of instructions for performing a specific task of the electronic device (100) or the processor (130) as an instruction set.

[0068] The memory (120) may store data, which is information in bit or byte units that can represent characters, numbers, images, etc. For example, the memory (120) may store mapping information. Here, the mapping information may be information in which at least one text and information about an image corresponding to at least one text are mapped.

[0069] The memory (120) is accessed by the processor (130), and reading / writing / modifying / deleting / updating instructions, instruction sets, or data can be performed by the processor (130).

[0070] The processor (130) controls the overall operation of the electronic device (100). Specifically, the processor (130) is connected to each component of the electronic device (100) and can control the overall operation of the electronic device (100). For example, the processor (130) is connected to components such as a microphone (110), a memory (120), a communication interface (e.g., see FIG. 3), and can control the operation of the electronic device (100).

[0071] The processor (130) may be one or more processors, and may include one or more of a CPU, a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit), a MIC (Many Integrated Core), an NPU (Neural Processing Unit), a hardware accelerator, or a machine learning accelerator. The one or more processors (130) may control one or any combination of other components of the electronic device (100), and may perform operations related to communication or data processing. The one or more processors (130) may individually or collectively execute one or more programs or instructions stored in the memory (120). For example, the one or more processors (130) may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in the memory (120).

[0072] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-specific processor).

[0073] One or more processors (130) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When one or more processors (130) are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include an internal processor memory, such as a cache memory or an on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.

[0074] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.

[0075] In embodiments of the present disclosure, one or more processors (130) may refer to a system on a chip (SoC) in which at least one processor and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto. However, for convenience of explanation, the operation of the electronic device (100) is described below using the expression processor (130).

[0076] The processor (130) can output a screen including a plurality of first images through the electronic device (100) or an external display device. When the processor (130) outputs the screen through the external display device, it can provide screen information to the external display device. For convenience of explanation, the electronic device (100) is described below as including a display and the screen is output through the display.

[0077] When a user's voice is received through a microphone (110) while a screen including a plurality of first images is output, the processor (130) can obtain text corresponding to the user's voice.

[0078] The processor (130) can identify a second image corresponding to the text corresponding to the user's voice among information about at least one text and an image corresponding to the at least one text stored in the memory (120). The memory (120) can map information about at least one text and an image corresponding to the at least one text and store the information as mapping information.

[0079] For example, the mapping information may be information in which at least one text obtained from a screen being output before the user's voice is received and information about an image corresponding to at least one text are mapped.

[0080] For example, when an application is executed, the processor (130) may receive a plurality of second images from a server corresponding to the executed application, and output a screen including a plurality of first images based on each of the plurality of second images. For example, each of the plurality of second images may have a resolution of 1280×720, and each of the plurality of first images may have a resolution of 180×60.

[0081] The processor (130) may obtain a plurality of texts from each of the plurality of second images, and obtain information about each of the plurality of texts and each of the second images corresponding to each of the plurality of texts. The processor (130) may store the obtained information in the memory (120). For example, the processor (130) may identify Title 1 representing Second Image 1 among the plurality of second images, map information about Title 1 and Second Image 1 to obtain mapping information, and store the obtained information in the memory (120). In this way, the processor (130) may obtain mapping information for all of the plurality of second images, and store the obtained information in the memory (120). That is, since the processor (130) obtains mapping information based on a plurality of second images that are originals of the plurality of first images and have a high resolution, rather than a plurality of first images included in the screen, text identification performance may be improved.

[0082] However, it is not limited thereto, and the mapping information may further include information obtained from the previous screen of the screen being output.

[0083] The processor (130) may change the resolution of a plurality of second images based on the resolution of each of the plurality of first images, obtain each of the plurality of second images with changed resolution as information about the second image corresponding to each of the plurality of texts, and store the obtained information in the memory (120). That is, during the text identification process, the second image with the original resolution is used, but after the text identification is completed, the second image with a lowered resolution may be stored in the memory (120) to save storage space.

[0084] However, it is not limited thereto, and the processor (130) may acquire mapping information by mapping information for each of the plurality of texts and the second image corresponding to each of the plurality of texts without changing the resolution, and store the acquired information in the memory (120).

[0085] The processor (130) obtains a compressed version of each of the plurality of second images with adjusted resolution as information about each of the plurality of second images corresponding to each of the plurality of texts, and the compressed version is compressed based on an encoder included in the electronic device (100), the resolution is adjusted, and the obtained information may be stored in the memory (120). That is, storage space can be further saved through downscaling and compression.

[0086] The processor (130) may identify text corresponding to each of the plurality of second images and at least one candidate text among the plurality of texts from each of the plurality of second images, obtain information about the text corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts, and store the obtained information in the memory (120).

[0087] In the above-described example, the processor (130) may further identify candidate title 1 as well as title 1 representing the second image 1 among the plurality of second images. The processor (130) may map title 1 and candidate title 1 to information about the second image 1 to obtain mapping information, and store the obtained information in the memory (120). Since errors may occur in the text identification process, the candidate text may be further stored and used in the subsequent process of identifying a location on the screen corresponding to the user's voice.

[0088] In the above, it is assumed that the processor (130) can access images related to a third-party application. However, this is not limited to the case, and the processor (130) may not be able to access images related to a third-party application.

[0089] For example, when an application is executed, the processor (130) may receive a screen from a server corresponding to the executed application and output the received screen. In this case, the processor (130) may obtain information about at least one text and a first image corresponding to each of the at least one text from a captured image corresponding to the screen, and store the obtained information in the memory (120).

[0090] When the second image is identified, the processor (130) can control an object included in an area corresponding to the second image identified in the captured image of the screen based on the user's voice.

[0091] The processor (130) can control an object based on a command input method supported by an application corresponding to the screen.

[0092] For example, if the command input method includes a touch input method, the processor (130) can control the object based on a command to touch a point in an area corresponding to the identified second image.

[0093] Alternatively, if the command input method does not include a touch input method, the processor (130) may control the object based on at least one first command for identifying a focus in a captured image and moving the focus to an area corresponding to the identified second image, and a second command for executing the object after the first command. Here, the processor (130) may move the focus and identify the current position of the focus by comparing the captured image and another captured image corresponding to the screen after the focus has been moved. However, the present invention is not limited thereto, and the processor (130) may also identify the position of the focus by analyzing the captured image itself.

[0094] FIG. 3 is a block diagram illustrating a detailed configuration of an electronic device (100) according to an embodiment of the present disclosure. The electronic device (100) may include a microphone (110), a memory (120), and a processor (130). In addition, according to FIG. 3, the electronic device (100) may further include a communication interface (140), a display (150), a user interface (160), a camera (170), and a speaker (180). For components illustrated in FIG. 3 that overlap with those illustrated in FIG. 2, a detailed description thereof will be omitted.

[0095] The communication interface (140) is a configuration that performs communication with various types of external devices according to various types of communication methods. For example, the electronic device (100) can perform communication with a server through the communication interface (140).

[0096] The communication interface (140) may include a Wi-Fi module, a Bluetooth module, an infrared communication module, and a wireless communication module. Here, each communication module may be implemented in the form of at least one hardware chip.

[0097] Wi-Fi modules and Bluetooth modules communicate via Wi-Fi and Bluetooth, respectively. When using a Wi-Fi or Bluetooth module, various connection information, such as the SSID and session key, is first transmitted and received. This information is then used to establish a communication connection before various other information can be transmitted and received. Infrared communication modules communicate using infrared data association (IrDA) technology, which uses infrared, a wavelength between visible light and millimeter waves, to wirelessly transmit data over short distances.

[0098] In addition to the above-described communication method, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards such as zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.

[0099] Alternatively, the communication interface (140) may include a wired communication interface such as HDMI, DP, Thunderbolt, USB, RGB, D-SUB, DVI, etc.

[0100] In addition, the communication interface (140) may include at least one of a LAN (Local Area Network) module, an Ethernet module, or a wired communication module that performs communication using a pair cable, a coaxial cable, or an optical fiber cable.

[0101] The display (150) is a configuration that outputs an image and can be implemented as a display of various forms such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, a PDP (Plasma Display Panel), etc. The display (150) may also include a driving circuit, a backlight unit, etc. that can be implemented as a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. The display (150) may be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, etc.

[0102] The user interface (160) may be implemented with buttons, a touch pad, a mouse, a keyboard, etc., or may be implemented with a touch screen capable of performing both display and operation input functions. Here, the buttons may be various types of buttons, such as mechanical buttons, touch pads, wheels, etc., formed on any area of ​​the front, side, or back of the main body of the electronic device (100).

[0103] The camera (170) is configured to capture still images or moving images. The camera (170) can capture still images at a specific point in time, but can also capture still images continuously.

[0104] The camera (170) can capture the front of the electronic device (100) to capture the actual environment in front of the electronic device (100). The processor (130) can also identify an area of ​​interest from an image captured by the camera (170).

[0105] The camera (170) includes a lens, a shutter, an aperture, a solid-state image sensor, an AFE (Analog Front End), and a TG (Timing Generator). The shutter controls the time at which light reflected from a subject enters the camera (170), and the aperture mechanically increases or decreases the size of the opening through which light enters to control the amount of light incident on the lens. When the solid-state image sensor accumulates light reflected from a subject as a photocharge, the image generated by the photocharge is output as an electrical signal. The TG outputs a timing signal for reading out pixel data of the solid-state image sensor, and the AFE samples and digitizes the electrical signal output from the solid-state image sensor.

[0106] The speaker (180) is a component that outputs various audio data processed by the processor (130) as well as various notification sounds and voice messages.

[0107] As described above, since the electronic device (100) stores mapping information in advance from the screen, the processing speed according to the user's voice can be improved. Furthermore, since the electronic device (100) acquires mapping information based on the original images of the images included on the screen, the accuracy can be improved, thereby improving the processing performance according to the user's voice.

[0108] In the above description, the application is described as being provided by a third party, but this is not limited thereto. For example, the application may be provided by the manufacturer of the electronic device (100). However, even if the application is provided by the manufacturer of the electronic device (100), if command input via voice recognition is not possible, the above disclosure may be applied. In addition, the electronic device (100) is a display device such as a TV, and the above disclosure may be applied to control a screen received from an external device such as an STB.

[0109] In addition, although the above is described as a hardware operation of the electronic device (100), the above disclosure may also be implemented in software. For example, the electronic device (100) may execute a voice control application in the background. In addition, when a user voice is received while an application different from the voice control application is executed, the voice control application may obtain a text corresponding to the user voice, identify a second image corresponding to the text among information about at least one text and an image corresponding to the at least one text stored in the electronic device (100), and control an object included in an area corresponding to the identified second image in a captured image corresponding to the screen based on the user voice. For such operations, the voice control application may be granted more authority than a general application. For example, the voice control application may be excluded from memory refresh, etc. even when executed in the background, and when a plurality of second images are received from the server (200) according to the execution of the application, the plurality of second images may be accessed.

[0110] The above has described the operation of the electronic device (100) storing information about at least one text and an image corresponding to at least one text as mapping information, but the present invention is not limited thereto. For example, when the electronic device (100) receives a plurality of second images from the server (200), the electronic device (100) may obtain a plurality of fingerprints from each of the plurality of second images, map information about each of the plurality of fingerprints and the second images corresponding to each of the plurality of fingerprints, and store the mapped information as mapping information. That is, the electronic device (100) may identify an image using a fingerprint rather than text corresponding to the image. In addition, any method that can identify and specify an image, as well as text and fingerprints, may be applied. For example, the electronic device (100) may obtain identification information about an image using a neural network model that generates identification information from an image, map the identification information to an image, and store the mapped information as mapping information.

[0111] Although the electronic device (100) has been described above as obtaining text corresponding to the user's voice from the user's voice, this is not a limitation. For example, when the electronic device (100) receives a user's voice, it may transmit the user's voice to a server and receive text corresponding to the user's voice from the server. Here, the server may include an STT (Speech to Text) server.

[0112] The functions related to artificial intelligence according to the present disclosure can be operated through a processor (130) and a memory (120).

[0113] The processor (130) may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, AP, DSP, etc., a graphics-only processor such as a GPU or VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU.

[0114] One or more processors are controlled to process input data according to predefined operating rules or artificial intelligence models stored in the memory (120). Alternatively, if one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model. The predefined operating rules or artificial intelligence models are characterized by being created through learning.

[0115] Here, "created through learning" means that a basic artificial intelligence model is learned using a learning algorithm using a plurality of learning data, thereby creating a predefined set of operating rules or an artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0116] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations by calculating the results of previous layers and the multiple weights. The multiple weights of the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated during the learning process to reduce or minimize the loss or cost values ​​obtained by the artificial intelligence model.

[0117] Artificial neural networks may include deep neural networks (DNNs), such as, but not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), or deep Q-networks.

[0118] Hereinafter, the operation of the electronic device (100) will be described in more detail with reference to FIGS. 4 to 8. For convenience of explanation, individual embodiments are described in FIGS. 4 to 8. However, the individual embodiments of FIGS. 4 to 8 may be implemented in any combination.

[0119] FIG. 4 is a block diagram illustrating the configuration of an electronic system (1000) according to an embodiment of the present disclosure. As illustrated in FIG. 4, the electronic system (1000) may include an electronic device (100) and a server (200).

[0120] When an application is executed, the electronic device (100) can receive a plurality of second images from a server (200) corresponding to the executed application, and output a screen including a plurality of first images based on each of the plurality of second images.

[0121] Alternatively, when an application is executed, the electronic device (100) may receive a screen from a server (200) corresponding to the executed application and output the screen.

[0122] The server (200) may be a device that stores data related to an application. For example, the server (200) may be a device that stores multiple contents provided by the application and multiple thumbnails (or images such as posters) corresponding to each of the multiple contents, and may be a desktop PC, laptop, smartphone, tablet PC, etc. However, the server (200) is not limited thereto, and any device capable of storing data related to the application may be used.

[0123] When an application is executed on an electronic device (100), the server (200) receives an image request signal from the electronic device (100) according to the execution of the application, and can provide a plurality of second images to the electronic device (100) or provide screen information corresponding to the application to the electronic device (100).

[0124] However, the present invention is not limited thereto, and the number of servers (200) may be plural for each application. In addition, the server (200) may be a device that reviews texts received from the electronic device (100). For example, the server (200) may provide a plurality of second images to the electronic device (100) at the request of the electronic device (100). The electronic device (100) may identify a plurality of texts from each of the plurality of second images and provide the plurality of texts to the server (200). The server (200) may correct the plurality of texts based on the stored data and provide the corrected plurality of texts to the electronic device (100). The electronic device (100) may map information about each of the plurality of corrected texts and the second images corresponding to each of the plurality of corrected texts to obtain mapping information, and store the obtained information in the memory (120). For example, the server (200) may obtain the second image 1 and the title AAA of the second image 1, and store the obtained information in the memory (120). At this time, if AAA' is received from the electronic device (100) as one of multiple texts, the server (200) can correct AAA' to AAA and provide it to the electronic device (100). Through this operation, errors in the text identification process can be reduced.

[0125] FIG. 5 is a flowchart illustrating an operation for storing mapping information according to one embodiment of the present disclosure. Since an actual application includes multiple posters (images such as thumbnails) on the screen, the terms "poster" and "image" are used interchangeably in FIG. 5 .

[0126] First, when an application is executed (S510), the processor (130) can identify whether the executed application is a controllable application without a remote control (S520).

[0127] The processor (130) can terminate the running application if it is an application that cannot be controlled without a remote control, and if it is an application that can be controlled without a remote control, it can identify whether to download individual posters as separate images (S530). For example, when the application is running, the processor (130) can receive multiple second images from a server corresponding to the application. Here, the second images may be posters.

[0128] The processor (130) may extract posters from captured images when individual posters are not downloaded as separate images (S560). For example, when the processor (130) receives pixel information about a screen itself including multiple first images from a server, the processor (130) may output the screen and extract posters from captured images corresponding to the screen. Alternatively, when individual posters are downloaded as separate images but are inaccessible, the processor (130) may output a screen including individual posters and extract posters from captured images corresponding to the screen.

[0129] The processor (130) can identify text in a poster (S570-1), map the text and the poster to obtain mapping information, and store the obtained information (S570-2).

[0130] Alternatively, the processor (130) may identify whether the downloaded image is accessible when downloading individual posters as separate images (S540), and if the downloaded image is accessible, determine whether the downloaded image is a new image (S550-1), and if a new image is found (S550-2), perform operation S570. Here, the operation of determining whether the image is new may be determined by checking whether an additional image exists compared to existing images.

[0131] However, the present invention is not limited thereto, and the processor (130) may extract a poster from the captured image at preset time intervals. Alternatively, the processor (130) may extract a poster from the captured image whenever the screen provided by the application is switched.

[0132] Although FIG. 5 describes that only posters and text are mapped and stored, the present invention is not limited thereto. For example, the processor (130) may identify the location of the poster before the user's voice is received. For example, the processor (130) may identify the location of each poster in the captured image after the operation of S570-2 is completed. In this case, when the user's voice is received, the processor (130) may obtain the text corresponding to the user's voice, identify the poster mapped to the text corresponding to the user's voice from the mapping information, and obtain the location of the identified poster from the stored information. In this operation, since the operation of comparing the captured image and the poster to identify the location of the poster is performed before, rather than after, the user's voice is received, the processing time after the user's voice is received can be shortened. This effect is described again in FIG. 6.

[0133] FIG. 6 is a flowchart illustrating an operation of processing an object corresponding to a user's voice based on mapping information according to an embodiment of the present disclosure.

[0134] First, the processor (130) receives a user voice (S610) and can identify whether the executed application is a voice-controllable application (S620).

[0135] The processor (130) can terminate the executed application if it is an application that cannot be controlled by voice, and can identify whether mapping information exists if the executed application is an application that can be controlled by voice (S630).

[0136] The processor (130) terminates if no mapping information is stored in the memory (120), and if mapping information is stored in the memory (120), searches for a text corresponding to the user's voice (S640) and can identify an image corresponding to the searched text (S650). For example, if mapping information such as (AAA, A image), (ABA, B image), and (CCC, C image) is stored in the memory (120), and if the text corresponding to the user's voice is ABA, the processor (130) can finally identify the B image based on (ABA, B image) among the mapping information.

[0137] The processor (130) can identify the location of the poster corresponding to the image identified in the captured image (S660). In the above-described example, the processor (130) can identify the location of the poster corresponding to the B image in the captured image.

[0138] The processor (130) can identify whether the executed application is an application capable of processing touch input (S670), and if it is capable of processing touch input, it can generate a virtual touch event at the poster location (S680), and if it is not capable of processing touch input, it can generate a remote control key sequence (S690).

[0139] The processor (130) may also identify the location of the poster before the user's voice is received, as described in FIG. 5. In this case, the processor (130) may identify location information corresponding to the searched text instead of performing operations S650 and S660. For example, the memory (120) may store mapping information such as (AAA, A image, output o, (x1, y1)), (ABA, B image, output o, (x2, y2)), (CCC, C image, output x, ( , )). Here, the mapping information may further store whether or not to output and the output coordinate values. In the case of the C image, it may not be output and thus may not have a coordinate value. If the text corresponding to the user's voice is ABA, the processor (130) may finally identify the output location (x2, y2) of the B image based on (ABA, B image, output o, (x2, y2)) among the mapping information. For convenience of explanation, the output location is indicated by coordinate values ​​such as (x2, y2) above, but it is not limited to this. For example, the coordinate values ​​may include values ​​for indicating an area such as (x2~x3, y2~y3).

[0140] The processor (130) can update mapping information as the page of the screen changes. For example, when images A and B are displayed on the first page of the screen, mapping information such as (AAA, image A, output o, (p1, x1, y1)), (ABA, image B, output o, (p1, x2, y2)), (CCC, image C, output x, ( , , )) may be stored in the memory (120). Thereafter, when the screen changes to the second page according to a user's operation and image C is displayed, the processor (130) can update mapping information such as (AAA, image A, output o, (p1, x1, y1)), (ABA, image B, output o, (p1, x2, y2)), (CCC, image C, output x, (p2, x3, y3)). Here, when a user voice such as “Run AAA of the previous page” is received, the processor (130) may identify the location of the A image based on the previous page and AAA.

[0141] FIG. 7 is a drawing for explaining an operation of identifying a position of focus according to one embodiment of the present disclosure.

[0142] The processor (130) can generate a virtual touch event at the poster location if the running application can process touch input, and can generate a remote control key sequence if the application cannot process touch input. In order to generate a remote control key sequence, the current focus position must first be identified.

[0143] For example, as illustrated in FIG. 7, the processor (130) can move the focus (S710) of the first position and identify the current position of the focus by comparing the captured image and another captured image corresponding to the screen including the focus (S720) of the second position.

[0144] However, this is not limited to this, and the method of moving the focus position may vary. In addition, the processor (130) may identify the focus position through image analysis without moving the focus position.

[0145] FIG. 8 is a diagram illustrating an operation of pre-identifying information about a poster before a user's voice is received according to an embodiment of the present disclosure.

[0146] The processor (130) can extract posters from the captured image (S810). For example, the processor (130) can extract posters from the captured image of the currently displayed screen even before the user's voice is received, based on at least one of the hardware performance or resource status of the electronic device (100). For example, the processor (130) can identify the location of each poster in the captured image.

[0147] The processor (130) can identify text in a poster (S820-1), map the text and the poster to obtain mapping information, and store the obtained information in the memory (120) (S820-2).

[0148] When a user's voice is received (S830), the processor (130) can identify the location of the poster corresponding to the user's voice based on the mapping information (S840).

[0149] As described in FIGS. 5 and 6, compared to the first embodiment in which poster information of the currently displayed screen is acquired after the user's voice is received, in the case of the second embodiment in which poster information of the currently displayed screen is acquired in advance before the user's voice is received, as shown in FIG. 8, the operation of comparing the poster included in the mapping information with the captured image can be omitted, so that the processing time after receiving the user's voice can be shortened.

[0150] FIG. 9 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.

[0151] First, when a user's voice is received while a screen including a plurality of first images is output, a text corresponding to the user's voice is acquired (S910). Then, a second image corresponding to the text corresponding to the user's voice is identified among information about at least one text and an image corresponding to at least one text stored in the electronic device (S920). Then, an object included in an area corresponding to the second image identified in the captured image corresponding to the screen based on the user's voice is controlled (S930).

[0152] In addition, the method may further include a step of outputting a screen including a plurality of first images based on a plurality of second images received from a server, and a step of obtaining information about a plurality of texts obtained from the plurality of second images and a second image corresponding to each of the plurality of texts.

[0153] And, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts may obtain each of the plurality of second images whose resolution has been changed based on the resolution of each of the plurality of first images as information about the second image corresponding to each of the plurality of texts.

[0154] In addition, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts may obtain each of the plurality of second images whose resolutions have been adjusted based on an encoder included in the electronic device as information about the second image corresponding to each of the plurality of texts.

[0155] And, the step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of texts may include identifying the text corresponding to each of the plurality of second images and at least one candidate text among the plurality of texts from each of the plurality of second images, and obtaining information about the text corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts.

[0156] In addition, the method may further include a step of outputting a screen received from a server and a step of obtaining information about at least one text obtained from a captured image and a first image corresponding to each of the at least one text.

[0157] And, the controlling step (S930) can control the object based on the command input method supported by the application corresponding to the screen.

[0158] In addition, if the control step (S930) includes a touch input method as the command input method, the object can be controlled based on a command to touch a point in an area corresponding to the identified second image.

[0159] And, in the controlling step (S930), if the command input method does not include a touch input method, the object can be controlled based on at least one first command for moving the focus included in the captured image to an area corresponding to the identified second image and a second command for executing the object after the first command.

[0160] Additionally, the step of identifying the focus can identify the current position of the focus by moving the focus and comparing the captured image and another captured image corresponding to the screen after the focus has been moved.

[0161] According to one or more embodiments of the present disclosure, the electronic device may improve processing speed according to user voice input because it stores mapping information in advance from the screen. Furthermore, because the electronic device acquires mapping information based on the original images included in the screen, accuracy may be improved, thereby improving processing performance according to user voice input.

[0162] According to an exemplary embodiment of the present disclosure, the various embodiments described above may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments, which is a device that can call instructions stored in the storage medium and operate according to the called instructions. When an instruction is executed by a processor, the processor may directly or under the control of the processor perform a function corresponding to the instruction using other components. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.

[0163] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0164] Furthermore, according to one embodiment of the present disclosure, the various embodiments described above may be implemented in a computer-readable recording medium or a similar device using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented by the processor itself. In a software implementation, embodiments such as the procedures and functions described herein may be implemented as separate software modules. Each of the software modules may perform one or more functions and operations described herein.

[0165] Computer instructions for performing processing operations of a device according to the various embodiments described above may be stored in a non-transitory computer-readable medium. The computer instructions stored in such a non-transitory computer-readable medium, when executed by a processor of a specific device, cause the specific device to perform processing operations of the device according to the various embodiments described above. A non-transitory computer-readable medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media may include a CD, DVD, hard disk, Blu-ray disk, USB, memory card, or ROM.

[0166] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0167] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person skilled in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

In electronic devices, mike; memory for storing one or more instructions; and one or more processors that individually or collectively execute one or more of the above instructions; , when the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, When a user's voice is received through the microphone while a screen including multiple first images is output, a text corresponding to the user's voice is obtained, Identifying a second image corresponding to the acquired text among information about at least one text and an image corresponding to the at least one text, and storing the information about the at least one text and the image corresponding to the at least one text in the memory, An electronic device that controls an object in an area of ​​the screen corresponding to the identified second image based on the user voice and the captured image corresponding to the screen. In the first paragraph, communication interface; and including display; When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, Controlling the display to output the screen based on a plurality of second images received from the server through the communication interface; An electronic device that obtains information about each of a plurality of texts obtained from the plurality of second images and information about a second image corresponding to each of the plurality of texts among the plurality of second images. In the second paragraph, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, An electronic device that obtains a version of a second image corresponding to each of the plurality of texts having a changed resolution as information about the second image corresponding to each of the plurality of texts. In the third paragraph, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, An electronic device that obtains a compressed version of each of the second images, each of which has its resolution adjusted, as information about the second image corresponding to each of the plurality of texts. In the second paragraph, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, Identifying at least one text corresponding to each of the plurality of second images among the plurality of texts, Identifying at least one candidate text among the one or more texts corresponding to each of the plurality of second images, An electronic device that obtains information about one or more texts corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts. In the first paragraph, communication interface; and including display; When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, Receive the above screen from the server through the above communication interface, Control the display to output the above screen, An electronic device that obtains information about at least one text obtained from the captured image and at least one first image corresponding to each of the at least one text. In the first paragraph, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, An electronic device that controls the object based on a command input method supported by an application corresponding to the above screen. In paragraph 7, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, An electronic device that controls the object based on a command corresponding to a touch at a point in an area corresponding to the identified second image, if the command input method includes a touch input method. In paragraph 7, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, An electronic device that controls the object based on at least one first command for moving the focus included in the captured image to an area corresponding to the identified second image and a second command for executing the object after the at least one first command, if the command input method does not include a touch input method. In paragraph 9, When the one or more instructions are individually or collectively executed by the one or more processors, the electronic device, Move the focus above, An electronic device that identifies the current position of the focus by comparing the captured image and another captured image corresponding to the screen after the focus has moved. In a method for controlling an electronic device, A step of obtaining text corresponding to a user voice when a user voice is received through a microphone of the electronic device while a screen including a plurality of first images is output; A step of identifying a second image corresponding to the text corresponding to the user voice among information about at least one text stored in the electronic device and an image corresponding to the at least one text; and A control method comprising: a step of controlling an object in an area of ​​the screen corresponding to the identified second image based on the user voice and the captured image corresponding to the screen. In Article 11, A step of outputting the screen based on a plurality of second images received from the server; and A control method further comprising: a step of obtaining information about each of a plurality of texts obtained from the plurality of second images and information about a second image corresponding to each of the plurality of texts among the plurality of second images. In paragraph 12, The step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of second images is as follows: A control method for obtaining a version of a second image corresponding to each of the plurality of texts having a changed resolution as information about the second image corresponding to each of the plurality of texts. In Article 13, The step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of second images is as follows: A control method for obtaining a compressed version of each second image with an adjusted resolution as information about the second image corresponding to each of the plurality of texts. In paragraph 12, The step of obtaining information about each of the plurality of texts and the second image corresponding to each of the plurality of second images is as follows: Identifying at least one text corresponding to each of the plurality of second images among the plurality of texts, Identifying at least one candidate text among the one or more texts corresponding to each of the plurality of second images, A control method for obtaining information about one or more texts corresponding to each of the plurality of second images, the at least one candidate text, and the second image corresponding to each of the plurality of texts.

Citation Information

Patent Citations

  • Mobile terminal and text correction method

    KR1020090123697A

  • Apparatus and Method for executing object using voice command

    KR1020140114519A

  • Watch type terminal

    KR1020160097913A

  • Secondary attery including welded parts and method of manufacturing the battery

    KR1020260022836A

  • KR20220013732A