Electronic device for providing analysis result of image, operation method thereof, and storage medium
The system in wearable devices dynamically selects cameras and processes user voice inputs to enhance image analysis and translation tasks, addressing inefficiencies in existing AI models by adapting to motion and environmental changes.
Patent Information
- Application Number
- PCT/KR2025/009342
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-02
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Existing AI models lack efficient methods for image analysis and translation in wearable devices, particularly when capturing images while in motion and integrating user voice inputs.
A system comprising a wearable electronic device with multiple cameras and processors that dynamically selects the appropriate camera based on image analysis and user voice inputs to perform tasks effectively, leveraging AI models for image processing and translation.
Enhances the accuracy and efficiency of image analysis and translation tasks in wearable devices by adaptively selecting cameras and processing user voice inputs, improving performance in dynamic environments.
Smart Images

Figure KR2025009342_08012026_PF_FP_ABST
Abstract
Description
Electronic device for providing image analysis results, method of operation thereof, and storage medium
[0001] The present disclosure relates to an electronic device for providing an analysis result of an image, an operating method thereof, and a storage medium.
[0002] Recently, the performance of artificial intelligence (AI) models has been advancing. For example, AI models can be trained to provide image analysis results. For example, an AI model can provide an image reading function as a result of the analysis. This image reading function can be a function that outputs speech for text (or Braille) contained in the image. For example, an AI model can provide a translation function as a result of the analysis. This translation function can provide a translation (or a corresponding speech) of text (or Braille) contained in the image into another language. For example, an AI model can provide a summary function as a result of the analysis. This summary function can provide a summary of the text (or Braille) contained in the image. There are no restrictions on the type of image analysis.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] An electronic device may include a communication device. The electronic device may include one or more processors.
[0005] An electronic device may include a memory that stores instructions.
[0006] The above instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to establish a communication connection between the wearable electronic device and the electronic device via the communication device.
[0007] The wearable electronic device includes a first camera and a second camera, and a focal length corresponding to a maximum resolution of the first camera may be smaller than a focal length corresponding to a maximum resolution of the second camera.
[0008] The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to: identify a plurality of first images acquired using the first camera while the wearable electronic device is moving based on the first camera being selected based on the analysis results of images captured by the first camera and / or the second camera, and / or the analysis results of the user's voice; and perform a first task identified based on the analysis results of the plurality of first images and the user's voice.
[0009] The instructions, when individually or collectively executed by the one or more processors, may cause the electronic device to: identify a second image acquired using the second camera based on the selection of the second camera based on the analysis results of the images captured by the first camera and / or the second camera, and / or the analysis results of the user's voice; and perform a second task identified based on the analysis results of the second image and the user's voice.
[0010] A method of operating an electronic device may include establishing a communication connection between a wearable electronic device and the electronic device.
[0011] The method of operating the electronic device may include an operation of checking a plurality of first images acquired using the first camera while the wearable electronic device is moving based on the analysis results of images captured by the first camera of the wearable electronic device and / or the second camera of the wearable electronic device, and / or the analysis results of the user's voice, based on the selection of the first camera of the wearable electronic device, and performing a first task checked based on the analysis results of the plurality of first images and the user's voice.
[0012] The method of operating the electronic device may include an operation of confirming a second image acquired using the second camera based on the second camera being selected based on an analysis result of an image captured by the first camera and / or the second camera, and / or an analysis result of a user's voice, and performing a second task confirmed based on the analysis result of the second image and the user's voice.
[0013] One or more non-transitory computer-readable storage media may be provided storing one or more programs comprising computer-executable instructions.
[0014] The above instructions, when executed individually or collectively by one or more processors of the electronic device, may cause the electronic device to perform operations.
[0015] The above actions may include actions of establishing a communication connection between a wearable electronic device and the electronic device.
[0016] The above operations may include an operation of confirming a plurality of first images acquired using the first camera while the wearable electronic device is moving based on the first camera of the wearable electronic device being selected based on an analysis result of an image captured by the first camera of the wearable electronic device and / or a second camera of the wearable electronic device, and / or an analysis result of a user's voice, and performing a first task confirmed based on the analysis result of the plurality of first images and the user's voice.
[0017] The above operations may include operations of confirming a second image acquired using the second camera based on the second camera being selected based on the analysis results of images captured by the first camera and / or the second camera, and / or the analysis results of the user's voice, and performing a second task confirmed based on the analysis results of the second image and the user's voice.
[0018] The wearable electronic device may include a first camera and a second camera.
[0019] The focal length corresponding to the maximum resolution of the first camera may be smaller than the focal length corresponding to the maximum resolution of the second camera.
[0020] The wearable electronic device may include one or more processors.
[0021] The wearable electronic device may include a memory for storing instructions.
[0022] The instructions, when individually or collectively executed by the one or more processors, may cause the wearable electronic device to: identify a plurality of first images acquired using the first camera while the wearable electronic device is moving based on the first camera of the wearable electronic device being selected based on the analysis results of images captured by the first camera and / or the second camera, and / or the analysis results of the user's voice; and perform a first task identified based on the analysis results of the plurality of first images and the user's voice.
[0023] The instructions, when individually or collectively executed by the one or more processors, may cause the wearable electronic device to: identify a second image acquired using the second camera based on the selection of the second camera based on the analysis results of the images captured by the first camera and / or the second camera, and / or the analysis results of the user's voice; and perform a second task identified based on the analysis results of the second image and the user's voice.
[0024] A method of operating a wearable electronic device may include an operation of checking a plurality of first images acquired using the first camera while the wearable electronic device is moving based on a result of analyzing an image captured by a first camera of the wearable electronic device and / or a second camera of the wearable electronic device, and / or a result of analyzing a user's voice, based on selection of the first camera of the wearable electronic device, and performing a first task checked based on a result of analyzing the plurality of first images and the user's voice.
[0025] The method of operating the wearable electronic device may include an operation of confirming a second image acquired using the second camera based on the second camera being selected based on an analysis result of an image captured by the first camera and / or the second camera, and / or an analysis result of a user's voice, and performing a second task confirmed based on the analysis result of the second image and the user's voice.
[0026] One or more non-transitory computer-readable storage media may be provided storing one or more programs comprising computer-executable instructions.
[0027] The above instructions, when executed individually or collectively by one or more processors of the electronic device, may cause the electronic device to perform operations.
[0028] The above operations may include an operation of confirming a plurality of first images acquired using the first camera while the wearable electronic device is moving based on the first camera of the wearable electronic device being selected based on an analysis result of an image captured by the first camera of the wearable electronic device and / or a second camera of the wearable electronic device, and / or an analysis result of a user's voice, and performing a first task confirmed based on the analysis result of the plurality of first images and the user's voice.
[0029] The above operations may include operations of confirming a second image acquired using the second camera based on the second camera being selected based on the analysis results of images captured by the first camera and / or the second camera, and / or the analysis results of the user's voice, and performing a second task confirmed based on the analysis results of the second image and the user's voice.
[0030] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.
[0031] FIG. 2A is a drawing for explaining an electronic device according to one embodiment.
[0032] FIG. 2b is a drawing for explaining the resolution according to the focal length of each of a plurality of cameras according to one embodiment.
[0033] FIG. 3A is a diagram for explaining an operation method of an electronic device and a wearable electronic device according to one embodiment.
[0034] FIG. 3b is a diagram for explaining an operation method of an electronic device and a wearable electronic device according to one embodiment.
[0035] FIG. 3c is a diagram for explaining an operation method of a wearable electronic device according to one embodiment.
[0036] FIG. 4a is a drawing for explaining images captured by a first camera according to one embodiment.
[0037] FIG. 4b is a drawing for explaining images captured by a second camera according to one embodiment.
[0038] FIG. 4c is a diagram illustrating the performance of a task related to an image captured by a first camera according to one embodiment.
[0039] FIG. 4D is a diagram illustrating the performance of a task related to an image captured by a second camera according to one embodiment.
[0040] FIG. 4e is a diagram illustrating the performance of a task related to an image captured by a second camera according to one embodiment.
[0041] FIG. 5A is a diagram for explaining the operation of an electronic device and a wearable electronic device according to one embodiment.
[0042] FIG. 5b is a diagram illustrating a method for verifying input data for an artificial intelligence model according to one embodiment.
[0043] FIG. 6a is a diagram for explaining training and inference of an artificial intelligence model according to one embodiment.
[0044] FIG. 6b is a diagram for explaining inference of an artificial intelligence model of multiple functions according to one embodiment.
[0045] FIG. 6c is a diagram for explaining various inferences of an artificial intelligence model of multiple functions according to one embodiment.
[0046] Figure 6d is a diagram for explaining the structure of the artificial intelligence model.
[0047] FIG. 7a is a diagram illustrating a task confirmation method based on image and user voice analysis results according to one embodiment.
[0048] Figure 7b is a diagram for explaining the generation of input data according to one embodiment.
[0049] FIG. 7c is a diagram illustrating a task confirmation method based on image and user voice analysis results according to one embodiment.
[0050] FIG. 8 is a diagram illustrating a task confirmation method based on image and user voice analysis results according to one embodiment.
[0051] Figure 9a is a diagram for explaining a training method according to one embodiment.
[0052] FIG. 9b is a diagram illustrating a personalized LLM according to one embodiment.
[0053] FIG. 10 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0054] FIG. 11A and FIG. 11B are drawings for explaining recognition modes according to one embodiment.
[0055] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0056] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0057] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with the electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication device (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0058] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication device (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0059] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication device (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication device (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing device) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0060] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0061] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0062] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0063] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0064] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0065] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0066] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0067] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0068] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0069] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0070] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0071] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0072] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0073] The communication device (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication device (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication device (190) may include a wireless communication device (192) (e.g., a cellular communication device, a short-range wireless communication device, or a global navigation satellite system (GNSS) communication device) or a wired communication device (194) (e.g., a local area network (LAN) communication device, or a power line communication device). Among these communication devices, the corresponding communication device can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication devices can be integrated into a single component (e.g., a single chip) or implemented as a plurality of separate components (e.g., multiple chips). The wireless communication device (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0074] The wireless communication device (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication device (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication device (192) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication device (192) may support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication device (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0075] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication device (190). A signal or power may be transmitted or received between the communication device (190) and the external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0076] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0077] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0078] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0079] FIG. 2A is a diagram for explaining an electronic device according to one embodiment. The embodiment of FIG. 2A will be described with reference to FIG. 2B. FIG. 2B is a diagram for explaining the resolution of each of a plurality of cameras according to a focal distance according to one embodiment. The wearable electronic device (201) of FIG. 2A may be implemented as, for example, an electronic device (102) connected to the electronic device (101) of FIG. 1, or may be implemented as the electronic device (101) of FIG. 1, and there are no limitations to the implementation thereof.
[0080] A wearable electronic device (201) may include a housing (200). The housing (200) may have a shape that can be worn on a user's finger, for example. For example, the housing (200) may be implemented to include a sub-housing that constitutes the outer surface of the wearable electronic device (201) and a sub-housing that has a shape that can be worn on a user's finger, but this is an example and is not limiting.
[0081] A wearable electronic device (201) may include a microphone (255). The wearable electronic device (201) may acquire a user's voice through the microphone (255). At least some of the multiple processing operations for the user's voice acquired through the microphone (255) may be performed by the wearable electronic device (201), by the electronic device (101), and / or by the server (108), and those skilled in the art will understand that there is no limitation on the subject of processing the user's voice.
[0082] A wearable electronic device (201) may include a PMIC (288), a charging interface (287), and / or a battery (289). For example, power provided from the battery (289) may be processed (e.g., voltage converted, but not limited to) by the PMIC (288) and provided to other components. Alternatively, the battery (289) may be charged under the control of the PMIC (288) (or a charger). The charging interface (287) may be, for example, a power line connecting between the PMIC (288) and the battery (289), but there is no limitation on its implementation form.
[0083] A wearable electronic device (201) may include a processor (220) and / or a memory (230). The memory (230) may store instructions. The instructions may be executed by the processor (220) and, when executed, cause the wearable electronic device (201) to perform at least one operation. The processor (220) may be executed individually or collectively by one or more entities (cores and / or processing units (e.g., but not limited to, a CPU, a GPU, an NPU, an FPGA, or an ASIC). Those skilled in the art will appreciate that the memory (230) may be implemented as one or more entities, and that the instructions may be stored in one entity or distributed across multiple entities.
[0084] The wearable electronic device (201) may include a communication device (290) and / or an antenna (291). The communication device (290) may support one or more communication schemes. For example, the communication device (290) may support short-range communication (e.g., but not limited to, Bluetooth, BLE (Bluetooth Low Energy), Zigbee, Thread, UWB (Ultra Wide Band), or Wi-Fi Direct). For example, the communication device (290) may support cellular communication or communication based on IEEE 802.11x (which may be referred to as WiFi communication). The antenna (291) may be implemented to correspond to the supported communication scheme.
[0085] A wearable electronic device (201) may include at least one sensor (276a) (for example, but not limited to, a photoplethysmography (PPG) sensor as a biometric sensor), at least one sensor (276b) (for example, but not limited to, an inertial sensor), and / or at least one sensor (276c) (for example, but not limited to, a temperature sensor). At least some of the at least one entities (220, 230, 276a, 276b, 276c, 281, 282, 288, 290) may be disposed on a flexible printed circuit board (FPCB) (299), but not limited to.
[0086] A wearable electronic device (201) may include a plurality of cameras (281, 282). Referring to FIG. 2B, for example, the resolution (211) of the first camera (281) according to its focal length may be at least partially different from the resolution (212) of the second camera (282) according to its focal length. For example, the focal length A corresponding to the maximum resolution of the first camera (281) may be smaller than the focal length B corresponding to the maximum resolution of the second camera (282). Accordingly, the first camera (281) may be referred to as a close-up camera, and the second camera (282) may be referred to as a long-distance camera. The focal length A corresponding to the maximum resolution of the first camera (281) may be, for example, 8 cm or less, but is not limited thereto. As described below with reference to FIG. 4a, a user can wear a wearable electronic device (201) on a finger and position the wearable electronic device (201) relatively close to a subject (for example, 8 cm or less, but not limited thereto), in which case at least one image can be acquired by a first camera (281) having a relatively high resolution at a close range. A focal length B corresponding to the maximum resolution of a second camera (282) can be, for example, 8 cm or more and 100 cm or less, but not limited thereto. As described below with reference to FIG. 4b, a user can wear a wearable electronic device (201) on a finger and position the wearable electronic device (201) relatively far from a subject (for example, 8 cm or more and 100 cm or less, but not limited thereto), in which case at least one image can be acquired by a second camera (282) having a relatively high resolution at a long range.For example, a camera to be used for shooting may be selected from among a plurality of cameras (281, 282) based on a selection by a user, for example, an analysis result of an image captured by at least some of the cameras (281, 282), sensing data of another sensor (for example, but not limited to, a proximity sensor), and / or an analysis result of the user's voice (for example, but not limited to, information indicating whether a task requires close-up shooting), but this is exemplary and there is no limitation on the selection method. Meanwhile, those skilled in the art will understand that there is no limitation on the number of cameras (281, 282).
[0087] FIG. 3A is a diagram illustrating an operating method of an electronic device and a wearable electronic device according to one embodiment. The embodiment of FIG. 3A will be described with reference to FIGS. 4A and 4B. FIG. 4A is a diagram illustrating images captured by a first camera according to one embodiment. FIG. 4B is a diagram illustrating images captured by a second camera according to one embodiment.
[0088] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0089] According to one embodiment, operations 301 to 315 may be understood to be performed in each processor (e.g., processor (120) of FIG. 1 and / or processor (290) of FIG. 2A) of the electronic device (101) and / or the wearable electronic device (201).
[0090] Referring to FIG. 3A, according to one embodiment, in operation 301, the electronic device (101) may establish a communication connection between the wearable electronic device (201) and the electronic device (101). The communication connection may be, for example, a BLE communication connection, but there is no limitation on the communication method. In operation 303, the wearable electronic device (201) may acquire at least one image captured by the first camera (281) and / or the second camera (282). In operation 305, the wearable electronic device (201) may provide the at least one image acquired to the electronic device (101) through the established communication connection. The electronic device (101), in operation 307, may select one of the first camera (281) and the second camera (282) based on the analysis result of at least one image and / or the analysis result of the user's voice obtained through a microphone (e.g., the microphone (255) of the wearable electronic device (201), the microphone of the electronic device (101), and / or another external electronic device operatively connected to the electronic device (101) (e.g., the true wireless stereo (TWS) earphones, but not limited to)). The electronic device (101), in operation 309, may provide the selection result to the wearable electronic device (201).
[0091] In one example, the wearable electronic device (201) may drive the second camera (282) to acquire an image. The electronic device (101) (or the wearable electronic device (201)) may select either the first camera (281) or the second camera (282) based on the analysis result of the image acquired by the second camera (282). For example, either of the cameras (281, 282) may be selected based on whether the image acquired by the second camera (282) is out of focus. As described with reference to FIG. 2B, the resolution of the second camera (282) may differ depending on the focal length. For example, when the distance from the second camera (282) to the subject is relatively short (for example, less than 8 cm, but there is no limitation), the resolution of the second camera (282) is relatively low, and the acquired image may be out of focus. The electronic device (101) (or the wearable electronic device (201)) may select the first camera (281) rather than the second camera (282) based on the image acquired by the second camera (282) being out of focus. Alternatively, the electronic device (101) (or the wearable electronic device (201)) may select the second camera (282) based on the image acquired by the second camera (282) not being out of focus. There is no limitation on the method or algorithm for determining whether or not an image is out of focus, and an integrated AI model described below may be used to determine whether or not an image is out of focus. Those skilled in the art will understand that the occurrence of out-of-focus can also be expressed in terms of whether a region of interest (ROI) is identified. Alternatively, or additionally, those skilled in the art will understand that conditions for determining whether an out-of-focus condition exceeds the camera's focusing capabilities can be used in camera selection.
[0092] Meanwhile, a single camera may be implemented to have resolution capabilities ranging from relatively close to relatively long distances. In this case, those skilled in the art will appreciate that the electronic device (101) may be implemented to control the focal length of a single camera, rather than selecting one of the multiple cameras.
[0093] Meanwhile, selecting one of the cameras (281, 282) based on the analysis result of the image acquired by the second camera (282) is exemplary. In another example, the electronic device (101) (or, wearable electronic device (201)) may select one of the cameras (281, 282) based on the analysis result of the image using the first camera (281). For example, the electronic device (101) (or, wearable electronic device (201)) may select the second camera (282) instead of the first camera (281) based on the determination that the image acquired using the first camera (281) is out of focus. For example, the electronic device (101) (or, wearable electronic device (201)) may select the first camera (281) based on determining that the image acquired using the first camera (281) is not out of focus.
[0094] In one example, the electronic device (101) may select either the first camera (281) or the second camera (282) based on the analysis result of the user's voice. For example, based on the fact that the text corresponding to the user's voice corresponds to (may be identical or similar, but not limited to) keywords corresponding to the first camera (281) (or may also be named as keywords corresponding to a close-up shooting task), the electronic device (101) may select the first camera (281). For example, based on the fact that the text corresponding to the user's voice corresponds to (may be identical or similar, but not limited to) keywords corresponding to the second camera (282) (or may also be named as keywords corresponding to a long-distance shooting task), the electronic device (101) may select the second camera (282). For example, based on the fact that “Braille” in “Read the Braille”, which is a text corresponding to the user’s voice, is a keyword corresponding to the first camera (281), the electronic device (101) may select the first camera (281). For example, either the first camera (281) or the second camera (282) may be selected based on the natural language understanding result or the artificial intelligence inference result of the text for the user’s voice. For example, if the natural language understanding result or the artificial intelligence inference result of the text corresponds to close-up photography, the first camera (281) may be selected, and if the natural language understanding result or the artificial intelligence inference result of the text corresponds to long-distance photography, the second camera (282) may be selected. For example, based on the fact that the natural language understanding result of “Read the Braille” is close-up photography, the electronic device (101) may select the first camera (281). For example, based on the artificial intelligence inference result for “read the Braille” being a close-up shot (or the first camera (281)), the electronic device (101) can select the first camera (281).Meanwhile, the aforementioned keyword comparison-based (or rule-based) selection, natural language understanding-based selection, and / or AI inference-based selection are exemplary, and there are no limitations to the method of camera selection based on user voice. For example, a camera may be selected solely based on user voice analysis results, in which case steps 303 and / or 305 may be omitted, as those skilled in the art will understand.
[0095] In operation 311, the wearable electronic device (201) may acquire a plurality of first images using the first camera (281) while the wearable electronic device (201) is moving based on whether the first camera (281) is selected, or may acquire a second image using the second camera (282) based on whether the second camera (282) is selected. In operation 313, the wearable electronic device (201) may provide the acquired image (e.g., a plurality of first images acquired using the first camera (281) or a second image acquired using the second camera (282)) to the electronic device (101). The electronic device (101) may, in operation 315, perform a first task identified based on the analysis results of a plurality of first images and the user voice analysis results, or perform a second task identified based on the analysis results of a second image and the user voice analysis results.
[0096] For example, referring to FIG. 4A, a wearable electronic device (201) may be positioned relatively close to a subject (410). The electronic device (101) may select a first camera (281) based on an analysis result of an image acquired for camera selection and / or an analysis result of a user voice (401). The electronic device (101) may provide a selection result (e.g., information indicating that the first camera (281) is selected) to the wearable electronic device (201). The wearable electronic device (201) may acquire a plurality of images (411, 412, 413, 414) of one side (401a) of the subject (410) while moving (418). For example, a user can move his / her hand in a certain direction while wearing a wearable electronic device (201), and accordingly, a plurality of images (411, 412, 413, 414) can be acquired through the first camera (281). Meanwhile, in the example of FIG. 4a, for convenience of explanation, the plurality of images (411, 412, 413, 414) are illustrated as not overlapping each other, but this is exemplary, and those skilled in the art will understand that, depending on the implementation, adjacent images acquired by the first camera (281) may overlap at least partially (or, adjacent images may be expressed as including objects corresponding to the same subject). For example, the text “If you ask me, we often go to church seeking answers the church can't give.” may be written on one side (401a). Accordingly, in one image, for example, image (411), there is only the text “If you ask me, we” in the first line and “give.” in the second line, and the text alone may provide inappropriate image analysis results.Accordingly, based on the analysis results of the plurality of images (411, 412, 413, 414) (or their composite images) and the user's voice (401), an accurate task can be confirmed. The electronic device (101) can obtain the plurality of images (411, 412, 413, 414) from the wearable electronic device (201). The electronic device (101) can perform the first task based on the analysis results of the plurality of images (411, 412, 413, 414) and the user's voice (401). For example, based on the plurality of images (411, 412, 413, 414), the text “If you ask me, we often go to church seeking answers the church can't give.” can be confirmed, and based on a “read” command for the text, a voice (402) corresponding to the text can be output. Verification of the first task can be performed by an electronic device (101) and / or a server (108), and those skilled in the art will understand that there is no limitation on the performing entity.
[0097] For example, referring to FIG. 4B, the wearable electronic device (201) may be positioned relatively far from the subject (420). The electronic device (101) may select the second camera (282) based on the analysis result of the image acquired for camera selection and / or the analysis result of the user's voice (431). The electronic device (101) may provide the selection result (e.g., information indicating that the second camera (282) is selected) to the wearable electronic device (201). The wearable electronic device (201) may acquire an image (421) of the subject (420). The electronic device (101) may acquire the image (421) from the wearable electronic device (201). The electronic device (101) may perform a second task based on the analysis result of the image (421) and the user's voice (431). For example, a summary result (432) for text included in an image (421) may be provided. Verification of the second task may be performed by an electronic device (101) and / or a server (108), and those skilled in the art will understand that there is no limitation on the performing entity.
[0098] FIG. 3b is a diagram for explaining an operation method of an electronic device and a wearable electronic device according to one embodiment.
[0099] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0100] According to one embodiment, operations 331 to 341 may be understood to be performed in each processor (e.g., processor (120) of FIG. 1 and / or processor (290) of FIG. 2A) of the electronic device (101) and / or the wearable electronic device (201).
[0101] The electronic device (101), in operation 331, may establish a communication connection between the wearable electronic device (201) and the electronic device (101). The wearable electronic device (201), in operation 333, may obtain at least one image captured by the first camera (281) and / or the second camera (282). The wearable electronic device (201), in operation 335, may select one of the first camera (281) and the second camera (282) based on the analysis result of the at least one image and / or the analysis result of the user's voice obtained through the microphone (255). In contrast to the embodiment of FIG. 3A in which the selection of the camera is performed by the electronic device (101), in the embodiment of FIG. 3B, the wearable electronic device (201) may select one of the cameras (281, 282). The selection of the camera has been described in detail with reference to FIG. 3A and will not be repeated here. In operation 337, the wearable electronic device (201) may acquire a plurality of first images using the first camera (281) while the wearable electronic device (201) is moving based on the selection of the first camera (281) as described with reference to FIG. 4A, or may acquire a second image using the second camera (282) based on the selection of the second camera (282) as described with reference to FIG. 4B. In operation 339, the wearable electronic device (201) may provide the acquired image (e.g., a plurality of first images acquired using the first camera (281) or a second image acquired using the second camera (282)) to the electronic device (101).The electronic device (101) may, in operation 341, perform a first task identified based on the analysis results of a plurality of first images and the user voice analysis results, or perform a second task identified based on the analysis results of a second image and the user voice analysis results.
[0102] FIG. 3c is a diagram for explaining an operation method of a wearable electronic device according to one embodiment.
[0103] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0104] According to one embodiment, operations 361 to 369 may be understood to be performed in each processor of the electronic device (201) (e.g., processor (290) of FIG. 2A).
[0105] In operation 361, the wearable electronic device (201) may select either the first camera (281) or the second camera (282) based on the analysis results of at least one image captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice acquired through the microphone (255). The selection of the camera has been described in detail with reference to FIG. 3A and is therefore not repeated here. Based on the selection of the first camera (281), the wearable electronic device (201) may, in operation 363, check a plurality of first images acquired using the first camera while the wearable electronic device (201) is moving, as described with reference to FIG. 4A. The wearable electronic device (201) may perform a first task identified based on the analysis results of a plurality of first images and the user's voice in operation 365. For example, the wearable electronic device (201) may be implemented to include a speaker, or may be operatively connected to an external electronic device including a speaker, and may output a voice corresponding to the reading function described in FIG. 4A directly or through the external electronic device. Based on the selection of the second camera (282), the wearable electronic device (201) may, in operation 367, identify a second image acquired using the second camera (282), as described with reference to FIG. 4B. The wearable electronic device (201) may, in operation 369, perform a second task identified based on the analysis results of the second image and the user's voice. For example, the wearable electronic device (201) may be implemented to include a speaker, or may be operatively connected to an external electronic device including a speaker, and may output voice as summarized in FIG. 4b directly or through the external electronic device.Alternatively, the wearable electronic device (201) may provide data for content expression to an external electronic device including a display device for displaying summarized content. Accordingly, the summarized content identified by the wearable electronic device (201) may be expressed on the external electronic device.
[0106] FIG. 4c is a diagram illustrating the performance of a task related to an image captured by a first camera according to one embodiment.
[0107] For example, referring to FIG. 4C, a wearable electronic device (201) may be positioned relatively close to a subject (410). The electronic device (101) (or, wearable electronic device (201)) may select a first camera (281) based on the analysis results of images acquired for camera selection and / or the analysis results of user voice (441). The wearable electronic device (201) may acquire a plurality of images of the subject (410). The wearable electronic device (201) may provide the acquired plurality of images to the electronic device (101). The electronic device (101) may perform a task identified based on the analysis results of the plurality of images and the user voice (441) received from the wearable electronic device (201). For example, a translation result (442) for text identified based on the plurality of images may be provided. Verification of the task can be performed by an electronic device (101) and / or a server (108), and those skilled in the art will understand that there is no limitation on the performing entity.
[0108] FIG. 4D is a diagram illustrating the performance of a task related to an image captured by a second camera according to one embodiment.
[0109] For example, referring to FIG. 4D, the wearable electronic device (201) may be positioned relatively far from the subject (440). The electronic device (101) (or the wearable electronic device (201)) may select the second camera (282) based on the analysis result of the image acquired for camera selection and / or the analysis result of the user's voice (443). The wearable electronic device (201) may acquire an image of the subject (440). The wearable electronic device (201) may provide the acquired image to the electronic device (101). The electronic device (101) may perform a task identified based on the analysis result of the image received from the wearable electronic device (201) and the user's voice (443). For example, storage of the image and / or a storage result (444) of the image (e.g., a representation of the stored image) may be performed. Verification of the task can be performed by an electronic device (101) and / or a server (108), and those skilled in the art will understand that there is no limitation on the performing entity.
[0110] FIG. 4e is a diagram illustrating the performance of a task related to an image captured by a second camera according to one embodiment.
[0111] For example, referring to FIG. 4E, the wearable electronic device (201) may be positioned relatively far from the subject (440). The electronic device (101) (or, the wearable electronic device (201)) may select the second camera (282) based on the analysis result of the image acquired for camera selection and / or the analysis result of the user's voice (451). The user may, for example, operate the wearable electronic device (201) worn on a finger to rotate and capture an image. The wearable electronic device (201) may acquire an image (452) of the subject. The wearable electronic device (201) may provide the acquired image (452) to the electronic device (101). The electronic device (101) can perform a task identified based on the analysis results of the image (452) and user voice (451) received from the wearable electronic device (201). For example, the storage of the image and / or the result of storing the image (e.g., the representation of the stored image) can be performed. The task identification can be performed by the electronic device (101) and / or the server (108), and those skilled in the art will understand that there is no limitation on the performing entity.
[0112] FIG. 5A is a diagram for explaining the operation of an electronic device and a wearable electronic device according to one embodiment.
[0113] According to one embodiment, a user may wear a wearable electronic device (201). The wearable electronic device (201) may acquire multiple images (501a) using, for example, a first camera (281). Alternatively, the wearable electronic device (201) may acquire an image (501b) using, for example, a second camera (282). The user (502) may speak a user voice (503) while capturing at least one image (501a or 501b). Acquisition of the user voice (503) may be performed by the electronic device (101), an external electronic device (504a, 504b), and / or the wearable electronic device (201), without limitation. The acquired user voice (503) may be converted into text through the ASR module (512).
[0114] The command module (511) can receive at least one image (501a or 501b) and converted text. The command module (511) can provide input data for an artificial intelligence model (520) based on at least one image (501a or 501b) and converted text. For example, the artificial intelligence model (520) can support text-to-speech (521), translation (522), braille / image-to-text (523), language identification (LID) detection (524), and / or region of interest (ROI) detection (525), but there is no limitation on the type and / or number of supported functions (or sub-artificial intelligence models). The process of generating input data for the artificial intelligence model (520) by the command module (511) will be described later.
[0115] The inference result of the artificial intelligence model (520) may be stored in the data storage (513). The inference result of the artificial intelligence model (520) may be utilized by a large language model (LLM) (526a), and as the utilization result is converted to text-to-speech (521), it may be output as a voice by, for example, an external electronic device (504b), but there is no limitation on the output entity. Depending on the function being performed, the LLM (526a) may be additionally applied, or the inference result of the artificial intelligence model (520) may be provided without the LLM (526a) being additionally applied. A fine-tuned LLM (526b) may be provided based on data stored in the data storage (513). Those skilled in the art will understand that the fine-tuned LLM (526b) may replace the existing LLM (526). Meanwhile, there is no limitation on the fine-tuning method. For example, LoRA (low-rank adaptation) can be used for fine-tuning. LoRA is a technique for adapting AI models, particularly in the field of natural language processing (NLP), to new tasks or domains by adjusting model parameters with a small number of additional parameters. Instead of directly modifying specific parameters of a pre-trained deep learning model, LoRA can utilize a small number of additional parameters to adjust model parameters. For example, weights can be fixed in a pre-trained model and a trainable rank decomposition matrix can be injected into each layer, but there are no limitations.
[0116] For example, the LLM (526a) may support a summary function. The summary function may refer to a function that summarizes a relatively large amount of input text. For example, based on a user's voice and a captured image, such as "What language is this text?", an inference result indicating that the language information is English may be provided. For example, the LLM (526a) may provide a response, "The language you mentioned is [English]," using the inference result of "English." In the above-described process, usage history may be used for fine-tuning. In addition, fine-tuning of the LLM (526a) may be performed based on usage patterns. Alternatively, sentiment analysis may be performed based on biometric information identified using a biometric sensor, and preference data may be classified based on the sentiment analysis and subsequently used for fine-tuning. Data stored in the data storage (513) may be used, for example, by an application (514), and the usage result (527) may be provided. Those skilled in the art will appreciate that editing functions for the usage results (527) may be supported by the application (514). The application (514) may, for example, support editing functions for texts generated by LLM (526a), summarized texts, and stored images.
[0117] FIG. 5B is a diagram illustrating a method for verifying input data for an artificial intelligence model according to one embodiment. Those skilled in the art will understand that the operations included in the method described in FIG. 5B may be performed, for example, by an electronic device (101), a wearable electronic device (201), and / or a server (108).
[0118] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0119] Referring to FIG. 5B, the method may include an operation (531) of acquiring a user voice. The method may include an operation (533) of processing the user voice and extracting features. The method may include an operation (535) of converting the user voice into text through ASR. The method may include an operation (537) of tokenizing the text and analyzing sentences. The method may include an operation (539) of verifying input data for an artificial intelligence model based on the analysis results. For example, tokenization and sentence analysis may provide an analysis result (which may be named a capsule command, but is not limited to it) based on tagging, keyword detection, intent identification, and / or morphological analysis, but is not limited to it. For example, based on keyword detection (or semantic analysis) of “text” and “Braille” in text output through ASR, the acquired image may be identified as a general image or a Braille image. The input data may include an item indicating the type of image, and information about the type of image identified according to the above-described method may be included in the input data. For example, based on keyword detection (or semantic analysis) of “text” and “Braille,” it may be determined that “perform image recognition” is required, and the image-to-text item included in the input data may have an activated value. For example, if the text includes “translate it,” “translate it into English,” and “translate it into French,” the translation item in the input data may have an activated value based on the keyword detection method and / or semantic analysis. The format of the input data will be described later.
[0120] FIG. 6a is a diagram for explaining training and inference of an artificial intelligence model according to one embodiment.
[0121] According to one embodiment, an artificial intelligence model (610) can be trained. The artificial intelligence model (610) can support multiple functions, and multiple types of training data can be used to train each of the multiple functions. For example, for training an image recognition function, training data consisting of pairs of images and correct text can be used. For example, for training a translation function, training data consisting of pairs of images and texts corresponding to the language after translation can be used. For example, for training a text-to-speech (TTS) function, training data consisting of pairs of texts and speech sounds can be used. For example, in FIG. 6A, training data of texts (603, 604) corresponding to various types of languages and corresponding speech sounds (601, 602) respectively can be used for training a TTS function and / or a translation function. Accordingly, the trained artificial intelligence model (610) may provide Korean text (606) as an inference result for English speech sounds (605).
[0122] FIG. 6b is a diagram for explaining inference of an artificial intelligence model of multiple functions according to one embodiment.
[0123] According to one embodiment, input data (680) for an artificial intelligence model (610) may be generated based on at least one image (621 or 622) acquired by a wearable device (201) and a user voice. The input data (680) may include, for example, information corresponding to each of a plurality of items (680a, 680b, 680c, 680d, 680e) as shown in FIG. 6B , but is not limited thereto. For example, information regarding whether TTS is performed may be expressed for an item (680a) of a TTS token. If it is determined that TTS is performed, a value of “TTS lang” may be expressed, and if it is determined that TTS is not performed, a value of “None” may be expressed. For example, information regarding whether translation is performed and / or at least one language associated with translation may be expressed for an item (680b) of a translation token. For example, if translation is not performed, a value of “None” may be expressed. For example, when translation is performed, the value “I2T English Korean” may be expressed, which may mean translation from English to Korean. For example, for the text item (680c), information about a portion of the source text input when TTS and / or translation is performed may be expressed. For the image token item (680d), information about the type of image may be expressed. For example, in the case of a general image, a value of “0” may be expressed, and in the case of a Braille image, a value of “1” may be expressed. For the image item (680e), information about the path where the image is stored may be expressed. The process of generating the above-described input data may be named prompting, and the input data may be named prompt.Assume that image (621) contains, for example, the text “Even the cat suddenly leapt off a root”, or that image (622) contains Braille corresponding to “Even the cat suddenly leapt off a root”.
[0124] The artificial intelligence model (610) can provide inference results (621, 622, 623, 624) for the input data (680). For example, the inference result (623) can be the text of “Even the cat suddenly leapt off a root,” which is a translation result corresponding to “Even the cat suddenly leapt off a root.” The inference result (622) can be the type of language before translation. The inference result (621) can include a speech sound (621). The speech sound (621) can be, for example, a spectrogram corresponding to the text of the translation result “Even the cat suddenly leapt off the roof,” but is not limited thereto. The inference result (624) can be information on a portion to be translated within the image (621 or 622).
[0125] FIG. 6c is a diagram for explaining various inferences of an artificial intelligence model of multiple functions according to one embodiment.
[0126] According to one embodiment, the artificial intelligence model (610) may support multiple functions as described above. For example, an inference result (631b) of an artificial intelligence model (610) of a prompt based on a Braille image (631a) may be provided. The inference result (631b) may include, for example, text corresponding to the Braille image (631a), ROI information, and information indicating Korean. For example, an inference result (632b) of an artificial intelligence model (610) of a prompt based on a general image (632a) may be provided. The inference result (632b) may include, for example, text corresponding to the general image (632a), ROI information, and information indicating English. For example, an inference result (633b) of an artificial intelligence model (610) of a prompt based on text (633a) may be provided. The inference result (633b) may be a speech sound of the text (633a). For example, after the primary inference result (632b) by the artificial intelligence model (610) is confirmed, a speech sound may be provided as a secondary additional inference result (633b) for the inference result (632b), and multiple inferences by the artificial intelligence model (610) may be possible.
[0127] Figure 6d is a diagram for explaining the structure of an artificial intelligence model.
[0128] According to one embodiment, the artificial intelligence model (610) may include, but is not limited to, a word embedding layer (642, 643), a feature extraction layer (644, 646), a casual CNN (645, 647), a concatenation layer (648), a transformer (649), a vocoder (650), and / or a dense layer (or a fully connected layer) (652, 654, 656). As described above, input data (641a, 641b, 641c, 641d, 641e, 641f) for the artificial intelligence model (610) may be generated. Input data (641a) corresponding to a TTS token and input data (641b) corresponding to a translation token may be input to the word embedding layer (642). Data (641c) corresponding to a text may be a word The word embedding layer (643) can be input to the word embedding layer (642, 643). The word embedding layer (642, 643) can extract text features, for example, and provide the extracted features to the connection layer (648). The features can be expressed as vectors, for example, but are not limited thereto. The input data (641d) corresponding to the image token can be provided to the connection layer (648). The input data (641e) corresponding to the image including text can be provided to the feature extraction layer (644), or the input data (641f) corresponding to the Braille image can be provided to the feature extraction layer (646). The features extracted by the feature extraction layer (644, 646) can be provided to the casual CNN (645, 647). The casual CNN (645, 647) can be configured to process, for example, one or more images (or can be expressed as streaming images), but are not limited thereto. The connection layer (648) receives the input. By connecting the data, it can be provided to the transformer (649).The transformer (649) may operate based on an attention mechanism, for example, and may include, but is not limited to, an encoder and / or a decoder. The transformer (649) may be trained to provide output data associated with a plurality of functions. For example, the transformer (649) may provide data for generating a speech sound (651) to the vocoder (650). The vocoder (650) may provide the speech sound (651). For example, the vocoder (650) may generate a spectrogram by aligning frames for an image output from the encoder in the form of an RNN and generated audio frames, but is not limited to. For example, the dense layers (652, 654, 656) may provide, as inference results, LID (653), data for text generation, and / or ROI information (657). Data for text generation can be converted into text and provided by a character decoder (655). The character decoder (655) can align text with vectors, for example, using connectionist temporal classification (CTC) loss, but is not limited thereto. Although not shown, those skilled in the art will understand that the inference results (651, 653, 655, 657) can also be used for training along with the correct data, and there is no limitation on the type of loss function used for training.
[0129] FIG. 7A is a diagram illustrating a task verification method based on image and user voice analysis results according to one embodiment. Those skilled in the art will understand that the operations included in the method described in FIG. 7A may be performed, for example, by an electronic device (101), a wearable electronic device (201), and / or a server (108).
[0130] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0131] According to one embodiment, the method may include an operation (operation 701) of confirming input data for an AI model based on the analysis results of at least one image captured by a selected camera and a user's voice. Since the selection of any one of the cameras (281, 282) of the wearable device (201) has been described in detail, the description will not be repeated here. Based on the analysis results of at least one image and the user's voice, input data composed of at least one item-specific piece of information may be confirmed. For example, the items constituting the input data for the AI model may be [TTS token], [Translation token], [text], [image token], and [image path]. Here, [TTS token] may mean a language in which TTS (text to speech) is to be performed, and may be expressed as a type of language (Korean, English), for example, but there is no limitation on the expression method. [Translation token] is information on a language for translation, and may be expressed as a type of language before translation and / or after translation, but there is no limitation on the expression method. [text] may include information related to the source text input when performing TTS and / or translation. [image token] may be expressed as 0 for a general image and 1 for a Braille image, but there is no limitation on the type of image and / or its expression method. [image path] may mean the storage path of at least one image. Accordingly, at least one image input to the AI model may be specified. The method may include an operation (operation 703) of confirming a task corresponding to an inference result of the AI model for the input data.The inference result of the AI model may include, but is not limited to, a TTS spectrogram, a language identification (LID), text, and / or a box position (or a position of a ROI). Unused items may be expressed as “none,” but are not limited to, either. For example, when a reading function (or TTS) is requested, the input value of the AI model may be expressed as “[TTS lang] [none][none][0 / 1][image path].” Here, TTS lang may be the language type for which TTS is to be performed. Here, [none] is for the Translation token, and since no translation is requested, it may be expressed as [none]. The input value of the AI model may be verified, for example, based on a rule. If the text, which is the ASR result of the user’s voice, contains a keyword related to “reading” but does not contain a keyword related to “translation,” the Translation token may be set to none. Meanwhile, rule-based input data verification based on keyword comparison is merely exemplary, and those skilled in the art will understand that NLU and / or AI model inference results may also be utilized. Furthermore, [image token] can be represented as 0 or 1. For example, if a translation function is requested, the input value for the AI model may be [TTS lang][I2T English Korean][none][0 / 1][image path] when used.
[0132] Figure 7b is a diagram for explaining the generation of input data according to one embodiment.
[0133] Referring to Fig. 7b, for example, based on the recognition result of the image being “text” (711a), the type of the image can be confirmed as “0” (711b). Based on the recognition result of the image being “Braille” (712a), the type of the image can be confirmed as “1” (712b). The data can be confirmed to correspond to an image-to-text function (720a) among the multiple task functions supported by the artificial intelligence model, and thus can be used to configure input data (721a, 722a) of the “image-to-text” among the multiple tasks. For example, based on the confirmation of the text “Translate this” (713a) from the user’s voice, it can be confirmed to correspond to a translation function (720b). The data can be used to configure input data (721b) of “translation.” Meanwhile, for the input data (721b) of "translation," for example, a first language (e.g., Korean (713b)) may be set as the default starting language, and a language (713c) extracted after translation may be used to configure the input data (721b). For example, based on the identification of "summary" (714a), "summary" (714b), question mark (714c), and LLM-related context (714d) from the user's speech, it may be determined that the LLM function (720c) is required. In this case, the LLM application result (722c) for the inference result of the artificial intelligence model may be provided. For example, based on the identification of “~hae” (715a), “~haejwo” (715b), “read me” (715c), “tell me” (715d), “tell me” (715e), or a question mark (715f) from the user’s voice, it can be determined that the TTS function (729d) is required. The data can be used to configure the input data (721d, 722d) of “TTS.” As described above, input data for an artificial intelligence model can be configured, but there is no limitation.
[0134] For example, a “reading” function may be performed based on the user’s voice confirmation of “read the text” or “read the Braille.” In this case, the keywords [text] and [read it] may be separated. To output a voice corresponding to an image corresponding to [text], the “image-to-text” function of the artificial intelligence model may be performed, and then the “TTS” function may be performed. For example, to perform the first “image-to-text” function, the first input data may be provided to the artificial intelligence model. Information for each item of the first input data may be as follows.
[0135] <Input data for the first inference of the AI model for "Read this text">
[0136] -[TTS:None]: TTS is not required when performing the first function, so the value of the TTS item can be “None”.
[0137] -[Translation: None]: No translation is required, so the value of the Translation field can be “None”.
[0138] -[Text: None]: There is no source text, so the value of the Text item can be “None”.
[0139] -[image: 0]: Based on the fact that the acquired image is a normal image, the value of the image item can be “0”.
[0140] -[image path: ~ / img]: “~ / img”, the path where the image is stored, can be expressed.
[0141] The function for this can be expressed as “ROI, text, LID, speech = MultiTaskModel[None, None, None, 0, ~ / img].” As the first inference result of the artificial intelligence model for this, ROI, text, and LID information can be provided. Here, since TTS is None, speech among the inference results of the artificial intelligence model can have the value “none.” The artificial intelligence model can generate input data for the second inference based on the first inference result, which can be as follows.
[0142] <Input data for the second inference of the AI model for "Read this text">
[0143] -[TTS:ko]: It can be confirmed based on the LID that the language for which TTS is requested is Korean, and accordingly, the value of the TTS item can be “ko”.
[0144] -[Translation: None]: No translation is required, so the value of the Translation field can be “None”.
[0145] -[Text: “Text of the first inference result”]: The text of the first inference result can be used as input data for the second inference.
[0146] -[image: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0147] -[image path: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0148] The function for this can be expressed as, “ROI, text, LID, speech = MultiTaskModel[None, None, “text of the first inference result”, None, None].” As the second inference result of the AI model for this, speech can be provided, and ROI, text, and LID can be expressed as the value “None.”
[0149] As described above, the “read” function may be performed based on two inferences of the artificial intelligence model, but those skilled in the art will understand that this is exemplary and it may be performed based on one inference, or three or more inferences.
[0150] For example, a “translation” function may be performed based on the user’s voice confirmation of “translate this text”, “translate it into Braille English”, or “translate it into Braille French”. In this case, the translation function may be confirmed based on the keyword detection of “translation”, but there is no limitation. In order to output voice corresponding to the image for translation, the “image-to-text” function of the artificial intelligence model may be performed, the “translation” function may be performed, and the “TTS” function may be performed. For example, in order to perform the first “image-to-text” function, the first input data may be provided to the artificial intelligence model. Information for each item of the first input data may be as follows.
[0151] <Input data for the first inference of the AI model for "translation">
[0152] -[TTS:None]: TTS is not required when performing the first function, so the value of the TTS item can be “None”.
[0153] -[Translation: None]: No translation is required, so the value of the Translation field can be “None”.
[0154] -[Text: None]: There is no source text, so the value of the Text item can be “None”.
[0155] -[image: 0]: Based on the fact that the acquired image is a normal image, the value of the image item can be “0”.
[0156] -[image path: ~ / img]: “~ / img”, the path where the image is stored, can be expressed.
[0157] The function for this can be expressed as “ROI, text, LID, speech = MultiTaskModel[None, None, None, 0, ~ / img].” As the first inference result of the artificial intelligence model for this, ROI, text, and LID information can be provided. Here, since TTS is None, speech among the inference results of the artificial intelligence model can have the value “none.” The artificial intelligence model can generate input data for the second inference based on the first inference result, which can be as follows.
[0158] <Input data for the second inference of the AI model for “translation”>
[0159] -[TTS:None]: Since the second function is translation and TTS is the third function, the value of the TTS item can be “None” at this stage.
[0160] -[Translation: I2T LID eng]: May contain information about the detected language (e.g., Korean) and the translated language (e.g., English).
[0161] -[Text: “Text of the first inference result”]: The text of the first inference result can be used as input data for the second inference.
[0162] -[image: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0163] -[image path: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0164] The function for this can be expressed as, "ROI, text, LID, speech = MultiTaskModel[None, I2T LID eng, "text of the first inference result", None, None]." As the second inference result of the AI model, a translation result can be provided. The AI model can generate input data for the third inference based on the second inference result, as follows.
[0165] <Input data for the second inference of the AI model for "translation">
[0166] -[TTS:LID]: May contain the LID of the language for which TTS is requested.
[0167] -[Translation: None]: Since translation has already been performed, no translation is required at this stage, and therefore the value of the Translation entry can be “None”.
[0168] -[Text: “Translation result of the text of the first inference result”]: The translation result, which is the second inference result, can be used as input data for the third inference.
[0169] -[image: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0170] -[image path: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0171] The function for this can be expressed as, “ROI, text, LID, speech = MultiTaskModel[LID, None, “Translation result of the text of the first inference result”, None, None].” As the third inference result of the artificial intelligence model for this, the speech corresponding to the translation result can be provided, and ROI, text, and LID can be expressed as the value “None.”
[0172] For example, a "Save" function may be performed based on the user's voice confirmation of "Capture image," "Save image," or "Save text." In this case, an "Image-to-Text" function may be performed, and the input data may be as follows.
[0173] <Input data for inference of artificial intelligence model for “storage”>
[0174] -[TTS:None]: TTS is not required, so the value of the TTS entry can be “None”.
[0175] -[Translation: None]: No translation is required, so the value of the Translation field can be “None”.
[0176] -[Text: None]: There is no source text, so the value of the Text item can be “None”.
[0177] -[image: 0]: Based on the fact that the acquired image is a normal image, the value of the image item can be “0”.
[0178] -[image path: ~ / img]: “~ / img”, the path where the image is stored, can be expressed.
[0179] The function for this can be expressed as, “ROI, text, LID, speech = MultiTaskModel[None, None, None, 0, ~ / img].” As the inference result of the artificial intelligence model for this, ROI, text, and LID information can be provided. The inference result of the artificial intelligence model can be stored.
[0180] As described above, multiple functions of an artificial intelligence model can be performed based on one or more inferences, and information on inferences for each function can be as shown in Table 1.
[0181] Task InferenceReadingFunction for first inference: image to textFunction for second inference: TTSTranslationFunction for first inference: image to textFunction for second inference: translationFunction for third inference: TTSSaveFunction for first inference: image to textSummaryFunction for first inference: image to textApply LLMFunction for second inference: TTSQueryFunction for first inference: image to textApply LLMFunction for second inference: TTS
[0182] Referring to Table 1, "Read" and "Translate" can be performed based on multiple inferences from the AI model as described above, while "Save" can be performed based on a single inference. Meanwhile, "Summary" and "Inquiry" may require additional application of LLM, as described with reference to Figure 7c. Whether tasks are required to be performed can be determined sequentially, for example, but there are no limitations.
[0183] FIG. 7C is a diagram illustrating a task verification method based on image and user voice analysis results according to one embodiment. Those skilled in the art will understand that the operations included in the method described in FIG. 7C may be performed by, for example, an electronic device (101), a wearable electronic device (201), and / or a server (108).
[0184] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0185] According to one embodiment, the method may include an operation (operation 761) of verifying input data based on the analysis results of at least one image captured by a selected camera and a user's voice. The method may include an operation (763) of verifying an inference result of an AI model for the input data. The method may include an operation (765) of verifying a task corresponding to a result of LLM processing for the inference result. For example, as described with reference to FIG. 7A, a task corresponding to the inference result of the AI model may be provided, and as shown in FIG. 7C, a task corresponding to a result of additional application of LLM for the inference result of the AI model may be provided. For example, a reading function may be a task corresponding to the inference result of the AI model. For example, a summarizing function may be a task requesting an additional LLM application result for text corresponding to the inference result of the AI model. For example, whether or not to additionally apply LLM may be determined depending on the type of task, but there is no limitation thereto. Meanwhile, those skilled in the art will understand that additional prompting may be further performed for the additional application of LLM.
[0186] For example, based on the user's voice confirmation of "What language is the text you are currently viewing?", an "Inquiry" (Q&A) function may be performed. To output voice corresponding to the image for the inquiry, the "Image-to-Text" function of the AI model may be performed, LLM may be applied, and the "TTS" function may be performed based on the result of the LLM application. For example, to perform the first "Image-to-Text" function, the first input data may be provided to the AI model. Information for each item of the first input data may be as follows.
[0187] <Input data for the first inference of the AI model for “Q&A”>
[0188] -[TTS:None]: TTS is not required when performing the first function, so the value of the TTS item can be “None”.
[0189] -[Translation: None]: No translation is required, so the value of the Translation field can be “None”.
[0190] -[Text: None]: There is no source text, so the value of the Text item can be “None”.
[0191] -[image: 0]: Based on the fact that the acquired image is a normal image, the value of the image item can be “0”.
[0192] -[image path: ~ / img]: “~ / img”, the path where the image is stored, can be expressed.
[0193] The function for this can be expressed as, “ROI, text, LID, speech = MultiTaskModel[None, None, None, 0, ~ / img].” As the first inference result of the artificial intelligence model for this, ROI, text, and LID information can be provided. The artificial intelligence model can apply the first inference result to LLM. For example, the prompt for applying LLM can be “What is the language type of [LID]?”, but there is no limitation. The prompt can be preset or generated based on the identified task, for example, and there is no limitation. The result of applying LLM to the prompt can be, for example, “The language is [LID].”
[0194] We can generate input data for the second inference, which may be as follows:
[0195] <Input data for the second inference of the AI model for "Q&A">
[0196] -[TTS:LID]: The value of the TTS entry can be the detected “LID”.
[0197] -[Translation: None]: This can be None, as no translation is required.
[0198] -[Text: “The language is [LID]”]: This could be the result of applying LLM, “The language is [LID]”.
[0199] -[image: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0200] -[image path: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0201] The function for this can be expressed as, "ROI, text, LID, speech = MultiTaskModel[LID, None, "The language is [LID]", None, None]." As the second inference result of the AI model, an answer to the question can be provided. "Q&A" can be performed based on the second inference result of the AI model and the application of LLM.
[0202] Based on the AI model and LLM trained in the above manner, answers can be provided for queries that do not have associated images. For example, if a user's voice is identified as "What role does Adams play in K-drama?", an image may not be obtained. In this case, LLM can be initially applied to the text corresponding to the user's voice, and an inference result by the AI model based on the application result can be provided. For example, a prompt for LLM application could be, "Answer the question 'What role does Adams play in K-drama?'", but there are no restrictions. The LLM's answer could be, for example, "Adams played the role of a king." Based on the LLM application result, input data for the AI model can be generated, for example. For example, the AI model may be required to perform a TTS function, and the input data may be as follows.
[0203] <Input data for inference of an AI model for “Q&A”>
[0204] -[TTS:ko]: The value of the TTS entry can be “ko”.
[0205] -[Translation: None]: This can be None, as no translation is required.
[0206] -[Text: “Adams played the role of the king”]: The result of applying LLM could be “Adams played the role of the king”.
[0207] -[image: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0208] -[image path: None]: This is a step for TTS where images are not used, so it can be expressed as “None”.
[0209] Accordingly, a voice corresponding to “Adams played the role of the king” can be output.
[0210] FIG. 8 is a diagram illustrating a task verification method based on image and user voice analysis results according to one embodiment. Those skilled in the art will understand that the operations included in the method described in FIG. 8 may be performed by, for example, an electronic device (101), a wearable electronic device (201), and / or a server (108).
[0211] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0212] According to one embodiment, the method may include an operation (operation 801) of confirming first input data based on an analysis result of at least one image captured by a selected camera and a user voice. The method may include an operation (operation 803) of confirming second input data based on an inference result of an AI model for the first input data. For example, items constituting the first input data may be [TTS token] [Translation token] [image token] [image path]. The AI model inference result for the first input data may include, for example, an LID, and the LID may be confirmed as the second input data. The method may include an operation (operation 805) of confirming a task based on the inference result of the AI model for the first input data of [TTS token] [Translation token] [image token] [image path] and the additionally confirmed second input data of [LID].
[0213] FIG. 9A is a diagram illustrating a training method according to one embodiment. The embodiment of FIG. 9A will be described with reference to FIG. 9B . FIG. 9B is a diagram illustrating a personalized LLM according to one embodiment. Those skilled in the art will understand that the operations included in the method described in FIG. 9A may be performed, for example, by an electronic device (101), a wearable electronic device (201), and / or a server (108).
[0214] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0215] According to one embodiment, the method may include an operation (operation 911) of identifying input data for an AI model based on an analysis result of at least one image captured by a selected camera and a user's voice. Since the selection of any one of the cameras (281, 282) of the wearable device (201) has been described in detail, the description will not be repeated here. Based on the analysis result of at least one image and the user's voice, input data consisting of at least one item-specific piece of information may be identified. The method may include an operation (operation 913) of identifying an inference result of an AI model for the input data. The method may include an operation (operation 915) of identifying a task corresponding to a result of LLM processing the inference result. As described above, the task may be identified based on an additional application result of LLM to the inference result of the AI model. The method may include an operation (operation 917) of training the LLM based on data related to the processing of the LLM. For example, LLM training can be performed based on the size of the pre-stored data exceeding a threshold size (e.g., 1 tera byte, but there is no limit), but there are no restrictions on the training execution conditions. Accordingly, as LLM is trained (or fine-tuned) based on personalized information, the application result of LLM can be customized to the user. For example, referring to FIG. 9B, the LLM application result (922) for a user's voice (921) can be provided. For example, there may be a history where the LLM application result for the AI model inference result for "What is this code?" and an image of the code is provided as "This is a phython code that extracts log-mel when writing an emotion speech recognition code."Based on the LLM application result, "This is a phython code that extracts log-mel when writing emotion speech recognition code," the LLM can be trained (fine-tuned). Thereafter, the results (922) of the fine-tuned LLM application to the user's voice (921) may be provided.
[0216] FIG. 10 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0217] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0218] According to one embodiment, operations 1001 to 1009 may be understood to be performed by a processor of an electronic device (101) (e.g., processor (120) of FIG. 1).
[0219] The electronic device (101) may, in operation 1001, determine whether it is in the multiple image recognition mode or the single image recognition mode. For example, the electronic device (101) may determine whether it is in the multiple image recognition mode or the single image recognition mode based on the analysis results of the captured image. For example, the electronic device (101) may determine whether it is in the multiple image recognition mode if the size of the text included in the image is less than a threshold size. For example, the electronic device (101) may determine whether it is in the single image recognition mode if the size of the text included in the image is greater than or equal to a threshold size. For example, the electronic device (101) may determine whether it is in the multiple image recognition mode based on determining that the texts included in the image fail to form a complete sentence. For example, the electronic device (101) may determine whether it is in the single image recognition mode based on determining that the texts included in the image form a complete sentence. Meanwhile, there is no limitation on the method for determining the multiple image recognition mode and the single image recognition mode.
[0220] When identified as a multiple image recognition mode, the electronic device (101) can, in operation 1003, identify multiple first images acquired while the electronic device (101) (or an external electronic device) is moved. In operation 1005, the electronic device (101) can perform a first task identified based on the analysis results of the multiple first images and the user's voice. When identified as a single image recognition mode, the electronic device (101) can, in operation 1007, identify a second image. In operation 1009, the electronic device (101) can perform a second task identified based on the analysis results of the second image and the user's voice. In one example, the electronic device (101) can identify multiple images identified by one of the multiple cameras based on the multiple image recognition mode, or can identify a single image identified by another of the multiple cameras based on the single image recognition mode. In another example, the electronic device (101) may be implemented to identify multiple images identified by one camera based on a multiple image recognition mode, or to identify a single image identified by one camera based on a single image recognition mode.
[0221] FIGS. 11A and 11B are diagrams illustrating recognition modes according to one embodiment.
[0222] Referring to FIG. 11A, an electronic device (1101) may be implemented in the form of, for example, glasses, and may include a camera capable of photographing the front of the electronic device (1101). The electronic device (1101) may be positioned relatively close to a subject (1110), for example. The electronic device (1101) may determine a recognition mode as a multiple image recognition mode based on, for example, a photographing result of the subject (1110). Based on the multiple image recognition mode, the electronic device (1101) may acquire multiple images (1121, 1122, 1123, 1124) of the subject (1110). For example, a user may rotate his / her head while wearing the electronic device (1101), and thus multiple images (1121, 1122, 1123, 1124) may be confirmed. The electronic device (1101) can acquire a user voice command, such as, for example, “Translate this.” The electronic device (1101) can provide a translation result for a text identified based on a plurality of images (1121, 1122, 1123, 1124) as a result of performing the task.
[0223] For example, referring to FIG. 11B, the electronic device (1101) may be positioned relatively far from, for example, a subject (1110). The electronic device (1101) may determine the recognition mode as a single image recognition mode based on, for example, a photographing result of the subject (1110). Based on the single image recognition mode, the electronic device (1101) may obtain single images (1131) of the subject (1110). The electronic device (1101) may obtain a user voice, for example, “Translate this.” The electronic device (1101) may provide a translation result for the text confirmed based on the single image (1131) as a result of performing the task. The camera used to acquire multiple images (1121, 1122, 1123, 1124) in FIG. 11a and the camera used to acquire a single image (1131) in FIG. 11b may be the same or different.
[0224] The electronic device (101) may include a communication device. The electronic device (101) may include one or more processors (120).
[0225] The electronic device (101) may include a memory (130) that stores instructions.
[0226] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to establish a communication connection between the wearable electronic device (201) and the electronic device (101) via the communication device.
[0227] The wearable electronic device (201) includes a first camera (281) and a second camera (282), and a focal length corresponding to the maximum resolution of the first camera (281) may be smaller than a focal length corresponding to the maximum resolution of the second camera (282).
[0228] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: identify a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moving based on the selection of the first camera (281) of the wearable electronic device (201) based on the analysis results of images captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice; and perform a first task identified based on the analysis results of the plurality of first images and the user's voice.
[0229] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to identify a second image acquired using the second camera (282) based on the selection of the second camera (282) based on the analysis results of the images captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice, and to perform a second task identified based on the analysis results of the second image and the user's voice.
[0230] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to perform the first task identified based on the analysis results of the plurality of first images and the user's voice, at least as a part of an operation of performing the first task identified based on the inference results of an AI model for data corresponding to the plurality of first images and the user's voice, and to perform the second task identified based on the inference results of the AI model for data corresponding to the second image and the user's voice, at least as a part of an operation of performing the second task identified based on the analysis results of the second image and the user's voice.
[0231] The above AI model may be configured to provide, as output data, data that causes the performance of at least some of a plurality of tasks including the first task and the second task.
[0232] The above AI model may be configured to include at least some of a plurality of sub-AI models associated with each of the plurality of tasks.
[0233] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to identify input data including information about a plurality of items set as input values of the AI model and the plurality of first images or the second images based on the analysis results of the user's voice.
[0234] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to identify information associated with TTS (text to speech) as at least a part of the information on the plurality of items, as at least a part of an operation of identifying input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice.
[0235] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to identify information associated with translation as at least a part of the information about the plurality of items, as at least a part of an operation of identifying input data including information about a plurality of items set as input values of the AI model based on the analysis results of the user's voice.
[0236] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to identify information associated with an image type as at least a part of the information on the plurality of items, as at least a part of an operation of identifying input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice.
[0237] The above instructions, when individually or collectively executed by the one or more processors (120), may cause information associated with a type of language to be identified as at least a part of the information on the plurality of items, as at least a part of the operation of identifying input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice.
[0238] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to output a voice corresponding to the inference result of the AI model, or to output a voice corresponding to the result of applying LLM to the inference result of the AI model, as at least part of an operation of performing the first task identified based on the inference result of the AI model.
[0239] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to output a voice corresponding to the inference result of the AI model, or to output a voice corresponding to the result of applying LLM to the inference result of the AI model, as at least part of an operation of performing the second task identified based on the inference result of the AI model.
[0240] If a region of interest (ROI) is identified based on an image acquired by the second camera (282), the second camera (282) may be selected, and if the ROI is not identified based on an image acquired by the second camera (282), the first camera (281) may be selected.
[0241] If a region of interest (ROI) is identified based on an image acquired by the first camera (281), the first camera (281) may be selected, and if the ROI is not identified based on an image acquired by the first camera (281), the second camera (282) may be selected.
[0242] Based on the analysis result of the user voice, the first camera (281) may be selected based on confirmation of a request to perform a proximity task, or the second camera (282) may be selected based on confirmation of a request to perform a distance task.
[0243] The method of operating the electronic device (101) may include an operation of establishing a communication connection between a wearable electronic device (201) and the electronic device (101).
[0244] The operating method of the electronic device (101) may include an operation of confirming a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moved based on the analysis result of an image captured by the first camera (281) of the wearable electronic device (201) and / or the second camera (282) of the wearable electronic device (201) and / or the analysis result of a user's voice, based on the selection of the first camera (281) of the wearable electronic device (201), and performing a first task confirmed based on the analysis result of the plurality of first images and the user's voice.
[0245] The method of operating the electronic device (101) may include an operation of confirming a second image acquired using the second camera (282) based on the selection of the second camera (282) based on the analysis result of an image captured by the first camera (281) and / or the second camera (282) and / or the analysis result of a user's voice, and performing a second task confirmed based on the analysis result of the second image and the user's voice.
[0246] One or more non-transitory computer-readable storage media may be provided storing one or more programs comprising computer-executable instructions.
[0247] The above instructions, when executed individually or collectively by one or more processors of the electronic device (101), may cause the electronic device (101) to perform operations.
[0248] The above operations may include operations for establishing a communication connection between a wearable electronic device (201) and the electronic device (101).
[0249] The above operations may include an operation of confirming a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moved based on the selection of the first camera (281) of the wearable electronic device (201) based on the analysis results of images captured by the first camera (281) of the wearable electronic device (201) and / or the second camera (282) of the wearable electronic device (201) and / or the analysis results of the user's voice, and performing a first task confirmed based on the analysis results of the plurality of first images and the user's voice.
[0250] The above operations may include operations of confirming a second image acquired using the second camera (282) based on the selection of the second camera (282) based on the analysis results of the images captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice, and performing a second task confirmed based on the analysis results of the second image and the user's voice.
[0251] A wearable electronic device (201) may include a first camera (281) and a second camera (282).
[0252] The focal length corresponding to the maximum resolution of the first camera (281) may be smaller than the focal length corresponding to the maximum resolution of the second camera (282).
[0253] The wearable electronic device (201) may include one or more processors (220).
[0254] The above wearable electronic device (201) may include a memory (230) that stores instructions.
[0255] The instructions, when individually or collectively executed by the one or more processors (220), may cause the wearable electronic device (201) to: identify a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moving based on the analysis results of images captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice, and perform a first task identified based on the analysis results of the plurality of first images and the user's voice.
[0256] The above instructions, when individually or collectively executed by the one or more processors (220), may cause the wearable electronic device (201) to: identify a second image acquired using the second camera (282) based on the selection of the second camera (282) based on the analysis results of the images captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice; and perform a second task identified based on the analysis results of the second image and the user's voice.
[0257] The instructions, when individually or collectively executed by the one or more processors (220), may cause the wearable electronic device (201) to perform, as at least part of an operation of performing the first task identified based on the analysis results of the plurality of first images and the user's voice, the first task identified based on the inference results of an AI model for data corresponding to the plurality of first images and the user's voice, and to perform, as at least part of an operation of performing the second task identified based on the analysis results of the second image and the user's voice, the second task identified based on the inference results of the AI model for data corresponding to the second image and the user's voice.
[0258] The above AI model may be configured to provide, as output data, data that causes the performance of at least some of a plurality of tasks including the first task and the second task.
[0259] The instructions, when individually or collectively executed by the one or more processors (220), may cause the wearable electronic device (201) to, as at least part of an operation of performing the first task identified based on the inference result of the AI model, output a voice corresponding to the inference result of the AI model, or output a voice corresponding to the result of applying LLM to the inference result of the AI model, and as at least part of an operation of performing the second task identified based on the inference result of the AI model, output a voice corresponding to the inference result of the AI model, or output a voice corresponding to the result of applying LLM to the inference result of the AI model.
[0260] If a region of interest (ROI) is identified based on an image acquired by the second camera (282), the second camera (282) may be selected, and if the ROI is not identified based on an image acquired by the second camera (282), the first camera (281) may be selected.
[0261] If a region of interest (ROI) is identified based on an image acquired by the first camera (281), the first camera (281) may be selected, and if the ROI is not identified based on an image acquired by the first camera (281), the second camera (282) may be selected.
[0262] Based on the analysis result of the user voice, the first camera (281) may be selected based on confirmation of a request to perform a proximity task, or the second camera (282) may be selected based on confirmation of a request to perform a distance task.
[0263] A method of operating a wearable electronic device (201) may include an operation of confirming a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moved based on a result of analyzing an image captured by the first camera (281) of the wearable electronic device (201) and / or the second camera (282) of the wearable electronic device (201) and / or a result of analyzing a user's voice, based on selection of the first camera (281) of the wearable electronic device (201), and performing a first task confirmed based on the result of analyzing the plurality of first images and the user's voice.
[0264] The operating method of the wearable electronic device (201) may include an operation of confirming a second image acquired using the second camera (282) based on the selection of the second camera (282) based on the analysis result of an image captured by the first camera (281) and / or the second camera (282) and / or the analysis result of a user's voice, and performing a second task confirmed based on the analysis result of the second image and the user's voice.
[0265] One or more non-transitory computer-readable storage media may be provided storing one or more programs comprising computer-executable instructions.
[0266] The above instructions, when executed individually or collectively by one or more processors of the electronic device (101), may cause the electronic device (101) to perform operations.
[0267] The above operations may include an operation of confirming a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moved based on the selection of the first camera (281) of the wearable electronic device (201) based on the analysis results of images captured by the first camera (281) of the wearable electronic device (201) and / or the second camera (282) of the wearable electronic device (201) and / or the analysis results of the user's voice, and performing a first task confirmed based on the analysis results of the plurality of first images and the user's voice.
[0268] The above operations may include operations of confirming a second image acquired using the second camera (282) based on the selection of the second camera (282) based on the analysis results of the images captured by the first camera (281) and / or the second camera (282) and / or the analysis results of the user's voice, and performing a second task confirmed based on the analysis results of the second image and the user's voice.
[0269] Electronic devices according to the embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.
[0270] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0271] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0272] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0273] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0274] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), Communication device (190); one or more processors (120); and Includes a memory (130) for storing instructions, The above instructions, when individually or collectively executed by one or more processors, cause the electronic device (101) to: Establishing a communication connection between a wearable electronic device (201) and the electronic device (101) through the communication device - the wearable electronic device (201) includes a first camera (281) and a second camera (282), and a focal length corresponding to the maximum resolution of the first camera (281) is smaller than a focal length corresponding to the maximum resolution of the second camera (282). Based on the analysis result of the image captured by the first camera (281) and / or the second camera (282), and / or the analysis result of the user's voice, based on which the first camera (281) of the wearable electronic device (201) is selected: Confirming a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moving, Performing a first task confirmed based on the analysis results of the plurality of first images and the user's voice, Based on the analysis results of the images captured by the first camera (281) and / or the second camera (282), and / or the analysis results of the user's voice, the second camera (282) is selected: Check the second image obtained using the second camera (282), An electronic device (101) that causes a second task to be performed based on the analysis results of the second image and the user's voice.
2. In paragraph 1, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: At least as part of an operation of performing the first task, which is identified based on the analysis results of the plurality of first images and the user's voice, Performing the first task confirmed based on the inference result of the AI model for the data corresponding to the plurality of first images and the user's voice, At least as part of the operation of performing the second task identified based on the analysis results of the second image and the user's voice, causing the second task identified based on the inference results of the AI model for data corresponding to the second image and the user's voice to be performed, The AI model is set to provide, as output data, data that causes the performance of at least some of a plurality of tasks including the first task and the second task, An electronic device (101) wherein the AI model is set to include at least a portion of a plurality of sub-AI models associated with each of the plurality of tasks.
3. In paragraph 1 or 2, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: An electronic device (101) that causes input data including information on a plurality of items set as input values of the AI model and the plurality of first images or the second images to be confirmed based on the analysis results of the user's voice.
4. In any one of paragraphs 1 to 3, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: As at least a part of an operation of checking input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice, An electronic device (101) that causes information associated with TTS (text to speech) to be identified as at least a portion of information about the plurality of items.
5. In any one of paragraphs 1 to 4, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: As at least a part of an operation of checking input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice, An electronic device (101) that causes information associated with a translation to be identified as at least some of the information about said plurality of items.
6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: As at least a part of an operation of checking input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice, An electronic device (101) that causes information associated with an image type to be identified as at least a portion of information about said plurality of items.
7. In any one of paragraphs 1 to 6, The above instructions, when individually or collectively executed by one or more processors (120), are at least part of an operation of verifying input data including information on a plurality of items set as input values of the AI model based on the analysis results of the user's voice. An electronic device (101) that causes information associated with a type of language to be identified as at least some of the information for said plurality of items.
8. In any one of paragraphs 1 to 7, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: As at least part of the operation of performing the first task confirmed based on the inference result of the AI model, outputting a voice corresponding to the inference result of the AI model, or outputting a voice corresponding to the result of applying LLM to the inference result of the AI model, An electronic device (101) that causes the device to output a voice corresponding to the inference result of the AI model, or to output a voice corresponding to the result of applying LLM to the inference result of the AI model, at least as part of an operation of performing the second task confirmed based on the inference result of the AI model.
9. In any one of paragraphs 1 to 8, An electronic device (101) in which the second camera (282) is selected when a region of interest (ROI) is identified based on an image acquired by the second camera (282), and the first camera (281) is selected when the ROI is not identified based on an image acquired by the second camera (282).
10. In any one of paragraphs 1 to 9, An electronic device (101) in which the first camera (281) is selected when a region of interest (ROI) is identified based on an image acquired by the first camera (281), and the second camera (282) is selected when the ROI is not identified based on an image acquired by the first camera (281).
11. In any one of paragraphs 1 to 10, An electronic device (101) in which the first camera (281) is selected based on a request for performing a proximity task confirmed based on the analysis result of the user voice, or the second camera (282) is selected based on a request for performing a distance task confirmed.
12. In a wearable electronic device (201), A first camera (281) and a second camera (282), wherein the focal length corresponding to the maximum resolution of the first camera (281) is smaller than the focal length corresponding to the maximum resolution of the second camera (282), one or more processors; and Includes a memory (230) for storing instructions, The above instructions, when individually or collectively executed by one or more processors (220), cause the wearable electronic device (201) to: Based on the analysis result of the image captured by the first camera (281) and / or the second camera (282), and / or the analysis result of the user's voice, based on which the first camera (281) of the wearable electronic device (201) is selected: Confirming a plurality of first images acquired using the first camera (281) while the wearable electronic device (201) is moving, Performing a first task confirmed based on the analysis results of the plurality of first images and the user's voice, Based on the analysis results of the images captured by the first camera (281) and / or the second camera (282), and / or the analysis results of the user's voice, the second camera (282) is selected: Check the second image obtained using the second camera (282), A wearable electronic device (201) that causes a second task to be performed based on the analysis results of the second image and the user's voice.
13. In paragraph 12, The above instructions, when individually or collectively executed by one or more processors (220), cause the wearable electronic device (201) to: At least as part of an operation of performing the first task, which is identified based on the analysis results of the plurality of first images and the user's voice, Performing the first task confirmed based on the inference result of the AI model for the data corresponding to the plurality of first images and the user's voice, At least as part of the operation of performing the second task identified based on the analysis results of the second image and the user's voice, causing the second task identified based on the inference results of the AI model for data corresponding to the second image and the user's voice to be performed. A wearable electronic device (201) configured to provide, as output data, data causing the performance of at least some of a plurality of tasks including the first task and the second task.
14. In either of paragraphs 12 or 13, The above instructions, when individually or collectively executed by one or more processors (220), cause the wearable electronic device (201) to: As at least part of the operation of performing the first task confirmed based on the inference result of the AI model, outputting a voice corresponding to the inference result of the AI model, or outputting a voice corresponding to the result of applying LLM to the inference result of the AI model, A wearable electronic device (201) that causes, as at least part of an operation of performing the second task identified based on the inference result of the AI model, to output a voice corresponding to the inference result of the AI model, or to output a voice corresponding to the result of applying LLM to the inference result of the AI model.
15. In any one of paragraphs 12 to 14, If a region of interest (ROI) is identified based on the image acquired by the second camera (282), the second camera (282) is selected, and if the ROI is not identified based on the image acquired by the second camera (282), the first camera (281) is selected. If a region of interest (ROI) is identified based on the image acquired by the first camera (281), the first camera (281) is selected, and if the ROI is not identified based on the image acquired by the first camera (281), the second camera (282) is selected. A wearable electronic device (201) in which the first camera (281) is selected based on a request for performing a proximity task confirmed based on the analysis result of the user voice, or the second camera (282) is selected based on a request for performing a distance task confirmed.
Citation Information
Patent Citations
Electronic apparatus for determining whether a card payment has been fraudulently canceled and operating method thereof
KR1020250157178A
Autonomous computing and telecommunications head-up displays glasses
US20140266988A1
User interface to select field of view of a camera in a smart glass
US20230031871A1
Wearable electronic device controlling camera module and operation method thereof
WO2024043438A1
Wearable electronic device including camera and operation method thereof
WO2024101747A1