Electronic device including multiple cameras and operation method of electronic device
The electronic device uses AI models to dynamically activate cameras based on environmental and user inputs, improving user experience through personalized and efficient camera operation.
Patent Information
- Application Number
- PCT/KR2025/008289
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-06-17
- Publication Date
- 2025-12-26
AI Technical Summary
Existing electronic devices with multiple cameras lack efficient methods for dynamically activating cameras based on environmental and user inputs, leading to suboptimal user experience and limited personalized functionality.
An electronic device equipped with multiple cameras, a processor, and AI models that analyze image and audio signals to determine which camera to activate, providing personalized and dynamic camera operation based on environmental and user inputs.
Enhances user convenience by automatically selecting appropriate camera modes and functionalities, offering personalized black box functions and efficient content collection.
Smart Images

Figure KR2025008289_26122025_PF_FP_ABST
Abstract
Description
Electronic device including multiple cameras and method of operating the electronic device
[0001] This document relates to an electronic device including a plurality of cameras and a method of operating the electronic device, for example, an electronic device that takes images through a plurality of cameras and a method of operating the electronic device.
[0002] This document relates to an AI-related electronic device capable of monitoring the surrounding environment through multiple input devices including a camera and audio, and a method for determining the activation of specific input / output devices based on trigger information from the user and the surrounding environment.
[0003] As the use of cameras in electronic devices such as smartphones has become more active in recent years, smartphones that include various types of cameras, such as ultra-wide-angle cameras, wide-angle cameras, and telephoto cameras, are being developed.
[0004] A technology has been proposed to automatically identify an object and apply camera option functions when shooting using multiple cameras in an electronic device including multiple cameras, and further advancement of this technology is required.
[0005] In one embodiment, the AI assistant application may use at least one sensor to obtain information about the subject's surroundings. For example, the AI assistant application may use a location sensor to obtain information such as the subject's current location and whether the subject is indoors or outdoors.
[0006] An AI assistant application can obtain the illuminance values surrounding a subject using a light sensor. Furthermore, the AI assistant application can use the light sensor to determine whether the surroundings of the subject are currently day or night. The AI assistant application can also obtain temperature information surrounding the subject using a temperature sensor, and humidity information surrounding the subject using a humidity sensor. The AI assistant application can also obtain environmental information using multiple image sensors (or multiple cameras).
[0007] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0008] The present invention provides a method for monitoring the surrounding environment through multiple input devices and determining the activation of specific input / output devices based on trigger information from the user and the surrounding environment. This improves user convenience, automatically collects content, and provides a personalized black box function.
[0009] An electronic device may include a plurality of cameras, a processor operatively connected to the plurality of cameras, and a memory storing instructions. The instructions, when executed by the processor, may cause the electronic device to perform object recognition on an image within an image based on receiving an image signal captured by at least some of the plurality of cameras, provide information about the captured image signal and the image on which object recognition was performed as input to an artificial intelligence (AI) model, analyze an audio signal based on receiving the audio signal, provide an analysis result and the audio signal as input to the AI model, determine a camera to be maintained in an activated state among the plurality of cameras using the AI model, and control an operation mode of the determined camera.
[0010] An operating method of an electronic device may include an operation of performing object recognition on an image in an image based on receiving an image signal captured by at least some of a plurality of cameras; an operation of providing information on the captured image signal and the image on which object recognition was performed as inputs to an AI model; an operation of analyzing an audio signal based on receiving an audio signal; an operation of providing an analysis result and the audio signal as inputs to the AI model; an operation of determining a camera to be maintained in an activated state among the plurality of cameras using the AI model, and an operation of controlling an operating mode of the determined camera.
[0011] According to various embodiments, the electronic device can determine whether a specific camera among a plurality of cameras can be used as a multimodal image input means, and thereby transmit data to be processed together with the user's voice to an AI (artificial intelligence) model.
[0012] According to various embodiments, the electronic device can determine whether a specific camera among a plurality of cameras can be utilized as a multimodal image input means and contribute to dynamic activation of the camera to provide convenience to the user.
[0013] According to various embodiments, the electronic device may determine whether a specific camera among a plurality of cameras can be utilized as a multimodal image input means, and may provide automatic content collection and personalized black box functions.
[0014] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0015] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0016] FIG. 2 is a block diagram illustrating an integrated intelligence system according to one embodiment.
[0017] FIG. 3 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment.
[0018] FIG. 4A is a block diagram of a generative artificial intelligence system according to one embodiment.
[0019] FIG. 4b illustrates a prompt to be input to an AI model according to one embodiment.
[0020] FIG. 5 is a block diagram of a prompt generation system according to one embodiment.
[0021] Figure 6 is a block diagram of an electronic device according to one embodiment.
[0022] FIG. 7 is a block diagram illustrating the operation of a processor and a large multimodal-model (LMM) according to various embodiments.
[0023] Figure 8 is a flowchart illustrating the operation of an electronic device periodically monitoring a user's voice.
[0024] Figure 9 is a flowchart illustrating the operation of an electronic device that automatically records a dangerous moment by monitoring both sound signals and sensor signals.
[0025] Figure 10 is a flowchart illustrating the operation of determining whether to activate the camera when an electronic device simultaneously detects a user's voice and a user input (e.g., a button input).
[0026] Fig. 11 is a flowchart illustrating a method for controlling camera operation of an electronic device according to one embodiment.
[0027] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0028] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0029] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0030] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0031] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0032] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0033] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0034] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0035] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0036] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0037] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0038] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0039] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0040] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0041] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0042] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0043] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0044] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0045] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0046] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0047] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0048] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0049] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0050] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0051] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0052] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0053] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0054] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0055] FIG. 2 is a block diagram illustrating an integrated intelligence system according to various embodiments.
[0056] Referring to FIG. 2, according to one embodiment, the integrated intelligence system may include an electronic device (210) (e.g., electronic device (101) of FIG. 1), an intelligent server (230) (e.g., server (108) of FIG. 1), and a service server (250) (e.g., server (108) of FIG. 1).
[0057] According to one embodiment, the electronic device (210) may be a terminal device (or electronic device) that can connect to the Internet, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a TV, white goods, a wearable device, an HMD, or a smart speaker.
[0058] According to the illustrated embodiment, the electronic device (210) may include a communication interface (213) (e.g., the interface (177) of FIG. 1), a microphone (212) (e.g., the input module (150) of FIG. 1), a speaker (216) (e.g., the audio output module (155) of FIG. 1), a display module (211) (e.g., the display module (160) of FIG. 1), a memory (215) (e.g., the memory (130) of FIG. 1), or a processor (214) (e.g., the processor (120) of FIG. 1). The components listed above may be operatively or electrically connected to each other. The electronic device (210) may include at least some of the configurations and / or functions of the electronic device (101) of FIG. 1.
[0059] In one embodiment, the communication interface (213) may be configured to connect to an external device and transmit and receive data. In one embodiment, the microphone (212) may receive sound (e.g., user speech) and convert it into an electrical signal. In one embodiment, the speaker (216) may output the electrical signal as sound (e.g., voice).
[0060] In one embodiment, the display module (211) may be configured to display an image or video. In one embodiment, the display module (211) may also display a graphical user interface (GUI) of a running app (or application program). In one embodiment, the display module (211) may receive a touch input via a touch sensor. For example, the display module (211) may receive a text input via a touch sensor in an on-screen keyboard area displayed within the display module (211).
[0061] According to one embodiment, the memory (215) may store a client module (218), a software development kit (SDK) (217), and a plurality of apps (219a, 219b). The client module (218) and the SDK (217) may constitute a framework (or solution program) for performing general-purpose functions. In addition, the client module (218) or the SDK (217) may constitute a framework for processing user input (e.g., voice input, text input, touch input).
[0062] According to one embodiment, the plurality of apps (219a, 219b) stored in the memory (215) may be programs for performing a specified function. According to one embodiment, the plurality of apps may include a first app (219a) and a second app (219b). According to one embodiment, each of the plurality of apps (219a, 219b) may include a plurality of operations for performing a specified function. For example, the apps (219a, 219b) may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the plurality of apps (219a, 219b) may be executed by the processor (214) to sequentially execute at least some of the plurality of operations.
[0063] According to one embodiment, the processor (214) can control the overall operation of the electronic device (210). For example, the processor (214) can be electrically connected to a communication interface (213), a microphone (212), a speaker (216), and a display module (211) to perform a specified operation.
[0064] According to one embodiment, the processor (214) may also execute a program stored in the memory (215) to perform a designated function. For example, the processor (214) may execute at least one of the client module (218) or the SDK (217) to perform the following operations for processing user input. The processor (214) may control the operations of a plurality of apps (219a, 219b), for example, through the SDK (217). The following operations described as operations of the client module (218) or the SDK (217) may be operations executed by the processor (214).
[0065] According to one embodiment, the client module (218) can receive user input. For example, the client module (218) can receive a voice signal corresponding to a user utterance detected through the microphone (212). Alternatively, the client module (218) can receive a touch input detected through the display module (211). Alternatively, the client module (218) can receive a text input detected through a keyboard or a visual keyboard. In addition, the client module (218) can receive various forms of user input detected through an input module included in the electronic device (210) or an input module connected to the electronic device (210). The client module (218) can transmit the received user input to the intelligent server (230). The client module (218) can transmit status information of the electronic device (210) together with the received user input to the intelligent server (230). The status information can be, for example, execution status information of an app.
[0066] In one embodiment, the client module (218) may receive a result corresponding to the received user input. For example, the client module (218) may receive a result corresponding to the received user input if the intelligent server (230) can produce a result corresponding to the received user input. The client module (218) may display the received result on the display module (211). Additionally, the client module (218) may output the received result as audio through the speaker (216).
[0067] According to one embodiment, the client module (218) can receive a plan corresponding to the received user input. The client module (218) can display the results of executing multiple operations of the app according to the plan on the display module (211). For example, the client module (218) can sequentially display the results of executing multiple operations on the display module (211) and output audio through the speaker (216). The electronic device (210) can, for another example, display only some results of executing multiple operations (e.g., the result of the last operation) on the display module (211) and output audio through the speaker (216).
[0068] In one embodiment, the client module (218) may receive a request from the intelligent server (230) to obtain information necessary to produce a result corresponding to the voice input. In one embodiment, the client module (218) may transmit the necessary information to the intelligent server (230) in response to the request.
[0069] According to one embodiment, the client module (218) may transmit result information of executing multiple operations according to a plan to the intelligent server (230). The intelligent server (230) may use the result information to confirm that the received user input has been processed correctly.
[0070] In one embodiment, the client module (218) may include a voice recognition module. In one embodiment, the client module (218) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (218) may execute an intelligent app that processes voice inputs to perform organic actions based on a specified input (e.g., "Wake up!").
[0071] According to one embodiment, the intelligent server (230) can receive information related to a user voice input from an electronic device (210) via a communication network. According to one embodiment, the intelligent server (230) can convert data related to the received voice input into text data. According to one embodiment, the intelligent server (230) can generate a plan for performing a task corresponding to the user voice input based on the text data.
[0072] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)), or a combination of the above or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI system may select at least one plan from a plurality of predefined plans.
[0073] According to one embodiment, the intelligent server (230) may transmit the results according to the generated plan to the electronic device (210), or transmit the generated plan to the electronic device (210). According to one embodiment, the electronic device (210) may display the results according to the plan on the display module (211). According to one embodiment, the electronic device (210) may display the results of executing an operation according to the plan on the display module (211).
[0074] According to one embodiment, the intelligent server (230) may include a front end (231), a natural language platform (232), a capsule database (238), an execution engine (233), an end user interface (234), a management platform (235), a big data platform (236), or an analytic platform (237).
[0075] According to one embodiment, the front end (231) can receive user input from the electronic device (210). The front end (231) can transmit a response corresponding to the user input.
[0076] According to one embodiment, the natural language platform (232) may include an automatic speech recognition module (ASR module) (232a), a natural language understanding module (NLU module) (232b), a planner module (232c), a natural language generator module (NLG module) (232d), or a text to speech module (TTS module) (232e).
[0077] According to one embodiment, the automatic speech recognition module (232a) can convert voice input received from the electronic device (210) into text data. According to one embodiment, the natural language understanding module (232b) can use the text data of the voice input to determine the user's intent. For example, the natural language understanding module (232b) can perform syntactic analysis or semantic analysis on user input in the form of text data to determine the user's intent. According to one embodiment, the natural language understanding module (232b) can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from the voice input, and can match the meaning of the determined words to the intent to determine the user's intent. The natural language understanding module (223b) can obtain intent information corresponding to the user's utterance. The intent information can be information indicating the user's intent determined by interpreting text data. The intent information can include information indicating an action or function that the user intends to execute using the device.
[0078] According to one embodiment, the planner module (232c) can generate a plan using the intent and parameters determined by the natural language understanding module (232b). According to one embodiment, the planner module (232c) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (232c) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (232c) can determine parameters necessary to execute the determined plurality of operations or result values output by the execution of the plurality of operations. The parameters and the result values can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and a plurality of concepts determined by the user's intent. The planner module (232c) can determine the relationships between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (232c) can determine the execution order of a plurality of actions based on the user's intention based on a plurality of concepts. In other words, the planner module (232c) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (232c) can generate a plan including association information (e.g., ontology) between the plurality of actions and the plurality of concepts. The planner module (232c) can generate the plan using information stored in a capsule database that stores a set of relationships between concepts and actions.
[0079] According to one embodiment, the natural language generation module (232d) can convert specified information into text format. The information converted into text format may be in the form of natural language speech. According to one embodiment, the text-to-speech module (232e) can convert text-to-speech information into speech information.
[0080] According to one embodiment, some or all of the functions of the natural language platform (232) may also be implemented in the electronic device (210).
[0081] The capsule database can store information about the relationships between multiple concepts and actions corresponding to multiple domains. According to one embodiment, the capsule can include multiple action objects (or action information) and concept objects (or concept information) included in the plan. According to one embodiment, the capsule database can store multiple capsules in the form of a concept action network (CAN). According to one embodiment, the multiple capsules can be stored in a function registry included in the capsule database.
[0082] The capsule database may include a strategy registry that stores strategy information necessary for determining a plan corresponding to a user input. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to the user input. According to one embodiment, the capsule database may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a given situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database may include a layout registry that stores layout information of information output through the electronic device (210). According to one embodiment, the capsule database may include a vocabulary registry that stores vocabulary information included in the capsule information. According to one embodiment, the capsule database may include a dialog registry that stores information on dialogue (or interaction) with the user. The capsule database may update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies that determine plans. The developer tool may include a dialog editor for creating conversations with users.The developer tool may include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on the currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database may also be implemented within the electronic device (210).
[0083] In one embodiment, the execution engine (233) can use the generated plan to produce a result. The end user interface (234) can transmit the produced result to the electronic device (210). Accordingly, the electronic device (210) can receive the result and provide the received result to the user. In one embodiment, the management platform (235) can manage information used in the intelligent server (230). In one embodiment, the big data platform (236) can collect user data. In one embodiment, the analysis platform (237) can manage the quality of service (QoS) of the intelligent server (230). For example, the analysis platform (237) can manage the components and processing speed (or efficiency) of the intelligent server (230).
[0084] According to one embodiment, the service server (250) may provide a service (e.g., food ordering or hotel reservation) specified to the electronic device (210). According to one embodiment, the service server (250) may be a server operated by a third party. According to one embodiment, the service server (250) may provide information for generating a plan corresponding to the received voice input to the intelligent server (230). The provided information may be stored in a capsule database. In addition, the service server (250) may provide result information according to the plan to the intelligent server (230). The service server (250) may include a plurality of service providers (e.g., CP Service A (251), CP Service B (252), CP Service C (253)), and each of the service providers (251, 252, 253) may provide a function for a domain associated with each capsule stored in the capsule database (238) of the intelligent server (230).
[0085] In the integrated intelligence system described above, the electronic device (210) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.
[0086] According to one embodiment, the electronic device (210) may provide a voice recognition service through an intelligent app (or voice recognition app) stored within the device. In this case, for example, the electronic device (210) may recognize a user utterance or voice input received through the microphone (212) and provide the user with a service corresponding to the recognized voice input.
[0087] According to one embodiment, the electronic device (210) may perform a designated operation based on the received voice input, either alone or together with the intelligent server (230) and / or the service server (250). For example, the electronic device (210) may execute an app corresponding to the received voice input and perform a designated operation through the executed app.
[0088] According to one embodiment, when an electronic device (210) provides a service together with an intelligent server (230) and / or a service server (250), the electronic device (210) may detect a user's speech using the microphone (212) and generate a signal (or voice data) corresponding to the detected user's speech. The electronic device (210) may transmit the voice data to the intelligent server (230) via a network (240) using a communication interface (213).
[0089] In one embodiment, an intelligent server (230) may generate a plan for performing a task corresponding to a voice input received from an electronic device (210), or a result of performing an operation according to the plan, in response to the voice input. The plan may include, for example, a plurality of operations for performing a task corresponding to a user's voice input, and a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations, or result values output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and the plurality of concepts.
[0090] According to one embodiment, the electronic device (210) can receive the response using the communication interface (213). The electronic device (210) can output a voice signal generated within the electronic device (210) to the outside using the speaker (216), or can output an image generated within the electronic device (210) to the outside using the display module (211).
[0091] Although FIG. 2 describes an example in which voice recognition of user input received from an electronic device (210), natural language understanding and generation, and calculation of results using a plan are performed on an intelligent server (230), the various embodiments of the present document are not limited thereto. For example, at least some components of the intelligent server (230) (e.g., natural language platform (232), execution engine (233), capsule database (238)) may be embedded in the electronic device (210) (or the electronic device (101) of FIG. 1), so that the operations may be performed by the electronic device (210).
[0092] FIG. 3 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to various embodiments.
[0093] According to one embodiment, a capsule database (e.g., capsule database (238) of FIG. 2) of an intelligent server (e.g., intelligent server (230) of FIG. 2) may store capsules in the form of a CAN (concept action network) (300). The capsule database may store operations for processing tasks corresponding to a user's voice input and parameters necessary for the operations in the form of a CAN (concept action network).
[0094] According to one embodiment, the capsule database may store a plurality of capsules (capsule (A) (310), capsule (B) (320)) corresponding to each of a plurality of domains (e.g., applications). According to one embodiment, one capsule (e.g., capsule (A) (310)) may correspond to one domain (e.g., location (geo), application). In addition, one capsule may correspond to at least one service provider (e.g., CP 1 (331) or CP 2 (332)) for performing a function for a domain related to the capsule. According to one embodiment, one capsule may include at least one operation (350) and at least one concept (360) for performing a specified function.
[0095] In one embodiment, a natural language platform (e.g., the natural language platform (232) of FIG. 2) can generate a plan for performing a task corresponding to a received speech input using capsules stored in a capsule database. For example, a planner module of the natural language platform (e.g., the planner module (232c) of FIG. 2) can generate a plan using capsules stored in a capsule database. For example, a plan can be generated using actions (311, 313) and concepts (312, 314) of capsule A (310) and actions (321) and concepts (322) of capsule B (320).
[0096] FIG. 4A is a block diagram of a generative artificial intelligence system according to one embodiment.
[0097] Referring to FIG. 4A, a generative artificial intelligence system (400) may include a generative AI model (450), an AI framework (440), a user query / response interface (410), an application / service component (430), and a knowledge repository (420). The AI framework (440) may include a prompt design component (442), an API / plugin management component (444), and an output modification component (446).
[0098] According to one embodiment, a user query / response interface (410) may receive a user's input. The user's input may be in the form of natural language, images, and / or videos. In addition, context information may also be transmitted when the user's input is transmitted. The context information may include various additional information at the time of the user's input. For example, the context information may include information related to the user or the electronic device, such as information about the application the user is currently using or information about the user's location. In addition, the user's input may also be in a form that mixes the aforementioned natural language, images, sounds, and context information. In addition, the user's input may also be in a non-natural language form, such as selecting a menu.
[0099] In one embodiment, the user question / response interface (410) may output the results of the generative artificial intelligence system (400) to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user.
[0100] According to one embodiment, the AI framework (440) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query.
[0101] In one embodiment, user input received from the user question / response interface (410) may be transmitted to a prompt design component (442). The prompt design component (442) may be used to generate prompts suitable for inputting the user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (442) may be an AI component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design component (442) may access a knowledge repository (420) containing user preference data, a prompt library, and prompt examples based on the user input to obtain and generate prompts, and may transmit the generated prompts to the LLM or LMM.
[0102] In one embodiment, the API / plug-in management component (444) may communicate with external information when there is a request for additional information when passing user input as input to the generative model. The API / plug-in management component (444) may establish a channel for communicating with the outside of the AI interface through the API, and may enable access to various data sources through the established channel. In addition, the API / plug-in management component (444) may request an action through the API that ultimately performs the user input, rather than an intermediate result, when the application or service needs to perform the action. Information obtained from the outside may be used to generate a prompt in the prompt design component (442) together with the user input, or may be passed as input to the generative model.
[0103] In one embodiment, the output modification component (446) (or refiner component) can fine-tune the output from the generative model. For example, the output modification component (446) can verify that the content generated through the LLM and / or LMM is not irrelevant, does not contain biased content, or does not contain harmful content. In addition, the output modification component (446) can determine to what extent the content matches the result desired by the user and, if necessary, can perform additional processing. The output modification component (446) can additionally configure and provide the user with hints to avoid undesired output.
[0104] According to one embodiment, a generative AI model (450) may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. The generative AI model (450) may include a model that generates images and / or a model that generates languages. Representative models for generating images include a generative adversarial network (GAN) and a variational auto encoder (VAE), and examples include a Diffusion-based generative model that uses a VAE and a Transformer structure. A model for generating languages is a model that is trained to statistically output the most appropriate output based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there are also LMMs (large multimodal models) that can recognize various types of data input such as text, images, and voices and generate new data corresponding thereto.
[0105] FIG. 4b illustrates a prompt to be input to an AI model according to one embodiment.
[0106] A generative AI model (e.g., the generative AI model (450) of FIG. 4a) may be an artificial intelligence model that learns various data to generate new information and sentences. In this document, a generative AI model may also be referred to as an AI model or a large language model (LLM).
[0107] In one embodiment, a user may input user input, including a query, to obtain desired information through an AI model. For example, the user may input the query via text input using a keypad or keyboard, or via voice input using a microphone. User input to the AI model may be input as a prompt. A prompt is a command for the AI model to generate a response, and may serve to guide the AI model to perform a desired action by the user.
[0108] In one embodiment, to obtain the desired response from an AI model, a user may need to create prompts that the AI model can understand and operate on. For example, parameters required for an AI model's response include information such as reader level, response length, respondent's perspective, output format, and output language. If these parameters are entered into the prompt through user input, the AI model can clearly understand the user's query and provide the desired response.
[0109] Referring to FIG. 4b, a prompt (490) created based on user input may include information related to the question, "What is the cutest cat?" and parameters, such as "reader level, length, perspective, format, answer me in English." When the prompt (490) with various parameters created in this way is input into an AI model (or LLM), the AI model can output a response in English, in a form of dialogue, within 500 characters at an elementary school level, from a marketer's perspective, as defined in the prompt.
[0110] Since AI models can be used by a variety of users, the ability / level of writing prompts for AI models may vary from person to person, making it difficult to write the optimal prompt for the AI model you want to use.
[0111] Below, we describe various embodiments for generating prompts to be input into an AI model from user input based on pre-stored prompt templates and user data, even when the user inputs only brief information.
[0112] FIG. 5 is a block diagram of a prompt generation system according to one embodiment.
[0113] According to one embodiment, the prompt generation system may include a voice assistant (590), a prompt manager (500), and an AI model (550).
[0114] According to one embodiment, a voice assistant (590) may include AI (artificial intelligence)-based hardware and / or software modules that understand the content of a voice input by a user and process actions according to the user's request. For example, the voice assistant (590) may analyze a voice signal using speech recognition technology (automatic speech recognition, ASR) to convert it into text, interpret the content of the converted text, and process various actions such as executing and controlling an application requested by the user, controlling device settings, and searching and providing information based on the interpretation of the text.
[0115] According to one embodiment, a user input to a voice assistant (590) may include a question (or query) requesting the AI model (550) to perform a task and respond, and the voice assistant (590) may transmit the user input including the question to the AI model (550), obtain a response to the question from the AI model (550), and provide the response to the question to the user.
[0116] According to one embodiment, the AI model (550) may include a generative artificial intelligence model (550) (e.g., the generative AI model (450) of FIG. 4A) or a large language model (LLM), which is an artificial intelligence model that learns various data to generate new information and sentences. For example, the AI model (550) may include an on-device AI model implemented on an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (210) of FIG. 2) and a server AI model implemented by an external server (e.g., the intelligent server (230) of FIG. 2), and the server AI model may be implemented and operated on a server of a manufacturer of the electronic device, or may be implemented and operated by a third party. The electronic device may select an AI model (550) set by the user among the available AI models (550) and transmit the user input, and / or select an AI model (550) that requests a response based on the type or content of the query of the user input and transmit the user input.
[0117] According to one embodiment, the prompt manager (500) can perform various operations to generate a prompt (or secondary prompt) from user input (or primary prompt). For example, the prompt manager (500) can receive a user input including a query for the AI model (550), verify the user input based on the number of parameters included in the user input, and if the user input fails the verification, generate a prompt based on a prompt template and pre-stored user data. The prompt generated by the prompt manager (500) includes various information compared to the initial user input, and since the included information is based on user data, when input to the AI model (550), the AI model (550) can provide more accurate and detailed information desired by the user as a response.
[0118] According to one embodiment, the prompt manager (500) may be implemented on an electronic device to generate a prompt in response to a user input and then provide the prompt to the AI model (550), or may be implemented on a server-type AI model to generate a prompt in response to a user input inputted to the AI model (550), and / or may be implemented on a separate server device to receive a user input over a network and then generate a prompt and then provide the prompt to the AI model (550).
[0119] Hereinafter, the prompt manager (500) that evaluates user input and generates a new prompt is described as being implemented on an electronic device that obtains user input; however, the various embodiments of this document are not limited thereto, and the embodiments described below may be provided by various devices such as a server device other than the electronic device.
[0120] Figure 6 is a block diagram of an electronic device according to one embodiment.
[0121] Referring to FIG. 6, the electronic device (600) may include a communication module (640), a display (630), a microphone (650), a processor (610), and a memory (620). In various embodiments of the present document, some of the illustrated components may be omitted or replaced. The electronic device (600) may include at least some of the components and / or functions of the electronic device (101) of FIG. 1 and / or the electronic device (210) of FIG. 2. At least some of the components of the illustrated (or not illustrated) electronic device (600) may be operatively, functionally, and / or electrically connected to each other.
[0122] In one embodiment, the hardwares of the electronic device (600) being operatively, functionally and / or electrically connected may mean that a direct connection, or an indirect connection, is established between the hardwares, either wired or wireless, such that the second hardware is controlled by the first hardware among the hardwares.
[0123] According to one embodiment, the display (630) can display various images provided from the processor (610). For example, the display (630) can be implemented as any one of a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a micro electro mechanical systems (MEMS) display, or an electronic paper display, but is not limited thereto. The display (630) can be configured as a touch screen that detects touch and / or proximity touch (or hovering) input using a part of a user's body (e.g., a finger) or an input device (e.g., a stylus pen). The display (630) can include at least some of the configurations and / or functions of the display module (160) of FIG. 1.
[0124] According to one embodiment, the display (630) may provide various screens provided by a voice assistant (e.g., the voice assistant (590) of FIG. 5) in the form of an interactive UI (user interface). For example, the display (630) may display a response (e.g., text, a photo) including a task execution result of an AI model (e.g., the AI model (550) of FIG. 5, the generative AI model (450) of FIG. 4A) in response to a user input, a prompt template (or sample prompt) selected in response to the user input, and / or a prompt generated by a prompt manager (e.g., the prompt manager (500) of FIG. 5). Alternatively, the electronic device (600) may provide a screen in the form of displaying a specific object on the screen or displaying a special effect in addition to the interactive UI.
[0125] According to one embodiment, the microphone (650) can pick up external sounds, such as a user's voice, and convert them into voice signals, which are digital data. According to one embodiment, the electronic device (600) may include a microphone in a part of a housing (not shown), or may receive voice signals picked up from an external microphone connected wired or wirelessly. According to one embodiment, the voice signal acquired from the microphone (650) is transmitted to the processor (610), and the processor (610) (or voice assistant) can convert the voice signal into text information through automatic speech recognition (ASR) and perform an operation corresponding to the text information.
[0126] According to one embodiment, the communication module (640) may support wireless communication with an external device using cellular wireless communication (e.g., 4G long term evolution (LTE), 5G new radio (NR)) and / or short-range wireless communication (e.g., WiFi). For example, the electronic device (600) may use the communication module (640) to communicate with an external server (e.g., the intelligent server (230) of FIG. 2) that provides a voice assistant function through a network. The communication module (640) may include at least some of the configurations and / or functions of the communication module (190) of FIG. 1 and / or the communication interface (213) of FIG. 2.
[0127] According to one embodiment, the memory (620) may temporarily or permanently store various data, including volatile memory and non-volatile memory. The memory (620) may include at least a portion of the configuration and / or function of the memory (130) of FIG. 1 and / or the memory (215) of FIG. 2, and may store the program (140) of FIG. 1. The memory (620) may store various applications (e.g., the first app (219a) and the second app (219b) of FIG. 2) and program modules supporting intelligent services (e.g., the client module (218) of FIG. 2).
[0128] According to one embodiment, the memory (620) may store various instructions that may be performed by the processor (610). Such instructions may include control commands such as arithmetic and logical operations, data movement, and / or input / output that may be recognized by the processor (610).
[0129] According to one embodiment, the processor (610) may be configured as one or more processors capable of performing calculations or data processing related to control and / or communication of each component of the electronic device (600). The processor (610) may include at least some of the configurations and / or functions of the processor (120) of FIG. 1 and / or the processor (214) of FIG. 2.
[0130] According to one embodiment, there is no limitation to the computational and data processing functions that the processor (610) may implement on the electronic device (600). However, this document will describe various embodiments that analyze ambient noise and / or user input to determine which camera to operate among multiple cameras. The operations of the processor (610) described below may be performed by loading instructions stored in the memory (620).
[0131] In this document, the description that the processor (610) can perform a certain operation (or function, work, task) may be interpreted to mean substantially the same as that an instruction (or command, computer program) that causes the electronic device (600) (or the processor (610)) to perform the corresponding operation is stored in the memory (620) (e.g., non-volatile memory, storage). In addition, the description that the processor (610) can perform a certain operation may be interpreted to mean substantially the same as that at least one processor (610) can perform the corresponding operation.
[0132] According to one embodiment, the memory (620) may store user data related to a user of the electronic device (600). For example, the user data may include personal information such as the user's name, gender, age, and occupation, location information, call / message records, schedules, application usage history, a user profile set by the user (e.g., name, gender, age, and occupation), and / or device information of the electronic device (600). According to one embodiment, the processor (610) may store a predetermined type of data among data obtained from each application or data input by the user in the memory (620) as user data, and may update the user data when the user data changes (e.g., when the location of the electronic device (600) moves).
[0133] According to one embodiment, the memory (620) may store various prompt templates. Here, the prompt templates define the formats of text provided as input when using the AI model and may include examples of information that enable the AI model to provide higher quality responses. The prompt templates may also be referred to as sample prompts. According to another embodiment, the prompt templates may be stored in an external server (e.g., the intelligent server (230) of FIG. 2), and the electronic device (600) may receive at least one of the prompt templates from the external server via the communication module (640).
[0134] Instructions for performing the operations of the electronic device (600) (or processor (610)) described above may be stored on a computer-readable recording medium. The recording medium may be tangible and non-transitory. The recording medium may store one or more computer programs including the instructions.
[0135] The electronic device (600) according to this document can monitor the surrounding environment through multiple input devices, including a camera and audio. The electronic device (600) according to this document can input information about the user and the surrounding environment into an AI model, and determine the activation of input and output devices based on the output results of the AI model.
[0136] The electronic device (600) may include a signal processing unit that processes voice, image, and composite signals, an LMM composite signal processing unit that determines an active function based on the collected signals, and an active function determination unit. The signal processing unit, the LMM composite signal processing unit, and the function determination unit may be separate components within the electronic device (600) or may be components included in the processor (610).
[0137] The electronic device (600) may include an image signal processing unit. The electronic device (600) may collect and analyze an image signal using at least one image sensor under the control of the processor (610).
[0138] Additionally, the electronic device (600) may include a voice signal processing unit. The electronic device (600) may acquire and analyze voice and sound signals surrounding the electronic device (600) using at least one sound receiving device. The sound receiving device may include, for example, any one of a speaker-oriented microphone, an external sound-oriented microphone, or an omnidirectional microphone.
[0139] The electronic device (600) can periodically monitor the user's speech and activate some or all of the cameras that are highly correlated with the user's speech received as an input signal.
[0140] The electronic device (600) can automatically save a preview image of the moment by taking an instant shot when a predefined keyword is detected in the user's speech. The electronic device (600) can evaluate the correlation between the preview images of all cameras and the user's speech and save the optimal preview image.
[0141] The electronic device (600) can automatically perform either image or video capture through an activatable camera when a loud noise is continuously picked up by an omnidirectional microphone for the purpose of collecting sounds in the surrounding environment.
[0142] The electronic device (600) may include a complex signal processing unit. The electronic device (600) may analyze speed values obtained using GPS information and / or a gyro sensor to determine whether to activate the camera. The electronic device (600) may include a server information receiving unit for obtaining information from an external server. The information obtained from the external server may include, for example, weather information.
[0143] The electronic device (600) can determine, through a pressure sensor during the preview activation phase, whether some cameras are in contact with the user or a specific object and thus cannot be photographed. If the electronic device (600) determines that the camera is in a state where photographing is impossible, the electronic device (600) may not activate the camera even when an input signal including a user's speech is detected.
[0144] The electronic device (600) can acquire information using an LMM composite signal processing unit, an audio signal processing unit, an image signal processing unit, and a composite signal processing unit. The electronic device (600) can analyze the acquired information to determine whether to activate some components including a camera.
[0145] The electronic device (600) can determine priorities among various input signals. For example, the electronic device (600) can receive various input signals including at least one of a voice signal, an image, a GPS signal, or a signal for a gyro sensor. When the movement speed of the electronic device (600) detected by the GPS signal or the signal for the gyro sensor exceeds a specified level, the electronic device (600) can determine a relatively higher priority for the GPS signal or the signal for the gyro sensor than the priority for the user's voice signal. When the movement speed is below the specified level, the electronic device (600) can determine that the priorities for the voice signal, the image, the GPS signal, or the signal for the gyro sensor are all equal.
[0146] The electronic device (600) may adjust the monitoring cycle of at least one input device among the microphone and the camera when a specific predefined situation is detected. The specific predefined situation may be defined by any one of the following factors: location, speed, or weather. The electronic device (600) may reduce power consumption by adjusting the monitoring cycle. For example, the electronic device (600) may reduce power consumption by increasing the monitoring cycle in a stable situation. A stable situation may mean, for example, a situation in which the moving speed of the electronic device (600) is below a specified level, or a situation in which the moving speed of surrounding objects is determined to be below a specified level based on an analysis of images of objects captured in the surroundings. Alternatively, the electronic device (600) may define a situation in which a noise exceeding a certain decibel level is detected as an urgent situation and adjust the monitoring cycle of the input device to a relatively short time.
[0147] FIG. 7 is a block diagram illustrating the operation of a processor and a large multimodal-model (LMM) according to various embodiments.
[0148] According to FIG. 7, a processor (e.g., processor (610) of FIG. 6, processor (120) of FIG. 1) may include a server information receiving unit (702) for receiving server information, a motion input unit (704) for receiving a motion by a user, a voice signal receiving unit (712), an audio signal receiving unit (714), and an image signal input unit (722). The server information receiving unit (702), the motion input unit (704) for receiving a motion by a user, the voice signal receiving unit (712), the audio signal receiving unit (714), and the image signal input unit (722) may be separate components within an electronic device (e.g., electronic device (600) of FIG. 6) or may be included in the processor (610). The operations described through the server information receiving unit (702), the operation input unit (704) for receiving an operation by the user, the voice signal receiving unit (712), the sound signal receiving unit (714), and the image signal input unit (722) may also be performed by the processor (610).
[0149] According to one embodiment, the processor (610) can process a composite signal using a server information receiving unit (702) and a motion input unit (704). The processor (610) can process a voice signal using a voice signal receiving unit (712) and an audio signal receiving unit (714). The processor (610) can process a video signal using a video signal input unit (722).
[0150] The processor (610) may perform at least one of the following operations: confirming context, recognizing an object of interest, recognizing a shooting mode, analyzing risk, or extracting keywords using the voice signal processing unit (720). The processor (610) may determine the meaning of a user's utterance by confirming the context.
[0151] The processor (610) can obtain and analyze preview images of multiple cameras using the image signal processing unit (730) to recognize objects in the preview images.
[0152] The processor (610) can analyze and process data including at least one of server information, user motion input (e.g., button input), voice signal, or video signal using a large multimodal-model (LMM) (740).
[0153] A large multimodal model (LMM) (740) can refer to an artificial intelligence model trained using large amounts of multimodal data. A large multimodal model (LMM) (740) can comprehensively process various types of data, including text, images, and audio. Compared to machine learning models that use only single-modal data, an LMM (740) can utilize richer information, enabling it to solve more accurate and diverse problems. For example, an LMM (740) can generate descriptions of objects within an image using both images and text, or perform speech recognition and natural language processing using audio and text.
[0154] In one embodiment, the processor (610) may determine that no additional information is needed based on whether a particular camera is activated in the large multimodal-model (LMM) (740). Conversely, the processor (610) may determine that additional information is needed based on whether a particular camera is not activated in the large multimodal-model (LMM) (740).
[0155] According to one embodiment, the processor (610) can receive at least one of a video or image using the video signal input unit (722).
[0156] The processor (610) can acquire ambient sound information using a voice signal processing unit (720), a user-oriented microphone, and an omnidirectional microphone. The processor (610) can obtain signals regarding user actions (e.g., button input) and server information (e.g., weather information) using a motion input unit (704) and a server information receiving unit (702).
[0157] The processor (610) can process each piece of information collected using an image signal processing unit, an audio signal processing unit, and a composite signal processing unit.
[0158] The processor (610) can process signals and server information about the user's actions using a complex signal processing unit. The processor (610) can convert the processed information into metadata input and provide it to the LMM (740). The LMM (740) can determine which camera to activate based on the metadata input.
[0159] The processor (610) can analyze the acquired image using an image signal processing unit. The processor (610) can perform object recognition within the image and provide image tagging information as input to the LMM (740). The processor (610) can also provide the acquired image itself as input to the LMM (740) without any separate processing.
[0160] The processor (610) can analyze the acquired sound signal using the voice signal processing unit (720). The processor (610) can perform at least one operation among object of interest, shooting mode inference, risk analysis, and keyword extraction. The processor (610) can provide the result obtained by performing the operation as input to the LMM (740). Alternatively, the processor (610) can provide the acquired sound signal itself as input to the LMM (740) without any separate processing.
[0161] In one embodiment, the processor (610) may determine whether to activate a camera based on the output of the LMM (740). If the processor (610) cannot determine which camera to activate from the LMM (740), the processor (610) may determine that insufficient information has been acquired and request additional information about the audio, video, or composite signal. If the processor (610) determines that periodic acquisition of audio, video, or composite signals is necessary, the processor (610) may request signal acquisition through periodic monitoring.
[0162] Figure 8 is a flowchart illustrating the operation of an electronic device periodically monitoring a user's voice.
[0163] The operations described through FIG. 8 may be implemented based on instructions that may be stored in a computer recording medium or memory (e.g., memory (130) of FIG. 1). The illustrated method (600) may be executed by the electronic device described above through FIGS. 1 to 6 (e.g., electronic device (101) of FIG. 1, electronic device (600) of FIG. 6), and the technical features described above will be omitted below. The order of each operation of FIG. 8 may be changed, some operations may be omitted, and some operations may be performed simultaneously.
[0164] In operation 802, an electronic device (e.g., electronic device (600) of FIG. 6, electronic device (101) of FIG. 1) can detect and analyze user speech under the control of a processor (e.g., processor (610) of FIG. 6, processor (120) of FIG. 1).
[0165] In operation 810, the electronic device (600) may perform analysis related to camera activation. The electronic device (600) may check the context to confirm the user's intent. The context may include the intent of the user's utterance. To confirm the context, the electronic device (600) may detect keywords related to camera activation set in advance in the user's utterance. If no keywords related to camera activation are detected, the electronic device (600) may wait until the user utters again or terminate the operation. Keywords related to camera activation may include, for example, words related to taking pictures or videos.
[0166] In operation 812, the electronic device (600) may temporarily activate the cameras based on the detection of a keyword related to camera activation. For example, the electronic device (600) may determine that a keyword related to camera activation has been detected based on the user utterance "Let's take a picture" and may temporarily activate the cameras.
[0167] In operation 814, the electronic device (600) activates the first camera and acquires a preview image. The first camera may refer to any one of a plurality of cameras. While the first camera and the second camera are described as examples herein, the number of cameras included in the electronic device (600) is not limited thereto.
[0168] In step 816, the electronic device (600) may activate the second camera and acquire a preview image. Steps 814 and 816 need not be sequential and may occur simultaneously. The positions of the first camera and the second camera may be different. For example, the first camera may be positioned on the front of the electronic device (600) relative to the display. The second camera may be positioned on the rear of the electronic device (600) relative to the display. Alternatively, the cameras may be positioned on the side of the electronic device (600).
[0169] In Fig. 8, the situation is explained assuming two cameras, but the number of cameras is not limited to this.
[0170] In operation 818, the electronic device (600) can analyze a preview image to recognize an object within the image. The preview image may refer to an image captured by the first camera and the second camera.
[0171] In operation 820, the electronic device (600) may perform a semantic visual search. Semantic visual search may refer to a technology for searching an image by utilizing semantic information including at least one of an object, scene, or relationship within the image. Semantic visual search may convert previously extracted semantic information into a high-dimensional feature representation in vector form. The electronic device (600) may quantify the semantic information within the image through semantic visual search and utilize it for search. The electronic device (600) may also use an AI model to understand a user's search query and match it with semantic information within the image to retrieve highly relevant images. The AI model may determine relevance based on semantic similarity rather than simple keyword matching. For example, if a user enters a search term such as "boat rowing in the sea," the electronic device (600) may understand this and match it with semantic information such as objects, scenes, and relationships within the image.
[0172] In one embodiment, the electronic device (600) may provide search results for images related to the matching result "a boat rowing in the sea," for example, images of "a boat on the sea" or "a boat with a rower on it." The AI model can determine relevance not only by matching keywords containing either "sea," "boat," or "oar," but also by considering the relationship between objects in the image and the context of the situation. The electronic device (600) may utilize semantic visual search to provide search results that better match the meaning desired by the user.
[0173] In operation 830, the electronic device (600) may determine which camera to activate. The electronic device (600) may use semantic visual search to determine the camera most relevant to the object in the image.
[0174] In operation 832, the electronic device (600) may keep the determined camera activated based on whether the camera to be activated has been determined. Alternatively, the electronic device (600) may terminate the operation or request additional information about the user's speech based on whether the camera to be activated has not been determined in operation 830.
[0175] In one embodiment, the electronic device (600) may transmit at least one of image data received from an activated camera or previous user speech data to the AI model. The electronic device (600) may receive a response from the AI model. The AI model may provide the response to the user using a voice assistant (e.g., voice assistant (590) of FIG. 5 ).
[0176] According to one embodiment, a voice assistant (e.g., voice assistant (590) of FIG. 5) may include AI (artificial intelligence)-based hardware and / or software modules that understand the content of a voice input by a user and process actions according to a user's request. For example, the voice assistant (590) may analyze a voice signal using speech recognition technology (automatic speck recognition, ASR) to convert it into text, interpret the content of the converted text, and process various actions such as executing and controlling an application requested by the user, controlling device settings, and searching and providing information based on the interpretation of the text.
[0177] According to one embodiment, a user input to a voice assistant (590) may include a question (or query) requesting the AI model (550) to perform a task and respond, and the voice assistant (590) may transmit the user input including the question to the AI model (550), obtain a response to the question from the AI model (550), and provide the response to the question to the user.
[0178] According to one embodiment, the electronic device (600) can periodically monitor the user's speech and perform speech analysis. The electronic device (600) can detect voice keywords corresponding to the user's object of interest through the analysis of the user's speech. Alternatively, the electronic device (600) can obtain information about the shooting mode in the camera activation stage through the analysis of the user's speech. If the electronic device (600) determines that it has obtained a voice keyword highly correlated with camera activation, it can temporarily activate multiple cameras (or one camera). The electronic device (600) can obtain preview information from the mounted cameras. The electronic device (600) can analyze each acquired preview image (or one preview image). The electronic device (600) can obtain information about object recognition within the image and the movement speed of the recognized object through the analysis of the preview information. The electronic device (600) can provide the analyzed preview information and the speech information from which the voice keyword is detected as input to the AI model. The electronic device (600) can perform semantic visual search using an AI model (e.g., LMM (740)). Based on the output of the semantic visual search, the electronic device (600) can determine a preview image with a high correlation with the utterance information. The electronic device (600) can identify the camera that captured the preview image with a high correlation. The electronic device (600) can determine to maintain the activated state for the identified camera.
[0179] Figure 9 is a flowchart illustrating the operation of an electronic device that automatically records a dangerous moment by monitoring both sound signals and sensor signals.
[0180] The operations described through FIG. 9 can be implemented based on instructions that can be stored in a computer recording medium or memory (e.g., memory (130) of FIG. 1). The illustrated method (600) can be executed by the electronic device described above through FIGS. 1 to 6 (e.g., electronic device (101) of FIG. 1, electronic device (600) of FIG. 6), and the technical features described above will be omitted below. The order of each operation of FIG. 9 can be changed, some operations can be omitted, and some operations can be performed simultaneously.
[0181] In operation 902, an electronic device (e.g., electronic device (600) of FIG. 6, electronic device (101) of FIG. 1) can recognize a sound source generating object and analyze the risk level under the control of a processor (e.g., processor (610) of FIG. 6, processor (120) of FIG. 1).
[0182] The sound source object may include, for example, a car. The electronic device (600) may detect a car within the preview image. Alternatively, the electronic device (600) may analyze ambient noise to determine whether the noise is related to a car and determine whether a car is located within the vicinity of the electronic device (600).
[0183] Additionally, the electronic device (600) can analyze the vehicle's moving speed by analyzing preview images at time intervals. If a vehicle is detected and its speed exceeds a specified level, the electronic device (600) can determine that the risk level is high. For example, the risk level can be determined to be high in proportion to the speed of vehicles around the electronic device (600). The electronic device (600) can determine that a dangerous situation exists based on the risk level exceeding a specified level.
[0184] In operation 910, the electronic device (600) can determine whether the user of the electronic device (600) is in a risk situation based on the risk analysis result.
[0185] In action 912, the electronic device (600) may temporarily activate the cameras based on what is determined to be a dangerous situation.
[0186] In operation 914, the electronic device (600) can activate the first camera and obtain a preview image. In operation 916, the electronic device (600) can activate the second camera and obtain a preview image. Operations 914 and 916 do not need to be sequential and can occur simultaneously. In FIG. 9, a situation with two cameras is described, but the number of cameras is not limited to this.
[0187] In operation 918, the electronic device (600) can analyze the preview image to recognize an object within the image.
[0188] In operation 920, the electronic device (600) can perform a semantic visual search. The semantic visual search has been described above in FIG. 8.
[0189] In operation 930, the electronic device (600) may determine whether a hazardous object is identified within the audio or video. The hazardous object may include, for example, a car. The electronic device (600) may detect a car within the preview image. Alternatively, the electronic device (600) may analyze ambient noise to determine whether the noise is related to a car and determine whether a car is located within the vicinity of the electronic device (600).
[0190] Additionally, the electronic device (600) can analyze the vehicle's moving speed by analyzing preview images at time intervals. If a vehicle is detected and its speed exceeds a specified level, the electronic device (600) can determine that the risk level is high. For example, the risk level can be determined to be high in proportion to the speed of vehicles around the electronic device (600). The electronic device (600) can determine that a dangerous situation exists based on the risk level exceeding a specified level.
[0191] In operation 932, the electronic device (600) may keep the camera that captured the dangerous object (e.g., a car) activated. Alternatively, in operation 930, the electronic device (600) may terminate the operation based on the absence of a dangerous object or may collect additional sound source information using a sensor.
[0192] According to one embodiment, the electronic device (600) may include an omnidirectional microphone capable of receiving ambient sounds. When a loud noise is detected in the user's surroundings, the electronic device (600) may analyze sensor information (e.g., GPS and / or gyro sensor) and the received sound information received from the sensor information input unit to perform a risk analysis.
[0193] The electronic device (600) can obtain information about the object that generated the sound source by analyzing sound information. The electronic device (600) can determine the level of risk based on the analyzed information. If the level of risk exceeds a preset level, the electronic device (600) can temporarily activate the installed cameras to obtain a preview. The electronic device (600) can analyze the obtained preview image using an AI model and recognize an object in the image. In addition, the electronic device (600) can perform analysis on the received sound and image using the AI model. If the electronic device (600) determines that a dangerous object (e.g., a car) exists in the image information, the electronic device (600) can activate the camera that has obtained the preview image related to the dangerous object. The electronic device (600) can activate the camera to obtain an image and / or a video.
[0194] According to one embodiment, the electronic device (600) may determine whether a sound exceeding a certain decibel is the voice of the user of the electronic device (600). The electronic device (600) may analyze the ambient noise based on the determination that the sound is not the voice of the user of the electronic device (600) but ambient noise. The electronic device (600) may determine that the user of the electronic device (600) is in a dangerous state based on the determination that the noise level exceeds a specified level or the change in the noise level per unit time exceeds a specified level. The electronic device (600) may control at least some of the cameras among the plurality of cameras to activate and acquire a preview image based on the determination that the user is in a dangerous state.
[0195] In one embodiment, the electronic device (600) may transmit at least one of image data received from an activated camera or previous user speech data to the AI model. The electronic device (600) may receive a response from the AI model. The AI model may provide the response to the user using a voice assistant (e.g., voice assistant (590) of FIG. 5 ).
[0196] According to one embodiment, a user input to a voice assistant (590) may include a question (or query) requesting the AI model (550) to perform a task and respond, and the voice assistant (590) may transmit the user input including the question to the AI model (550), obtain a response to the question from the AI model (550), and provide the response to the question to the user.
[0197] Figure 10 is a flowchart illustrating the operation of determining whether to activate the camera when an electronic device simultaneously detects a user's voice and a user input (e.g., a button input).
[0198] The operations described through FIG. 10 may be implemented based on instructions that may be stored in a computer recording medium or memory (e.g., memory (130) of FIG. 1). The illustrated method (600) may be executed by the electronic device described above through FIGS. 1 to 6 (e.g., electronic device (101) of FIG. 1, electronic device (600) of FIG. 6), and the technical features described above will be omitted below. The order of each operation of FIG. 10 may be changed, some operations may be omitted, and some operations may be performed simultaneously.
[0199] In operation 1010, an electronic device (e.g., electronic device (600) of FIG. 6, electronic device (101) of FIG. 1) may detect a user input (e.g., a button input) under the control of a processor (e.g., processor (610) of FIG. 6, processor (120) of FIG. 1). The user input may include a touch input to a display in addition to a button input.
[0200] In operation 1012, the electronic device (600) may temporarily activate the cameras based on detecting a user input (e.g., a button press).
[0201] In operation 1014, the electronic device (600) can activate the first camera and obtain a preview image. In operation 1016, the electronic device (600) can activate the second camera and obtain a preview image. Operations 1014 and 1016 do not need to be sequential and can occur simultaneously. In FIG. 10, a situation with two cameras is described, but the number of cameras is not limited to this.
[0202] In operation 1018, the electronic device (600) can analyze the preview image to recognize an object within the image.
[0203] In operation 1020, the electronic device (600) can perform a semantic visual search. The semantic visual search has been described above in FIG. 8.
[0204] In operation 1030, the electronic device (600) can detect and analyze user speech. Operations 1010 and 1030 are not sequential and may occur simultaneously or one operation may be performed first.
[0205] In operation 1032, the electronic device (600) may verify the context related to camera activation. The electronic device (600) may determine the meaning of the user utterance by verifying the context. The context may include the intent of the user utterance. To verify the context, the electronic device (600) may detect keywords related to camera activation set in advance in the user utterance. If keywords related to camera activation are not detected, the electronic device (600) may wait until the user utters again or terminate the operation.
[0206] In operation 1034, the electronic device (600) can identify keywords for objects of interest preset by the user. The objects of interest preset by the user may vary depending on the settings. The objects of interest may be preset by the user or may be preset during the development phase of the electronic device (600).
[0207] In operation 1036, the electronic device (600) may check a keyword for a shooting mode. The shooting mode may include, for example, an image shooting mode or a video shooting mode.
[0208] The electronic device (600) can identify keywords for a preset object of interest and provide related information as input to the AI model. The electronic device (600) can identify keywords for a shooting mode and provide related information as input to the AI model.
[0209] The electronic device (600) may perform a semantic visual search using an AI model in operation 1020 and determine which camera to activate in operation 1022.
[0210] At subsequent operation 1024, the electronic device (600) can determine the shooting mode of the activated camera.
[0211] According to one embodiment, the electronic device (600) may determine that a user's voice utterance and a user input (e.g., a button press) are detected together as a starting point for taking a photo or video. The electronic device (600) may determine the camera in the direction in which the touch input is detected or the camera in the opposite direction as the priority activation target.
[0212] The electronic device (600) can acquire multiple preview images by activating multiple cameras when a user input is detected on the display or a specific button. The electronic device (600) can determine the most relevant preview image among the multiple preview images based on speech information acquired by analyzing the user's voice. The electronic device (600) can maintain the activation status of the camera that captured the determined preview image.
[0213] In one embodiment, the electronic device (600) may transmit at least one of image data received from an activated camera or previous user speech data to the AI model. The electronic device (600) may receive a response from the AI model. The AI model may provide the response to the user using a voice assistant (e.g., voice assistant (590) of FIG. 5 ).
[0214] According to one embodiment, a user input to a voice assistant (590) may include a question (or query) requesting the AI model (550) to perform a task and respond, and the voice assistant (590) may transmit the user input including the question to the AI model (550), obtain a response to the question from the AI model (550), and provide the response to the question to the user.
[0215] Fig. 11 is a flowchart illustrating a method for controlling camera operation of an electronic device according to one embodiment.
[0216] The operations described through FIG. 11 may be implemented based on instructions that may be stored in a computer recording medium or memory (e.g., memory (130) of FIG. 1). The illustrated method (600) may be executed by the electronic device described above through FIGS. 1 to 10 (e.g., electronic device (101) of FIG. 1, electronic device (600) of FIG. 6), and the technical features described above will be omitted below. The order of each operation of FIG. 11 may be changed, some operations may be omitted, and some operations may be performed simultaneously.
[0217] In operation 1110, the electronic device (600) may perform object recognition on an image in a video under the control of a processor (e.g., processor (610) of FIG. 6, processor (120) of FIG. 1) and provide it as input to an AI model.
[0218] According to one embodiment, the electronic device (600) may perform object recognition on an image in an image based on receiving an image signal captured by at least some of the plurality of cameras. The electronic device (600) may perform object recognition on an image in an image using a machine learning model or a large multimodal-model (LMM) (e.g., LMM (740) of FIG. 7). The LMM (740) may refer to an artificial intelligence model learned by utilizing a large amount of multimodal data. The LMM (740) may generate a description of an object in an image by using an image and text together, or may perform speech recognition and natural language processing by utilizing audio and text.
[0219] According to one embodiment, the electronic device (600) can provide information about the captured image signal and the image on which object recognition was performed as input to the AI model.
[0220] In operation 1120, the electronic device (600) may analyze a voice signal and provide the analysis result as input to an AI model. The electronic device (600) may analyze the voice signal using an LMM (740). The electronic device (600) may provide the analysis result and the original data of the voice signal as input to the AI model (e.g., the LMM (740)).
[0221] In operation 1130, the electronic device (600) can use the AI model to determine which camera to keep activated and control the operating mode of the activated camera.
[0222] According to one embodiment, the electronic device (600) may determine whether a sound exceeding a certain decibel is the voice of the user of the electronic device (600). The electronic device (600) may analyze the ambient noise based on the determination that the sound is not the voice of the user of the electronic device (600) but ambient noise. The electronic device (600) may determine that the user of the electronic device (600) is in a dangerous state based on the determination that the noise level exceeds a specified level or the change in the noise level per unit time exceeds a specified level. The electronic device (600) may control at least some of the cameras among the plurality of cameras to activate and acquire a preview image based on the determination that the user is in a dangerous state.
[0223] According to one embodiment, the electronic device (600) can detect an object causing the noise among preview images acquired using a plurality of cameras. The electronic device (600) can control the activation of at least some of the cameras that captured the object causing the noise.
[0224] According to one embodiment, the electronic device (600) can analyze an object on the acquired preview image. The electronic device (600) can activate the camera that acquired the preview image to capture an image and / or video based on whether the object analyzed on the acquired preview image matches an object determined as a noise source from the analysis information on ambient noise.
[0225] According to one embodiment, the electronic device (600) can determine the movement speed of the electronic device based on at least one of a global positioning system (GPS), a gyro sensor, and an acceleration sensor. The electronic device (600) can determine that the user of the electronic device (600) is in a dangerous state based on whether the movement speed of the electronic device (600) exceeds a specified level or whether an increase in the movement speed of the electronic device (600) exceeds a specified level. Based on whether the user is determined to be in a dangerous state, the electronic device (600) can control at least some of the cameras among the plurality of cameras to activate and acquire a preview image.
[0226] According to one embodiment, the electronic device (600) can detect the movement of external objects using a plurality of cameras. The electronic device (600) can determine that the user of the electronic device (600) is in a dangerous state based on detecting an object among the external objects moving at a specified speed or an increase in the speed of the external object per unit time exceeding a specified level. Based on determining that the user is in a dangerous state, the electronic device (600) can control activation of at least some of the plurality of cameras to acquire a preview image.
[0227] In one embodiment, the electronic device (600) may identify a camera among multiple cameras that detects an object moving at a specified speed. Alternatively, the electronic device (600) may identify a camera that detects an object whose speed increase per unit time exceeds a specified level. The electronic device (600) may control the activation of the identified cameras.
[0228] According to one embodiment, the electronic device (600) may use an AI model to output information about a camera to be maintained in an activated state among the plurality of cameras. Based on the inability of the AI model to determine a camera to be maintained in an activated state among the plurality of cameras, the electronic device (600) may determine that additional information related to a video signal, an audio signal, and a sensor signal for at least one of a global positioning system (GPS), a gyro sensor, or an acceleration sensor is required. The electronic device (600) may additionally obtain a signal for at least one of the video signal, the audio signal, or the sensor signal.
[0229] According to one embodiment, the electronic device (600) may activate a microphone for receiving a user's voice based on detecting a user input on the display or a specific button. The electronic device (600) may activate multiple cameras to acquire multiple preview images at the time the user input on the display or a specific button is detected. The electronic device (600) may determine a preview image with the highest relevance among the multiple preview images based on speech information acquired by analyzing the user's voice. The electronic device (600) may maintain an activated state for the camera that captured the determined preview image.
[0230] According to one embodiment, the electronic device (600) can activate a plurality of cameras based on detection of a user's voice to acquire a plurality of preview images. The electronic device (600) can compare tag information present in the plurality of preview images with tag information acquired from the user's voice. The electronic device (600) can store the tag information in the memory (130) based on a match between the tag information present in the plurality of preview images and the tag information acquired from the user's voice. The electronic device (600) can maintain an activated state for a camera that captures a preview image having matching tag information.
[0231] The embodiments of this document disclosed in this specification and drawings are merely specific examples to easily explain the technical contents according to the embodiments of this document and to help understand the embodiments of this document, and are not intended to limit the scope of the embodiments of this document. Therefore, the scope of one embodiment of this document should be interpreted to include all changes or modified forms derived based on the technical idea of one embodiment of this document, in addition to the embodiments disclosed herein.
Claims
1. In an electronic device (101), Multiple cameras (180); At least one processor (120) operatively connected to the plurality of cameras; and Includes a memory (130) for storing instructions, The above instructions, when executed by the at least one processor, cause the electronic device to Perform object recognition on an image within a video based on receiving video signals captured by at least some of a plurality of cameras, Provide information about the captured video signal and the image on which object recognition was performed as input to the AI (artificial intelligence) model, Analyzing the voice signal based on receiving the voice signal, The analysis results and the above voice signal are provided as input to the AI model, Using the above AI model, determine which camera among the multiple cameras will remain active, An electronic device that controls the operating mode of a determined camera.
2. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to Determine whether a sound exceeding a certain decibel is received and whether it is the voice of a user of the electronic device, Analyze the ambient noise based on the determination that it is ambient noise and not the user's voice of the electronic device, The user of the electronic device is determined to be in a dangerous state based on whether the noise level exceeds a specified level or the change in the noise level per unit time exceeds a specified level, An electronic device that controls at least some of the plurality of cameras to be activated to obtain a preview image based on the determination that the user is in a dangerous state.
3. In paragraph 2, The above instructions, when executed by the at least one processor, cause the electronic device to Detecting an object causing the noise among the preview images acquired using the above multiple cameras, An electronic device that controls at least some of the cameras that are photographed by an object that causes the above noise to be activated.
4. In paragraph 2, The above instructions, when executed by the at least one processor, cause the electronic device to Analyze the objects in the acquired preview image, An electronic device that activates a camera that has acquired a preview image to capture an image and / or video based on whether an object analyzed in the acquired preview image matches an object determined to be a noise source from analysis information about ambient noise.
5. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to Determining the moving speed of the electronic device based on at least one of a global positioning system (GPS), a gyro sensor, or an acceleration sensor; If the movement speed of the electronic device exceeds the specified level, Determining that the user of the electronic device is in a dangerous state based on an increase in the movement speed of the electronic device exceeding a specified level, An electronic device that controls at least some of the plurality of cameras to be activated to obtain a preview image based on the determination that the user is in a dangerous state.
6. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to Detecting the movement of external objects using multiple cameras, The user of the electronic device is determined to be in a dangerous state based on the detection of an object among external objects moving at a speed exceeding a specified speed or an increase in the speed of an external object per unit time exceeding a specified level, An electronic device that controls at least some of the plurality of cameras to be activated to obtain a preview image based on the determination that the user is in a dangerous state.
7. In paragraph 6, The above instructions, when executed by the at least one processor, cause the electronic device to Among the above multiple cameras, identify a camera that detects an object moving at a speed exceeding a specified speed or an object whose speed increase per unit time exceeds a specified level, Activate the confirmed cameras, An electronic device that controls data captured from a camera determined to be in an activated state and user voice data to be provided as input to the AI model.
8. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to Using the above AI model, information about a camera to be maintained in an activated state among the plurality of cameras is provided as output. Based on the inability to determine which camera among the plurality of cameras will remain active in the above AI model, it is determined that additional information related to a video signal, an audio signal, and a sensor signal for at least one of a global positioning system (GPS), a gyro sensor, or an acceleration sensor is required, An electronic device that controls to additionally acquire at least one signal among a video signal, an audio signal, or a sensor signal.
9. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to Activate the microphone to receive the user's voice based on detecting user input on the display or a specific button; Acquire multiple preview images by activating multiple cameras at the time when user input for the display or a specific button is detected, Determine the most relevant preview image among multiple preview images based on the speech information obtained by analyzing the user's voice, An electronic device that controls the camera that captured the determined preview image to remain active.
10. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to Activate multiple cameras based on the detection of user voice to acquire multiple preview images, Compare the tag information present in the above multiple preview images with the tag information obtained from the user's voice, Store the tag information in the memory based on the tag information existing in the plurality of preview images and the tag information obtained from the user's voice matching, An electronic device that controls a camera to remain active while capturing preview images having matching tag information.
11. In the method of operating an electronic device, An operation of performing object recognition on an image within an image based on receiving an image signal captured through at least some of a plurality of cameras; An action that provides information about the captured video signal and the image on which object recognition has been performed as input to the AI model; An operation of analyzing a voice signal based on receiving the voice signal; An operation of providing the analysis results and the above voice signal as input to the AI model; An operation of determining a camera to be maintained in an activated state among the plurality of cameras using the AI model; and A method comprising an operation for controlling the operating mode of a determined camera.
12. In paragraph 11, An action to determine whether a sound exceeding a certain decibel is received and whether it is the voice of a user of the electronic device; An action of analyzing ambient noise based on determining that the ambient noise is not the user's voice of the electronic device; An action for determining that the user of the electronic device is in a dangerous state based on whether the noise level exceeds a specified level or whether the change in the noise level per unit time exceeds a specified level; and A method further comprising an action of controlling at least some of the plurality of cameras to be activated to obtain a preview image based on the determination that the user is in a risk state.
13. In paragraph 12, An operation of detecting an object that causes the noise among the preview images acquired using the plurality of cameras; and A method further comprising activating at least some of the cameras that capture the object causing the noise.
14. In paragraph 12, An action to analyze an object on an acquired preview image; and A method further comprising an action of controlling a camera that acquired the preview image to capture an image and / or video based on whether an object analyzed in the acquired preview image matches an object determined as a noise source from analysis information about ambient noise.
15. In paragraph 11, An operation of determining a moving speed of the electronic device based on at least one of a global positioning system (GPS), a gyro sensor, or an acceleration sensor; An action of determining that a user of the electronic device is in a dangerous state based on whether the movement speed of the electronic device exceeds a specified level or whether an increase in the movement speed of the electronic device exceeds a specified level; and A method further comprising an action of controlling at least some of the plurality of cameras to be activated to obtain a preview image based on the determination that the user is in a risk state.
Citation Information
Patent Citations
Method for emergency diagnosis having nonlinguistic speech recognition function and apparatus thereof
KR1020170018140A
Emergency rescue request method using smart device
KR1020170067251A
Ice maker and refrigerator
KR1020230018503A
A method for monitoring the environment of a parked car including an asynchronous camera.
KR102564994B1
Personal monitoring apparatus and methods
US20230413014A1