Image recognition method and electronic device for supporting same

The electronic device addresses resource wastage by using speech and image data to identify multiple objects, enhanced by motion data from a secondary device, for efficient information provision.

WO2026005331A1PCT designated stage Publication Date: 2026-01-02SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/007732
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-07
Filing Date
2025-06-05
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Electronic devices face resource wastage and difficulty in simultaneously identifying and providing information about multiple objects within image data due to continuous monitoring of user gaze or finger direction.

Method used

An electronic device acquires speech data and image data to identify at least two objects corresponding to a gesture, providing associated information through a processor that utilizes motion data from a secondary device to enhance object identification.

Benefits of technology

Reduces resource waste and enables efficient selection and information provision for multiple objects within image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007732_02012026_PF_FP_ABST
    Figure KR2025007732_02012026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to various embodiments may comprise at least one processor, a camera, a microphone, and a memory (313) which is operatively connected to the at least one processor, the camera, and the microphone and stores at least one instruction, wherein the at least one instruction is configured to cause, when executed by the at least one processor, the electronic device to: acquire image data through the camera while speech data including a designated word related to designation of an object is being acquired through the microphone; identify, among a plurality of objects included in the image data, at least two objects corresponding to a gesture identified in the image data at a point in time when the designated word is spoken; and provide association information for the at least two objects.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition method and electronic device supporting the same

[0001] Embodiments disclosed in this document relate to an image recognition method and an electronic device supporting the same.

[0002] With the advancement of digital technology, a variety of electronic devices capable of communicating and processing personal information on the go, such as mobile terminals, electronic organizers, smartphones, tablet PCs, and wearable devices, have been released. These devices have evolved from simple voice calls and messaging to include video calls, electronic organizers, document processing, email, internet access, and photography.

[0003] Additionally, electronic devices can provide image recognition capabilities that recognize objects within image data and apply them to various services. For example, the electronic device can provide information (e.g., object attribute information) about a target object indicated by a user among multiple objects contained in image data.

[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0005] In general, an electronic device can identify the direction in which a user's gaze or finger is pointed, and identify an object corresponding to the direction in which the user's gaze or finger is pointed in image data as a target object.

[0006] However, electronic devices have the problem of having to continuously monitor the user's gaze or the direction of their finger to identify the target object. This monitoring can waste the electronic device's resources.

[0007] Additionally, electronic devices can identify a single object as the target object, whether the user's gaze is focused on it or the user's finger is pointing at it. For example, electronic devices have difficulty simultaneously identifying multiple target objects within image data and providing relevant information about them.

[0008] The technical problems to be achieved in this document are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.

[0009] According to various embodiments, an electronic device includes at least one processor, a camera, a microphone, and a memory (313) operatively connected to the at least one processor, the camera, and the microphone and storing at least one command, wherein the at least one command, when individually or collectively executed by the at least one processor, causes the electronic device to: acquire image data through the camera while speech data including a designated word related to designation of an object is acquired through the microphone; identify, among a plurality of objects included in the image data, at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered; and provide associated information for the at least two objects.

[0010] An operating method of an electronic device according to various embodiments may include an operation of acquiring speech data including a designated word related to designation of an object, an operation of acquiring image data while the speech data is acquired, an operation of identifying at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and an operation of providing associated information about the at least two objects.

[0011] A computer-readable recording medium according to various embodiments may be configured such that when executed by an electronic device, the electronic device obtains speech data including a designated word related to designation of an object, obtains image data while the speech data is being obtained, identifies at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and provides associated information about the at least two objects.

[0012] An image recognition system according to various embodiments includes a first electronic device and a second electronic device, wherein the first electronic device is configured to acquire image data through a camera while acquiring speech data including a designated word related to designation of an object through a microphone, identify at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and provide associated information about the at least two objects, and the second electronic device is configured to provide information related to a posture of the second electronic device acquired through a sensor to the first electronic device, and the first electronic device may be configured to identify the at least two objects corresponding to a gesture identified in the image data when the information related to the posture of the second electronic device satisfies a designated condition.

[0013] Electronic devices according to various embodiments disclosed in this document can reduce resource waste and enable selection of multiple objects within image data and provision of associated information therefor.

[0014] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.

[0015] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0016] FIG. 1 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0017] FIG. 2A is a drawing for explaining an image recognition function of an electronic device according to various embodiments.

[0018] FIG. 2b is a diagram for explaining information provided through an image recognition function according to various embodiments.

[0019] FIG. 2c is a drawing for explaining an image recognition function of an electronic device according to various embodiments.

[0020] FIG. 2D is a diagram illustrating a configuration for an electronic device according to various embodiments to determine whether a target object is suitable as a comparison target.

[0021] FIG. 2e is a diagram illustrating a configuration for guiding an electronic device to reselect a comparison target according to various embodiments.

[0022] FIG. 2f is a diagram illustrating an operation of providing information about a target object according to various embodiments.

[0023] FIG. 3a is a diagram illustrating an image recognition system according to various embodiments.

[0024] FIG. 3b is a diagram schematically illustrating the configuration of an image recognition system according to various embodiments.

[0025] Figure 3c is a drawing for explaining motion data used for object identification.

[0026] FIG. 4 is a drawing for explaining an image recognition function of an electronic device according to various embodiments.

[0027] FIG. 5A is a drawing for explaining an image recognition function of a first electronic device according to various embodiments.

[0028] FIG. 5b is a drawing for explaining an operation of recognizing a user's gesture in a first electronic device according to various embodiments.

[0029] FIG. 5c is a diagram illustrating the operation of a first electronic device controlled based on gestures according to various embodiments.

[0030] Figure 6 is a diagram illustrating the configuration of an information provision model according to various embodiments.

[0031] FIG. 7 is a diagram illustrating a procedure for processing input data of an object identification model according to various embodiments.

[0032] FIG. 8A is a diagram illustrating an image recognition system according to various embodiments.

[0033] FIG. 8b is a diagram illustrating an image recognition system according to various embodiments.

[0034] FIG. 9A is a flowchart illustrating the operation of an electronic device according to various embodiments.

[0035] FIG. 9b is a diagram for explaining related information according to various embodiments.

[0036] FIG. 10 is a flowchart illustrating a motion data acquisition operation of an electronic device according to various embodiments.

[0037] FIG. 11 is a flowchart illustrating a target object identification operation of an electronic device according to various embodiments.

[0038] FIG. 12 is a flowchart illustrating a field of view correction operation of an electronic device according to various embodiments.

[0039] Figure 13 is a drawing for explaining a field of view correction process according to various embodiments.

[0040] FIG. 14 is a flowchart illustrating an operation of providing related information in an electronic device according to various embodiments.

[0041] FIG. 15 is another flowchart illustrating the operation of an electronic device according to various embodiments.

[0042] FIGS. 16A to 16C are drawings for explaining the operation of the first electronic device in various embodiments.

[0043] Hereinafter, various embodiments of this document will be described with reference to the attached drawings. However, this is not intended to limit the technology described in this document to specific embodiments, and it should be understood that various modifications, equivalents, and / or alternatives of the embodiments of this document are included. In connection with the description of the drawings, similar reference numerals may be used for similar components.

[0044]

[0045] FIG. 1 is a block diagram of an exemplary electronic device (100) capable of performing the operations described in this document.

[0046] Referring to FIG. 1, the electronic device (100) may be one of various forms of electronic devices, such as a notebook (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable) type smartphone (191-3)), a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 1 are exemplary only and do not limit the implementations described or claimed in this document. The electronic device (100) may be referred to as a mobile device, a user device, a multi-function device, a portable device, or a server.

[0047] The electronic device (100) may include components including at least one processor (110) (hereinafter referred to as processor (110)), at least one memory (120) (hereinafter referred to as memory (120)), at least one display (140) (hereinafter referred to as display (140)), at least one image sensor (150) (hereinafter referred to as image sensor (150)), at least one communication circuit (160) (hereinafter referred to as communication circuit (160)), and / or at least one sensor (170) (hereinafter referred to as sensor (170)). The above components are merely exemplary. For example, the electronic device (100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input / output interface). For example, some components may be omitted from the electronic device (100). For example, some components may be integrated into one component.

[0048] The processor (110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (110) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data) stored in the memory (120). The processor (110) may include a processor assembly including one or more processing circuits. The processor (110) may include any processing circuit operative to control the performance and operations of one or more components (e.g., the memory (120), the display (140), the image sensor (150), the communication circuit (160), and / or the sensor (170)) of the electronic device (100). For example, the processor (110) (e.g., the application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or a chipset). For example, the processor (110) may be implemented with multiple cores (or at least one core circuit), multiple chips, or multiple chipsets. For example, the processor (110) may include one or more processing circuits. For example, the processor (110) may include one or more processing circuits configured to individually and / or collectively perform various functions of the present disclosure. As a non-limiting example, at least a portion of the processor (110) may be included in a first chip of the electronic device (100), and at least another portion of the processor (110) may be included in a second chip of the electronic device (100) that is different from the first chip of the electronic device (100).

[0049] For example, the processor (110) may include a central processing unit (CPU) (111), a graphics processing unit (GPU) (112), a neural processing unit (NPU) (113), an image signal processor (ISP) (114), a display controller (115), a memory controller (116), a storage controller (117), a communication processor (CP) (118), and / or a sensor interface (119). These components of the processor (110) are merely exemplary. For example, the processor (110) may further include other components. For example, some components of the processor (110) may be omitted from the processor (110). For example, some components of the processor (110) may be included as separate components of the electronic device (100) outside the processor (110). For example, some components of the processor (110) (e.g., a memory controller (116)) may be included within other components (e.g., at least a portion of the memory (120), an interface (e.g., available for connection to at least one component of the electronic device (100)), a display (140) and / or an image sensor (150)).

[0050] The processor (110) can control the operations of the electronic device (100) by executing instructions stored in the memory (120). For example, the processor (110) can correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.

[0051] The processor (110) may cause other components of the electronic device (100) to perform various operations by executing instructions stored in the memory (120). The CPU (111) (or central processing circuit) may be configured to control components of the processor (110) based on the execution of instructions stored in the memory (120) (e.g., volatile memory (121) and / or non-volatile memory (122)). The GPU (112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (113) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). The ISP (114) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (150) into a format suitable for a component within the electronic device (100) or a component of the processor (110). The display controller (115) (or display control circuit, or display processing unit (DPU)) may be configured to process an image acquired from the CPU (111), the GPU (112), the ISP (114), or the memory (120) (e.g., the volatile memory (121)) into a format suitable for the display (140). The memory controller (116) (or memory control circuit) may be configured to control reading data from the volatile memory (121) and writing data to the volatile memory (121). The storage controller (117) (or storage control circuit) may be configured to control reading data from the nonvolatile memory (122) and writing data to the nonvolatile memory (122).The CP (118) (communication processing circuit) may be configured to process data acquired from a component of the processor (110) into a format suitable for transmitting to another electronic device via the communication circuit (160), or to process data acquired from another electronic device via the communication circuit (160) into a format suitable for processing by the component of the processor (110). For example, the communication circuit (160) may include one or more communication circuits. The sensor interface (119) (or sensing data processing circuit, sensor hub) may be configured to process data about the state of the electronic device (100) and / or the state of the surroundings of the electronic device (100), acquired via the sensor (170), into a format suitable for the component of the processor (110).

[0052] The memory (120) may include one or more storage media (or one or more storage devices). For example, the memory (120) may include a memory assembly including one or more storage media. For example, the one or more storage media may include permanent memory (e.g., non-volatile memory (122)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (121)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof. The memory (120) may include cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (100). As a non-limiting example, the cache memory may be included within the processor (110). The memory (120) may be fixedly embedded within the electronic device (100) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and / or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (100).

[0053] For example, the memory (120) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (110). For example, the memory (120) may store instructions callable by an application programming interface (API). For example, the memory (120) may store instructions within a library.

[0054]

[0055] According to various embodiments, the aforementioned electronic device (100) may provide an image recognition function that recognizes (or identifies) objects within image data and applies the recognition function to various services. This will be described in detail with reference to FIGS. 2A to 15 below. Furthermore, at least one of the various embodiments described with reference to FIGS. 2A to 15 below may be combined with other embodiments.

[0056]

[0057] FIG. 2A is a diagram for explaining an image recognition function of an electronic device according to various embodiments. FIG. 2B is a diagram for explaining information provided through an image recognition function according to various embodiments. FIG. 2C is a diagram for explaining an image recognition function of an electronic device according to various embodiments. FIG. 2D is a diagram for explaining a configuration for an electronic device according to various embodiments to determine whether a target object is appropriate as a comparison target. FIG. 2E is a diagram for explaining a configuration for guiding an electronic device according to various embodiments to reselect a comparison target. FIG. 2F is a diagram for explaining an operation for providing information on a target object according to various embodiments.

[0058] 200 of FIG. 2a represents a situation in which a user speaks while sequentially pointing to a plurality of specific objects with an index finger (e.g., repeating a pointing gesture), and 230 of FIG. 2a represents image data (210) acquired by an electronic device (100) while the user speaks.

[0059] Referring to FIG. 2A, an electronic device (100) according to various embodiments may identify a target object (or comparison target) indicated by a user among a plurality of objects included in the image data (210) based on speech data (or speech input) (201) and image data (210). According to one embodiment, the electronic device (100) may identify a plurality of target objects in the image data (210) based on speech data (201) and provide related information thereon. An image recognition function according to various embodiments related thereto will be described in more detail below.

[0060] According to various embodiments, the electronic device (100) can obtain speech data (201) and identify a first part, a second part, and a third part thereof.

[0061] According to one embodiment, the first part (and the second part) of the utterance data (201) may correspond to designated words (e.g., demonstrative pronouns indicating a single object such as this, that, here, there, he, she, you) that designate the first target object (and the second target object), and the third part of the utterance data (201) may correspond to a designated query.

[0062] For example, the electronic device (100) may identify a first word (e.g., this child) (201-1) designating a first target object (or one object) as a first part, a second word (e.g., that child) (201-2) designating a second target object (e.g., another object) as a second part, and a sentence corresponding to the query (e.g., which child will be taller when they grow up?) (201-3) as a third part for the utterance data (201).

[0063] According to various embodiments, the electronic device (100) may acquire image data (210) while acquiring speech data (201). According to one embodiment, the electronic device (100) may recognize (or extract) an object included in the image data (210) prior to identifying the target object.

[0064] For example, the electronic device (100) can recognize a first object (211), a second object (213), a third object (215), and a fourth object (217) based on feature data (e.g., feature points) extracted from image data (210). In this regard, the electronic device (100) can extract feature data using various known techniques such as histogram of oriented gradient (HOG), scale invariant feature transform (SIFT), local binary pattern (LBP), and modified census transform (MCT).

[0065] According to various embodiments, the electronic device (100) can identify a target object indicated by a user among recognized objects (e.g., a first object (211) to a fourth object (217)) based on the first part and the second part identified from the speech data (201).

[0066] According to an embodiment, the electronic device (100) may identify an object (e.g., a second object (213)) corresponding to a first location (221) of an index finger identified in image data (210) as a first target object at a first time point (or after the first time point) when a first portion (e.g., a first word (201-1)) is identified, as illustrated in 230 of FIG. 2A. For example, the electronic device (100) may identify a second object (213) existing within a certain range (or a certain distance) based on the first location (221) as the first target object.

[0067] Similarly, the electronic device (100) can identify an object (e.g., a third object (215)) corresponding to the second position (223) of the index finger identified in the image data (210) as the second target object at a second time point (or after the second time point) when the second part (e.g., the second word (201-2)) is identified. For example, when the index finger located at the first position (221) moves to the second position (223) within a certain range based on the third object (215) at the second time point, the electronic device (100) can identify the third object (215) as the second target object.

[0068] According to various embodiments, the electronic device (100) may provide associated information about identified target objects (e.g., a first target object (e.g., a second object (213)) and a second target object (e.g., a third object (215))) based on the third portion (201-3) of the speech data (201).

[0069] According to one embodiment, as illustrated in FIG. 2B, information (250) about an identified target object may include at least one of first information (251) related to a first target object (e.g., a poodle), second information (253) related to a second target object (e.g., a Welsh corgi), and third information (255) which is detailed information about them. The related information may include related information (e.g., comparison information) about the first target object and the second target objects. In this regard, the electronic device (100) may use the first target object, the second target object, and the third part (201-3) of the data (201) as a search term (e.g., which child will grow up to be bigger, a poodle or a Welsh corgi?), and obtain a search result therefor from an external source (e.g., a search server). However, this is merely an example, and various embodiments are not limited thereto. For example, the associated information may include information about each of the first target object and the second target object. In this regard, the electronic device (100) may use queries inquiring about the types of the first target object, the second target object, and each target object as search terms (e.g., "Describe the types of poodles and Welsh corgis"), and may obtain search results (e.g., "Description of poodles" and "Description of Welsh corgis") from an external source.

[0070] According to various embodiments, the electronic device (100) may output the associated information (250) regarding the identified target object through the electronic device (100) or may output it through another electronic device. According to one embodiment, the associated information (250) regarding the identified target object may be output as visual information through the display (140) of the electronic device (100). However, this is merely exemplary, and various embodiments are not limited thereto. For example, the associated information (250) regarding the identified target object may be output as auditory information through the speaker of the electronic device (100), as illustrated in 295 of FIG. 2F, or may be output through another electronic device (e.g., wireless earphones (298), a smartwatch, another smart phone, or an AI speaker) that is connected to the electronic device (100) through communication, as illustrated in 297 of FIG. 2F.

[0071] Additionally or optionally, the electronic device (100) according to various embodiments may also identify a target object from a plurality of image data.

[0072] According to an embodiment, the electronic device (100) may obtain first image data (260) corresponding to a first field of view (e.g., 11 o'clock direction) at a first point in time when a first portion (e.g., a first word (201-1)) is identified, as illustrated in FIG. 2C. In addition, the electronic device (100) may obtain second image data (270) corresponding to a second field of view (e.g., 1 o'clock direction) different from the first field of view at a second point in time when a second portion (e.g., a second word (201-2)) is identified.

[0073] In this regard, the electronic device (100) may identify an object (e.g., a second object (213)) corresponding to a first position (261) of an index finger identified in the first image data (260) as a first target object at a first time point when the first portion is identified. In addition, the electronic device (100) may identify an object (e.g., a third object (215)) corresponding to a second position (271) of an index finger identified in the second image data (270) as a second target object at a second time point when the second portion is identified.

[0074] Additionally or optionally, the electronic device (100) according to various embodiments may compare attributes (e.g., type) of the first target object with attributes of the second target object in providing association information (e.g., comparison information) for identified target objects (e.g., first target object and second target object).

[0075] According to one embodiment, the electronic device (100) can determine whether a first target object and a second target object are appropriate as comparison objects through attribute comparison. For example, as illustrated in 208-1 of FIG. 2D, when a first target object (e.g., a first smart phone) (281) and a second target object (e.g., a second smart phone) (282) indicated by a user are identified, the electronic device (100) can determine whether the first target object (281) and the second target object (282) have attributes that are related to each other, and determine that the first target object (281) and the second target object (282) having attributes that are related to each other (e.g., a mobile device) are appropriate as comparison objects. In addition, the electronic device (100) can provide association information (250) for target objects (e.g., the first target object and the second target object) that are determined to be appropriate as comparison objects.

[0076] For example, as illustrated in 208-2 of FIG. 2d, when a second target object (e.g., a second smart phone) (282) and a third target object (e.g., a laptop) (283) indicated by the user are confirmed, the electronic device (100) may check an attribute (e.g., a mobile device) related to the second target object (282) and an attribute (e.g., a computer device) related to the third target object (283), and may determine that the second target object (282) and the third target object (283) having different attributes are not appropriate as comparison objects. In addition, when the electronic device (100) determines that the target objects (e.g., the second target object (282) and the third target object (283)) are not appropriate as comparison objects, the electronic device (100) may provide guide information (291) for reselecting a comparison object or guide information for re-inputting a query, as illustrated in FIG. 2e. According to an embodiment, the electronic device (100) may also provide information related to the cause of the inappropriateness of the comparison target (e.g., the properties of the specified target objects are different from each other) as at least part of the guide information (291).

[0077] As described above, the image recognition function according to various embodiments may be provided through the independent operation of the electronic device (100). Depending on the embodiment, the electronic device (100) may also provide the image recognition function through collaboration with at least one other device. This will be described in detail with reference to FIGS. 3A to 3C below.

[0078]

[0079] FIG. 3a is a diagram illustrating an image recognition system according to various embodiments. FIG. 3b is a diagram schematically illustrating the configuration of an image recognition system according to various embodiments, and FIG. 3c is a diagram illustrating motion data utilized for object identification.

[0080] Referring to FIGS. 3A to 3C, an image recognition system (300) according to various embodiments may be configured with a first electronic device (310) and a second electronic device (320). According to one embodiment, the first electronic device (310) and the second electronic device (320) may communicate with each other via a network (e.g., a short-range communication network or a long-range communication network).

[0081] According to an embodiment, the first electronic device (310) may be referred to as a smartphone, and the second electronic device (320) may be referred to as a wearable device (e.g., a ring-shaped electronic device) worn on a part of the body (e.g., an index finger).

[0082] However, this is merely an example, and various embodiments are not limited thereto. For example, the first electronic device (310) may be referred to as a foldable type smartphone (191-2) that can be attached to clothing worn by a user in a folded state (e.g., an outside pocket of clothing), or a pin (or clip) type wearable device that can be attached to clothing (e.g., a shirt) or accessories (e.g., a tie, a necklace) worn by the user. Depending on the embodiment, the first electronic device (310) and the second electronic device (320) may be referred to as a smartphone, and at least one of the first electronic device (310) and the second electronic device (320) may be referred to as a wearable device worn on another part of the body (e.g., a watch-type electronic device wearable on the wrist, a necklace-type electronic device wearable on the neck, eyeglass-type electronic device wearable on the face, or an augmented reality or virtual reality device wearable on the head).

[0083] According to various embodiments, the first electronic device (310) can provide an image recognition function through collaboration with the second electronic device (320).

[0084] According to one embodiment, the first electronic device (310) may acquire (or collect) speech data (201) and image data (210), and the second electronic device (320) may acquire motion data (341) related to the posture of the second electronic device (320) while it is worn on the body (e.g., motion data for a part of the body on which the second electronic device is worn (e.g., an index finger)). In addition, the first electronic device (310) may identify a target object indicated by a user among a plurality of objects (e.g., a first object (211) to a fourth object (217)) included in the image data (210) based on the speech data (201) and image data (210) acquired by the first electronic device (310) and the motion data (341) acquired by the second electronic device (320).

[0085] The configuration of the first electronic device (310) according to various embodiments related thereto will be described in more detail below.

[0086] Referring to FIG. 3B, a first electronic device (310) according to various embodiments may be configured with at least one first processor (311) (hereinafter referred to as a first processor (311)), at least one first communication circuit (312) (hereinafter referred to as a first communication circuit (312)), at least one first memory (313) (hereinafter referred to as a first memory (313)), at least one image sensor (314) (or at least one camera) (hereinafter referred to as an image sensor (314)), at least one microphone (315) (hereinafter referred to as a microphone (315)), and at least one output device (316) (hereinafter referred to as an output device (316)).

[0087] Depending on the embodiment, the first electronic device (310) may be implemented to have more or fewer components than the aforementioned components. For example, the first electronic device (310) may correspond to the electronic device (100) illustrated in FIG. 1.

[0088] According to various embodiments, the first communication circuit (312) may support wireless communication with the second electronic device (320). According to one embodiment, the first communication circuit (312) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first electronic device (310) and the second electronic device (320).

[0089] For example, the first communication circuit (312) may communicate with the second electronic device (320) via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)).

[0090] According to various embodiments, the first memory (313) may store commands or data related to components of the first electronic device (310). For example, the first memory (313) may store programs, algorithms, routines, and commands related to image recognition.

[0091] According to various embodiments, the image sensor (314) may acquire (or capture) image data (210). According to one embodiment, the image data (210) may be stored in the first memory (313). In addition, if the first electronic device (310) has an output device (316) configured to output visual information, the image data (210) may also be output through the output device (316).

[0092] According to various embodiments, the microphone (315) may be configured to convert external sounds into electrical audio signals and output them. According to one embodiment, the microphone (315) may be configured to acquire speech data (201). Depending on the embodiment, the microphone (315) may be applied with various noise removal algorithms to remove noise generated during the process of receiving external sounds.

[0093] According to various embodiments, the output device (316) may provide various information related to the operation of the first electronic device (310). At least some of the various information may be related to the image recognition function. According to one embodiment, the output device (316) may include a display configured to provide visual information (e.g., text, images, videos, icons, or symbols) to the user and receive user input (e.g., touch input). However, this is merely an example, and various embodiments are not limited thereto. For example, according to an embodiment, the output device (316) may be configured to provide auditory information, tactile information, or a combination thereof.

[0094] According to various embodiments, the first processor (311) may be operatively connected to a first communication circuit (312), a first memory (313), an image sensor (314), a microphone (315), and an output device (316), and may control various components (e.g., hardware or software components) of the first electronic device (310).

[0095] According to one embodiment, the first processor (311) can identify a target object indicated by a user in image data (210) and provide related information (250) about the identified target object. The target object may be part or all of a plurality of objects included in the image data (210).

[0096] In this regard, the first processor (311) may acquire image data (210) while acquiring speech data (201). For example, the first processor (311) may identify a target object (e.g., the first target object and the second target object of FIG. 2A) among objects included in the image data (210) based on a first portion (e.g., a designated first word (201-1) of FIG. 2A) and a second portion (e.g., a designated second word (201-2) of FIG. 2A) identified from the speech data (201). In addition, the first processor (311) may provide related information (250) about the target object based on a third portion (e.g., a sentence (201-3) corresponding to a query of FIG. 2A) identified from the speech data (201). A specific description of the first processor (311) related to this may refer to the operation of the electronic device (100) described through FIGS. 2a and 2b.

[0097] Additionally, the first processor (311) according to various embodiments may utilize motion data (341) acquired through the second electronic device (320) to identify a target object. In this regard, the first processor (311) may acquire motion data (341) from the second electronic device (320) while acquiring speech data (201) (and / or while acquiring image data (210)). For example, the motion data (341) acquired from the second electronic device (320) may include position information, velocity information, acceleration information, direction information, or a combination thereof.

[0098] According to one embodiment, the first processor (311) may obtain first motion data acquired through the second electronic device (320) at a first time point when a first portion of the speech data (201) is identified, and utilize the same to identify a target object (e.g., a first target object). In addition, the first processor (311) may obtain second motion data acquired through the second electronic device (320) at a second time point when a second portion of the speech data (201) is identified, and utilize the same to identify another target object (e.g., a second target object).

[0099] For example, as illustrated in FIG. 3C, the first processor (311) may acquire (or extract) a portion (341-1) corresponding to a first time point (t1) at which a first portion (e.g., a designated first word (e.g., this child) (201-1)) of the speech data (201) is identified, as first motion data, from the motion data (341) acquired from the second electronic device (320). Similarly, the first processor (311) may acquire another portion (341-2) of the motion data (341) corresponding to a second time point (t2) at which a second portion (e.g., a designated second word (e.g., that child) (201-2)) of the speech data (201) is identified, as second motion data.

[0100] According to one embodiment, the first processor (311) may determine whether the first motion data satisfies a specified condition when identifying the first target object. For example, if the first motion data acquired at the first time point (t1) satisfies the specified condition, the first processor (311) may identify an object corresponding to the first position of a body (e.g., an index finger) identified in the image data (210) as the first target object.

[0101] Similarly, the first processor (311) can determine whether the second motion data satisfies a specified condition when identifying the second target object. For example, if the second motion data acquired at the second point in time satisfies the specified condition, the first processor (311) can identify an object corresponding to the second position of the body identified in the image data (210) as the second target object.

[0102] As described above, the first processor (311) according to various embodiments can identify a target object in image data (210) when motion data (341) obtained through the second electronic device (320) satisfies a specified condition.

[0103] In this regard, the first processor (311) may determine that a specified condition is satisfied if the first motion data (and / or the second motion data) corresponds to a specified gesture. For example, the specified gesture may include a gesture in which the user indicates a specific object using a part of the body (e.g., a pointing gesture). Depending on the embodiment, the specified gesture may also occur when the second electronic device (320) (e.g., a touch sensor) is at least partially touched by a part of the body.

[0104] However, this is merely an example, and various embodiments are not limited thereto. For example, various gestures that can indicate a specific object (e.g., drawing a circle with the index finger, drawing with the hand while pointing to an object with the finger, forming a finger frame, or forming a circle using the thumb and index finger) can be used under specified conditions.

[0105] Additionally or optionally, the first processor (311) may acquire second motion data if a second portion of the speech data (201) is identified within a predetermined time (e.g., 5 seconds) after the first portion of the speech data (201) is identified. For example, the first processor (311) may exclude a second portion that is identified after a predetermined time has elapsed since the first portion was identified from the identification of the second target object.

[0106] As described above, the first electronic device (310) can determine whether motion data (341) satisfying a specified condition is obtained from the second electronic device (320). However, in this case, a synchronization problem or a transmission speed problem for the motion data (341) may occur.

[0107] In this regard, according to one embodiment, the second processor (321) may determine whether the first motion data satisfies a specified condition when identifying the first target object. The first processor (311) may also obtain motion data (341) that satisfies the specified condition from the second electronic device (320). For example, the second electronic device (320) may obtain motion data (341) related to a posture of the second electronic device (320) at the time of identifying the first part (and / or the second part) of the speech data (201) or receiving information representing the first part (and / or the second part) of the speech data (201) from the first electronic device (310). In addition, the second electronic device (320) may provide the motion data (341) to the first electronic device (310) when the obtained motion data (341) satisfies the specified condition.

[0108] Additionally or optionally, the second electronic device (320) may extract feature data for the motion data (341) and provide it to the first electronic device (310) instead of the motion data (341) that satisfies the specified conditions. In this case, the first electronic device (310) may identify the target object in the image data (210) based on the feature data provided from the second electronic device (320).

[0109] According to an embodiment, the second electronic device (320) may extract feature data related to a motion feature of a body part (e.g., a finger) (or the second electronic device (320)). For example, the feature data may include at least one of feature data corresponding to a moving body part, feature data corresponding to a stationary body part, feature data corresponding to a body part that has stopped moving, and feature data corresponding to a body part that has stopped moving.

[0110] As described above, the first processor (311) can obtain motion data (341) used to identify a target object through the second electronic device (320). The configuration of an exemplary second electronic device (320) related thereto will be described in more detail below.

[0111] According to various embodiments, a second electronic device (320) may be configured with at least one second processor (321) (hereinafter referred to as a second processor (321)), a second communication circuit (322) (hereinafter referred to as a second communication circuit (322)), at least one second memory (323) (hereinafter referred to as a second memory (323)), and at least one sensor (324) (hereinafter referred to as a sensor (224)), as illustrated in FIG. 3B. Depending on the embodiment, the second electronic device (320) may be implemented to have more or fewer components than the aforementioned components.

[0112] The second communication circuit (322) and the second memory (323) described above may be similar to or identical to the first communication circuit (312) and the first memory (313) of the first electronic device (210) described above, and thus a detailed description thereof may be omitted.

[0113] According to various embodiments, the second communication circuit (322) may support wireless communication with the first electronic device (310). According to one embodiment, the second communication circuit (322) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first electronic device (310) and the second electronic device (320).

[0114] According to various embodiments, the second memory (323) may store commands or data related to at least one other component of the second electronic device (320).

[0115] According to various embodiments, the sensor (324) may be configured to obtain motion data related to the posture of the second electronic device (320). According to one embodiment, the sensor (324) may include at least one of an acceleration sensor, a gyro sensor, a gesture sensor, or a barometric pressure sensor.

[0116] According to various embodiments, the second processor (321) may be operatively connected to the second communication circuit (322), the second memory (323), and the sensor (324), and may control various components (e.g., hardware or software components) of the second electronic device (320).

[0117] According to one embodiment, the second processor (321) may provide motion data collected through the sensor (324) to the first electronic device (310). In this regard, the second processor (321) may control the sensor (324) so ​​that motion data (341) is collected while the first electronic device (310) and the second electronic device (320) are connected to each other through communication.

[0118] According to an embodiment, the second processor (321) may control the sensor (324) to collect motion data (341) based on the occurrence of a specified event. For example, if the second electronic device (320) is configured to receive a user input (e.g., a touch input) (e.g., if a touch sensor is provided in the second electronic device (320), the first processor (321) may process the motion data (341) to be collected after (or while) the user input is detected.

[0119] As described above, the first electronic device (310) according to various embodiments can identify a target object in the image data (210) based on the speech data (201) and image data (210) collected by the first electronic device (310) and the motion data (314) collected by the second electronic device (320).

[0120] According to an embodiment, the first electronic device (310) may perform an operation of scanning a beam of an optical pointer, such as a laser pointer, to identify a target object. In this regard, the first electronic device (310) or the second electronic device (320) may include a light emitting unit configured to irradiate a laser point, and the first electronic device (310) may identify an object corresponding to the direction in which the laser point is pointed among a plurality of objects as a target object, thereby further improving object identification accuracy.

[0121]

[0122] FIG. 4 is a drawing for explaining an image recognition function of an electronic device according to various embodiments.

[0123] The image recognition method illustrated in FIG. 4 differs from the image recognition methods illustrated in FIGS. 2a and 2b in that it utilizes a gesture that simultaneously designates multiple specific objects. The image recognition functions according to various related embodiments will be described in more detail below.

[0124] Referring to FIG. 4, a first electronic device (310) according to various embodiments can obtain utterance data (411) and identify a fourth part and a third part thereof.

[0125] According to one embodiment, the fourth part of the utterance data (411) may correspond to a designated word (e.g., a demonstrative pronoun indicating multiple objects such as these, those, we, they, you, and she) that designates multiple target objects (e.g., a first target object and a second target object) at once, and the third part of the utterance data (411) may correspond to a designated query.

[0126] For example, the first electronic device (310) can identify a word (e.g., these) designating multiple target objects as a fourth part for the speech data (411) (e.g., which of these is the tallest when it grows up?), and can identify a sentence corresponding to the query (e.g., which is taller when it grows up?) as a third part.

[0127] According to various embodiments, the first electronic device (310) can identify a target object in the image data (210) based on the fourth part of the speech data (411).

[0128] In this regard, the first electronic device (310) can recognize (or extract) objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) after acquiring the image data (210) while the speech data (411) is acquired. In addition, the first electronic device (310) can identify a plurality of target objects indicated by the user among the recognized objects (e.g., the first object (211) to the fourth object (217)) based on the fourth part of the speech data (411).

[0129] According to an embodiment, the first electronic device (310) can identify a target object in the image data (210) when the motion data (341) obtained through the second electronic device (320) satisfies a specified condition.

[0130] For example, the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) corresponding to the gesture (413) identified in the image data (210) as target objects at a fourth time point when the fourth portion is identified, as illustrated. For example, the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) included in a circle (413-1) drawn by the gesture (413) as target objects.

[0131] According to various embodiments, the first electronic device (310) may provide association information (250) for a plurality of target objects based on the third part of the speech data (411).

[0132]

[0133] FIG. 5A is a diagram illustrating an image recognition function of a first electronic device according to various embodiments. FIG. 5B is a diagram illustrating an operation of recognizing a user's gesture in a first electronic device according to various embodiments. FIG. 5C is a diagram illustrating an operation of a first electronic device controlled based on a gesture according to various embodiments.

[0134] The image recognition method illustrated in FIG. 5A differs from the image recognition method illustrated in FIG. 4 in that motion data is acquired through a plurality of second electronic devices (320-1 and 320-2). For example, one (320-1) of the plurality of second electronic devices (320-1 and 320-2) may be worn on a first part of the body (e.g., the left hand (511-1)), and the other (320-2) of the plurality of second electronic devices (320-1 and 320-2) may be worn on a second part of the body (e.g., the right hand (511-2)). The image recognition function according to various embodiments related thereto will be described in more detail below.

[0135] Referring to FIG. 5A, a first electronic device (310) according to various embodiments can obtain utterance data (513) and identify a fourth part and a third part thereof.

[0136] According to one embodiment, the fourth part of the utterance data (513) may correspond to a designated word (e.g., a demonstrative pronoun indicating multiple objects such as these, those, we, they, you, and she) that designates multiple target objects (e.g., a first target object and a second target object) at once, and the third part of the utterance data (513) may correspond to a designated query.

[0137] For example, the first electronic device (310) can identify a word (these) designating multiple target objects as a fourth part for the user's speech data (513) (e.g., "Which of these is the tallest?") and identify a sentence corresponding to the query (e.g., "Who is taller?") as a third part.

[0138] According to various embodiments, the first electronic device (310) can identify a target object in the image data (210) based on the fourth part of the speech data (411).

[0139] In this regard, the first electronic device (310) can recognize (or extract) objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) after acquiring the image data (210) while acquiring the speech data (513). In addition, the first electronic device (310) can identify a plurality of target objects indicated by the user among the recognized objects (e.g., the first object (211) to the fourth object (217)) based on the fourth part of the speech data (411).

[0140] According to an embodiment, the first electronic device (310) can identify a target object in the image data (210) when the motion data (341) obtained through the plurality of second electronic devices (320-1 and 320-2) satisfies a specified condition.

[0141] For example, the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) corresponding to a gesture (e.g., a hand frame gesture) identified in the image data (210) as target objects at a fourth time point when the fourth part is identified, as illustrated. For example, the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) included in a hand frame formed by a gesture (511) using the first part (511-1) of the body and the second part (511-2) of the body as target objects.

[0142] According to one embodiment, the first electronic device (310) may utilize predetermined signals transmitted and received with a plurality of second electronic devices (320-1 and 320-2) via wireless communication (e.g., ultra wideband (UWB) communication) for gesture identification. For example, as illustrated in FIG. 5B, the first electronic device (310) may determine a first distance (B) with one second electronic device (320-1), a second distance (C) with another second electronic device (320-2), and a third distance (A) between the plurality of second electronic devices (320-1 and 320-2) based on the transmitted and received predetermined signals. According to an embodiment, the first electronic device (310) may identify a gesture in which the third distance (A) between the plurality of second electronic devices (320-1 and 320-2) falls within a specified range (e.g., within 15 cm).

[0143] In this regard, the first electronic device (310) may transmit and receive a predetermined signal with a plurality of second electronic devices (320-1 and 320-2) through an algorithm related to ranging. For example, the algorithm related to ranging may include at least one of ToF (time of flight), TWR (two way ranging), DS-TWR (double sided-TWR), SS-TWR (single sided-TWR), TDoA (time difference of arrival), or AoA (angle of arrival).

[0144] According to various embodiments, the first electronic device (310) may provide association information (250) for a plurality of target objects based on the third portion (201-3) of the speech data (513).

[0145] As described above, the first electronic device (310) according to various embodiments can identify a target object from the image data (210) based on speech data (201, 411, 513) collected by the first electronic device (310), image data (210), and motion data (341) collected by the second electronic device (320).

[0146] According to an embodiment, the first electronic device (310) may utilize an artificial intelligence model generated through machine learning to identify a target object. In this regard, the first electronic device (310) (e.g., the first memory (313)) may store an artificial intelligence model (e.g., an information provision model (3131)) configured to identify a target object from image data (210) and provide related information (e.g., comparison information) about the target object. A more detailed description thereof will be provided with reference to FIG. 6 below.

[0147] Additionally or alternatively, the first electronic device (310) according to various embodiments may perform a designated action based on a gesture (e.g., a hand frame gesture) identified in the image data (210). According to one embodiment, the first electronic device (310) may activate an image sensor that is in an inactive state in response to the gesture identification. For example, the first electronic device (310) may obtain an image corresponding to a field of view of the first electronic device (310) (e.g., an image sensor), as illustrated in 540 of FIG. 5C . According to an embodiment, the first electronic device (310) may obtain an image that includes only target objects (e.g., a second object (213) and a third object (215)) identified by the gesture (e.g., a hand frame gesture), as illustrated in 550 of FIG.

[0148]

[0149] Figure 6 is a diagram illustrating the configuration of an information provision model according to various embodiments.

[0150] The information providing model (3131) according to various embodiments may be an artificial intelligence model generated through machine learning. According to an embodiment, the information providing model (3131) may input motion data (610) (e.g., motion data (341)), image data (620) (e.g., image data (210)), and speech data (630) (e.g., speech data (201, 411, 513)), and output related information (e.g., comparison information) regarding a target object identified in the image data (620).

[0151] According to one embodiment, the information providing model (3131) may include a plurality of artificial neural network layers. The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, a transformer network, or a combination of two or more thereof, but is not limited to the examples described above. In addition to the software structure, the information providing model (3131) may additionally or alternatively include a hardware structure. The configuration of the information providing model (3131) according to various embodiments related thereto will be described in more detail below.

[0152] Referring to FIG. 6, an information provision model (3131) according to various embodiments may include an object identification model (605), an information retrieval model (607), and an information generation model (609).

[0153] According to various embodiments, the object identification model (605) can identify a target object from image data (620). According to one embodiment, the object identification model (605) can receive motion data (610), image data (620), and speech data (630) as inputs, and output identification information (615) about the target object identified in the image data (620). For example, the object identification model (605) can identify the target object based on a gesture identified in the image data (620) while speech data (630) is input.

[0154] According to various embodiments, the information retrieval model (607) can retrieve information related to a target object identified in image data (620). According to one embodiment, the information retrieval model (607) can receive speech data (630) and information about the target object (e.g., identification result (615)) as inputs, and output a search result (617). For example, the information retrieval model (607) can generate a search term based on the speech data (630) and the identification result (615), and obtain a search result (617) based on the search term from an external source (e.g., a search server).

[0155] According to various embodiments, the information generation model (609) can convert the search result (617) of the information retrieval model (607) into a form that can be recognized by the user and output information (619) (e.g., related information (250)) about the target object. According to one embodiment, the information generation model (609) can convert the text-type search result (617) output by the information retrieval model (607) into an auditory form. Depending on the embodiment, the information (619) about the target object may also be output in a visual form.

[0156] Additionally or optionally, the information generation model (609) may include a generative model configured to generate new output data (e.g., image data) based on the search results (617). However, this is merely exemplary, and various embodiments are not limited thereto. For example, the information generation model (609) may be comprised of various types of models other than the generative model.

[0157] Depending on the embodiment, the information provision model (3131) may be configured with fewer components than the aforementioned components. For example, the information generation model (609) may be omitted from the information provision model (3131). In such a case, the search results (617) obtained by the information retrieval model (607) may be provided as information (619) regarding the target object.

[0158] According to an embodiment, the information provision model (3131) may be configured to have more components than the aforementioned components. For example, a pattern recognition model (603) may be included in the information provision model (3131). In this case, the object identification model (605) may identify the user's gesture based on pattern information (613) of motion data (610) output by the pattern recognition model (603). The configuration of the pattern recognition model (603) according to various embodiments related thereto will be described in more detail below.

[0159] According to various embodiments, the pattern recognition model (603) can identify pattern information (613) based on motion data (610). The pattern information can be related to the repetition of motion, the direction of motion, the speed of motion, or the magnitude of motion. According to one embodiment, the pattern recognition model (603) can receive motion data (610) as input and identify a specific pattern in the motion data (610). In addition, the pattern recognition model (603) can output pattern information (613) related to the identified specific pattern. This pattern information (613) is input to the object identification model (605), and the object identification model (605) can utilize the pattern information (613) to identify a user's gesture.

[0160] Additionally or optionally, the pattern recognition model (603) may identify a specific pattern in the motion data (610) based on feature data (611) extracted from the motion data (610). In this case, the information provision model (3131) may further include a feature extraction model (601). The configuration of the feature extraction model (601) according to various related embodiments will be described in more detail below.

[0161] According to various embodiments, the feature extraction model (601) can extract feature data (611) from motion data (610). The feature data (611) may be data that can be utilized to identify a specific pattern in the motion data (610). According to one embodiment, the feature extraction model (601) can extract feature data (611) related to appearance features and motion features of a body part (e.g., a finger) from the motion data (610). This feature data (611) is input to a pattern recognition model (603), and the pattern recognition model (603) can utilize the feature data (611) to identify the pattern.

[0162]

[0163] FIG. 7 is a diagram illustrating a procedure for processing input data of an object identification model according to various embodiments.

[0164] According to various embodiments, the object identification model (605) receives image data (620) and speech data (630) as inputs and can identify a target object from the image data (620).

[0165] According to an embodiment, the information provision model (3131) can input the image data (620) and speech data (630) into the object identification model (605).

[0166] Referring to FIG. 7, the information provision model (3131) can perform an operation of converting one-dimensional data into two-dimensional data when merging (650) image data (620) and speech data (630).

[0167] For example, the information provision model (3131) can obtain refined image data by preprocessing (621) the image data (620) in a manner such as filtering and sampling, and convert it into two-dimensional image data expressed in terms of the relationship between time and frequency (623).

[0168] Similarly, the information provision model (3131) can obtain refined speech input by preprocessing (631) the speech input (630) in a manner such as filtering and sampling, and convert it into two-dimensional speech data expressed in terms of the relationship between time and frequency (633).

[0169] These two-dimensional image data (623) and speech data (633) are merged (650) and provided as input to an object identification model (605), and the object identification model (605) can output identification information (615) for the target object based on the merged image data (623) and speech data (633).

[0170] However, this is merely an example, and various embodiments are not limited thereto. For example, as described above, in identifying a target object, motion data (610) (or feature data (611) or pattern information (613)) may be further utilized in addition to image data (620) and speech data (630). In this regard, the information provision model (3131) may convert motion data (610) into two-dimensional motion data, merge it with two-dimensional image data (623) and speech data (633), and output it as an object identification model (605).

[0171]

[0172] FIG. 8A is a diagram illustrating an image recognition system according to various embodiments.

[0173] Referring to FIG. 8A, an image recognition system (81) according to various embodiments may be composed of a first electronic device (810), a second electronic device (820), and a third electronic device (830). According to one embodiment, the first electronic device (810) may communicate with the second electronic device (820) and the third electronic device (830) through a network (e.g., a short-range communication network or a long-range communication network).

[0174] According to various embodiments, the first electronic device (810) may provide an image recognition function through collaboration with the second electronic device (820) and the third electronic device (830).

[0175] According to one embodiment, a first electronic device (810) may obtain (or collect) image data (210), a second electronic device (820) may obtain motion data (341) related to a posture of the second electronic device (820) while it is worn on a body, and a third electronic device (830) may obtain speech data (201). In addition, the first electronic device (810) may identify a target object indicated by a user among a plurality of objects (e.g., a first object (211) to a fourth object (217)) included in the image data (210) based on the image data (210) obtained by the first electronic device (810), the motion data (341) obtained by the second electronic device (820), and the speech data (201) obtained by the third electronic device (830).

[0176] The configuration of the first electronic device (810) according to various embodiments related thereto will be described in more detail below.

[0177] A first electronic device (810) according to various embodiments may be composed of a first processor (811), a first communication circuit (812), a first memory (813), an image sensor (814), and an output device (816).

[0178] According to an embodiment, the configurations of the first electronic device (810) illustrated in FIG. 8a may be similar or identical to the configurations of the first electronic device (320) described above through FIG. 3b, and thus a detailed description thereof may be omitted.

[0179] According to various embodiments, the first processor (811) can identify a target object indicated by a user from image data (210) acquired through an image sensor (814). For example, the first processor (811) can utilize motion data (341) acquired by a second electronic device (820) and speech data (201) acquired by a third electronic device (830) to identify the target object.

[0180] The configuration of the second electronic device (820) and the configuration of the third electronic device (830) according to various embodiments related thereto will be described in more detail below.

[0181] A second electronic device (820) according to various embodiments may be composed of a second processor (821), a second communication circuit (822), a second memory (823), and a sensor (824).

[0182] According to an embodiment, the configurations of the second electronic device (820) illustrated in FIG. 8a may be similar or identical to the configurations of the second electronic device (320) described above through FIG. 3b, and thus a detailed description thereof may be omitted.

[0183] According to one embodiment, the second processor (821) may provide motion data (341) collected through a sensor (824) to the first electronic device (810).

[0184] A third electronic device (830) according to various embodiments may be composed of a third processor (831), a third communication circuit (832), a third memory (833), and a microphone (834).

[0185] According to an embodiment, the configurations of the third electronic device (830) illustrated in FIG. 8a may be similar or identical to the configurations of the first electronic device (810) or the second electronic device (820) described above through FIG. 3b, and thus a detailed description thereof may be omitted.

[0186] According to one embodiment, the third processor (831) may provide speech data (201) collected through a microphone (834) to the first electronic device (810).

[0187]

[0188] FIG. 8b is a diagram illustrating an image recognition system according to various embodiments.

[0189] Referring to FIG. 8B, an image recognition system (82) according to various embodiments may be composed of a first electronic device (810), a second electronic device (820), a third electronic device (830), and a fourth electronic device (840). According to one embodiment, the first electronic device (810) may communicate with the second electronic device (820), the third electronic device (830), and the fourth electronic device (840) through a network (e.g., a short-range communication network or a long-range communication network).

[0190] According to various embodiments, the first electronic device (810) can provide an image recognition function through collaboration with the second electronic device (820), the third electronic device (830), and the fourth electronic device (840).

[0191] According to one embodiment, the first electronic device (810) can identify a target object indicated by a user among a plurality of objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) based on motion data (341) acquired by the second electronic device (820), speech data (201) acquired by the third electronic device (830), and image data (210) acquired by the fourth electronic device (840).

[0192] The configuration of the first electronic device (810) according to various embodiments related thereto will be described in more detail below.

[0193] A first electronic device (810) according to various embodiments may be composed of a first processor (811), a first communication circuit (812), a first memory (813), and an output device (816).

[0194] According to one embodiment, the first processor (811) can utilize motion data (341) acquired by the second electronic device (820), speech data (201) acquired by the third electronic device (830), and image data (210) acquired by the fourth electronic device (840) to identify a target object.

[0195] The configuration of the second electronic device (820) to the fourth electronic device (840) according to various embodiments related thereto will be described in more detail below.

[0196] A second electronic device (820) according to various embodiments may be composed of a second processor (821), a second communication circuit (822), a second memory (823), and a sensor (824).

[0197] According to one embodiment, the second processor (821) may provide motion data (341) collected through a sensor (824) to the first electronic device (810).

[0198] A third electronic device (830) according to various embodiments may be composed of a third processor (831), a third communication circuit (832), a third memory (833), and a microphone (834).

[0199] According to one embodiment, at least one third processor (831) may provide speech data (201) collected via a microphone (834) to the first electronic device (810).

[0200] A fourth electronic device (840) according to various embodiments may be composed of a fourth processor (841), a fourth communication circuit (842), a fourth memory (843), and an image sensor (844).

[0201] According to one embodiment, the fourth processor (841) may provide image data (210) collected through the image sensor (844) to the first electronic device (810).

[0202] The image recognition system (81, 82) described through the aforementioned FIGS. 8A and 8B is merely exemplary, and various embodiments are not limited thereto. For example, an image recognition system according to various embodiments may be configured with a first electronic device (810) configured to acquire image data (210) and a second electronic device (820) configured to collect motion data (341) and speech data (201). In addition, an image recognition system according to various embodiments may be configured with a first electronic device (810) configured to acquire speech data (201) and a second electronic device (820) configured to collect motion data (341) and image data (210).

[0203]

[0204] An electronic device (310) according to various embodiments may include at least one processor (311), a camera (e.g., an image sensor) (314), a microphone (315), and a memory (313) operatively connected to the at least one processor, the camera, and the microphone and storing at least one command. According to one embodiment, the at least one command, when individually or collectively executed by the at least one processor (311), may cause the electronic device (310) to: acquire image data (210) through the camera (314) while utterance data (201) including a designated word related to the designation of an object is acquired through the microphone (315), identify at least two objects corresponding to a gesture identified in the image data at the time when the designated word is uttered among a plurality of objects (211 to 217) included in the image data, and provide associated information (250) for the at least two objects.

[0205] According to various embodiments, the utterance data may further include a query (201-3). According to one embodiment, the at least one command, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: provide the associated information related to the query.

[0206] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: generate a search term based on the at least two objects and at least a portion of the query, and provide the associated information related to the generated search term.

[0207] According to various embodiments, the electronic device (310) may further include a communication circuit (312). According to one embodiment, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: obtain the associated information from an external device (320) via the communication circuit (312).

[0208] According to various embodiments, the designated word may include a first word designating a plurality of objects. According to one embodiment, the at least one command, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: identify, among a plurality of objects included in the image data, a first object and a second object corresponding to a first gesture (413, 511) identified in the image data at the time when the first word is uttered, and provide the associated information for the first object and the second object.

[0209] According to various embodiments, the designated word may include a second word (201-1) and a third word (201-2) designating a single object. According to one embodiment, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: identify, among a plurality of objects included in the image data, a first object corresponding to a second gesture (221) identified in the image data at a time when the second word is uttered; identify, among a plurality of objects included in the image data, a second object corresponding to a third gesture (223) identified in the image data at a time when the third word is uttered; and provide the association information for the first object and the second object.

[0210] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: identify a second object in the image data if the third word is uttered within a predetermined time after the second word is uttered.

[0211] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: acquire first image data (260) and second image data (270) corresponding to different fields of view through the camera (314), identify a first object corresponding to a second gesture (221) identified in the first image data at a time when the second word is uttered among a plurality of objects included in the first image data, and identify a second object corresponding to a third gesture (223) identified in the second image data at a time when the third word is uttered among a plurality of objects included in the second image data.

[0212] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: acquire motion data (341) from an external device (320) via the communication circuit (312) while the speech data is acquired, and identify the at least two objects if the motion data satisfies a specified condition.

[0213] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: input the speech data and the image data into an artificial intelligence model (3131) stored in the electronic device (310) (e.g., memory (313)), and identify the at least two objects based on an output of the artificial intelligence model.

[0214]

[0215] Figure 9a is a flowchart illustrating the operation of an electronic device according to various embodiments. Furthermore, Figure 9b is a diagram for explaining related information according to various embodiments. Depending on the embodiment, the operations in the following embodiments may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.

[0216] According to one embodiment, operations 910 to 950 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).

[0217] Referring to FIG. 9A, an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may obtain image data (210) in operation 910. According to one embodiment, the image data (210) may be obtained through an image sensor (e.g., an image sensor (314)) of the electronic device (100) (e.g., the first electronic device (310)) or may be obtained from an external source (e.g., a second electronic device (320)).

[0218] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may obtain motion data (341) in operation 920. According to one embodiment, the electronic device (100) may obtain motion data (341) while obtaining image data (210). For example, the motion data (341) may be obtained from an external source (e.g., a second electronic device (320)).

[0219] According to various embodiments, an electronic device (100) (e.g., a first electronic device (310)) may obtain a speech input (or speech data) (201) in operation 930. According to one embodiment, the speech input (201) may be obtained through a microphone (e.g., a microphone (315)) of the electronic device (100) or may be obtained externally. The speech may be a speech requesting related information (e.g., comparison information) about an object. However, this is merely an example, and various embodiments are not limited thereto. For example, the speech may be a speech requesting information about each object, and according to an embodiment, the speech may be a speech indicating various requests, such as taking pictures of the object.

[0220] An electronic device (100) according to various embodiments (e.g., a first electronic device (310)) may, at operation 940, identify at least two objects (e.g., two target objects) included in image data (210) based on speech input (201) and motion data (341).

[0221] According to one embodiment, the electronic device (100) can identify a target object based on the location of a body part (e.g., an index finger) identified in the image data (210) when motion data satisfying a specified condition is acquired while a speech input is acquired.

[0222] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may provide association information on a target object corresponding to a speech input at operation 950. According to one embodiment, the electronic device (100) may generate a search term based on the speech input (201) and the identified target object, and obtain search results based on the search term from an external source (e.g., a search server). For example, the association information may include at least one of the visual association information described above with reference to FIG. 2b or the auditory association information described above with reference to FIG. 2f.

[0223] According to an embodiment, when the electronic device (100) obtains a gesture indicating a plurality of objects (971, 972) and a user utterance requesting a comparison of the plurality of objects (971, 972) (e.g., "Tell me where to buy the cheaper object among the two objects") (970), as illustrated in FIG. 9b, the electronic device (100) may provide the price comparison result and information related to the purchase location (e.g., the seller's homepage address) to the plurality of objects (971, 972) (980). According to an embodiment, the gestures indicating the plurality of objects (971, 972) may be the same or different from each other.

[0224] For example, one object (971) may be designated by a first gesture (e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing at an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and index finger), and another object (972) may also be designated by the first gesture.

[0225] In an embodiment, one object (971) may be designated by a first gesture (e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing to an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and an index finger), and another object (972) may be designated by a second gesture different from the first gesture (e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing to an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and an index finger).

[0226]

[0227] FIG. 10 is a flowchart illustrating motion data acquisition operations of an electronic device according to various embodiments. The operations of FIG. 10 described below may represent various embodiments of operation 910 of FIG. 9.

[0228] Referring to FIG. 10, an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may obtain motion data (341) in operation 1010. According to one embodiment, the electronic device (100) may obtain motion data (341) in a state in which the operation of an image sensor (e.g., an image sensor (314)) that consumes relatively much power is deactivated.

[0229] According to various embodiments, an electronic device (100) (e.g., a first electronic device (310)) may determine, in operation 1020, whether motion data (341) corresponding to a specified gesture is acquired. According to one embodiment, the electronic device (100) may determine whether motion data (341) collected by another electronic device (e.g., a second electronic device (320)) that is connected to the electronic device (100) through communication corresponds to the specified gesture. For example, the specified gesture may be a gesture of drawing a specified shape using a body on which another electronic device is worn (e.g., a gesture of repeatedly drawing a circle a certain number of times). According to an embodiment, the specified gesture may also be a gesture of generating a specified input (e.g., a touch input) to another electronic device worn on the body using another part of the body.

[0230] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may repeatedly perform operations 1010 and 1020 if motion data (341) corresponding to a specified gesture is not obtained.

[0231] According to various embodiments, an electronic device (100) (e.g., a first electronic device (310)) may activate an image sensor in operation 1030 when motion data (341) corresponding to a specified gesture is obtained. According to one embodiment, the electronic device (100) may reduce power consumption of the electronic device (100) by controlling the operation of the image sensor (314) using a sensor (e.g., a motion sensor) that consumes relatively little power.

[0232] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may obtain image data (210) in operation 1040. As described in operation 910 of FIG. 9, the image data (210) may be obtained through an image sensor of the electronic device (100) or may be obtained from an external source (e.g., a second electronic device (320)).

[0233]

[0234] FIG. 11 is a flowchart illustrating a target object identification operation of an electronic device according to various embodiments. The operations of FIG. 10 described below may represent various embodiments of operation 940 of FIG. 9.

[0235] Referring to FIG. 11, an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments can, in operation 1110, check first motion data corresponding to a first portion of a speech input (201).

[0236] According to one embodiment, the first part of the speech input (201) may correspond to a designated word designating a first target object (e.g., a demonstrative pronoun indicating a single object such as this, that, here, there, he, she, you), as described above through FIG. 2a.

[0237] In this regard, the electronic device (100) can acquire, as described above through FIG. 3c, some motion data corresponding to the first time point (t1) at which the first part of the speech data (201) is identified from the acquired motion data (341), as first motion data.

[0238] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1120, identify second motion data corresponding to a second portion of a speech input.

[0239] According to one embodiment, the second part of the speech input (201) may correspond to a designated word designating a second target object, as described above with reference to FIG. 2a.

[0240] In this regard, the electronic device (100) can acquire, as described above through FIG. 3c, some motion data corresponding to the second time point (t2) at which the second part of the speech data (201) is identified from the acquired motion data (341), as second motion data.

[0241] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1130, identify a first object (e.g., a first target object) corresponding to first motion data among a plurality of objects included in image data (210).

[0242] According to one embodiment, the electronic device (100) may determine whether the first motion data satisfies a specified condition when identifying the first object. For example, if the first motion data acquired at the first time point (t1) satisfies the specified condition, the electronic device (100) may identify an object corresponding to the position of a body (e.g., an index finger) identified in the image data (210) as the first target object.

[0243] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1140, identify a second object (e.g., a second target object) corresponding to second motion data among a plurality of objects included in image data (210).

[0244] According to one embodiment, the electronic device (100) may determine whether the second motion data satisfies a specified condition when identifying the second object. For example, if the second motion data acquired at the second time point (t2) satisfies the specified condition, the electronic device (100) may identify an object corresponding to the position of a body (e.g., an index finger) identified in the image data (210) as the second target object.

[0245] As described above, the electronic device (100) according to various embodiments can identify a target object by recognizing a user's gesture that sequentially points to a plurality of specific objects in image data (210). However, if the image data (210) corresponding to the user's field of view is not acquired, the target object identified by the electronic device (100) and the object indicated by the user may not match each other. In this regard, the electronic device (100) according to various embodiments can acquire the image data (210) corresponding to the user's field of view and improve the identification performance for the target object by matching the acquisition range of the image data (210) to the user's field of view. This will be described in more detail with reference to FIGS. 12 and 13 below.

[0246]

[0247] FIG. 12 is a flowchart illustrating a field of view correction operation of an electronic device according to various embodiments. FIG. 13 is a diagram for explaining a field of view correction process according to various embodiments. In addition, each operation in the following embodiments may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. In addition, at least one of the above-described operations may be omitted depending on the embodiment.

[0248] According to one embodiment, operations 1210 to 1250 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).

[0249] Referring to FIGS. 12 and 13, an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may obtain image data (1301) in operation 1210. According to one embodiment, the image data (1301) may be obtained through an image sensor (e.g., an image sensor (314)) of the electronic device (100) or may be obtained from an external source (e.g., a second electronic device (320)).

[0250] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1220, select a reference object from image data (1301). The reference object may include at least one of the objects included in the image data (1301).

[0251] According to one embodiment, the electronic device (100) may recognize a first object (e.g., a sofa) (1303), a second object (e.g., a table) (1305), a third object (e.g., a vase) (1307), and a fourth object (e.g., a picture frame) (1309) based on feature data (e.g., feature points) extracted from image data (1301), and select at least one of the recognized objects (e.g., the fourth object (1309)) as a reference object.

[0252] According to various embodiments, an electronic device (100) (e.g., a first electronic device (310)) may output guide information that guides selection of a reference object in operation 1230. For example, the electronic device (100) may output guide information in an auditory form (e.g., “Point to the picture frame,” “Point to the right corner of the picture frame”) that guides a user to point to the reference object. However, this is merely an example, and various embodiments are not limited thereto. For example, the electronic device (100) may provide guide information in a visual form, and according to an embodiment, may provide guide information in an auditory form and guide information in a visual form together.

[0253] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1240, identify a first location (1321) corresponding to a gesture related to selection of a reference object in image data (1301).

[0254] An electronic device (100) according to various embodiments (e.g., a first electronic device (310)) may, in operation 1250, correct a field of view of an image sensor (314) based on a first location (1321) and a second location corresponding to a reference object. For example, the electronic device (100) may correct a field of view of the image sensor based on a distance and direction difference (1323) between the first location and the second location.

[0255] Additionally or optionally, the electronic device (100) according to various embodiments may select an object (e.g., a third object (e.g., a vase) (1307)) having a size smaller than a specified size among a plurality of objects (e.g., a first object (e.g., a sofa) (1303), a second object (e.g., a table) (1305), a third object (e.g., a vase) (1307), and a fourth object (e.g., a picture frame) (1309)) recognized from image data (1301) as a reference object in order to improve the field of view correction performance of the image sensor.

[0256]

[0257] FIG. 14 is a flowchart illustrating operations for providing related information in an electronic device according to various embodiments. The operations of FIG. 14 described below may represent various embodiments of operation 950 of FIG. 9.

[0258] Referring to FIG. 14, an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments can check a third part of a speech input (201) in operation 1410.

[0259] According to one embodiment, the third part of the speech input (201) may correspond to a specified query, as described above with reference to FIG. 2a.

[0260] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may perform a search operation in operations 1420 and 1430.

[0261] According to one embodiment, the electronic device may perform a search operation using the third part of the speech data (201), the first target object, and the second target object as search words (operation 1420). In this regard, the electronic device (100) may generate a search word based on at least a portion of the third part of the speech data (201), the first target object, and the second target object. In addition, the electronic device (100) may obtain a search result based on the generated search word from an external source (e.g., a search server) (operation 1430).

[0262]

[0263] Figure 15 is another flowchart illustrating the operation of an electronic device according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.

[0264] According to one embodiment, operations 1510 to 1580 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).

[0265] Referring to FIG. 15, an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1510, acquire image data (210) through an image sensor (e.g., an image sensor (314)) of the electronic device (100) or may acquire it from the outside (e.g., a second electronic device (320)).

[0266] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may detect a designated first gesture in operation 1520. According to one embodiment, the designated first gesture may be a gesture in which a user points to a specific object with an index finger.

[0267] An electronic device (100) according to various embodiments (e.g., a first electronic device (310)) may determine, in operation 1530, whether a first utterance for object designation is detected. The first utterance may correspond to a designated word designating an object (e.g., a demonstrative pronoun indicating a single object such as this, that, here, there, he, she, you).

[0268] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may repeatedly perform operations related to operations 1510 to 1530 if the first ignition is not detected.

[0269] According to various embodiments, an electronic device (100) (e.g., a first electronic device (310)) may, when a first utterance is detected, identify a first object in image data (210) in operation 1540. According to one embodiment, the electronic device (100) may identify a first object corresponding to a first gesture among a plurality of objects included in the image data (210).

[0270] An electronic device (100) according to various embodiments (e.g., a first electronic device (310)) may detect a designated second gesture at operation 1550. The designated second gesture may be similar to the first gesture described above. According to one embodiment, after detecting the first gesture, the electronic device (100) may detect a second gesture indicating another object.

[0271] An electronic device (100) according to various embodiments (e.g., a first electronic device (310)) may determine, at operation 1560, whether a second utterance for object designation is detected. The second utterance may be similar to the first utterance described above. According to one embodiment, the electronic device (100) may detect a second utterance indicating another object after detecting the first utterance.

[0272] An electronic device (100) according to various embodiments (e.g., a first electronic device (310)) may repeatedly perform operations related to operations 1510 to 1560 if a second ignition is not detected.

[0273] According to various embodiments, an electronic device (100) (e.g., a first electronic device (310)) may, when a second utterance is detected, identify a second object in the image data (210) at operation 1570. According to one embodiment, the electronic device (100) may identify a second object corresponding to the second gesture among a plurality of objects included in the image data (210).

[0274] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may detect an information request utterance at operation 1580. For example, the information request utterance may be an utterance requesting related information (e.g., comparison information) about a first object and a second object. For example, the information request utterance may be an utterance requesting information about each of the first object and the second object.

[0275] An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may, in operation 1590, provide related information about a first object and a second object based on an information request utterance.

[0276] According to one embodiment, when an information request utterance requesting related information about a first object and a second object is detected, the electronic device (100) can obtain related information about the first object and the second object from the outside.

[0277] According to one embodiment, when an information request utterance requesting information about each of a first object and a second object is detected, the electronic device (100) can obtain information about the first object and information about the second object from the outside.

[0278]

[0279] FIGS. 16A to 16C are drawings for explaining the operation of the first electronic device in various embodiments.

[0280] Referring to 1610 of FIG. 16A, a first electronic device (310) (e.g., electronic device (100)) according to various embodiments may acquire a first gesture (1603) and a first utterance (1605) while acquiring image data (1601). For example, the first gesture (1603) may be a gesture identified in the image data (1601) that indicates at least one specific object included in the image data (1601) (e.g., a gesture of making a finger frame). For example, the first utterance (1605) may be an utterance that indicates processing (e.g., capturing or storing) at least a portion of the image data (1601) (e.g., capturing here).

[0281] Referring to 1620 of FIG. 16A, a first electronic device (310) (e.g., electronic device (100)) according to various embodiments may acquire a second gesture (1623) and a second utterance (1625) after acquiring a first gesture (1603) and a first utterance (1605). For example, the second gesture (1623) may be another gesture identified in the image data (1601) that indicates at least one other specific object included in the image data (1601) (e.g., a gesture pointing to a target with a finger). For example, the second utterance may be an utterance that indicates another processing (e.g., a search) for at least a portion of the image data (e.g., “What is this? Translate it into Korean, make an album name, and save the photos I took.”). According to one embodiment, the first electronic device (310) can acquire a second gesture (1623) and a second utterance (1625) within a certain time after the first gesture (1603) and the first utterance (1605) are acquired.

[0282] Referring to 1630 of FIG. 16B, a first electronic device (310) (e.g., electronic device (100)) according to various embodiments may perform first processing on image data (1601) based on a first gesture (1603) and a first utterance (1605). According to one embodiment, the first electronic device (310) may designate the entire image data (1601) as a first processing target based on the first gesture (1603), and store (1633) the image data (1601) designated as the first processing target based on the first utterance (1605).

[0283] Additionally, the first electronic device (310) (e.g., the electronic device (100)) according to various embodiments may perform second processing on the image data (1601) based on the second gesture (1623) and the second utterance (1625). According to one embodiment, the first electronic device (310) may designate a portion of the image data (1601) as a second processing target based on the second gesture (1623), perform a search operation on the second processing target designated as the processing target based on the second utterance (1625), and use at least a portion of the search result to create a storage album (or folder) for the image data (1601). For example, as illustrated, the first electronic device (310) may set at least a portion of the result as a storage album name (1631) based on the second utterance (1625).

[0284] Additionally or alternatively, the first electronic device (310) (e.g., electronic device (100)) according to various embodiments may provide a processing result for at least one of the first processing and the second processing while performing the first processing (e.g., storing the image data (1601)) and the second processing (e.g., setting a name for a storage album after a search) on the image data (1601). For example, the first electronic device (310) may provide visual information (1635) indicating a search result (e.g., search content) for the second processing target (e.g., Gwanghwamun. Gwanghwamun is the main gate of Gyeongbokgung Palace, and so on). However, this is merely an example, and various embodiments are not limited thereto. For example, the processing result for at least one of the first processing and the second processing may be output as auditory information through an external electronic device (e.g., wireless earphones (1641)) connected to the first electronic device (310) through communication, as illustrated in 1640 of FIG. 16c, or may be output as auditory information through a speaker of the first electronic device (310), as illustrated in 1650 of FIG. 16c.

[0285] As described above, the first electronic device (310) according to various embodiments may designate the entire image data (1601) as a processing target based on the first gesture (1603). However, this is merely an example, and various embodiments are not limited thereto. For example, as illustrated in 1635 of FIG. 16B , the first electronic device (310) may designate only a specific object included in the hand frame formed by the first gesture (1603) as a processing target. For example, the first electronic device (310) may store a portion (1634) of the image data (1601) corresponding to the first gesture (1603) stored based on the first utterance (1605) in a storage album set based on the second utterance (1625).

[0286]

[0287] An operating method of an electronic device (310) according to various embodiments may include an operation of acquiring speech data (201) including a designated word related to designation of an object, an operation of acquiring image data (210) while the speech data is being acquired, an operation of identifying at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects (211 to 217) included in the image data, and an operation of providing associated information (250) for the at least two objects.

[0288] According to various embodiments, the utterance data may further include a query (201-3). According to one embodiment, the method of operating the electronic device (310) may include an operation of providing the related information related to the query.

[0289] According to various embodiments, the method of operating the electronic device (310) may include generating a search term based on the at least two objects and at least a portion of the query, and providing the associated information related to the generated search term.

[0290] According to various embodiments, the method of operating the electronic device (310) may include an operation of obtaining the related information from an external device (320).

[0291] According to various embodiments, the designated word may include a first word designating a plurality of objects. According to one embodiment, the operating method of the electronic device (310) may include an operation of identifying a first object and a second object corresponding to a first gesture (413, 511) identified in the image data at a time when the first word is uttered, among a plurality of objects included in the image data, and an operation of providing the associated information for the first object and the second object.

[0292] According to various embodiments, the designated word may include a second word (201-1) and a third word (201-2) designating a single object. According to one embodiment, the operating method of the electronic device (310) may include an operation of identifying a first object corresponding to a second gesture (221) identified in the image data at a time when the second word is uttered among a plurality of objects included in the image data, an operation of identifying a second object corresponding to a third gesture (223) identified in the image data at a time when the third word is uttered among a plurality of objects included in the image data, and an operation of providing the association information for the first object and the second object.

[0293] According to various embodiments, the operating method of the electronic device (310) may include an operation of identifying a second object in the image data when the third word is uttered within a certain time after the second word is uttered.

[0294] According to various embodiments, the operating method of the electronic device (310) may include an operation of acquiring first image data (260) and second image data (270) corresponding to different fields of view, an operation of identifying a first object corresponding to a second gesture (221) identified in the first image data at a time when the second word is uttered among a plurality of objects included in the first image data, and an operation of identifying a second object corresponding to a third gesture (223) identified in the second image data at a time when the third word is uttered among a plurality of objects included in the second image data.

[0295] According to various embodiments, the operating method of the electronic device (310) may include an operation of acquiring motion data (341) from an external device (320) while the speech data is acquired, and an operation of identifying at least two objects if the motion data satisfies a specified condition.

[0296] According to various embodiments, the operating method of the electronic device (310) may include an operation of inputting the speech data and the image data into an artificial intelligence model (3131) stored in the electronic device (310) (e.g., memory (313)) and an operation of identifying the at least two objects based on an output of the artificial intelligence model.

[0297]

[0298] A computer-readable recording medium according to various embodiments may be configured such that when executed by an electronic device, the electronic device obtains speech data including a designated word related to designation of an object, obtains image data while the speech data is being obtained, identifies at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and provides associated information about the at least two objects.

[0299]

[0300] An image recognition system according to various embodiments may include a first electronic device (310) and a second electronic device (320).

[0301] According to one embodiment, the first electronic device (310) includes at least one first processor (311), a camera (314), a microphone (315), and a first memory (313) operatively connected to the at least one first processor, the camera, and the microphone and storing at least one first command, wherein the at least one first command, when executed by the at least one first processor, causes the first electronic device to: acquire image data (210) through the camera while utterance data (201) including a designated word related to designation of an object is acquired through the microphone; identify at least two objects corresponding to a gesture identified in the image data at the time when the designated word is uttered, among a plurality of objects (211 to 217) included in the image data; and provide associated information (250) for the at least two objects.

[0302] According to one embodiment, the second electronic device (320) includes at least one second processor (321), a sensor (324), and a second memory (323) operatively connected to the at least one second processor and the sensor and storing at least one second instruction, wherein the at least one second instruction, when executed by the at least one second processor, is configured to cause the second electronic device to provide information related to a posture of the second electronic device obtained through the sensor to the first electronic device.

[0303] According to various embodiments, the at least one first instruction, when executed by the at least one first processor, may be configured to cause the first electronic device to: identify at least two objects corresponding to a gesture identified in the image data if information related to a posture of the second electronic device provided from the second electronic device satisfies a specified condition.

Claims

1. In an electronic device (310), At least one processor (311); Camera (314); Mike (315); and A memory (313) operatively connected to at least one processor (311), the camera (314) and the microphone (315) and storing at least one command, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: While speech data (201) containing a designated word related to the designation of an object is acquired through the microphone (315), image data (210) is acquired through the camera (314), Among the plurality of objects (211 to 217) included in the above image data, at least two objects corresponding to the gesture identified in the image data at the time when the specified word is uttered are identified, An electronic device configured to provide association information (250) for at least two objects.

2. In paragraph 1, The above utterance data further includes a query (201-3), The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: An electronic device configured to provide said associated information related to said query.

3. In paragraph 2, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: Generate a search term based on at least two objects and at least a part of the query, An electronic device configured to provide the above-mentioned related information related to the above-mentioned generated search word.

4. In paragraph 1, It further includes a communication circuit (312), The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: An electronic device configured to obtain the related information from an external device (320) through the communication circuit (312).

5. In paragraph 1, The above-mentioned words include a first word designating multiple objects, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: Among the plurality of objects included in the above image data, the first object and the second object corresponding to the first gesture (413, 511) identified in the image data at the time when the first word is uttered are identified, An electronic device configured to provide the associated information for the first object and the second object.

6. In paragraph 1, The above-mentioned words include a second word (201-1) and a third word (201-2) that designate a single object, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: Among the plurality of objects included in the above image data, a first object corresponding to a second gesture (221) identified in the image data at the time when the second word is uttered is identified, Among the plurality of objects included in the above image data, a second object corresponding to a third gesture (223) identified in the image data at the time when the third word is uttered is identified, An electronic device configured to provide the associated information for the first object and the second object.

7. In paragraph 6, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: An electronic device configured to identify a second object in the image data when the third word is uttered within a certain time after the second word is uttered.

8. In paragraph 6, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: Obtain first image data (250) and second image data (270) corresponding to different fields of view through the above camera (314), Among the plurality of objects included in the first image data, a first object corresponding to a second gesture (221) identified in the first image data at the time when the second word is spoken is identified, An electronic device configured to identify a second object corresponding to a third gesture (223) identified in the second image data at the time when the third word is uttered, among a plurality of objects included in the second image data.

9. In paragraph 1, It further includes a communication circuit (312), The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: While the above speech data is being acquired, motion data (341) is acquired from an external device (320) through the communication circuit (312), An electronic device configured to identify at least two objects when the motion data satisfies a specified condition.

10. In any one of paragraphs 1 to 9, The at least one instruction, when individually or collectively executed by the at least one processor (311), causes the electronic device (310) to: Input the above speech data and the above image data into the artificial intelligence model (3131) stored in the electronic device (310), An electronic device configured to identify at least two objects based on the output of the artificial intelligence model.

11. In the operating method of the electronic device (310), An action of obtaining utterance data (201) containing a specified word related to the designation of an object; An operation of acquiring image data (210) while the above speech data is being acquired; An operation of identifying at least two objects corresponding to a gesture identified in the image data at the time when the specified word is uttered among a plurality of objects (211 to 217) included in the image data; and A method comprising an action of providing association information (250) for at least two objects.

12. In paragraph 11, The above utterance data further includes a query (201-3), A method comprising an action of providing the above-mentioned related information related to the above-mentioned query.

13. In paragraph 12, An operation of generating a search term based on at least two objects and at least a portion of the query; and A method comprising an action of providing the above-mentioned related information related to the above-mentioned generated search word.

14. In paragraph 11, A method comprising an operation of obtaining the above-mentioned related information from an external device (320).

15. In any one of paragraphs 11 to 14, An operation of inputting the above speech data and the above image data into an artificial intelligence model (3131) stored in the electronic device (310); and A method comprising an action of identifying at least two objects based on the output of the artificial intelligence model.

Citation Information

Patent Citations

  • Information Processing Systems

    JP7373068B2

  • Display device and light emitting element

    KR1020250036301A

  • Methods and systems for displaying virtual objects from an augmented reality environment on a multimedia device

    US20230169744A1

  • Data processing system, data processing method, and information providing system

    US20240037956A1

  • KR20190013390A