Circuitry, electronic device and method

The described circuitry system improves object localization by using user commands and image segment data to generate precise localization data through machine learning, addressing inefficiencies in existing localization methods.

WO2025176839A1PCT designated stage Publication Date: 2025-08-28SONY GROUP CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/054716
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-21
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing object localization techniques, such as those using Apple AirTags, lack efficiency and precision in determining the exact location of misplaced items within an environment.

Method used

A circuitry system that utilizes user command data and image segment data to determine the localization of an object by identifying image segments corresponding to the user's query and generating data based on the relations between these segments, employing machine learning models like neural networks for image segmentation and recognition.

Benefits of technology

Enhances the accuracy and efficiency of locating misplaced items by leveraging image segment data and user commands, providing precise localization results through intelligent vision sensors and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025054716_28082025_PF_FP_ABST
    Figure EP2025054716_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure pertains to circuitry for localizing an object, wherein the circuitry is configured to obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CIRCUITRY, ELECTRONIC DEVICE AND METHOD

[0002] TECHNICAL FIELD

[0003] The present disclosure generally pertains to circuitry, an electronic device and a method.

[0004] TECHNICAL BACKGROUND

[0005] It is generally known to localize an object

[0006] In some instances, an Apple AirTag is attached to an object. The AirTag communicates wirelessly with nearby user devices such as smartphones. The user devices notify a cloud service of a position at which they have received a signal from the AirTag. Thus, a user looking for the object can query the cloud service for the position of the AirTag to localize the object.

[0007] Although there exist techniques for localizing an object, it is generally desirable to provide improved circuitry for localizing an object, an improved electronic device, and an improved method for localizing an object.

[0008] SUMMARY

[0009] According to a first aspect, the disclosure provides circuitry for localizing an object, wherein the circuitry is configured to obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

[0010] According to a second aspect, the disclosure provides an electronic device, that includes: circuitry for localizing an object, wherein the circuitry is configured to obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments; and a camera configured to acquire the image data.

[0011] According to a third aspect, the disclosure provides a method for localizing an object, wherein the method includes obtaining localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

[0012] Further aspects are set forth in the dependent claims, the drawings and the following description.

[0013] BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Embodiments are explained by way of example with respect to the accompanying drawings, in which:

[0015] Fig. 1 illustrates a first embodiment of an electronic device;

[0016] Fig. 2 illustrates a second embodiment of an electronic device;

[0017] Fig. 3 illustrates an embodiment of a method for generating image segment data;

[0018] Fig. 4 illustrates an embodiment of a method for localizing an object;

[0019] Fig. 5 illustrates an embodiment of localizing an object; and

[0020] Fig. 6 illustrates an embodiment of a general-purpose computer.

[0021] DETAILED DESCRIPTION OF EMBODIMENTS

[0022] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.

[0023] Some embodiments of the present disclosure pertain to circuitry for localizing an object, wherein the circuitry is configured to obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

[0024] The circuitry may include a programmed microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or the like. The circuitry may include a central processing unit (CPU) for executing instructions (e.g., software, firmware). The circuitry may include a graphics processing unit (GPU) and / or a tensor processing unit (TPU) for executing a machine learning model, e.g., an artificial neural network. The circuitry may include an intelligent vision sensor with artificial intelligence (Al) processing functionality, e.g., the IMX500 by Sony. The Al processing functionality may include executing a machine learning model such as an artificial neural network. The intelligent vision sensor may be included in the circuitry in addition or alternatively to any one of a CPU, a GPU and a TPU. The circuitry may include the general-purpose computer described below with respect to Fig. 6.

[0025] The circuitry may be used for localizing an object. For example, a user may have misplaced an object (e.g., a television (TV) remote, a keyring, a bag etc.). When looking for the misplaced object, the user may input a user command to the circuitry for identifying a location of the misplaced object. For example, as the user command, the user may utter a voice command, type a text message and / or perform a gesture. The voice command or text message may indicate an object for which the user is looking, e.g.: “Where is my TV remote?” or “Find my keyring!” or “Tell me where I put my bag.” The circuitry may then obtain localization data that are associated with the object. For example, the localization data may indicate a location of the object.

[0026] Based on the user command, the circuitry may generate user command data that indicate a label (object label) of the object. The object label may correspond to the designation which the user used in his user command for identifying the object (e.g., “TV remote”, “keyring”, “bag” etc.), and / or may represent a general concept in which a machine learning model executed by the circuitry may subsume the object identified in the user command. For example, the object label may be formatted as text data (e.g., a sequence of one or more characters according to a character encoding such as American Standard Code for Information Interchange (ASCII), Windows-1252, Unicode Transformation Format - 8-bit (UTF-8), UTF-16 or the like). For example, the object label may be formatted as audio data (e.g., a sequence of sound samples in the time domain and / or in the frequency domain). For example, the object label may be formatted as an embedding (e.g., a sequence of values indicating a vector in a vector space that may allow a machine learning model to analyze a relationship with another concept represented as a vector in the vector space).

[0027] The circuitry obtains the localization data based on image segment data. The image segment data may be based on image data. For example, a camera may acquire the image data by imaging an environment (e.g., one or more rooms, an apartment, an office, a garden, a street section, a park or the like). For example, the user may carry or wear the camera or an electronic device that may include the camera, such that the camera may image an environment in the surroundings of the user. Thus, the image data may represent one or more images of the environment. The image data may include the image(s) as grayscale image(s), as color image(s) (e.g., red-green-blue (RGB), YUV, hue-saturation-value (HSV), and / or hue-saturation-lightness (HSL)) or the like.

[0028] For generating the image segment data, an image segmentation may be performed on the image data, wherein for each detected object shown in an image of the image data, an image segment that may include the detected object may be extracted from the image. Thus, the set of image segments may represent objects detected in the image(s) of the image data. An object recognition may be performed, which may determine one or more labels (image segment labels) for each image segment of the set of image segments. The image segment label(s) may be associated with the image segment. For example, the image segment label(s) may represent a category, a designation, a description or the like of the respective detected object. Further, indications of relations between image segments of the set of image segments may be stored in the image segment data. In some embodiments, the image segment data also include portions of the image(s) of the image data and / or information that may allow to at least partially reconstruct portions of the image(s) of the image data.

[0029] The image segment labels may be stored in a format that may correspond to a format in which the user command label may be formatted, such that the circuitry may be able to compare the user command label with the image segment labels. For example, the image segment labels may be stored as text data, audio data and / or embeddings like the user command labels. Thus, it may be possible to determine a similarity (e.g., a difference and / or a distance in a vector space) of the user command label to the image segment labels.

[0030] For obtaining the localization data, the circuitry may compare the user command label with the image segment labels. For example, after generating the user command data, the circuitry may iterate through the image segment labels and compare each of them with the user command label. The image segment labels may be stored in a data structure that may facilitate the comparison with the user command label, e.g., in a tree data structure such as a binary search tree, a B-tree, a B+-tree or the like. The image segment labels may be stored in a database (e.g., in a relational database), and the comparison may be based on an index of the database. The image segment labels may be stored in one or more files (e.g., Comma-Separated Values (CSV), JavaScript Object Notation (JSON), Extensible Markup Language (XML), Hierarchical Data Format (HDF), Python pickle, Excel spreadsheet (XLS, XLSX), or the like). The circuitry may select an image segment label that has a highest similarity (e.g., smallest difference, shortest distance) to the user command label, and may determine that the image segment associated with the selected image segment label corresponds to the object which the user is looking for (e.g., which is indicated by the user command from the user).

[0031] The circuitry may then generate the localization data based on the relation of the image segment that is associated with the selected image segment label to another image segment of the set of image segments. For example, the other image segment may correspond to a portion of an image of the image data that may show another object. The other object may be related (e g., positionally) to the object the user is looking for. For example, the relation indicated by the image segment data may indicate that the other object may be located close to (e.g., below, above, beside, surrounding, (partially) covering or the like) the object the user is looking for. The circuitry may select, as the other image segment, an image segment of the set of image segments which the user may easily find. For example, the circuitry may select, as the other object shown by the image portion of the other image segment, a largest object in a surrounding of the object the user is looking for. For example, the circuitry may select, as the other object, an object that may be easily described, e.g., an object with a unique shape, unique color and / or unique pattern among objects detected in the set of image segments, and / or an object represented by an image segment, of the set of image segments, with a unique associated image segment label among the image segment labels of the image segment data. The skilled person may find further suitable criteria for selecting the other image segment.

[0032] The localization data may indicate a location of the object the user is looking for relative to the other object represented by the other image segment. For example, in a case where the user is looking for a key, and an image of the image data shows the key located on a table, the circuitry may select, as the image segment whose associated image segment label corresponds to the user command label, an image segment of the set of image segments which represents the key, and may select, as the other image segment, an image segment of the set of image segments which represents the table, because the circuitry may determine that the table is a large object in close proximity to the key the user is looking for. The circuitry may then generate the localization data such that the localization data indicate that the table is located below the key the user is looking for, and that, therefore, the key the user is looking for is located on the table.

[0033] In a case where the similarity between the selected image segment label and the user command label does not satisfy a predefined threshold (e.g., a difference or distance between the selected image segment label and the user command label is larger than the predefined threshold), the circuitry may determine that the object which the user is looking for cannot be found in the image data and, thus, that the circuitry cannot obtain corresponding localization data. In such a case, the circuitry may cause a notification to be issued to the user, which may indicate that the object cannot be found.

[0034] The image segment data may be generated and stored before the user starts looking for the object (and, thus, before the user inputs the user command, e g , utters the voice command, writes the text message or performs the gesture). For example, the image data may be acquired and the image segment data generated in advance and stored, such that the circuitry may be able to generate localization data based on the image segment data as soon as the circuitry receives the user command from the user. The generated image segment data may be stored for a predefined time (e.g., an hour, a day, a week or the like, without limiting the disclosure to these durations or to this time range), and may be deleted or overwritten after the predefined time has elapsed and / or a predefined amount of storage is occupied with newer image segment data.

[0035] The image data may be acquired by a camera that may be provided in a same electronic device as the circuitry. The image segment data may be generated by the circuitry. For example, in some embodiments where the circuitry includes an intelligent vision sensor (e.g., IMX500 by Sony), the image data and / or the image segment data may be generated by the intelligent vision sensor. The circuitry may also include a storage unit (e.g., based on flash memory, on Double Data Rate Synchronous Dynamic Random-Access Memory (DDR-SDRAM), on magnetic storage or the like) and may store the image segment data on the storage unit.

[0036] However, the disclosure is not limited to the circuitry acquiring the image data, generating the image segment data and / or storing the image segment data. For example, the image data may be acquired by a camera that may be provided separately from the circuitry and / or from an electronic device that includes the circuitry. For example, the image segment data may be generated by another device than the circuitry, e.g., by a (separate) camera acquiring the image data and / or by a computing device, e.g., a server in a cloud For example, the image segment data may be stored by another device than the circuitry, e.g., by a server in a cloud, by a network-attached storage (NAS) or the like. For example, the image segment data may be stored by a database server, and the circuitry may determine the image segment whose associated image segment label corresponds to the user command label by querying the database server based on the user command data and receiving, from the database server, a response which may indicate the image segment to be selected as well as its associated image segment label and its relation to one or more other image segments of the set of image segments.

[0037] In some embodiments, the user command data are based on a voice command inputted by a user.

[0038] As mentioned, the user may utter the voice command and may mention, in the voice command, the object the user is looking for. The voice command may be recorded by a microphone and may be analyzed based on speech recognition. The analyzing of the voice command may include determining that the user is indicating an object the user is looking for, and generating, based on the voice command, user command data that indicate a label of the object the user is looking for.

[0039] The microphone that may record the voice command may be included in an electronic device that may also include the circuitry, and the circuitry may perform the analyzing of the recorded voice command and the generating of the user command data. However, the disclosure is not limited to embodiments in which the circuitry analyzes the recorded voice command and generates the user command data. The analyzing of the voice command and the generating of the user command data may as well be performed by another device separate from the circuitry (and an electronic device in which the circuitry may be included).

[0040] In some embodiments, the circuitry is further configured to cause a machine learning model to identify the label in the voice command and generate the user command data in accordance with the identified label.

[0041] The machine learning model may include an artificial neural network (e g., a convolutional neural network (CNN), a recurrent neural network (RNN), a transformer etc.) and may be based on long short-term memory (LSTM), attention, or the like. For example, the machine learning model may include or may be based on Whisper by OpenAI, on Bidirectional Encoder Representations from Transformers (BERT) by Google, or the like.

[0042] The machine learning model may identify the user command label in the voice command by determining, in the voice command, one or more words that indicate the object the user is looking for, and converting these determined word(s) to a format specified for the user command label. The machine learning model may identify the user command label in the voice command by generating, based on the voice command, a value that may be formatted according to the format specified for the user command label. For example, the user command label may be formatted as text data, audio data and / or an embedding, as described above. The machine learning model may output the user command label as the user command data. The user command data may include further information in addition to the user command label.

[0043] The circuitry may execute the machine learning model, e.g., with its GPU, TPU, intelligent vision sensor or the like.

[0044] In some embodiments, the generating of the user command data includes converting at least a portion of the voice command into text data.

[0045] For example, the machine learning model (e.g., Whisper, BERT or the like) may convert the voice command (or a portion of the voice command that includes one or more words which indicate the object the user is looking for) to text data. The text data may correspond to the user command label, or the machine learning model may generate the user command label based on the text data.

[0046] In some embodiments, the set of image segments is based on an output of a machine learning model configured to extract the set of image segments from the image data by image segmentation.

[0047] The machine learning model for extracting the set of image segments from the image data may include an artificial neural network (e.g., including or based on Segment Anything Model (SAM) by Meta Al) that may be configured to segment any detected object in an image. For example, the machine learning model may receive one or more images of the image data, extract an image segment for each object it detects in the image(s), and output the extracted image segments. The machine learning model may output the image segments as images that show a content of the respective image segments, as data indicating a border of the respective image segments in the image(s), as data (e.g., a bitmap) that indicate which pixels of the image(s) correspond to the respective image segments, and / or as embeddings that may by suited to be inputted to a machine learning model configured to determine labels associated with the respective image segments.

[0048] After the image segment labels associated with the extracted image segments have been determined, the image segments and / or the image data may be deleted and / or may be included into the image segment data.

[0049] The circuitry may execute the machine learning model, e.g., with its GPU, TPU, intelligent vision sensor or the like, or the machine learning model may be executed by another device separate from the circuitry. In some embodiments, the image segment data are based on an output of a machine learning model configured to determine the labels associated with the set of image segments based on image recognition.

[0050] The machine learning model for determining the image segment labels may include an artificial neural network (e.g., including or based on Contrastive Language-Image Pre-training (CLIP) by OpenAI) that may be configured to generate a label for an image (e.g., for an image segment or a portion, of an image of the image data, that corresponds to the image segment) inputted to the machine learning model. For example, the machine learning model may be trained to classify the inputted image and output, as the label, one or more terms that may correspond to one or more best matching classes that have been determined by the classifying. The machine learning model may output the labels formatted as text data, as audio data and / or as embeddings.

[0051] The circuitry may execute the machine learning model, e.g., with its GPU, TPU, intelligent vision sensor or the like, or the machine learning model may be executed by another device separate from the circuitry.

[0052] In some embodiments, the image segment data include embeddings of the labels associated with the set of image segments; and the determining of the image segment whose associated label corresponds to the label indicated by the user command data is based on a difference between an embedding of the label indicated by the user command data and the embeddings of the labels associated with the set of image segments.

[0053] The embeddings may include a sequence (e.g., list, set) of values that indicate a vector in a vector space. The vector space may allow a machine learning model to analyze a relationship with another concept represented as a vector in the vector space. The relationship may be quantified as a similarity, difference and / or distance between vectors / embeddings. For example, a measure for quantifying the relationship may include a cosine distance, a Euclidean distance, a Manhattan distance or the like. The values of the vector may be integer values, floating point values or the like, with a suitable bit-length (e g., 8 bit, 16 bit, 32 bit, 64 bit or the like, without limiting the bit-length to these values, to this range of bit-lengths, or to bit-lengths that are multiples of 8).

[0054] The user command label may be generated as an embedding in a same vector space as the image segment labels, such that a similarity between the user command label and the image segment labels may correspond to a distance of the respective embeddings in the vector space.

[0055] Formatting the user command label and the image segment labels as embeddings may allow for robust comparison of the user command label and the image segment labels. For example, if the user is looking for an object that can be designated with any one of several different words, the embeddings that correspond to these different words may have a high similarity (e.g., a small distance in the corresponding vector space), such that the circuitry may be able to determine the image segment whose associated image segment label corresponds to the user command label irrespective of which of these different words the user uses in his user command (e g., voice command or text message) to identify the object.

[0056] Further, formatting the user command label and the image segment labels as embeddings may allow saving computing resources and / or electrical energy when generating and / or comparing the labels.

[0057] In some embodiments, the relation of the image segment to the other image segment is based on a position of the image segment with respect to a position of the other image segment.

[0058] As mentioned, the image segment data may indicate that the other object is located close to, below, above, beside, surrounding, (partially) covering or the like the object the user is looking for. For example, the positions of the image segment and of the other image segment may correspond to a coordinate in a pixel grid of an image of the image data, and the relation may indicate a distance in pixels between the positions, a direction between the positions, a coordinate vector from a first one of the positions to a second one of the positions, or the like. For example, the positions of the image segment and the other image segment may correspond to three-dimensional positions (e.g., determined based on Simultaneous Localization and Mapping (SLAM) or the like) of detected objects represented by the respective image segments in a three- dimensional space, and the relation indicated by the image segment data may indicate a distance, direction and / or coordinate vector between the object the user is looking for and the other object in the three-dimensional space.

[0059] Thus, the circuitry may be able to determine localization data that indicate a relation between the object the user is looking for and the other object.

[0060] In some embodiments, the circuitry is further configured to output a localization result according to the localization data.

[0061] The circuitry may output, as the localization result, an indication of a position of the object the user is looking for relative to the other object, such that the user may be able to look for the object in a proximity to the other object.

[0062] For example, if the user is looking for a key, and it is determined that the key is lying on a table, the localization data may indicate that a position of the key is above a position of the table. The circuitry may then output, as the localization result, an indication that the key is located on the table, such that the user can look on the table and find the key.

[0063] The circuitry may output the localization result as a voice notification through a loudspeaker (e g., a loudspeaker of an electronic device that includes the circuitry, or a loudspeaker wirelessly controlled by the circuitry, e.g., via Bluetooth or Wi-Fi). The voice notification may be generated by a machine learning model (e g., an artificial neural network, for example based on a generative large language model). The circuitry may output the localization result as a text message and / or an image (e.g., on a screen of an electronic device that includes the circuitry). The circuitry may output the localization result as annotations (e.g., arrows, lines, text and / or an image that may represent the object the user is looking for) via augmented reality (e g., on smartglasses or on a head-mounted display). The skilled person may find further suitable ways of outputting the localization result to the user.

[0064] Some embodiments pertain to an electronic device that includes: the circuitry of any one of the embodiments described above; and a camera configured to acquire the image data.

[0065] The electronic device may include a home assistant device (e.g., Amazon Echo), a smartphone, a laptop, a tablet, a smartwatch, smartglasses, a head-mounted display (HMD), a camera or the like.

[0066] The camera may include an image sensor and an optical element (e.g., lens, mirror). The optical element may focus incident light on the image sensor. The image sensor may include a plurality of photosensitive elements (e.g., photodiodes) that may generate an electrical signal corresponding to an amount of incident light. The image sensor may include circuitry that may generate image data based on the electrical signal. In some embodiments, the electronic device includes an intelligent vision sensor (e.g., IMX500 by Sony). The intelligent vision sensor may include the image sensor of the camera and / or the circuitry of any one of the embodiments described above.

[0067] The electronic device (e.g., its circuitry) may extract the set of image segments from one or more images of the image data acquired by the camera, determine image segment labels that are associated with the set of image segments, and generate the image segment data based on the image segment labels and on a relation between image segments of the set of image segments. In some embodiments, the electronic device may transmit the image data to another device (e.g., to a server in a cloud or, in a case where the electronic device is configured as a wearable device (e g., smartwatch, smartglasses, HMD), to a smartphone, laptop or tablet to which the electronic device may be coupled) and the other device may generate the image segment data based on the image data.

[0068] In some embodiments, the electronic device further includes a storage unit; and the circuitry is further configured to: obtain the image data from the camera; generate the image segment data based on the image data; and store the generated image segment data in the storage unit.

[0069] The storage unit may be based on flash memory, on DDR-SDRAM, on magnetic storage or the like, and may be included in the circuitry.

[0070] For generating the image segment data based on the image data, the circuitry may execute a machine learning model that may extract the set of image segments from the image data by image segmentation and / or execute a machine learning model that may determine the labels associated with the set of image segments based on image recognition, as described above.

[0071] The camera of the electronic device may acquire images of a surrounding of the electronic device, e.g., while a user is holding and / or wearing the electronic device, and may provide image data that include the images to the circuitry. The camera may acquire the images at a predefined rate (e.g., every second, every 30 seconds, every minute, every 5 minutes, every 10 minutes, every 30 minutes, or the like, without limiting the disclosure to these values or to this range) and / or if a change in the surrounding of the electronic device is detected (e.g., if objects are moved in the surrounding of the electronic device, if the camera is turned to another direction, if the user takes the electronic device to another room etc.). The camera may acquire the images even if no user command (e.g., voice command, text message or gesture), which indicates an object the user is looking for, has been received.

[0072] The circuitry may generate the image segment data based on the image data even if no user command, which indicates an object the user is looking for, has been received, and may store the generated image segment data in the storage unit. As mentioned, the image segment data may be deleted from the storage unit or overwritten when a predefined time (e.g., an hour, a day, a week or the like, without limiting the disclosure to these durations or to this time range) is elapsed and / or when a predefined amount of storage is occupied with newer image segment data.

[0073] Thus, the storage unit may store historical image segment data of surroundings of the user, such that, when the user inputs a user command (e.g., voice command, text message or gesture) that indicates an object the user is looking for, the circuitry may be able to obtain localization data based on the historical image segment data stored in the storage unit without having to wait for the camera to acquire image data. Therefore, even if the object the user is looking for is located in another room or building than a room or building in which the electronic device (and the user) is located when the user is inputting the user command, the circuitry may obtain localization data based on the historical (stored) image segment data without having to acquire image data of the other room or building as long as the storage unit stores image segment data from a previous time when the user (with the electronic device) was in the other room or building.

[0074] In embodiments where the circuitry (and / or the electronic device that includes the circuitry) acquires the image data, generates the image segment data based on the image data, stores the image segment data, generates the user command data, and obtains the localization data without transmitting data to an external device that is separate from the circuitry or electronic device, a privacy of the user may be protected because the image data may be kept in the circuitry or electronic device such that others may not be able to abuse the image data. Further, generating and storing the image segment data locally may allow to obtain localization data independently of any cloud access, e.g., without requiring access to a communication network (e g., the internet) for accessing the cloud, and / or without registering a user account for a cloud service.

[0075] As mentioned, in embodiments where the circuitry includes an intelligent vision sensor (e.g., IMX500 by Sony), the intelligent vision sensor may acquire the image data, generate the image segment data, generate the user command data, and obtain the localization data. For example, the intelligent vision sensor may execute one or more machine learning models (e.g., artificial neural networks) for extracting the set of image segments from the image data, determining the image segment labels associated with the set of image segments, and / or generating the user command data. The intelligent vision sensor may also determine, from the set of image segments, the image segment whose associated image segment label corresponds to the user command label indicated by the user command data.

[0076] The one or more machine learning models may be adapted to the intelligent vision sensor, e.g., to reduce a storage required to store the machine learning model(s), to reduce a power required for executing the machine learning model(s), to reduce an execution time of the machine learning model(s) on the intelligent vision sensor etc. For example, if the machine learning model(s) is / are based on an existing pre-trained artificial neural network, portions of the artificial neural network that are not required for obtaining the localization data may be removed in order to save storage. For example, the machine learning model(s) may be configured to support only input and output data formats used by the intelligent vision sensor, and support for further input or output formats may be dropped in order to save storage and / or increase an efficiency of the machine learning model(s). For example, a first machine learning model may be merged with a second machine learning model to a combined machine learning model if the second machine learning model uses an output of the first machine learning model as input (e g., if a machine learning model generates image segment labels based on a set of image segments outputted from another machine learning model), such that unnecessary output and / or input operation as well as unnecessary output and / or input layers may be avoided. The skilled person may find further ways of optimizing one or more machine learning models for an intelligent vision sensor.

[0077] As mentioned, the machine learning model(s) for extracting the set of image segments from the image data, for determining the image segment labels associated with the set of image segments and / or for determining the user command label may include an artificial neural network. The artificial neural network may include a Feed-Forward Network, a Residual Network (ResNet), a Recurrent Neural Network (RNN), a Convolutional Neural Network (CNN), a Generative Adversarial Network (GAN), a Transformer Neural Network and / or any other suitable neural network architecture. The skilled person may find a suitable architecture for the artificial neural network based on his expert knowledge. For example, the artificial neural network may include or may be configured similar to SAM, CLIP, Whisper, a large language model (LLM) such as BERT, generative pre-trained transformer (GPT), ChatGPT, Llama, or the like.

[0078] Some embodiments pertain to a method for localizing an object, wherein the method includes: obtaining localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

[0079] The circuitry described above may be configured to perform the method, and the method may be performed by the circuitry described above. Any embodiment of the circuitry described above may correspond to an embodiment of the method with corresponding features.

[0080] The features described above with respect to the circuitry, the electronic device and / or the method may be combined in any suitable way to form a corresponding embodiment of the circuitry, electronic device and / or method with the combined features.

[0081] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.

[0082] Returning to Fig. 1, Fig. 1 illustrates a first embodiment of an electronic device 10. The electronic device 10 includes an intelligent vision sensor 11, a storage unit 12, a microphone 13, a loudspeaker 14 and a lens 15.

[0083] The intelligent vision sensor 11 includes an array of photosensitive elements as well as a processing section. The lens 15 focuses incident light on the array of photosensitive elements, which generates electrical signals that correspond to the incident light. The processing section generates image data based on the electrical signals from the photosensitive elements. Therefore, the intelligent vision sensor 11 and the lens 15 constitute a camera 16. The processing section of the intelligent vision sensor 11 further executes instructions (software, firmware) as well as a machine learning model.

[0084] The storage unit 12 stores digital data that are generated and / or used by the intelligent vision sensor 11. For example, the digital data stored in the storage unit 12 include software instructions for the intelligent vision sensor 11, parameters of a machine learning model executed by the intelligent vision sensor 11, and image segment data.

[0085] The microphone 13 records sound (e.g., a voice command from a user of the electronic device 10) and generates sound data that represent the recorded sound. The sound data is provided to the intelligent vision sensor 11 for further processing.

[0086] The loudspeaker 14 emits sound (e g., a voice notification).

[0087] The intelligent vision sensor 11 and the storage unit 12 are an example of circuitry. The electronic device 10 and the circuitry is configured to perform any one of the methods described below with respect to Fig. 3 to 5. The electronic device 10 may further include a processor for controlling a function of the electronic device 10.

[0088] Fig. 2 illustrates a second embodiment of an electronic device 20. The electronic device 20 includes a CPU 21, an artificial intelligence (Al) unit 22, a camera 23, a storage unit 24, a microphone 25 and a loudspeaker 26.

[0089] The CPU 21 executes software instructions and controls a general functionality of the electronic device 20. The Al unit 22 includes a GPU and executes a machine learning model. The camera 23 includes a lens 23a and acquires image data, as described for the camera 16 of Fig. 1. The storage unit 24 stores digital data generated and / or used by the CPU 21 and / or the Al unit 22, e.g., software instructions for the CPU 21, parameters of a machine learning model executed by the Al unit 22, and image segment data.

[0090] The microphone 25 and the loudspeaker 26 provide functionality corresponding to the functionality described with respect to the microphone 13 and the loudspeaker 14, respectively, of Fig. 1.

[0091] The CPU 21, the Al unit 22 and the storage unit 24 are an example of circuitry. The electronic device 20 and the circuitry is configured to perform any one of the methods described below with respect to Fig. 3 to 5.

[0092] The electronic devices 10 and 20 may be configured as a home assistant device (e.g., Amazon Echo), a smartphone, a laptop, a tablet, a smartwatch, smartglasses, an HMD, a camera or the like. The electronic devices 10 and 20 may store a virtual assistant software (e.g., Amazon Alexa, Google Assistant, Siri by Apple, Bixby by Samsung), and the virtual assistant software may cause the electronic devices 10 and 20 to perform any one of the methods described with respect to Fig. 3 to 5.

[0093] Fig. 3 illustrates an embodiment of a method 30 for generating image segment data. The method 30 is an example of a method performed by the electronic device 10 or 20 described with respect to Fig. 1 or 2, respectively.

[0094] At 31, a camera (e.g., the camera 16 or 23) acquires image data.

[0095] At 32, circuitry (e g., the intelligent vision sensor 11 or the CPU 21) obtains the image data acquired at 31 from the camera.

[0096] At 33, the circuitry generates image segment data based on the image data obtained at 32.

[0097] The generating of the image segment data includes causing, at 34, an artificial neural network 35 (an example of a machine learning model) to extract a set of image segments from the image data by image segmentation. Therefore, the set of image segments is based on an output of the artificial neural network 35.

[0098] The generating of the image segment data further includes causing, at 36, the artificial neural network 35 to determine, based on image recognition, labels (image segment labels) associated with the set of image segments that has been extracted at 34. The artificial neural network 35 outputs embeddings of the image segment labels, and the circuitry includes the embeddings of the image segment labels in the image segment data such that the image segment data include the embeddings of the image segment labels. The generated image segment data indicate the image segment labels, which have been determined at 36, and indicate relations, in the image data, between image segments of the set of image segments. The relations between image segments of the set of image segments are based on coordinates of the image segments of the set of image segments within a pixel grid of an image of the image data, and are based on (relative) positions of the image segments with respect to each other.

[0099] Accordingly, the image segment data are based on an output of the artificial neural network 35.

[0100] At 37, the circuitry stores the image segment data generated at 33 in a storage unit (e g., the storage unit 12 or 24).

[0101] The artificial neural network 35 is, for example, executed by the intelligent vision sensor 11 or by the Al unit 22. It is noted that, in some embodiments, the extracting of the set of image segments at 34 and the determining of the image segment labels at 36 are performed by different machine learning models. Further, in some embodiments, the method 30 is performed by circuitry separate from the electronic devices 10 and 20.

[0102] Fig. 4 illustrates an embodiment of a method 40 for localizing an object. The method 40 is an example of a method performed by the electronic device 10 or 20 described with respect to Fig. 1 or 2, respectively.

[0103] At 41, circuitry (e g., the intelligent vision sensor 11 or the CPU 21) generates user command data.

[0104] The generating of the user command data includes obtaining, at 42, sound data from a microphone 43 (e.g., the microphone 13 or 25). The microphone 43 records, in the sound data, a voice command inputted by a user.

[0105] The generating of the user command data includes converting, at 44, the voice command into text data. It is noted that, in some embodiments, the whole voice command is converted into text data, and in some embodiments, only a portion of the voice command in which a mention of an object is detected is converted into text data.

[0106] The converting of (at least the portion of) the voice command into text data at 44 is performed by an artificial neural network 45 (an example of a machine learning model) configured to convert a voice command into text data, and the circuitry causes the artificial neural network 45 to convert (at least the portion of) the voice command into text data. The artificial neural network 45 is further configured to identify a label in text data, and the generating of the user command data includes causing the artificial neural network 45 to identify a label (user command label) in the voice command that has been converted into text data at 44.

[0107] The generating of the user command data further includes generating, at 46, the user command data in accordance with the user command label identified at 44. The generating of the user command data includes inserting the user command label into the user command data. In some embodiments, the user command label is inserted into the user command data in a text format, and in some embodiments, an embedding of the user command label is inserted into the user command data.

[0108] Accordingly, the user command data are based on the voice command inputted by the user and indicate the user command label of the object.

[0109] At 47, the circuitry obtains localization data associated with the object based on the user command data from 41 and on image segment data (e.g., the image segment data generated and stored by the method 30 of Fig. 3).

[0110] The obtaining of the localization data includes determining, at 48 and based on the image segment data, an image segment (of the set of image segments extracted at 34 of Fig. 3 and indicated by the image segment data) whose associated image segment label corresponds to the user command label indicated by the user command data from 41 (“corresponding image segment”).

[0111] The determining of the corresponding image segment at 48 includes determining, at 49, a difference (e g., a distance in a vector space) between an embedding of the user command label and the embeddings of the image segment labels. The user command label is indicated by the user command data (e.g., included in the user command data and / or generated based on the user command data). The image segment labels are determined at 36 of Fig. 3 and are indicated by the image segment data.

[0112] The determining of the corresponding image segment at 48 further includes selecting, at 50, as the corresponding image segment, an image segment of the set of image segments whose associated image segment label has an embedding with a smallest difference, among the set of image segments, to the embedding of the user command label.

[0113] Therefore, the determining of the corresponding image segment at 48 is based on a difference between an embedding of the user command label and the embeddings of the image segment labels. The obtaining of the localization data at 47 further includes generating, at 51, the localization data based on a relation of the (corresponding and selected at 50) image segment to another image segment of the set of image segments. As mentioned, the relation of the image segment to the other image segment is based on a position of the image segment with respect to a position of the other image segment.

[0114] At 52, the circuitry outputs a localization result according to the localization data through a loudspeaker 53 (e.g., the loudspeaker 14 or 26) as a voice notification, such that the user can hear the voice notification and find the object based on the localization result.

[0115] The artificial neural network 45 is, for example, executed by the intelligent vision sensor 11 or by the Al unit 22. It is noted that, in some embodiments, the converting of the voice command into text data at 44 and the identifying of the user command label at 46 are performed by different machine learning models. In some embodiments, the user command label is identified at 46 based on the voice command in an audio format (e.g., independent of any converting the voice command to text data), and the converting of the voice command to text data at 44 may be omitted.

[0116] Further, the method 30 of Fig. 3 (acquiring the image data, generating and storing the image segment data) may be performed before and / or after the generating of the user command data at 41. If the user command label from 41 is already known when the image segment data is generated at 33, each image segment label generated at 36 may directly be compared to the user command label according to 49 (difference between embeddings) such that storing the image segment labels at 37 may be unnecessary and may be omitted in such a case.

[0117] It is noted that the voice command is provided as an example of a basis on which to determine the user command label; and in some embodiments, the user command label is instead or in addition determined based on a text message from the user (e.g., entry in a webform, chat message in a conversation with a chatbot, entry in a command line prompt etc.) and / or based on a gesture of the user. The skilled persons may find further ways of generating the user command label.

[0118] Fig. 5 illustrates an embodiment of localizing an object. The localizing of the object is performed by a home assistant 60 (e.g., the electronic device 10 or 20 of Fig. 10 or 20, respectively). In the following, a case is described where the home assistant 60 corresponds to the electronic device 10 of Fig. 1 and includes, as the intelligent vision sensor 11, an IMX500 by Sony.

[0119] The home assistant 60 has stored three artificial neural networks 60a, 60b and 60c, which are examples of machine learning models, and executes the artificial neural networks 60a, 60b and 60c on the IMX500. The artificial neural network 60a is based on SAM by Meta Al and is configured (e.g., trained) to extract image segments from an image. The artificial neural network 60b is based on CLIP by OpenAI and is configured (e.g., trained) to determine labels (image segment labels) associated with images (e.g., image segments). The artificial neural network 60c is based on Whisper by OpenAI and is configured (e.g., trained) to convert voice input (audio data) into text data.

[0120] Image data 61, which include images, are acquired by the camera 16. The home assistant 60 generates image segment data 62 based on the image data 61. For generating the image segment data 62, the home assistant 60 provides the image data 61 to SAM 60a. SAM 60a extracts a set of image segments 62a to 62f from the image data 61. The home assistant 60 provides the set of image segments 62a to 62f to CLIP 60b. CLIP 60b determines an image segment label for each of the image segments 62a to 62f.

[0121] In the case illustrated in Fig. 5, the image segment label of the image segment 62a corresponds to a TV set, the image segment label of the image segment 62b corresponds to a lamp, the image segment label of the image segment 62c corresponds to a table, the image segment label of the image segment 62d corresponds to a TV remote, and the image segment labels of the image segments 62e and 62f, respectively, correspond to a chair. The image segment data 62 associate the set of image segments 62a to 62f with their respective image segment labels. As illustrated by lines connecting the image segments 62a to 62f, the image segment data 62 further indicate relations between image segments of the set of image segments 62a to 62f, wherein the relations are based on (relative) positions of the image segments 62a to 62f with respect to each other in an image of the image data 61. The image segment data 62 also indicate a size (in pixels) of the respective image segments 62a to 62f.

[0122] The home assistant 60 stores the image segment data 62 in a local database 63. The local database 63 is stored in the home assistant 60 (e.g., in the storage unit 12). Accordingly, the home assistant 60 performs the method 30 of Fig. 3 for generating the image segment data 62.

[0123] If a user utters a voice command 64 (e.g., “Where is my TV remote?”), the home assistant 60 records the voice command 64 with its microphone 13 and performs the method 40 of Fig 4 for localizing the TV remote.

[0124] The home assistant 60 causes Whisper 60c to convert the voice command 64 into text data. The home assistant 60 then identifies a user command label (which corresponds to “TV remote”) in text data of the voice command 64. The home assistant 60 queries the local database 63 for the user command label and determines that the image segment label of the image segment 62d corresponds to the user command label. The home assistant 60 also determines that a position of the image segment 62d is above a position of the image segment 62f, whose associated image segment label corresponds to a chair, and that the position of the image segment 62d is below a position of the image segment 62c, whose associated image segment label corresponds to a table.

[0125] Accordingly, the home assistant 60 generates localization data that indicate that the TV remote is located under the table on the chair. The home assistant 60 generates a voice notification 65 (e g., “The TV remote is located under the table on the chair.”), which indicates a localization result according to the localization data, and outputs the voice notification 65 through its loudspeaker 14.

[0126] It is noted that, although SAM 60a, CLIP 60b and Whisper 60c are mentioned as examples of the machine learning models 60a to 60c, the disclosure is not limited to these models. The skilled person may find other suitable machine learning models.

[0127] Fig. 6 illustrates an embodiment of a general -purpose computer 150. The general -purpose computer 150 can be implemented such that it can basically function as any type of electronic device (e.g., the electronic device 10 of Fig. 1 or the electronic device 20 of Fig. 2), for example, a home assistant device, a smartphone, smartglasses, a head-mounted display (HMD), a smartwatch, a mobile phone, a mobile tablet, a laptop / notebook, a camera, a terminal device or the like. The general -purpose computer 150 is an example of an electronic device that includes circuitry that is configured to perform the method according to the present technology (e.g., the method of Fig. 4 and / or 3). The computer has components 151 to 161, which can form a circuitry, such as any one of the units 11, 21, 22 or the like, as described herein.

[0128] Embodiments which use software, firmware, programs or the like for performing the methods as described herein can be installed on computer 150, which is then configured to be suitable for the concrete embodiment.

[0129] The computer 150 has a CPU 151 (Central Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 152, stored in a storage 157 and loaded into a random-access memory (RAM) 153, stored on a medium 160 which can be inserted in a respective drive 159, etc.

[0130] Furthermore, the computer 150 includes an artificial intelligence (Al) processor 151a (e.g., the intelligent vision sensor 11 of Fig. 1 or the Al unit 22 of Fig. 2). The Al processor 151a may include a graphics processing unit (GPU), a tensor processing unit (TPU) and / or an intelligent vision sensor (e.g., IMX500 by Sony). The Al processor 151a may be configured to execute an Al model (e.g., an artificial neural network), for example, the artificial neural network 35 of Fig. 3 and / or the artificial neural network 45 of Fig. 4.

[0131] The CPU 151, the ROM 152 and the RAM 153 are connected with a bus 161, which in turn is connected to an input / output interface 154. The number of CPUs, memories and storages is only exemplary, and the skilled person will appreciate that the computer 150 can be adapted and configured accordingly for meeting specific requirements which arise when it functions as an information processing apparatus according to the present technology.

[0132] At the input / output interface 154, several components are connected: an input 155, an output 156, the storage 157, a communication interface 158 and the drive 159, into which a medium 160 (compact disc (CD), digital video disc (DVD), universal serial bus (USB) flash drive, secure digital (SD) card, CompactFlash (CF) memory, or the like) can be inserted.

[0133] The input 155 can be a pointer device (mouse, graphic table, or the like), a keyboard, a microphone, a camera, a touchscreen, an eye-tracking unit etc.

[0134] The output 156 can have a display (liquid crystal display (LCD), cathode ray tube (CRT) display, light-emitting diode (LED) display, electronic paper, etc.; e.g., included in a touchscreen), loudspeakers, etc.

[0135] The storage 157 can have a hard disk drive (HDD), a solid-state drive (SSD), a flash drive and the like.

[0136] The communication interface 158 can be adapted to communicate, for example, via universal serial bus (USB), a serial port (RS-232), parallel port (IEEE 1284), a local area network (LAN; e.g., ethemet), wireless local area network (WLAN; e.g., Wi-Fi, IEEE 802.11), mobile telecommunications system (GSM, UMTS, LTE, NR etc ), Bluetooth, near-field communication (NFC), ZigBee, infrared, etc.

[0137] It should be noted that the description above only pertains to an example configuration of computer 150. Alternative configurations may be implemented with additional or other sensors, storage devices, interfaces or the like. For example, the communication interface 158 may support other radio access technologies than the mentioned UMTS, LTE and NR.

[0138] In some instances, in contemporary fast-paced lifestyles, a common struggle of locating misplaced belongings within a home remains a persistent challenge, often leading to valuable time lost and unnecessary frustration. Existing solutions, such as the Apple AirTag may require users to purchase separate tags for each individual item, leading to increased costs and resource waste. Moreover, in some instances, these solutions rely on external data processing, which may raise concerns about data security and compromising user privacy.

[0139] Some embodiments of the present disclosure aim to address these limitations by providing a (in some cases cost-effective and / or private) method for swiftly retrieving lost items within a home environment, offering an advanced alternative that may prioritize user privacy while ensuring a seamless recovery of items.

[0140] In some embodiments, a home assistant leverages the power of Sony’s IMX500 sensor to offer a camera-based solution for locating misplaced household items. A system that includes the home assistant may operate as an edge device and may enhance user privacy and convenience through voice-commanded retrieval of lost items, without a need for cloud processing.

[0141] In some embodiments, a proposed solution is a camera-based home assistant that includes a camera equipped with the IMX500 intelligent sensor, which provides the computational resources for the Segment Anything Model (SAM) and for the Contrastive Language-Image Pretraining (CLIP) model. This home assistant may run continuously to record scenes in a home, and all data may be saved locally without a need for internet-based services. The data processing may solely happen on the home assistant device using the IMX500 sensor.

[0142] An embodiment of a process of localizing a misplaced item may include four steps.

[0143] A first step of the process may include a voice-activated query. Users may initiate the process by issuing a simple voice command to the camera-based home assistant, prompting the system to begin a search for the misplaced item within a household, e g. “Where did I put my TV remote?”

[0144] A second step of the process may include voice command processing. A system equipped with the advanced IMX500 sensor may process the command entirely on an edge device (e.g., on the home assistant), ensuring that all data remain within secure boundaries of the user’s home. A general-purpose speech recognition model (e.g., Whisper) may be used to convert raw audio signals to spoken text.

[0145] A third step of the process may include object retrieval in historical data. Utilizing SAM and the CLIP framework, which may seamlessly integrate language and image features, the system may efficiently translate semantic information embedded in the user’s voice command to identify specific objects based on segmented results of SAM. This may enable the system to scan through historical video records captured by the camera-based home assistant device, pinpointing most recent occurrences of the requested item within the user’ s home environment. A fourth step of the process may include item localization. Through comprehensive analysis, a segmentation model based on SAM and CLIP may precisely pinpoint a last known location of the requested item and may, e.g., present the user with a clear visual representation on TV of where the item was last seen within the home or with a voice description of the location.

[0146] An advantage of some embodiments may lie in a segmentation technology, which may enable the system to effectively identify and locate any object within the household, regardless of the user’s query. By harnessing the SAM model in conjunction with the CLIP framework, the system may seamlessly integrate language and image features, ensuring precise object recognition. This approach may empower users to effortlessly retrieve a wide range of misplaced items through simple voice commands, providing a comprehensive and versatile solution that may cater to diverse user needs and preferences.

[0147] Furthermore, the system’s on-device processing capability, powered by the IMX500 sensor, may enhance user privacy and security. By performing all processing tasks locally, the system may eliminate the need for external data transfer, ensuring that sensitive user information may remain within the secure boundaries of the user’s home. This on-device processing may not only guarantee data privacy but may also streamline a retrieval process, enabling users to swiftly locate their belongings without compromising their personal data.

[0148] As mentioned, the methods disclosed herein may be performed by a home assistant device (like Amazon Echo or the like) which, however, may include a camera and may not require an internet access. In some embodiments, a product may contain a simple camera with an IMX500 sensor may be equipped with a microphone.

[0149] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. Changes of the ordering of method steps may be apparent to the skilled person.

[0150] Please note that the division of the electronic device 10 into units 11 to 15 (and of its circuitry into units 11 and 12) as well as the division of the electronic device 20 into units 21 to 26 (and of its circuitry into units 21, 22 and 24) is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. For instance, the units 11 and 12 as well as the units 21, 22 and 24 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.

[0151] A method for controlling an electronic device, such as the electronic devices 10 and 20 discussed above, is described above and under reference of Fig. 3, 4 and 5. The method can also be implemented as a computer program causing a computer and / or a processor, such as processors 11 and / or 21 discussed above, to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.

[0152] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.

[0153] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.

[0154] Note that the present technology can also be configured as described below.

[0155] (1) Circuitry for localizing an object, the circuitry being configured to: obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

[0156] (2) The circuitry of (1), wherein the user command data are based on a voice command inputted by a user.

[0157] (3) The circuitry of (1) or (2), further configured to: cause a machine learning model to identify the label in the voice command and generate the user command data in accordance with the identified label. (4) The circuitry of (3), wherein the generating of the user command data includes converting at least a portion of the voice command into text data.

[0158] (5) The circuitry of any one of (1) to (4), wherein the set of image segments is based on an output of a machine learning model configured to extract the set of image segments from the image data by image segmentation.

[0159] (6) The circuitry of any one of (1) to (5), wherein the image segment data are based on an output of a machine learning model configured to determine the labels associated with the set of image segments based on image recognition.

[0160] (7) The circuitry of any one of (1) to (6), wherein the image segment data include embeddings of the labels associated with the set of image segments; and wherein the determining of the image segment whose associated label corresponds to the label indicated by the user command data is based on a difference between an embedding of the label indicated by the user command data and the embeddings of the labels associated with the set of image segments.

[0161] (8) The circuitry of any one of (1) to (7), wherein the relation of the image segment to the other image segment is based on a position of the image segment with respect to a position of the other image segment.

[0162] (9) The circuitry of any one of (1) to (8), further configured to: output a localization result according to the localization data.

[0163] (10) An electronic device, comprising: the circuitry of any one of (1) to (9); and a camera configured to acquire the image data.

[0164] (11) The electronic device of (10), further comprising a storage unit; wherein the circuitry is further configured to: obtain the image data from the camera; generate the image segment data based on the image data; and store the generated image segment data in the storage unit. (12) A method for localizing an object, the method comprising: obtaining localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

[0165] (13) The method of (12), wherein the user command data are based on a voice command inputted by a user.

[0166] (14) The method of (12) or (13), further comprising: causing a machine learning model to identify the label in the voice command and generate the user command data in accordance with the identified label.

[0167] (15) The method of ( 14), wherein the generating of the user command data includes converting at least a portion of the voice command into text data.

[0168] (16) The method of any one of (12) to (15), wherein the set of image segments is based on an output of a machine learning model configured to extract the set of image segments from the image data by image segmentation.

[0169] (17) The method of any one of (12) to (16), wherein the image segment data are based on an output of a machine learning model configured to determine the labels associated with the set of image segments based on image recognition.

[0170] (18) The method of any one of (12) to (17), wherein the image segment data include embeddings of the labels associated with the set of image segments; and wherein the determining of the image segment whose associated label corresponds to the label indicated by the user command data is based on a difference between an embedding of the label indicated by the user command data and the embeddings of the labels associated with the set of image segments.

[0171] (19) The method of any one of (12) to (18), wherein the relation of the image segment to the other image segment is based on a position of the image segment with respect to a position of the other image segment.

[0172] (20) The method of any one of (12) to (19), further comprising: outputting a localization result according to the localization data.

[0173] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (12) to (20), when being carried out on a computer. (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (12) to (20) to be performed.

Claims

CLAIMS1. Circuitry for localizing an object, the circuitry being configured to: obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

2. The circuitry of claim 1, wherein the user command data are based on a voice command inputted by a user.

3. The circuitry of claim 1, further configured to: cause a machine learning model to identify the label in the voice command and generate the user command data in accordance with the identified label.

4. The circuitry of claim 3, wherein the generating of the user command data includes converting at least a portion of the voice command into text data.

5. The circuitry of claim 1, wherein the set of image segments is based on an output of a machine learning model configured to extract the set of image segments from the image data by image segmentation.

6. The circuitry of claim 1, wherein the image segment data are based on an output of a machine learning model configured to determine the labels associated with the set of image segments based on image recognition.

7. The circuitry of claim 1, wherein the image segment data include embeddings of the labels associated with the set of image segments; and wherein the determining of the image segment whose associated label corresponds to thelabel indicated by the user command data is based on a difference between an embedding of the label indicated by the user command data and the embeddings of the labels associated with the set of image segments.

8. The circuitry of claim 1, wherein the relation of the image segment to the other image segment is based on a position of the image segment with respect to a position of the other image segment.

9. The circuitry of claim 1, further configured to: output a localization result according to the localization data.

10. An electronic device, comprising: circuitry for localizing an object, the circuitry being configured to: obtain localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segments extracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments; and a camera configured to acquire the image data.

11. The electronic device of claim 10, further comprising a storage unit; wherein the circuitry is further configured to: obtain the image data from the camera; generate the image segment data based on the image data; and store the generated image segment data in the storage unit.

12. A method for localizing an object, the method comprising: obtaining localization data associated with the object based on user command data and on image segment data; wherein the user command data indicate a label of the object; wherein the image segment data indicate labels associated with a set of image segmentsextracted from image data and indicate relations, in the image data, between image segments of the set of image segments; and wherein the obtaining of the localization data includes: determining, based on the image segment data, an image segment whose associated label corresponds to the label indicated by the user command data; and generating the localization data based on a relation of the image segment to another image segment of the set of image segments.

13. The method of claim 12, wherein the user command data are based on a voice command inputted by a user.

14. The method of claim 12, further comprising: causing a machine learning model to identify the label in the voice command and generate the user command data in accordance with the identified label.

15. The method of claim 14, wherein the generating of the user command data includes converting at least a portion of the voice command into text data.

16. The method of claim 12, wherein the set of image segments is based on an output of a machine learning model configured to extract the set of image segments from the image data by image segmentation.

17. The method of claim 12, wherein the image segment data are based on an output of a machine learning model configured to determine the labels associated with the set of image segments based on image recognition.

18. The method of claim 12, wherein the image segment data include embeddings of the labels associated with the set of image segments; and wherein the determining of the image segment whose associated label corresponds to the label indicated by the user command data is based on a difference between an embedding of the label indicated by the user command data and the embeddings of the labels associated with the set of image segments.

19. The method of claim 12, wherein the relation of the image segment to the other image segment is based on a position of the image segment with respect to a position of the other image segment.

20. The method of claim 12, further comprising: outputting a localization result according to the localization data.

Citation Information

Patent Citations

  • Object searching method and system for visually impaired people

    CN115218903A

  • Wearable eyeglasses for providing social and environmental awareness

    US9922236B2