Device, method and computer program
A device with image segmentation and zero-shot classification aids users in locating objects in real-world environments by providing feedback, enhancing object identification and robot assistance.
Patent Information
- Application Number
- PCT/EP2025/054529
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-02-20
- Publication Date
- 2025-09-11
AI Technical Summary
Individuals, especially those with visual impairments or in unfamiliar environments, face challenges in identifying and locating items within their surroundings.
A device equipped with a camera, processor, and zero-shot classification model performs image segmentation and object classification on a real-world video stream, allowing users to provide inputs for feedback on the object's location using audible, tactile, or visual cues.
Enables efficient object identification and location in real-world scenes without requiring internet connectivity, aiding users in unfamiliar environments and assisting robots in object retrieval.
Smart Images

Figure EP2025054529_12092025_PF_FP_ABST
Abstract
Description
[0001] DEVICE, METHOD AND COMPUTER PROGRAM
[0002] BACKGROUND
[0003] Field of the Disclosure
[0004] The present technique relates to a device, computer program and method.
[0005] Description of the Related Art
[0006] The “background” description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in the background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly or impliedly admitted as prior art against the present technique.
[0007] People in unfamiliar environments often find it difficult identifying and locating items within the environment. For example, a person who is visually impaired may find an environment such as a supermarket or street challenging. Similarly, a person who is not visually impaired but is looking for an item in an unfamiliar environment may find it difficult.
[0008] It is an aim of the disclosure to at least address this issue.
[0009] SUMMARY
[0010] According to one aspect of the disclosure, there is provided a device comprising circuitry configured to: receive a video of a real-world scene containing an object; perform image segmentation on each frame in the video; classify the segmented object using a zero-shot classification model; receive a user input and compare the user input with the classified object; and in the event of a positive comparison, provide feedback to indicate the location of the object in the real-world scene.
[0011] The foregoing paragraphs have been provided by way of general introduction, and are not intended to limit the scope of the following claims. The described embodiments, together with further advantages, will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein: Figure 1 shows a pair of Augmented Reality (AR) glasses 100 containing a device 200 which is connected to a camera 110 in the AR glasses;
[0014] Figure 2 shows a device 200 according to embodiments of the disclosure;
[0015] Figure 3 A shows a real-life scene 305 A captured as a video by the camera 110 in the AR glasses 100 worn by the user;
[0016] Figure 3B shows the real life scene captured by the camera 110 after image segmentation has been performed;
[0017] Figure 4 shows feedback according to embodiments;
[0018] Figure 5 shows a flowchart 500 explaining embodiments of the disclosure; and
[0019] Figure 6 shows a robot 600 according to embodiments with a device 200 according to embodiments embedded therein.
[0020] DESCRIPTION OF THE EMBODIMENTS
[0021] Referring now to the drawings, wherein like reference numerals designate identical or corresponding parts throughout the several views.
[0022] Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims, the disclosure may be practiced otherwise than as specifically described herein.
[0023] Figure 1 shows a pair of Augmented Reality (AR) glasses 100 containing a device 200 which is connected to a camera 110 in the AR glasses 100. The camera 110 is positioned on the bridge of the AR glasses 100 and captures images in front of the user of the AR glasses 100. In embodiments, the camera 110 captures a video stream of everything the user of the AR glasses sees. The device 200 receives the video stream of the real-world scene from the camera 110.
[0024] Although the foregoing describes the device 200 being located in AR glasses, the disclosure is in no way limited to this. For example, the device 200 could be provided in a mobile telephone, video camera, laptop computer or the like. Indeed, in embodiments, the device 200 could be a server or other computer equipment such as a tablet computer which receives a video of a real world scene from a camera.
[0025] Figure 2 shows a device 200 according to embodiments of the disclosure. The device 200 comprises a processor 210 which is embodied as circuitry and may be any solid state circuitry such as circuitry controlled by software or an application specific integrated circuit. The processor 210 is connected to the camera 110. This connection may be over a wired or wireless connection. The processor 210 is also connected to storage 220. In embodiments, the storage 220 is solid-state storage, but is not limited and may be optically readable storage or the like. Moreover, the storage 220 may be located remote to the device 200. In embodiments, the storage 220 contains computer readable instructions which, when loaded onto the processor 210, configures the device 200 to perform a method or methods according to embodiments of the disclosure.
[0026] In embodiments, the device 200 receives a user input and provides feedback to the user. Accordingly, although not shown, the processor 210 of the device 200 is connected to a user input which may be a microphone, a keyboard or other mechanism that allows a user to provide an input and a feedback mechanism such as a speaker, haptic feedback mechanism or a display or the like that allows the user to receive feedback.
[0027] Figure 3 A shows a real-life scene 305 A captured as a video by the camera 110 in the AR glasses 100 worn by the user. In the embodiment of Figure 3 A, the real-life scene is a plate containing several food items include salad leaves, chick-peas and beetroot chunks. The video stream is provided to the device 200 so that the device 200 receives the video of the real-world scene containing the plate and various food items. Specifically, the video is received by the processor 210 located within the device 200.
[0028] The processor 210 within the device 200 performs image segmentation on each image in the video. In other words, the processor 210 performs image segmentation on each frame of video to identify where one or more objects are within each image of the video. In embodiments, the processor 210 uses the Segment Anything Model (SAM) created by Meta Al ® to segment the image into objects and to identify the location of each object within each image of the video. The purpose of using SAM is that it can be carried out in the device 200 and there is no requirement for the device 200 to be connected to the internet. Of course, although the processor 210 in embodiments uses SAM, the disclosure is not so limited and the use of any appropriate image segmentation algorithm is envisaged such as Mockups by Glorify ® or ruttl or the like is envisaged.
[0029] Figure 3B shows the real life scene captured by the camera 110 after image segmentation has been performed. In other words, Figure 3B shows an object segmented real life scene 305B. As will be evident from Figure 3B, where the outline of each object has been highlighted, each individual object in the image has been isolated.
[0030] The processor 210 then generates a segmentation map where the location of each isolated object in the image is stored in association with an identifier for the isolated object. The location may be a pixel position of the isolated object or may be any kind of position information that will enable the position of the object in the real-world scene relative to the position of the user capturing the image to be determined.
[0031] The processor 210 then classifies each isolated object using a zero-shot classification model such as Contrastive Language Image Pre-training (CLIP) developed by OpenAI ®. The benefit of using a zeroshot classification model to classify each object is that the model does not need to be specifically trained on known objects and so is much better at classifying objects on which the model has not been specifically trained. The processor 210 assigns a label to each object to uniquely identify each object in the image.
[0032] Accordingly, the processor 210 creates, for each image in the video, an inventory (or table) of identified objects, the classification and their respective positions within the frame. An example of the inventory is provided in Table 1 below:
[0033] Table 1
[0034] The table is stored in storage 220.
[0035] In the embodiments of table 1, the label is based upon the classification provided by the zero-shot classification model. Accordingly, there is no need for a separate classification column as the classification can be derived from the object identifier. This save storage space within storage 220.
[0036] After the table has been created and stored in storage 220, a user input is received by device 200. Specifically, the user input is received by the processor 210. In embodiments, the user input is spoken by the user and so is received via a microphone. However, the disclosure is not so limited and the user input may be textual or may be through a mouse or gesture recognition or the like.
[0037] In embodiments, the user input relates to one or more of the objects isolated from the image of video. In other words, in embodiments, when the user provides a user input, the processor 210 compares the user input with the classified object and if there is a positive comparison between the user input and the classified object, the processor 210 provides feedback to the user indicating the location of the classified object within the real-life scene.
[0038] In embodiments, the user input may be a verbal command such as “does the plate include beetroot?”. A speech recognition toolkit such as Vosk is provided in the storage 220 which is used by the processor 210 to extract the command in the user input. The processor 210 compares the command to the stored classified objects and in the event of a positive comparison, provides feedback to the user.
[0039] Feedback according to embodiments is described with reference to Figure 4. In Figure 4, a first image 410A in a video is captured and image segmentation and classification is performed and the information in table 1 is populated. The user then issues the verbal command “does this plate include beetroot?”. The processor 210 interrogates the table and identifies that beetroot is a classified object in the image and so provides feedback to the user to indicate the location of beetroot in the real-world. This is shown in 420A.
[0040] In embodiments, the processor 210 guides the user’s hand to the location of the classified object. However, in the first image 410A, no hand is seen. Accordingly, the feedback 420A to the user is that the user’s hand is not in the image and that the user’s hand needs to be placed in the field of view of the camera 110.
[0041] The user then moves their hand 405 into the camera field of view in the second image 410B. This is segmented and classified as an object as described above. Accordingly, as the processor 210 knows the location of the user’s hand, the processor 210 determines the direction the user must move their hand to locate the object in the real-world scene. Specifically, the processor 210 provides feedback 420B that the user must move their hand to the left.
[0042] The user has moved their hand in the third image 410C. Again, the user’s hand is segmented and classified as an object and so the processor 210 knows the location of the user’s hand and the direction the user must move their hand to locate the object in the real-world scene. The processor 210 provides feedback 420C that the user must move their hand forwards.
[0043] The user has moved their hand in the fourth image 410D. The user’s hand is segmented and classified as an object and so the processor 210 knows the location of the user’s hand and the direction the user must move their hand to locate the object in the real-world scene. In this case, the user’s hand is above the object. Therefore, the processor 210 provides feedback 420D that the user’s hand is above the object in the real-life scene.
[0044] In embodiments, the processor 210 provides audible feedback via a speaker (such as audio or voice instructions) directing the user’s hand to the object. In embodiments, the processor 210 provides tactile feedback via haptic feedback located on the user’s body or face or on the device 200 or the like. This haptic feedback directs the user’s hand to the object by applying a vibration indicating the direction the hand needs to move or other feedback. Of course, other types of feedback is envisaged such as visual feedback for example lights or a message on a display or any combination of feedback mechanisms such as haptic feedback and audible feedback or audible feedback and visual feedback is envisaged.
[0045] Although the above feedback directs the user’s hand to the object by recognising the hand and directing the hand, the disclosure is not so limited. In embodiments, the feedback may simply indicate that location of the object in the real-life scene. For example, in the embodiments of Figure 4, rather than directing the user’s hand, the processor 210 may simply provide the feedback that “the beetroot is on the upper left side of the plate”. This feedback may be provided audibly or via a tactile feedback mechanism.
[0046] By performing image segmentation on each image in the video to locate the object in each image and then classify the object(s) using a zero-shot classification model, it is possible to provide feedback to the user so that the user can locate the object in the real world scene. Moreover, by using a zero-shot classification model, two distinct advantages are provided. Firstly, no active internet connection is required which provides the user with the freedom to use the device 200 in areas with little or no internet coverage and secondly, zero-shot classification models are adept at classifying objects not previously detected and trained on by the model. This is especially advantageous where the object(s) is / are in a new environment for the user.
[0047] There are numerous applications for a device 200 according to embodiments. The device 200 may be used by a user who is in an unfamiliar environment such as a supermarket. The user may provide a shopping list to the device 200 and as the device 200 classifies objects from the shopping list in its stored table as the user walks around the supermarket, the device 200 may provide feedback to the user directing the user to each item in the shopping list.
[0048] Similarly, there may be objects which are desired by the user but whose presence is difficult to establish due to a plethora of other objects being in the same location. One example would be a forest where the user wishes for a particular mushroom to be identified. The user can provide a command such as “please tell me when I see a button mushroom” and in the event that the button mushroom is captured by the camera and so becomes a classified object in the list, the device will provide feedback indicating the presence of the button mushroom indicating the directions to the object. This allows a user to identify objects that would be normally difficult to see due to a plethora of other objects.
[0049] The device 200 may be used to assist a user in identifying an object which is unfamiliar to the user in a more timely manner. Typically, it takes longer for a user to identify an object that is unfamiliar to that user compared with an object that is familiar to that user, even if the user is aware of the shape of the unfamiliar object. This is because the user takes longer to review all the objects in a particular scene to identify the unfamiliar object. One example would be in a warehouse environment where a user must identify an unfamiliar product from many hundreds of other products. The user can provide a command such as “please tell me when I see a pair of grape scissors”. The user can walk around the warehouse at their normal pace and when the device identifies the grape scissors, the device will provide feedback indicating the presence of the grape scissors indicating the directions to the object. Similarly, in an agricultural setting, a user who is unfamiliar with various plants may list the desired plant before walking into an area and in the event that the desired plant is identified, the device will provide feedback to the user to direct the user to the desired plant. Figure 5 shows a flowchart 500 explaining embodiments of the disclosure. The process starts in step 505 and then moves to step 510. In step 510 the processor 210 receives a video of a real-world scene containing an object. This is received, in embodiments, from camera 110 but may be received from a different camera or from storage or the like. The process moves to step 515 where image segmentation is performed on each image in the video to locate the object in each frame. As noted above, in embodiments, this is SAM. The process moves to step 520 where the segmented object is classified using a zero-shot classification model. The process moves to step 525 where a user input is received and the user input is compared with the classified object. The process moves to step 530 where in the event of a positive comparison, feedback is provided to indicate the location of the object in the real -world scene. The process moves to step 535 where the process ends.
[0050] Although the foregoing has been described with reference to providing feedback to a human, the disclosure is not so limited. In particular, a user may give a command to a robot into which the device 200 according to embodiments is embedded. The user’s command to the robot may be “please find a red pair of socks in my sock drawer”. Rather than giving the feedback to the user, however, the feedback is provided to the robot so that the robot’s movement system can move the robot. The device 200 may be integrated, in embodiments, into the robot’s vision system and the feedback will be provided to one or more robot actuator to manoeuvre the robot to pick up the identified object by controlling the robot’s movement system. In embodiments, therefore, the feedback will be comprised of movement instructions for the actuators controlling a robot’s moving mechanism such as limbs or wheels rather than a tactile or audible feedback or in addition to tactile or audible feedback.
[0051] Figure 6 shows a robot 600 according to embodiments with a device 200 according to embodiments embedded therein. In embodiments, the robot 600 may be aibo ® developed by Sony Corporation ® or a humanoid shaped robot. In these embodiments, the robot’s moving mechanism are limbs, but the disclosure is not so limited and may be wheels or tracks or a combination of any of these.
[0052] Accordingly, the feedback in these embodiments is provided to an actuator which controls the robot’s moving mechanism.
[0053] In so far as embodiments of the disclosure have been described as being implemented, at least in part, by software-controlled data processing apparatus, it will be appreciated that a non-transitory machine- readable medium carrying such software, such as an optical disk, a magnetic disk, semiconductor memory or the like, is also considered to represent an embodiment of the present disclosure.
[0054] It will be appreciated that the above description for clarity has described embodiments with reference to different functional units, circuitry and / or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuitry and / or processors may be used without detracting from the embodiments. Described embodiments may be implemented in any suitable form including hardware, software, firmware or any combination of these. Described embodiments may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of any embodiment may be physically, functionally and logically implemented in any suitable way. Indeed the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the disclosed embodiments may be implemented in a single unit or may be physically and functionally distributed between different units, circuitry and / or processors.
[0055] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in any manner suitable to implement the technique.
[0056] Embodiments of the present technique can generally described by the following numbered clauses:
[0057] 1. A device comprising circuitry configured to: receive a video of a real-world scene containing an object; perform image segmentation on each frame in the video; classify the segmented object using a zero-shot classification model; receive a user input and compare the user input with the classified object; and in the event of a positive comparison, provide feedback to indicate the location of the object in the real-world scene.
[0058] 2. A device according to clause 1, wherein the circuitry is configured to: provide feedback to the user.
[0059] 3. A device according to clause 2, wherein the circuitry is configured to: provide the feedback to the user to indicate the location of the object in the image.
[0060] 4. A device according to any preceding clause, wherein the feedback is audible feedback.
[0061] 5. A device according to any one of clauses 1 to 3, wherein the feedback is tactile feedback.
[0062] 6. A robot comprising a moving mechanism, an actuator controlling the moving mechanism and a device according to any preceding clause, wherein the feedback is provided to the actuator to control the moving mechanism.
[0063] 7. A method performed in a device, the method comprising: receiving a video of a real-world scene containing an object; performing image segmentation on each frame in the video; classifying the segmented object using a zero-shot classification model; receiving a user input and compare the user input with the classified object; and in the event of a positive comparison, providing feedback to indicate the location of the object in the real-world scene.
[0064] 8. A method according to clause 7, wherein the providing step comprises: providing feedback to the user.
[0065] 9. A method according to clause 8, wherein the providing step further comprises: providing the feedback to the user to indicate the location of the object in the image.
[0066] 10. A method according to any one of clauses 7 to 9, wherein the feedback is audible feedback.
[0067] 11. A method according to any one of clauses 7 to 9, wherein the feedback is tactile feedback.
[0068] 12. A computer program product comprising computer readable instructions which, when loaded onto a computer, configures the computer to perform a method according to any one of clauses 7 to 11.
Claims
CLAIMS1. A device comprising circuitry configured to: receive a video of a real-world scene containing an object; perform image segmentation on each frame in the video; classify the segmented object using a zero-shot classification model; receive a user input and compare the user input with the classified object; and in the event of a positive comparison, provide feedback to indicate the location of the object in the real-world scene.
2. A device according to claim 1, wherein the circuitry is configured to: provide feedback to the user.
3. A device according to claim 2, wherein the circuitry is configured to: provide the feedback to the user to indicate the location of the object in the image.
4. A device according to claim 1, wherein the feedback is audible feedback.
5. A device according to claim 1, wherein the feedback is tactile feedback.
6. A robot comprising a moving mechanism, an actuator controlling the moving mechanism and a device according to claim 1, wherein the feedback is provided to the actuator to control the moving mechanism.
7. A method performed in a device, the method comprising: receiving a video of a real-world scene containing an object; performing image segmentation on each frame in the video; classifying the segmented object using a zero-shot classification model; receiving a user input and compare the user input with the classified object; and in the event of a positive comparison, providing feedback to indicate the location of the object in the real-world scene.
8. A method according to claim 7, wherein the providing step comprises: providing feedback to the user.9, A method according to claim 8, wherein the providing step further comprises: providing the feedback to the user to indicate the location of the object in the image.
10. A method according to claim 7, wherein the feedback is audible feedback.
11. A method according to claim 7, wherein the feedback is tactile feedback.
12. A computer program product comprising computer readable instructions which, when loaded onto a computer, configures the computer to perform a method according to claim 7.
Citation Information
Patent Citations
Image processing techniques to quickly find a desired object among other objects from a captured video scene
US20220343647A1