Wearable device and article searching method

Through the image collector and graphics of the wearable device, the large model recognizes and stores the position relationship of items, and generates prompt voice, solving the problem that existing devices cannot quickly find common items, and achieving efficient and accurate item positioning.

CN120451254APending Publication Date: 2025-08-08HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377070.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing wearable devices such as smart watches and smart glasses can only achieve a single interaction function with the mobile phone, and cannot help users find commonly used items quickly and accurately. The search methods for existing items are inefficient and prone to errors, especially in large spaces or messy environments.

Method used

Video is recorded in real time through the wearable device's image collector, and the graphics and text understanding model is used to identify the relative position relationship between key items and other items, save the location information to the storage list, and generate a prompt voice to help users find items.

Benefits of technology

It realizes the rapid and accurate search of a variety of daily life items, reduces resource usage, supports natural voice command interaction, and improves the efficiency and accuracy of item search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451254A_ABST
    Figure CN120451254A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses wearable equipment and an article searching method. The method comprises the following steps: acquiring a video recorded by an image collector; key frames in the video are input into an image-text understanding large model, position information of the key article is obtained, the position information comprises a relative position relation between the key article and a first article, and the first article comprises articles except the key article; storing the position information of the key article in a storage list; after an instruction of searching for the first key article input by a user is received, position information of the first key article is obtained from the storage list; and generating prompt voice based on the position information of the first key article, and controlling a loudspeaker to play the prompt voice. According to the application, the image collector of the wearable device records the video and the image-text in real time to understand the recognition capability of the large model on the video image, the related position relationship between the required article and other articles can be recognized and stored, and the position relationship can help a user to quickly search for various daily life articles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of device control technology, and in particular to a wearable device and an object search method. Background Art

[0002] In daily life, people often encounter the problem of not being able to find commonly used items, such as keys, mobile phones, and remote controls. These small items are often difficult to find when needed because they are used frequently and are often placed casually. Existing methods of finding items mostly rely on manual memory or simple markings, such as hanging keys on fixed hooks or setting a specific storage location for remote controls. However, this method is not only inefficient but also prone to errors. In large spaces or environments where items are placed in a disorderly manner, it becomes even more difficult to find the required items quickly and accurately.

[0003] Existing wearable devices, such as smartwatches and smart glasses, are mostly limited in functionality, capable of interacting only with specific devices like mobile phones. For example, these devices typically pair with a user's smartphone via Bluetooth or Wi-Fi, enabling a single function like "Find My Phone." If a user can't find their phone, simply tapping a button on their smartwatch triggers a tone on the phone, helping them quickly locate it.

[0004] However, the original intention of designing this function was only to meet the needs in the scenario of lost mobile phone, and it does not cover the search for other common items. Summary of the Invention

[0005] Some embodiments of the present application provide a wearable device and an item search method. By using the image collector of the wearable device to record video in real time and the large-scale model of graphic and text understanding to recognize video images, the relative positional relationship between the required item and other items can be identified and stored. This positional relationship can help users quickly find a variety of daily life items.

[0006] In a first aspect, some embodiments of the present application provide a wearable device, including:

[0007] an image collector configured to record video;

[0008] The controller is configured as:

[0009] Get the video recorded by the image collector;

[0010] Inputting key frames in the video into the large image-text understanding model to obtain the location information of the key items, the location information including the relative position relationship between the key items and the first items, and the first items including items other than the key items;

[0011] Save the location information of key items to the storage list;

[0012] After receiving a user input instruction to search for a first key item, obtaining location information of the first key item from a storage list;

[0013] A prompt voice is generated based on the location information of the first key item, and the speaker is controlled to play the prompt voice.

[0014] The above technical solution has the following advantages or beneficial effects: through the real-time video recording by the image collector of the wearable device and the recognition ability of the large model of graphic understanding of the video image, the relative positional relationship between the required items and other items can be identified and stored, and this positional relationship can help users quickly find a variety of daily life items.

[0015] In some embodiments, the controller is further configured to:

[0016] If the key item to be stored is already in the storage list and the location information of the key item to be stored is different from the location information of the key item in the storage list, obtaining a first storage time of the key item in the storage list and a second storage time of the key item to be stored;

[0017] When the difference between the second storage time and the first storage time is less than the preset time length, the location information of the key item to be stored is added to the location information corresponding to the key item in the storage list.

[0018] The above technical solution has the following advantages or beneficial effects: when the location information recognized by the large model of graphic and text understanding changes due to different angles, but the actual position does not change, the embodiment of the present application can recognize multiple location information of the same key item in a short period of time and save the multiple location information so that the user can obtain the location information of the key item more accurately.

[0019] In some embodiments, after obtaining the storage time of the key items in the storage list, the controller is further configured to:

[0020] When the difference between the second storage time and the first storage time is greater than or equal to the preset time length, the location information corresponding to the key item in the storage list is replaced with the location information of the key item to be stored.

[0021] The above technical solution has the following advantages or beneficial effects: when different location information of the same key item is identified after a certain period of time, the embodiment of the present application can replace the saved location information with the latest location information, so that the user can obtain the location information of the key item more accurately.

[0022] In some embodiments, after saving the location information of the key items to the storage list, the controller is further configured to:

[0023] When a key item corresponds to at least two location information, the broadcast priority is set according to the size of the space occupied by the first item in the location information, or the broadcast priority is set according to the storage time of the location information. The broadcast priority is used to represent the broadcast order of the prompt voice corresponding to the location information.

[0024] The above technical solution has the following advantages or beneficial effects: In this embodiment, the announcement priority can be set based on the space occupied by the first item, so that items with a smaller space are announced first. Items with a smaller space occupation indicate that the search range is smaller, which facilitates users to quickly locate the desired key item. The announcement priority can also be set based on the storage time, so that items with a recent storage time are announced first. The closer the storage time, the higher the search accuracy, which facilitates users to accurately locate the desired key item.

[0025] In some embodiments, the controller generates a prompt voice based on the location information of the first key item, and is further configured to:

[0026] If the first item in the location information of the first key item is a key item, obtaining the location information of the first item from the storage list;

[0027] A prompt voice is generated based on the first key item and the location information of the first item.

[0028] The above technical solution has the following advantages or beneficial effects: When searching for a specific key item, if the relative position relationship with another key item is obtained from the stored list, the location information of the other key item needs to be obtained. Given that key items are often relatively small and difficult to find, the solution of the embodiment of the present application can help users quickly find the required items by using the positional relationships of more items.

[0029] In some embodiments, the location information of the first key item includes first location information and second location information, and the controller generates a prompt voice based on the location information of the first key item and controls the speaker to play the prompt voice, and is further configured to:

[0030] generating a first prompt voice based on the first location information, and generating a second prompt voice based on the second location information;

[0031] When the space occupied by the first object in the first location information is smaller than that occupied by the first object in the second location information, the speaker is controlled to play the first prompt voice first.

[0032] The above technical solution has the following advantages or beneficial effects: in the embodiment of the present application, items that occupy a small space are given priority, and items that occupy a small space have a smaller search range, which makes it easier for users to quickly locate the key items they need.

[0033] In some embodiments, the location information of the first key item includes first location information and second location information, and the controller generates a prompt voice based on the location information of the first key item and controls the speaker to play the prompt voice, and is further configured to:

[0034] generating a first prompt voice based on the first location information, and generating a second prompt voice based on the second location information;

[0035] When the storage time of the first location information is later than the storage time of the second location information, the speaker is controlled to play the first prompt voice first.

[0036] The above technical solution has the following advantages or beneficial effects: items with a recent storage time are reported first, and the more recent the storage time, the higher the accuracy of the search, which facilitates users to accurately locate the key items they need.

[0037] In some embodiments, after receiving a user input instruction to search for a first key item, the controller obtains the location information of the first key item from the storage list, and is further configured to:

[0038] Control the sound collector to collect ambient voice, perform voice recognition on the ambient voice, and obtain voice text;

[0039] When the voice text includes the first key item and the search intention, location information of the first key item is obtained from the storage list.

[0040] The above technical solution has the following advantages or beneficial effects: voice commands allow users to complete various tasks without manually operating the device, and users can communicate with the device in a more natural way without having to learn complex commands or interfaces.

[0041] In some embodiments, the controller inputs key frames in the video into a large image-text understanding model to obtain location information of key items, and is further configured to:

[0042] Get the key frames in the video;

[0043] When the similarity between a key frame and its previous key frame is less than a preset value, the key frame is input into the large image-text understanding model to obtain the location information of the key item.

[0044] The above technical solution has the following advantages or beneficial effects: When multiple keyframes are highly similar, only one keyframe is input into the large image-text understanding model, reducing resource usage. When multiple consecutive keyframes are different, all keyframes are input into the large image-text understanding model, preventing information about individual key items from being missed and not stored.

[0045] In a second aspect, some embodiments of the present application provide an item search method, comprising:

[0046] Get the video recorded by the image collector;

[0047] Inputting key frames in the video into the large image-text understanding model to obtain the location information of the key items, the location information including the relative position relationship between the key items and the first items, and the first items including items other than the key items;

[0048] Save the location information of key items to the storage list;

[0049] After receiving a user input instruction to search for a first key item, obtaining location information of the first key item from a storage list;

[0050] A prompt voice is generated based on the location information of the first key item, and the speaker is controlled to play the prompt voice.

[0051] The above technical solution has the following advantages or beneficial effects: through the real-time video recording by the image collector of the wearable device and the recognition ability of the large model of graphic understanding of the video image, the relative positional relationship between the required items and other items can be identified and stored, and this positional relationship can help users quickly find a variety of daily life items.

[0052] In the technical solution provided by the embodiment of the present application, the video recorded by the image collector is obtained in real time, and then the key frames in the video are input into the large model of image-text understanding to obtain the location information of the key items. The location information includes the relative position relationship between the key item and the first item, and the first item includes items other than the key item. The location information of the key item is saved in a storage list. After receiving the user input instruction to find the first key item, the location information of the first key item is obtained from the storage list, and then a prompt voice is generated based on the location information of the first key item, and the speaker is controlled to play the prompt voice. Through the real-time video recording of the image collector of the wearable device and the recognition ability of the large model of image-text understanding for video images, the relative position relationship between the required item and other items can be identified and stored. This position relationship can help users quickly find a variety of daily life items. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0054] Figure 1 A flowchart of an item search method provided in some embodiments of the present application;

[0055] Figure 2 A flowchart of a key frame processing method provided in some embodiments of the present application;

[0056] Figure 3 A schematic diagram of a key frame provided in some embodiments of the present application;

[0057] Figure 4 A schematic diagram of another key frame provided in some embodiments of the present application;

[0058] Figure 5 A schematic diagram of another key frame provided in some embodiments of the present application;

[0059] Figure 6 A flowchart of a method for generating a prompt voice according to some embodiments of the present application;

[0060] Figure 7 A schematic diagram of the appearance of smart glasses provided in some embodiments of the present application;

[0061] Figure 8 A schematic diagram of an internal module of smart glasses provided in some embodiments of the present application;

[0062] Figure 9 A flowchart of another item search method provided in some embodiments of the present application;

[0063] Figure 10 A timing diagram of an item search method provided in some embodiments of the present application. DETAILED DESCRIPTION

[0064] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.

[0065] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0066] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," and "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0067] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0068] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.

[0069] In the embodiments of this application, wearable devices refer to electronic devices that can be worn directly on the body and not only have computing capabilities but also provide various services through an internet connection. Wearable devices include, but are not limited to, smart watches, fitness trackers, smart glasses, smart clothing, smart jewelry, ear-worn devices, and virtual reality (VR) / augmented reality (AR) headsets.

[0070] In some embodiments, the wearable device may include at least one of a communication device, a detector, a controller, a display, an image processing module, an audio processor, an audio output device, a memory, a power supply, and a user input interface.

[0071] In some embodiments, the communication device is a component used to communicate with an external device or server 400 according to various communication protocol types. The wearable device can be provided with multiple communication devices depending on the communication methods supported. For example, when the wearable device supports wireless network communication, the wearable device can be provided with a communication device including WiFi functionality. When the wearable device supports Bluetooth connection communication, the wearable device needs to be provided with a communication device including Bluetooth functionality.

[0072] The communication device can connect the wearable device to an external device or server via wireless or wired connections. Wired connections can connect the wearable device to an external device through components such as data cables and interfaces. Wireless connections can connect the wearable device to an external device via wireless signals or wireless networks. A wearable device can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.

[0073] In some embodiments, the detector is used to collect signals from the external environment or external interactions. For example, the detector includes a light receiver, a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting images of the external environment, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds.

[0074] In an embodiment of the present application, the image collector can record video in real time and send the video to the image processing module, which recognizes the image in the video to obtain the relative position relationship between the key item and other items.

[0075] In an embodiment of the present application, the sound collector can collect ambient sound in real time and send the ambient sound to the audio processor. The audio processor can recognize the ambient sound and obtain voice text.

[0076] In some embodiments, the controller may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller controls the operation of the wearable device and responds to user operations through various software control programs stored in a memory. The controller controls the overall operation of the wearable device.

[0077] In some embodiments, the display includes a display component for presenting an image and a driver component for driving the image display. The display is configured to receive image signals output from the controller for display. For example, the display can be configured to display video content, image content, menu control interface components, and user control UI interfaces.

[0078] In some embodiments, a user may input a user command through a graphical user interface (GUI) displayed on a display, and the user input interface receives the user input command through the graphical user interface (GUI).

[0079] In some embodiments, the image processing module may have a built-in AI (Artificial Intelligence) module. The AI module may use a trained large-scale image and text understanding model to identify the captured images, i.e., the key frames of the video, and output the relative position relationship between the key items and other items.

[0080] In some embodiments, the image processing module may include an embedded AI module. This AI module can invoke the server's image processing interface via a communication device, thereby transmitting the captured image, i.e., the key frames of the video, to the server via the communication device. The server then uses a trained large-scale image and text understanding model to recognize the image and output the relative positional relationship between the key item and other items, which is then sent to the AI module.

[0081] In some embodiments, the audio processor is used to use automatic speech recognition (ASR) to recognize the collected voice data to obtain voice recognition text, and use the acoustic model and language model to perform semantic understanding on the voice recognition text to obtain semantic parsing results.

[0082] In some embodiments, the audio processor calls the audio processing interface of the server through the communication device, that is, uses automatic speech recognition to recognize the collected voice data to obtain speech recognition text, and uses the acoustic model and language model to perform semantic understanding of the speech recognition text, and sends the semantic analysis results to the audio processor.

[0083] In some embodiments, the audio output device may be a speaker on the wearable device or an external audio output device connected to the wearable device. For the external audio output device connected to the wearable device, the wearable device may also be provided with an external audio output terminal, through which the audio output device may be connected to the wearable device to output the sound of the wearable device.

[0084] In some embodiments, the user input interface can be used to receive instructions from the user. Exemplarily, the user input interface can also receive voice instructions input by the user.

[0085] In daily life, people often encounter the problem of not being able to find commonly used items, such as keys, mobile phones, and remote controls. These small items are often difficult to find when needed because they are used frequently and are often placed casually. Existing methods of finding items mostly rely on manual memory or simple markings, such as hanging keys on fixed hooks or setting a specific storage location for remote controls. However, this method is not only inefficient but also prone to errors. In large spaces or environments where items are placed in a disorderly manner, it becomes even more difficult to find the required items quickly and accurately.

[0086] Existing wearable devices, such as smartwatches and smart glasses, are mostly limited in functionality, capable of interacting only with specific devices like mobile phones. For example, these devices typically pair with a user's smartphone via Bluetooth or Wi-Fi, enabling a single function like "Find My Phone." If a user can't find their phone, simply tapping a button on their smartwatch triggers a tone on the phone, helping them quickly locate it.

[0087] However, the original intention of designing this function was only to meet the needs in the scenario of lost mobile phone, and it does not cover the search for other common items.

[0088] In order to help users quickly find a variety of daily life items, the embodiment of the present application provides a wearable device. The structure and functions of each part of the wearable device can refer to the above embodiment. In addition, based on the wearable device shown in the above embodiment, this embodiment further improves some functions of the wearable device. Figure 1 As shown, the controller is configured to perform the following steps:

[0089] Step S101: Acquire the video recorded by the image collector.

[0090] The image collector is set on the wearable device. After the user wears the wearable device and starts the wearable device, the video can be recorded in real time through the image collector.

[0091] In some embodiments, a capture start button can be provided on the wearable device. After receiving a user's press operation on the capture start button, the image collector can be activated and a video of the current environment can be recorded through the image collector. After receiving a user's press operation on the capture start button again, the image collector can be turned off to stop the image collector from recording the current environment. The setting of the capture start button allows the user to flexibly choose whether to record the video according to their needs, meeting the needs of different scenarios.

[0092] In some embodiments, the image collector continuously collects image data of the surrounding environment to form a coherent video stream, and sends the collected video stream data to the image processing module in real time. The embodiments of the present application can promptly process the collected image data and promptly save the relative positions of key items identified in a storage list, making it easier for users to quickly find the items they need.

[0093] In other embodiments, the image collector records the video in real time, and the duration of each video recording is set to a target duration. During the recording process, the image collector continuously collects image information of the surrounding environment to form a coherent video stream. When the recording of a video of a target duration is completed, the video is transmitted to the image processing module for processing, and then the image collector immediately starts recording the next target duration video, and so on. The embodiment of the present application can centrally process the collected image data. Centralized processing can facilitate the screening and processing of image data, and there is no need to perform image analysis and recognition on each frame of video data, which helps to reduce resource usage.

[0094] Step S102: Input the key frames in the video into the image-text understanding model to obtain the location information of key items.

[0095] Among them, the location information includes the relative position relationship between the key item and the first item, and the first item includes items other than the key item. A key frame refers to a complete image frame in a video sequence, which contains all the information of the frame and can be independently decoded and displayed without relying on other frames. Key items include but are not limited to keys, mobile phones, remote controls, wallets, chargers, headphones, rings, scissors, ID cards, thermoses, etc. The first item can include key items, and can also include items such as sofas, tables, sinks, chairs and beds. Relative position relationships include up and down, left and right, front and back, inside and outside, etc. For example, the location information of a mobile phone can be above the table, to the right of a water cup, or inside a handbag.

[0096] In some embodiments, the step of inputting key frames from a video into a large model for image-text comprehension to obtain location information of key items includes: obtaining a preset number of key frames from a video at a preset time interval, inputting the obtained key frames into the large model for image-text comprehension, and obtaining location information of key items. Exemplarily, 5 key frames are taken from each second of video. The number of key frames is related to the processing power of the image processing module. The stronger the processing power of the image processing module, the more key frames can be obtained from each second of video. Embodiments of the present application can reduce the recognition of image data in a video and reduce resource usage.

[0097] In other embodiments, the step of inputting key frames from a video into a large model for image-text comprehension to obtain location information of key items includes: acquiring key frames at preset time intervals, inputting the acquired key frames into the large model for image-text comprehension, and obtaining location information of key items. Exemplarily, the most recent key frame is acquired every 10 ms. Alternatively, a key frame is acquired every interval of a preset number of key frames. Exemplarily, after acquiring the key frames in the video, a key frame is acquired every three key frames. Embodiments of the present application can reduce the recognition of image data in the video and reduce resource usage.

[0098] In some other embodiments, Figure 2 As shown in the figure, the steps of inputting the key frames in the video into the large image and text understanding model to obtain the location information of key items include:

[0099] Step S201: Acquire key frames in the video.

[0100] Step S202: determining whether the similarity between the key frame and the previous key frame is less than a preset value.

[0101] Similarity can be calculated by comparing the differences between corresponding pixels in two images. It can also be done by statistically analyzing the color distribution of the images to generate a color histogram, and then using different distance metrics to compare the similarity between the two histograms. It is also possible to use pre-trained convolutional neural networks to extract image features and calculate similarity based on these features.

[0102] When the similarity between the key frame and the previous key frame is less than a preset value, step S203 is executed: the key frame is input into the image-text comprehension model to obtain the location information of the key item.

[0103] When the similarity between the key frame and the previous key frame is greater than or equal to the preset value, step S204 is executed: the key frame is deleted, and the key frame is not input into the large model of image-text comprehension.

[0104] The embodiment of the present application can input only one key frame into the large model of image-text understanding when multiple key frames are highly similar, thereby reducing resource usage. For example, when the user is stationary, the key frames in the captured video are all the same, and the present application does not need to input all the key frames at this time into the large model of image-text understanding. The embodiment of the present application can also input multiple key frames into the large model of image-text understanding when multiple key frames are different in a row. For example, when the user is in a fast-moving state, the present application can input image data into the large model of image-text understanding for recognition, thereby avoiding the failure to recognize and store the information of individual key items.

[0105] The training method of the large model of image and text understanding may include: first, a large amount of image data needs to be collected as a training set. Each image data includes at least two objects, and the image data in the training set is as diverse as possible, that is, it covers common and commonly used objects to improve the generalization ability of the model. Then the image data in the training set are labeled, and the labels are the relative position relationship of at least two objects. Among them, in order to increase the amount and diversity of data, a series of transformation operations such as rotation, scaling, cropping, flipping, color adjustment, etc. can be performed on the original image. The basic model of image recognition is trained using labeled image data. The basic model can be a convolutional neural network. A part of the training set is divided as a validation set to evaluate the performance of the large model of image and text understanding on unseen data. Once the model training is completed and all evaluation criteria are passed, the large model of image and text understanding can be obtained.

[0106] If the large image-text understanding model identifies a key item from a key frame, the location information of the key item is output. If the large image-text understanding model identifies at least two key items from a key frame, the location information of all key items is output. If the large image-text understanding model does not identify a key item from a key frame, the location information of the items does not need to be output.

[0107] For example, the key frames in the collected video are as follows: Figure 3 As shown in the figure, the key frame is input into the image and text understanding model to obtain the location information of the remote control, which is above the sofa. Figure 4 As shown, the key frame is input into the image and text understanding model to obtain the position information of the remote control and the mobile phone. The position information of the remote control is above the sofa and to the left of the mobile phone, and the position information of the mobile phone is above the sofa and to the right of the remote control. Figure 5 As shown, the key frame is input into the large model of image and text understanding, and the output result is that the key item is not recognized.

[0108] Step S103: Save the location information of the key items into a storage list.

[0109] The storage list includes item name, location information, storage time, etc.

[0110] In some embodiments, the step of saving the location information of the key item to a storage list includes: determining whether the key item to be stored is in the storage list; if the key item to be stored is in the storage list, replacing the location information corresponding to the key item in the storage list with the location information of the key item to be stored; and if the key item to be stored is not in the storage list, adding the location information of the key item to be stored to the storage list.

[0111] For example, if the key item to be stored is a mobile phone, and its corresponding location information is above the sofa, if the mobile phone is already in the storage list and its location information is to the right of the sink, the location information of the mobile phone in the storage list is replaced with above the sofa. If the mobile phone is not in the storage list, the mobile phone and its location information above the sofa are directly added to the storage list.

[0112] In other embodiments, the step of saving the location information of the key item to the storage list includes: determining whether the key item to be stored is in the storage list; if the key item to be stored is already in the storage list, determining whether the location information of the key item to be stored is the same as the location information of the key item in the storage list.

[0113] When the location information of the key item to be stored is the same as that of the key item in the storage list, the first storage time corresponding to the key item in the storage list is replaced with the second storage time corresponding to the key item to be stored. The second storage time may be the current time.

[0114] If the location information of the key item to be stored is different from the location information of the key items in the storage list, the first storage time of the key item in the storage list and the second storage time of the key item to be stored are obtained, and it is determined whether the difference between the second storage time and the first storage time is less than a preset time period. If the difference between the second storage time and the first storage time is less than the preset time period, the location information of the key item to be stored is added to the location information corresponding to the key items in the storage list, and the second storage time corresponding to the key item to be stored is saved.

[0115] When the difference between the second storage time and the first storage time is greater than or equal to the preset time length, the location information corresponding to the key item in the storage list is replaced with the location information of the key item to be stored, and the first storage time corresponding to the key item in the storage list is replaced with the second storage time corresponding to the key item to be stored.

[0116] If the key item to be stored is not in the storage list, the location information of the key item to be stored is added to the storage list.

[0117] For example, the storage list is shown in Table 1.

[0118] Item Name Location information Storage time cell phone Left side of the remote control 9:00 AM, x / x key Above the table 8:00 AM, x / x …… …… ……

[0119] The preset duration is 2 minutes. If the key item to be stored is a mobile phone and its corresponding location information is above the sofa, if the current time is 9:01, the mobile phone's location information will be updated to the left side of the remote control and above the sofa, and the storage times will be 9:00 and 9:01 respectively. If the current time is 9:10, the mobile phone's location information will be updated to above the sofa, and the storage time will be 9:10. The mobile phone's location information in the storage list will be replaced with above the sofa. If the key item to be stored is a wallet and its corresponding location information is above the sofa, the wallet and its location information above the sofa will be directly added to the storage list.

[0120] Wearable devices will capture images at different angles as the user moves. The different angles may cause the location information recognized by the large model of image and text understanding to change, but the actual location does not change. For example, when the image collector of the wearable device is close to the mobile phone, the captured image can only recognize that the mobile phone is on the right side of the remote control. When the image collector of the wearable device is far away from the mobile phone and the remote control is not captured, the captured image can only recognize that the mobile phone is above the sofa, and the location information corresponding to the mobile phone needs to be changed. When the embodiment of the present application recognizes multiple location information of the same key item in a short period of time, all the multiple location information can be saved so that the user can obtain the location information of the key item more accurately.

[0121] In some embodiments, after the location information of key items is saved in a storage list, if a key item corresponds to one location information, there is no need to set a priority for the announcement. If a key item corresponds to at least two location information, the announcement priority is set based on the space occupied by the first item in the location information. The announcement priority is used to indicate the order in which the prompt voice messages corresponding to the location information are announced.

[0122] A first item priority mapping table may be pre-set, wherein the mapping table includes a correspondence between first items and priorities. The smaller the space occupied by the first item, the higher the priority. The higher the priority value, the higher the priority.

[0123] For example, the mobile phone's location information is to the left of the remote control and above the sofa. The priority values corresponding to the remote control and the sofa in the first item priority mapping table are obtained as 50 and 10, respectively. When the mobile phone's location information needs to be broadcast, the location information of the left side of the remote control can be sent to the speech synthesis module first.

[0124] In this embodiment of the present application, the priority of the announcement can be set based on the space occupied by the first item, so that the items with the smallest space are announced first. The items with the smallest space occupying a smaller space have a smaller search range, which makes it easier for users to quickly locate the key items they need. The items with the smallest space may also be moved, so other location information can be directly or indirectly announced to help users locate the actual location of the items they need.

[0125] In other embodiments, after the location information of the key item is saved in the storage list, if the key item corresponds to at least two pieces of location information, the priority of the announcement is set according to the storage time of the location information, wherein the more recent the storage time, the higher the priority.

[0126] For example, the location information of the mobile phone is the left side of the remote control and above the sofa, and the storage time is 9:00 and 9:01 respectively. When the location information of the mobile phone needs to be broadcast, the location information above the sofa can be sent to the speech synthesis module first.

[0127] In this embodiment of the present application, the priority of the broadcast can be set according to the storage time, so that the items with the latest storage time are broadcast first. The closer the storage time, the higher the accuracy of the search, which helps users accurately locate the key items they need. If the item is not found at the location, other location information can be broadcast directly or indirectly to help users determine the actual location of the required item.

[0128] When saving the location information of key items to the storage list, multiple key items of the same type may appear. For example, a user may have two mobile phones, a wallet, and keys. If both mobile phones are stored under the same item name, the location information may be mixed up.

[0129] In some embodiments, to accurately store the location information of multiple key items of the same type, characteristic information can be added to the storage list. Different characteristic information can be set for different types of objects. For example, a mobile phone can have its border color set as characteristic information, and a wallet can have its length or color set as characteristic information.

[0130] When saving the location information of key items to the storage list, in addition to verifying the item names, you must also verify the item's feature information. If both the item names and feature information are identical, they are considered the same item and their location information can be updated to the corresponding location information in the storage list. If the item names are identical but the feature information is different, they should be treated as separate items and their location information added to the storage list.

[0131] For example, the key item to be stored is a mobile phone (red border), and its corresponding location information is above the sofa. If a mobile phone with a red border already exists in the storage list and its location information is to the right of the sink, the location information of the mobile phone (red border) in the storage list will be replaced with above the sofa. If a mobile phone already exists in the storage list but has a black border, the mobile phone (black border) and its location information above the sofa will be directly added to the storage list.

[0132] In some embodiments, the image collector of the wearable device sends the recorded video to the image processing module, which inputs the key frames in the video into the large image and text understanding model to obtain the location information of the key items. In the embodiment of the present application, the large image and text understanding model can be deployed on the terminal side, without sending the key frames to the server, which can better protect user privacy and security. In addition, it can continue to run and provide services even when the network connection is unstable or there is no network.

[0133] In other embodiments, the image collector of the wearable device sends the recorded video to the image processing module, and the image processing module calls the image processing interface of the server to process the key frames in the video, that is, the image processing module sends the key frames in the video to the server, and the server inputs the key frames into the large model of image and text understanding to obtain the location information of the key items, and sends the location information of the key items to the image processing module. The embodiment of the present application can deploy the large model of image and text understanding on the server side. The server is equipped with high-performance hardware resources, has faster processing speed and higher efficiency, is easy to update and maintain, and even devices with limited computing power can take advantage of the most advanced model services.

[0134] Step S104: receiving a user input instruction to search for a first key item.

[0135] In some embodiments, a voice command input by a user to search for a first key item may be received, specifically comprising: controlling a sound collector to collect ambient voice, and determining whether the volume of the ambient voice is greater than a preset volume. If the volume of the ambient voice is less than or equal to the preset volume, then the ambient voice collection continues. If the volume of the ambient voice is greater than the preset volume, then voice recognition is performed on the ambient voice to obtain a voice text. Then, it is determined whether the voice text includes the first key item and the search intent. It may be determined first whether the voice text includes the first key item. If the voice text does not include the first key item, then the ambient voice collection continues. If the voice text includes the first key item, then semantic analysis is performed on the voice text to obtain the intent corresponding to the voice text. If the intent corresponding to the voice text is a search intent, then it may be determined that a voice command input by the user to search for the first key item has been received. If the intent corresponding to the voice text is not a search intent, then it may be determined that a voice command input by the user to search for the first key item has not been received.

[0136] In some embodiments, a voice button is provided on the wearable device. Upon receiving an input instruction from a user pressing the voice button, a sound collector is controlled to collect the user's voice. Voice recognition is performed on the ambient voice to obtain a voice text. If the voice text includes the first key item and the search intent, it is determined that a voice instruction input by the user to search for the first key item has been received.

[0137] Step S105: Obtain the location information of the first key item from the storage list.

[0138] Step S106: Generate a prompt voice based on the location information of the first key item and control the speaker to play the prompt voice.

[0139] The voice synthesis module composes a prompt text with the name of the first key item, the location information of the first key item, and a conjunction, then uses voice synthesis technology to convert the prompt text into a prompt voice, and then controls the speaker to play the prompt voice. Among them, the conjunctions include words such as "at" and "located at".

[0140] Exemplarily, if the first key item is a thermos cup and the location information of the first key item is above the table, the prompt text is "The thermos cup is above the table".

[0141] In some embodiments, as Figure 6 shown, after obtaining the location information of the first key item from the storage list, step S601 can be executed: Determine whether the first item in the location information of the first key item is a key item.

[0142] If the first item in the location information of the first key item is a key item, that is, the second key item, then step S602 is executed: Obtain the location information of the first item in the location information of the first key item, that is, the second key item, from the storage list.

[0143] In some embodiments, after step S602, a prompt voice can be generated based on the location information of the first key item and the first item, that is, the second key item.

[0144] In some embodiments, after step S602, step S603 is executed: Determine whether the first item in the location information of the second key item is a key item;

[0145] If the first item in the location information of the second key item is not a key item, then step S604 is executed: Generate a prompt voice based on the location information of the first key item and the first item in the location information of the first key item.

[0146] If the first item in the location information of the second key item is a key item, that is, the third key item, continue to obtain the location information of the third key item from the storage list. Among them, the third key item is different from the first key item. If the third key item is the same as the first key item, there is no need to obtain the location information of the third key item, and a prompt voice is directly generated based on the location information of the first key item.

[0147] If the first item in the location information of the first key item is not a key item, step S605 is executed: generating a prompt voice based on the location information of the first key item.

[0148] For example, the first key item is a mobile phone. The location information of the mobile phone obtained from the storage list is to the left of the remote control. Since the remote control is also a key item, the location information of the remote control is also obtained from the storage list. The location information of the remote control is above the sofa, and the prompt voice message "The mobile phone is to the left of the remote control, and the remote control is on the sofa" is generated.

[0149] In the embodiment of the present application, when searching for a specified key item, if the relative position relationship with another key item is obtained from the storage list, it is necessary to obtain the location information of another key item until the position relationship with the non-key item is obtained. Key items are often relatively small and difficult to find. The prompt voice provides the position relationship with another key item, which still cannot help the user quickly locate the location of the required item. The solution of the embodiment of the present application can help users quickly find the required item through the position relationship with non-key items.

[0150] In some embodiments, in addition to sending the location information of key items to the voice synthesis module in sequence according to the priority, the image processing module can also directly send multiple location information of key items to the voice synthesis module together. After obtaining multiple location information of the first key item, namely the first location information and the second location information, the voice synthesis module can generate a first prompt voice based on the first location information and a second prompt voice based on the second location information, and then determine whether the first item in the first location information occupies less space than the first item in the second location information. If the first item in the first location information occupies less space than the first item in the second location information, the first prompt voice is played first. If the first item in the first location information occupies more space than the first item in the second location information, the second prompt voice is played first.

[0151] In other embodiments, after the speech synthesis module generates the first prompt voice based on the first location information and the second prompt voice based on the second location information, it may further determine whether the storage time of the first location information is later than the storage time of the second location information. If the storage time of the first location information is later than the storage time of the second location information, the first prompt voice is played first. If the storage time of the first location information is earlier than the storage time of the second location information, the second prompt voice is played first.

[0152] In some embodiments, prioritizing the first location information to play the prompt voice is to first play the first location information prompt voice, and then immediately play the second location information prompt voice. For example, the phone is played to the left of the remote control and the phone is above the sofa.

[0153] In other embodiments, prioritizing the voice prompt corresponding to the first location information means first playing the voice prompt corresponding to the first location information. After receiving the user input command "describe the first key item in detail," the voice prompt corresponding to the second location information is played. For example, the voice prompt indicating that the phone is on the left side of the remote control is played first. If the user does not find the item and asks again, the voice prompt indicating that the phone is above the sofa is played.

[0154] In some embodiments, the speech synthesis module may obtain the storage time corresponding to the location information of the first key item while obtaining the location information of the first key item, and may form a prompt text with the name of the first key item, the location information of the first key item, a conjunction, and the storage time. For example, the mobile phone was on the table an hour ago.

[0155] In other embodiments, the speech synthesis module, while obtaining the location information of the first key item, may also obtain the storage time corresponding to the location information and then determine whether the difference between the storage time and the current time is greater than a preset difference. If the difference between the storage time and the current time is greater than the preset difference, the storage time is added to the prompt text. If the difference between the storage time and the current time is less than or equal to the preset difference, the storage time does not need to be added to the prompt text. The preset difference may be 24 hours.

[0156] If the location information of the first key item is not obtained in the storage list, the speaker is controlled to play a prompt voice that the first key item is not found. The prompt voice can also prompt the user whether to start searching for the first key item immediately. If the user inputs an instruction to immediately search for the first key item, the user is prompted to move with the wearable device, and the image collector is controlled to collect environmental image data in real time. The environmental image data is input into the image-text comprehension model to determine whether the environmental image data includes the first key item. After the first key item is identified, the prompt voice of the location information of the first key item is played.

[0157] For example, the wearable device may be smart glasses. Figure 7 As shown, the smart glasses look similar to ordinary glasses, but they are equipped with an image collector 61. The image collector 61 is used to record video, that is, environmental image data. The smart glasses are also equipped with an image processing module, an audio module, and a speaker.

[0158] like Figure 8As shown, the image acquisition module and audio module are respectively connected to the image processing module, responsible for transmitting the collected video and audio information to the image processing module for processing. The image processing module is connected to the speaker module, which transmits the processed object location information to the speaker module for voice broadcast. Speech synthesis functionality can be integrated into either the image processing module or the speaker module. These modules work together to realize the object search function.

[0159] like Figure 9 As shown, the image acquisition module records 10 minutes of video each time, which is then transferred to the image processing module for processing. The next video is then recorded, providing the image processing module with environmental image data for object identification. The audio module is responsible for real-time audio reception and voice recognition. Upon detecting a voice command containing the name of a key item and followed by "where is it?" or "where is...?", it immediately notifies the image processing module. The image processing module receives the video from the image acquisition module and uses a large-scale image-text comprehension model to identify key items and their relative positions to other items, storing the latest location information of the key items. Upon receiving notification from the audio module, the key item's location is obtained and notified to the speaker module. The speaker module receives the item location information sent by the image processing module and plays it back as audio, informing the user of the item's specific location.

[0160] The image collector and audio module work together to record real-time video, with each recording duration set to 10 minutes. During recording, the image collector continuously captures image information of the surrounding environment, forming a coherent video stream. Once a 10-minute video is recorded, it is transmitted to the image processing module for processing, and the image collector immediately begins recording the next 10-minute video, repeating the cycle. This process ensures continuous monitoring of the surrounding environment, providing a rich data foundation for subsequent object identification.

[0161] After receiving the video from the image acquisition device, the image processing module extracts five key frames per second and uses a large model to analyze and identify the video content. Trained with extensive data, the large model accurately identifies key items in the video and determines their specific locations within the scene. For example, "remote control + right side of yellow sofa" or "cell phone + left side of white sink." After identification, the image processing module stores and retains the key items and their locations, retaining only the most recent location information to ensure data timeliness and accuracy. This ensures that the image processing module always provides the most up-to-date location information, regardless of any changes in the item's position.

[0162] The audio module is responsible for real-time audio reception and collecting voice information from the surrounding environment. During this process, the audio module continuously performs voice recognition. When it detects a voice command containing the name of a key item followed by the phrase "where is it?" or "where is it?"—for example, "Where's my phone?" or "Where's my remote control?"—the audio module immediately notifies the image processing module of the item search request. This feature allows users to quickly initiate the item search process with a simple voice query, improving operational convenience.

[0163] When the image processing module receives the notification from the audio module, it uses the stored key item location information to determine the specific location of the item the user is searching for. It then transmits this location information to the speaker module. The speaker module then plays the information as a voice message, such as "The remote control is to the right of the yellow sofa," allowing the user to quickly locate the item and complete the search.

[0164] Suppose a user can't find their remote control at home and is wearing smart glasses. The user says, "Where's the remote?" The audio module receives and recognizes the voice command in real time, immediately notifying the image processing module. After receiving the notification, the image processing module retrieves the remote's location from stored location information—assuming it's "to the right of the yellow sofa." The image processing module then sends this location information to the speaker module, which plays "The remote control is to the right of the yellow sofa." The user then hears the voice prompt and can quickly find the remote.

[0165] In some embodiments, the timing diagram of the item search method can be as follows: Figure 10 As shown. The controller includes an image processing module, a speech recognition and analysis module, and a speech synthesis module. The image collector records the video and sends the video to the image processing module. The image processing module inputs the key frames of the video into the image-text understanding model to obtain the location information of the key items. The location information includes the relative position relationship between the key items and other items, and saves the location information of the key items to a storage list. The sound collector obtains the ambient voice and sends the ambient voice to the speech recognition and analysis module. The speech recognition and analysis module recognizes and parses the ambient voice. When the recognized text includes the key items and the search intention, a search request is sent to the image processing module, wherein the search request includes the name of the key items. The image processing module sends the location information of the key items to the speech synthesis module. The speech synthesis module combines the name and location information of the key items into a prompt text, synthesizes the prompt text into a prompt voice, and sends the prompt voice to the speaker. The speaker plays the prompt voice.

[0166] The technical solution provided in the embodiment of the present application can identify and store the relative positional relationship between the required items and other items through the real-time video recording by the image collector of the wearable device and the recognition ability of the large model of graphic and text understanding of the video image. This positional relationship can help users quickly find a variety of daily life items.

[0167] Some embodiments of the present application further provide a computer-readable storage medium that can store a program. When the computer storage medium is configured in a wearable device or a server, the program, when executed, can include the program steps involved in the terminal device control method in the above embodiment. The computer storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0168] An embodiment of the present application provides an electronic device comprising: a processor and a memory for storing processor-executable instructions, wherein the processor is configured to read the executable instructions from the memory and execute the instructions to implement the terminal device control method in the above embodiment.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

[0170] For ease of explanation, the above description has been presented in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments have been selected and described to better explain the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various variations of the embodiments suitable for specific use considerations.

Claims

1. A wearable device, characterized in that: include: an image collector configured to record video; The controller is configured as: Obtaining the video recorded by the image collector; Inputting key frames in the video into a large model for image-text comprehension to obtain location information of key items, wherein the location information includes a relative positional relationship between the key items and a first item, wherein the first item includes items other than the key items; Saving the location information of the key items to a storage list; After receiving a user input instruction to search for a first key item, obtaining location information of the first key item from the storage list; A prompt voice is generated based on the location information of the first key item, and a speaker is controlled to play the prompt voice.

2. The wearable device according to claim 1, wherein: The controller executes the step of saving the location information of the key item into a storage list, and is further configured to: If the key item to be stored is already in the storage list and the location information of the key item to be stored is different from the location information of the key item in the storage list, obtaining a first storage time of the key item in the storage list and a second storage time of the key item to be stored; When the difference between the second storage time and the first storage time is less than a preset time period, the location information of the key item to be stored is added to the location information corresponding to the key item in the storage list.

3. The wearable device according to claim 2, wherein: After obtaining the storage time of the key items in the storage list, the controller is further configured to: When the difference between the second storage time and the first storage time is greater than or equal to a preset time period, the location information corresponding to the key item in the storage list is replaced with the location information of the key item to be stored.

4. The wearable device according to claim 1, wherein: After executing the step of saving the location information of the key item into the storage list, the controller is further configured to: In the case where the key item corresponds to at least two location information, the broadcast priority is set according to the size of the space occupied by the first item in the location information, or the broadcast priority is set according to the storage time of the location information. The broadcast priority is used to represent the broadcast order of the prompt voice corresponding to the location information.

5. The wearable device according to claim 1, wherein: The controller generates a prompt voice based on the location information of the first key item, and is further configured to: If the first item in the location information of the first key item is a key item, acquiring the location information of the first item from the storage list; A prompt voice is generated based on the first key item and the location information of the first item.

6. The wearable device according to claim 1, wherein: The location information of the first key item includes first location information and second location information, and the controller generates a prompt voice based on the location information of the first key item and controls the speaker to play the prompt voice, and is further configured to: generating a first prompt voice based on the first location information, and generating a second prompt voice based on the second location information; When the space occupied by the first object in the first location information is smaller than that occupied by the first object in the second location information, the speaker is controlled to play the first prompt voice preferentially.

7. The wearable device according to claim 1, wherein: The location information of the first key item includes first location information and second location information, and the controller generates a prompt voice based on the location information of the first key item and controls the speaker to play the prompt voice, and is further configured to: generating a first prompt voice based on the first location information, and generating a second prompt voice based on the second location information; When the storage time of the first location information is later than the storage time of the second location information, the speaker is controlled to play the first prompt voice first.

8. The wearable device according to claim 1, wherein: The controller is further configured to: after receiving a user input instruction to search for a first key item, obtain the location information of the first key item from the storage list; Control the sound collector to collect ambient voice, perform voice recognition on the ambient voice, and obtain voice text; When the voice text includes the first key item and the search intention, the location information of the first key item is acquired from the storage list.

9. The wearable device according to claim 1, wherein: The controller inputs the key frames in the video into the image-text comprehension model to obtain the location information of the key items, and is further configured to: Obtaining key frames in the video; When the similarity between the key frame and the previous key frame of the key frame is less than a preset value, the key frame is input into the image-text comprehension model to obtain the location information of the key item.

10. A method for searching for an item, characterized in that: include: Get the video recorded by the image collector; Inputting key frames in the video into a large model for image-text comprehension to obtain location information of key items, wherein the location information includes a relative positional relationship between the key items and a first item, wherein the first item includes items other than the key items; Saving the location information of the key items to a storage list; After receiving a user input instruction to search for a first key item, obtaining location information of the first key item from the storage list; A prompt voice is generated based on the location information of the first key item, and a speaker is controlled to play the prompt voice.