Article position retrieval method and device, computer equipment, medium and product

By combining a central control screen and a camera, voice commands are parsed and image data is recognized, and the location of items is associated in a multimodal manner. This solves the problem of low efficiency in recording and searching for items in existing technologies, and achieves efficient and accurate item management.

CN121979855APending Publication Date: 2026-05-05GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREE ELECTRIC APPLIANCE INC OF ZHUHAI
Filing Date
2025-12-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing voice assistants cannot directly link to the actual location of items and lack efficient audio and visual retrieval mechanisms, resulting in low efficiency in recording and searching for item placement.

Method used

The system wakes up the central control screen and receives voice commands. It combines image data collected by the camera, parses audio files and performs object recognition, performs multimodal association matching of item categories and location information, stores the data in the database, and outputs search results when querying.

Benefits of technology

It enables accurate, intuitive, and efficient recording and retrieval of item locations, improving the efficiency and accuracy of information management and providing a natural and convenient multimodal item management experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979855A_ABST
    Figure CN121979855A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an article position retrieval method and device, computer equipment, a medium and a product, and the method comprises the steps: responding to a triggering operation of a user on a central control screen, waking up the central control screen, receiving a voice instruction which is input by the user and contains article placement information, converting the voice instruction into an audio file, and associating the audio file with a timestamp; controlling a camera to collect image data including an article placement scene; analyzing the audio file to extract a keyword, and performing object recognition on the image data to obtain object category and position information; performing multi-modal association matching on the keyword and the item category and position information to determine and record the storage position of the target item; the audio file, the image data, the keyword and the storage position of the target object are stored in a database in an associated mode; and receiving an article query instruction input by a user on the central control screen, performing retrieval in the database based on the article query instruction, and outputting a retrieval result associated with the storage position of the target article.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of application services, and in particular to a method, apparatus, computer equipment, medium, and product for retrieving the location of an item. Background Technology

[0002] With the rapid development of smart home and voice interaction technologies, users' demand for voice assistants to help manage items in daily life is increasing. Currently, users often lack effective means of recording where they place items (such as keys, remote controls, etc.), leading to difficulties in finding them later.

[0003] Although some existing voice assistants have voice memo or voice command recording functions, these functions can only record the text information entered by the user through voice recognition. They cannot be directly linked to the actual location of the item, nor can they provide intuitive audio or visual prompts when the user needs them.

[0004] Furthermore, traditional voice assistants lack efficient mechanisms for retrieving historical audio, resulting in low information retrieval efficiency. Therefore, existing technologies have significant shortcomings in areas such as item placement recording, voice interaction guidance, and audio retrieval. Summary of the Invention

[0005] In view of this, embodiments of this application provide at least one method, apparatus, computer device, medium, and product for retrieving the location of an item.

[0006] The technical solution of this application embodiment is implemented as follows: On one hand, embodiments of this application provide a method for retrieving the location of an item, the method comprising: In response to the user's trigger operation on the central control screen, the central control screen is woken up and the user's voice command containing the placement information of the item is received. The voice command is converted into an audio file and associated with a timestamp. Control the camera to capture image data of the scene containing the placement of objects; The audio file is parsed to extract keywords, and the image data is used for object recognition to obtain item category and location information; The keywords are matched with the item category and the location information using a multimodal association to determine and record the storage location of the target item. The audio file, the image data, the keywords, and the storage location of the target item are associated and stored in the database; The central control screen receives item query commands input by the user, performs a search in the database based on the item query commands, and outputs search results associated with the storage location of the target item.

[0007] On the other hand, embodiments of this application provide another device for retrieving the location of an item, the method comprising: The recording module is used to respond to the user's trigger operation on the central control screen, wake up the central control screen and receive the user's voice command containing the placement information of the item, convert the voice command into an audio file and associate it with a timestamp; Control the camera to capture image data of the scene containing the placement of objects; The audio file is parsed to extract keywords, and the image data is used for object recognition to obtain item category and location information; The keywords are matched with the item category and the location information using a multimodal association to determine and record the storage location of the target item. The audio file, the image data, the keywords, and the storage location of the target item are associated and stored in the database; The retrieval module is used to receive item query instructions input by the user on the central control screen, perform a retrieval in the database based on the item query instructions, and output retrieval results associated with the storage location of the target item.

[0008] In another aspect, embodiments of this application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method for retrieving the location of an item.

[0009] In another aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method for retrieving the location of an item.

[0010] In another aspect, embodiments of this application provide a computer program including computer-readable code. When the computer-readable code is run in a computer device, the processor in the computer device executes some or all of the steps in the above-described method for retrieving the location of an item.

[0011] In another aspect, embodiments of this application provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-mentioned method for retrieving the location of an item.

[0012] This application embodiment achieves accurate, intuitive, and efficient recording and retrieval of item placement information through response triggering, receiving voice commands, image acquisition, parsing and recognition, multimodal matching and association, structured storage, and intelligent retrieval output. It overcomes the shortcomings of single voice recording lacking spatial intuitiveness, utilizes images to provide visual anchors, and greatly improves the efficiency and accuracy of information management through the associated storage of voice and images and keyword-based retrieval. It provides users with a natural, convenient, and reliable multimodal item location management experience, and is particularly helpful in solving the problem of difficulty in finding forgotten items in daily life.

[0013] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this application. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0015] Figure 1 A schematic diagram illustrating the implementation process of an item location retrieval method provided in this application embodiment; Figure 2 One of the schematic diagrams illustrating the principle of an item location retrieval method provided in this application embodiment; Figure 3 A second schematic diagram illustrating the principle of an item location retrieval method provided in this application embodiment; Figure 4 A schematic diagram illustrating the structural composition of an item location retrieval device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0018] The terms “first / second / third” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first / second / third” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.

[0020] This application provides a method for retrieving the location of an item, which can be executed by a processor of a computer device. The computer device can refer to a server, laptop computer, tablet computer, desktop computer, smart TV, set-top box, mobile device (such as a mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or other similar computer equipment. Figure 1 This is a schematic diagram illustrating the implementation process of an item location retrieval method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes: Step 101: In response to the user's trigger operation on the central control screen, wake up the central control screen and receive the voice command input by the user containing the item placement information, convert the voice command into an audio file and associate it with a timestamp.

[0021] In this embodiment, the central control screen is an intelligent terminal device integrating display, touch, and voice interaction functions, typically serving as the control center of a smart home system. An audio file is a digital audio data storage unit. A timestamp is precise time data recording the exact moment a specific event occurs.

[0022] The system responds to physical triggering operations by the user on the central control screen, such as touch or gesture, activating the screen's standby state. The system receives voice commands from the user, including the item name and placement location, via the microphone array integrated into the central control screen. The system then calls the built-in audio encoding module to convert the received analog voice signal into a digital audio file. Simultaneously, the system calls the clock module to generate a unique timestamp representing the current time and associates this timestamp with the generated digital audio file using metadata.

[0023] Step 102: Control the camera to acquire image data containing the scene of the object placement.

[0024] In this embodiment, the camera is an electronic device used to capture optical images and convert them into digital image data. It can be integrated into the central control screen or set up separately; no limitation is made here. Image data is digital information composed of an array of pixels, used to characterize a visual scene.

[0025] After waking up the central control screen, the system automatically activates the network camera integrated with or connected to the central control screen. The system sends control commands to the camera, driving its lens to adjust its focus and angle to cover the physical environment area where the user placed the item. The camera performs image acquisition operations according to the commands, generating one or more static or dynamic image data streams representing the scene where the item was placed, and transmits the image data to the system's image processing buffer.

[0026] Step 103: Parse the audio file to extract keywords, and perform object recognition on the image data to obtain item category and location information.

[0027] In this embodiment, keywords are lexical units extracted from speech content that represent core semantics. Object recognition is a technology in the field of computer vision, which refers to the process of analyzing image data using algorithms to identify objects contained within it and determine their categories and boundaries.

[0028] The system invokes a speech recognition engine to decode and transcribe the generated audio file, converting continuous speech signals into discrete text information. The system then performs natural language processing on this text, applying named entity recognition and part-of-speech tagging algorithms to extract a list of keywords representing item names and locations. Simultaneously, the system invokes a pre-trained object detection model to process the acquired image data. This model analyzes image pixels, locates the contours of all identifiable objects in the image, and outputs a preset category label and the object's position coordinates in the image coordinate system for each detected object.

[0029] Step 104: Perform multimodal association matching between the keywords, the item category, and the location information to determine and record the storage location of the target item.

[0030] In the embodiments of this application, multimodal association matching is an information fusion technology that refers to the process of aligning and associating data from different perceptual modalities (such as auditory speech and visual images) to form a unified semantic understanding.

[0031] The system establishes a mapping relationship between voice keywords and image recognition results. The system compares the item name keywords extracted from the voice command with all item category labels output by the image recognition. When a completely matching and unique item category is found, the system determines the item's location coordinates in the image as the target item's storage location. If multiple similar items exist in the image, the system further compares the location description keywords in the voice command or initiates an interactive confirmation process. The system calculates semantic similarity or guides the user to confirm via touch on the central control screen, ultimately identifying and recording a unique target item and its corresponding physical storage location information.

[0032] Step 105: Associate and store the audio file, the image data, the keywords, and the storage location of the target item in the database.

[0033] In this embodiment, the database is a repository that organizes, stores, and manages data according to a specific data structure, supporting efficient data retrieval, updating, and management operations.

[0034] The system constructs a structured data record. This record uses the timestamp generated in step 101 as a unique index identifier. The system packages and encapsulates the timestamp, the corresponding audio file, the collected image data, the keyword list parsed from the audio, and the target item storage location information finally determined in step 104. The system writes this encapsulated complete data record to the local storage or the database of a remote cloud server through a data access interface, ensuring that the relationships between the data elements are correctly established and persistently saved.

[0035] Step 106: Receive the item query command input by the user on the central control screen, perform a search in the database based on the item query command, and output the search results associated with the storage location of the target item.

[0036] In this embodiment, the item search command is a voice or touch command issued by the user to the system to find a specific item. The search result is a set of relevant information retrieved from the database by the system based on the search criteria.

[0037] The system achieves accurate, intuitive, and efficient recording and retrieval of item location information through response triggers, receiving voice commands, acquiring images, parsing and recognizing them, multimodal matching and association, structured storage, and intelligent retrieval output. It overcomes the shortcomings of single voice recording lacking spatial intuitiveness, utilizes images to provide visual anchors, and greatly improves the efficiency and accuracy of information management through the associated storage of voice and images and keyword-based retrieval. It provides users with a natural, convenient, and reliable multimodal item location management experience, and is particularly helpful in solving the problem of finding forgotten items in daily life.

[0038] This application embodiment simultaneously collects contextual information about the placement of items through both voice and visual channels, and accurately establishes the association between items and their locations using multimodal fusion technology. Subsequently, the system persistently stores the structured association information. When a user needs to find an item, the system can quickly locate it through an efficient retrieval mechanism and intuitively present the historical records. This achieves efficient management of the entire lifecycle of household item locations, from recording and storage to retrieval, significantly improving the efficiency and experience of finding items for users and enhancing its applicability to users with different needs.

[0039] In some embodiments, step 101 includes: Step 1011: Receive a wake-up signal from the user via voice command or physical action to activate the recording function of the voice control screen.

[0040] In this embodiment, a voice command refers to a verbal command issued by a user to the system by speaking specific words or sentences. A physical action refers to a user interacting with the device through physical contact or non-contact actions such as touch or gestures. A wake-up signal refers to a specific instruction or action pattern used to trigger the system to enter a working state from a standby or low-power state. A voice control screen is an intelligent interactive terminal integrating a display screen, audio input / output devices, a camera, and a computing processing unit, used to receive and process user voice and touch commands and provide feedback. The recording function refers to the system's ability to activate its audio input device and prepare to collect and record sound signals.

[0041] The system continuously monitors its audio input channel or its touch / gesture sensors. When the system detects a voice command matching a preset wake-up word pattern (e.g., the user says "record the location of an object") or a preset physical action (e.g., the user taps a specific area of ​​the screen or makes a specific gesture), the system determines that a valid wake-up signal has been received. Subsequently, the system switches from standby or monitoring state to active working state and activates its audio processing module, ready to begin collecting and recording the user's voice input.

[0042] Step 1012: Collect the voice signal of the user describing the name of the item and its location through the audio input device of the voice control screen.

[0043] In this embodiment, an audio input device refers to a hardware component used to capture sound and convert it into an electrical signal, such as a microphone array. A voice signal refers to an analog or digital sound wave signal generated by the user's voice and captured by the audio input device. An item name refers to an identifying designation given by the user to the item to be recorded, such as a key or remote control. Placement location refers to the specific location or orientation information of where the item is placed, as described by the user, such as on the coffee table in the living room.

[0044] Once the recording function of the voice control screen is activated, the system captures the user's voice through its built-in microphone array. The system guides the user (e.g., by displaying prompts on the screen or playing guiding voice) to speak a statement containing the name of the item to be recorded and a description of its location. The microphone converts the captured sound waves into analog electrical signals, which are then converted into a digital audio data stream via an analog-to-digital converter for use by subsequent processing modules.

[0045] Step 1013: Convert the speech signal into a digital audio file and add a timestamp representing the recording time to the audio file.

[0046] In this application embodiment, digital format refers to converting continuous analog signals (such as sound waveforms) into a signal form represented by discrete numerical values ​​(usually binary numbers) through sampling and quantization processes. Audio file refers to a computer file that stores digital audio data according to a specific encoding format (such as WAV, MP3). A timestamp is a marker recorded by the system clock according to a specific time unit (such as year, month, day, hour, minute, second, millisecond) when a piece of data (here, an audio file) is created or modified, used to identify the moment the data was generated.

[0047] The system compresses and encapsulates the acquired digital voice data stream according to a preset audio encoding standard to generate a structured audio file. Simultaneously, the system obtains the current standard time from its internal clock or network time protocol server and embeds this time information as metadata into the file header of the audio file or the associated database record, thus completing the timestamp marking of the audio file.

[0048] Step 1014: If the voice signal does not contain preset key intent information, the user is prompted to supplement the key intent information through the audio output device of the voice control screen.

[0049] In this embodiment, the preset key intent information refers to a set of core information elements predefined by the system to clarify the user's recording intent, typically including the item name, placement location, etc. Audio output device refers to a hardware component used to convert electrical signals back into sound, such as a speaker.

[0050] The system performs real-time or near-real-time speech recognition and natural language understanding analysis on the collected user speech signals, extracting semantic components from the sentences. The system compares the analysis results with preset key intent information templates. If the analysis finds that the user's speech description is missing necessary key intent information (for example, only saying "placed here" without specifying what the item is), the system plays a pre-recorded or speech-synthesized prompt through its built-in speaker to guide the user to fill in the missing information.

[0051] This application embodiment activates the system by receiving diverse wake-up signals, guides the user to make a voice description, automatically converts the voice into a digital record with a timestamp, and actively provides interactive prompts when the information is incomplete. This process achieves convenient and structured capture of the user's intention to place items.

[0052] In some embodiments, step 102 includes: Step 1021: When the voice control screen is woken up, control the camera to capture a first panoramic image with the first preset parameters.

[0053] In this embodiment, the voice control screen is an intelligent terminal device integrating voice interaction, image acquisition, and display functions. The camera is an image sensor component integrated into the voice control screen, used to acquire optical images and convert them into digital signals. The first preset parameters are a set of camera operating parameters pre-set by the system, suitable for capturing images covering a wide field of view, typically including specific focal length, aperture, and sensor settings. The first panoramic image refers to a wide-field digital image captured by the camera using the first preset parameters, capable of covering the main monitoring area in front of the voice control screen.

[0054] Upon detecting that the voice control screen has been activated by a user through voice command or physical action, the system immediately sends a control command to the camera integrated into the screen. The system instructs the camera to load and apply the first preset parameter configuration. After the camera completes the optical parameter adjustment according to the first preset parameters, the system triggers it to perform an image acquisition. The system stores and marks the acquired wide-field-of-view digital image, which covers a broad scene in front, as the first panoramic image, serving as the baseline scene data for subsequent image analysis.

[0055] Step 1022: During the process of receiving the voice command, dynamically adjust the angle and focal length of the camera according to the content of the voice command.

[0056] In this embodiment, a voice command is an audio signal containing the user's operational intention, issued by the user to the voice control screen. Angle is a spatial orientation parameter of the camera's optical axis relative to its initial mounting reference direction. Focal length is an optical parameter in the camera's optical system that determines the size of the imaging field of view and the size of the object image.

[0057] While receiving the user's voice command audio stream through the microphone array of the voice control screen, the system simultaneously activates a real-time speech recognition engine to parse the audio stream. From the parsed text information, the system extracts directional or regional descriptive keywords related to the object's placement. Based on these keywords, the system generates corresponding camera pan-tilt control commands and lens zoom control commands. These control commands are sent to the camera's drive mechanism in real time, which dynamically adjusts the camera's shooting angle and the optical lens's focal length accordingly, gradually aligning and focusing the camera's field of view on the target area described by the user's voice.

[0058] Step 1023: Take a second partial image containing the area where the target item is placed using the adjusted second preset parameters.

[0059] In this embodiment, the second preset parameters are a set of working parameters suitable for capturing a specific local area, after the system dynamically adjusts them according to the voice command in step 1022, resulting in a stable camera position. The target item placement area is the specific local spatial range within which the user intends to place the item. The second local image refers to a clear, focused, and fully contained close-up digital image of the target item placement area captured by the camera using the second preset parameters.

[0060] Once the system determines that the camera's angle and focus have been adjusted and stabilized according to the voice command, the system sends an image acquisition command to the camera. The camera then acquires an image using the currently adjusted second preset parameters. The system stores and labels the acquired close-up digital image, which clearly shows the details of the area where the target item is placed, as the second local image.

[0061] Step 1024: Save the first panoramic image and the second partial image together as the image data.

[0062] In this embodiment, the system creates a new data record entry to associate and store all information related to the current item placement event. The system associates the first panoramic image and the second partial image together under this data record entry, constituting the image data portion of this record. The system associates and binds this image data with synchronously recorded audio files of voice commands, parsed text information, timestamps, and other information, and performs data persistence operations to save it to local storage or upload it to a cloud database.

[0063] This application embodiment provides a global environmental context for the placement of an item through a first panoramic image, while a second local image provides precise visual details of the placement point. The combined image data, strictly correlated with voice commands in time and space, together construct a multimodal and precise record of the item's location. In some embodiments, step 103 includes: Step 1031: Perform speech recognition on the audio file to convert the speech content into text information.

[0064] In this embodiment, speech recognition is a technology that converts human speech signals into corresponding text information. Audio files are digital files that store sound information. Text information is readable data composed of character sequences.

[0065] The system invokes its built-in speech recognition engine to load the audio file to be processed. The speech recognition engine analyzes the sound waveform in the audio file, segmenting the continuous speech signal into discrete text units using acoustic and language models, ultimately generating complete text information corresponding to the speech content. This text information serves as input data for subsequent processing.

[0066] Step 1032: Extract keywords representing item attributes and location attributes from the text information.

[0067] In this embodiment, text information is readable data composed of character sequences. Keywords are words or phrases extracted from the text that represent core semantics. Item attributes are attributes that describe the characteristics of an item, such as name and category. Location attributes are attributes that describe the location or orientation of an item.

[0068] The system performs natural language processing on the generated text information. It applies named entity recognition and dependency parsing techniques to scan the text, identify and extract noun entities related to items as item attribute keywords, and simultaneously identify and extract nouns or phrases related to location and place as location attribute keywords. These extracted keywords are stored in a structured manner to represent the core intent of the user's voice commands.

[0069] Step 1033: Upload the image data to a preset recognition server, and use an object detection model to identify multiple items in the image data.

[0070] In this embodiment, image data is digital data that records visual information. The recognition server is a remote computer configured with a specific algorithm model for processing recognition tasks. The object detection model is a trained algorithm model capable of identifying specific object categories in an image and locating their positions.

[0071] The system transmits locally stored image data to a pre-configured remote recognition server via a network connection. Upon receiving the image data, the recognition server invokes the object detection model deployed on it. This model performs forward inference calculations on the input image data, analyzes image features, identifies the contours of multiple independent objects within the image, and outputs information for each identified object.

[0072] Step 1034: Generate location information for each identified item, including category label and coordinates in the image.

[0073] In this embodiment, the category label is a text identifier used to identify the category to which an item belongs. Coordinate position is a pair of numerical values ​​used to precisely locate a point in a two-dimensional plane of the image, typically including a horizontal coordinate and a vertical coordinate.

[0074] The system receives object detection results from the recognition server. For each item identified by the model, the system assigns it a most probable category name as a category label based on the classification confidence score output by the model. Simultaneously, the system reads the bounding box coordinates of the item output by the model; these coordinates, based on image pixels, contain the item's specific location and extent within the image. The system binds each item's category label to its corresponding coordinates, forming a complete record of item location information.

[0075] This application embodiment converts voice content into analyzable text and extracts key attributes. At the same time, it utilizes the server's powerful computing capabilities to perform accurate object recognition and positioning on images. This lays the data foundation for the subsequent accurate association and matching of voice commands with specific items in the visual scene, thereby supporting efficient and intuitive item location recording and retrieval functions.

[0076] In some embodiments, step 104 includes: Step 1041: Perform similarity matching between the item names in the keywords and the item categories.

[0077] In this embodiment, keywords refer to words or phrases extracted by the system from user voice commands that represent the user's intent, such as "key" or "desk." Item names are words within keywords that refer to specific items, such as "key." Item categories are labels obtained by the system after classifying objects detected in an image using image recognition technology; for example, a key might be categorized as a small metal item or a more specific key category. Similarity matching refers to the system using algorithms to calculate the semantic or character-level similarity between two text strings (such as item names and item category labels) to determine whether they refer to the same thing.

[0078] First, the system extracts the keyword "item name" from the parsed user voice command. Simultaneously, it obtains a list of all item category tags derived from image recognition analysis of the current scene image. Next, the system calls a pre-defined text similarity matching algorithm to compare the extracted item name with each item category tag obtained from image recognition. This algorithm evaluates the semantic relevance between the name and the category tag and outputs a similarity score. The system then determines whether the match is successful based on a preset threshold, thereby filtering out items in the image that may correspond to the user's stated item name.

[0079] Step 1042: If there is an item whose name and category are successfully matched in terms of similarity, mark the coordinates of the item in the image data as the storage location of the target item.

[0080] In this embodiment, image data refers to digital image information captured and stored by the system through a camera, specifically a photograph recording the environment in which an item is placed. Coordinate position refers to a two-dimensional numerical pair (usually X and Y coordinates) used to precisely locate a pixel in the digital image, here used to identify the specific location of the item in the image. Target item refers to the specific item that the user intends to place and whose position needs to be recorded. Storage location, in this context, refers to the physical location of the target item, which the system records by marking its coordinates in the image data.

[0081] After the system completes similarity matching in step 1041, it checks the number of successfully matched items. If, after calculation, only one item's category label is highly similar to the item name in the user's speech (i.e., a successful and unique match), the system proceeds to step 1042. The system retrieves the pixel region occupied by the uniquely matching item in the image data from the image recognition results and calculates the coordinates of a representative point (such as the center point) within that region. Subsequently, the system overlays a visual marker (such as an icon or highlight box) at the calculated coordinate point on the image data. This marked coordinate position is recorded by the system and associated with the storage location of the target item, completing the precise association from speech information to image location.

[0082] Step 1043: If there are multiple items whose names and categories are successfully matched for similarity, match the location description in the keywords with the location information.

[0083] In this embodiment, location description refers to words or phrases extracted from keywords in the user's voice commands that describe the location of the item, such as "on the desk" or "in the left drawer." Location information is a textual description of the environment in which the identified item is located, generated by the system through image analysis (such as combining panoramic and partial images) or through spatial relationship analysis of objects in image recognition, such as the left side of the screen or the desktop area.

[0084] In the matching results of step 1041, if the system finds multiple items that successfully match the similarity of the item name described by the user (e.g., multiple sets of keys identified in the image), the system proceeds to step 1043. At this point, the system needs further information to distinguish these similar items. The system extracts potential location description information (e.g., placed on a TV cabinet) from the previously parsed voice command keywords. Simultaneously, the system obtains the location information text descriptions generated for each identified item during the image recognition stage. Next, the system performs text matching or semantic comparison between the extracted location description keywords and the location information of each of these multiple candidate items to find a more accurate correspondence.

[0085] Step 1044: If there is an item whose location description in the keyword successfully matches the location information, mark the coordinate position of the item in the image data as the storage location of the target item.

[0086] In this embodiment, the system judges the matching result between the location description and the location information. If, after matching, only one item's location information successfully corresponds to the location description in the user's voice among multiple similar items (for example, the user says "keys on the TV cabinet," and only one set of keys in the image is described as being in the TV cabinet area), then the system executes step 1044. The system determines that this uniquely matched item is the target item pointed to by the user. Subsequently, the system obtains the coordinate position of the item in the image data and visually marks this coordinate point. This marked position is officially recorded by the system as the storage location of the target item.

[0087] Step 1045: If there are multiple items whose location descriptions in the keywords successfully match the location information, mark multiple candidate locations in the image data for the user to select and confirm.

[0088] In this embodiment, a candidate location refers to multiple possible location points identified in the image data when the system cannot uniquely determine the target item through automatic matching, allowing the user to make further selections. Here, "user" refers to the operator interacting with the voice control screen.

[0089] If, in step 1043, the system finds that after a successful match between the location description and location information, there are still multiple corresponding items (for example, the user says "keys on the table," but the image recognition shows two sets of keys on the table), then the system proceeds to step 1045. At this point, the system cannot make a decision autonomously. The system will visually mark the coordinates corresponding to these multiple matching items (for example, using dotted lines with different numbers). These marked points together constitute multiple candidate locations. The system displays these image data with candidate location markings on the interactive interface of the central control screen, clearly prompting the user to select from these options to confirm the final correct location.

[0090] Step 1046: Determine the storage location of the target item based on the user's selection results for the multiple candidate locations.

[0091] In this embodiment of the application, the selection result refers to the input information that the user selects one of the multiple candidate positions provided by the system as the final answer through touch screen click, voice confirmation or other interactive methods.

[0092] After the system displays an image with multiple candidate locations in step 1045, it awaits user input. The user taps one of the marked candidate locations on the central control screen. The system captures this touch event and parses the coordinates of the tapped point on the image. The system compares these coordinates with the coordinates of each candidate location generated in step 1045 to determine the specific location selected by the user. Subsequently, the system determines this user-confirmed coordinate location as the storage location of the target item and completes data association and storage.

[0093] This application's embodiments, through the sequential execution of similarity matching, location description matching, and final user interaction confirmation, can progressively approach and ultimately pinpoint the exact location of the item pointed to by the user in the image. This effectively solves the ambiguity problem of automatic system recognition when multiple similar items exist in the scene or when voice commands are not precise enough. It ensures the accuracy of item location recording, transforming vague voice descriptions into precise coordinates in the image, providing a reliable data foundation for subsequent efficient retrieval, thereby comprehensively improving the intelligence and practicality of the item location management system.

[0094] In some embodiments, step 105 includes: Step 1051: Establish an association index between the audio file, the image data, the keywords, and the storage location of the target item.

[0095] In this embodiment, the association index is a data structure used to establish and store explicit pointing relationships between different types of data items, enabling the system to quickly locate all other related data items through one of the data items.

[0096] After completing audio file recording, image data acquisition, keyword extraction, and determining the storage location of the target item, the system creates an associated index. This index logically binds the unique identifier of the audio file, the storage path of the image data, the set of keywords extracted from the user's speech, and the specific storage location of the item confirmed through image recognition or user interaction (e.g., coordinate information or area description). This binding relationship allows the system to subsequently retrieve all other related information (such as the corresponding audio, images, and location) based on any information (such as keywords).

[0097] Step 1052: Store the associated index and corresponding file data in a local database or a cloud database.

[0098] In this embodiment, the local database is a storage system deployed inside a user terminal device (such as a voice control screen); the cloud database is a networked storage system deployed on a remote server and accessible via the Internet.

[0099] The system generates a related index, along with the original audio files, image data, and other file data that the index points to, and stores them in a designated data storage system. Depending on the configuration or network conditions, the system chooses to store the data in the device's local embedded database or, after encryption, store it on a remote cloud database server via network transmission. This storage process ensures the persistent preservation of all recorded information, providing a data foundation for subsequent retrieval operations.

[0100] Step 1053: Perform intelligent summarization processing on the audio file to extract key information fragments and generate summary text or summary audio.

[0101] In this application embodiment, intelligent summarization processing is a technical process that automatically analyzes long audio content through algorithms, identifies and extracts core information points, and its output is a summary text or audio segment.

[0102] The system applies audio analysis algorithms to stored audio files. These algorithms analyze the speech content, identifying phrases describing key information such as item names, locations, and times. The system then extracts these key phrases from the original audio and combines them into a shorter summary audio file. Simultaneously, the system can also use speech-to-text technology to convert this key information into a condensed summary text. This process aims to extract the core content of the audio recording, reducing the time users need to review the complete recording.

[0103] Step 1054: Receive the custom tag added by the user for the target item, and bind and store the custom tag with the associated index.

[0104] In this embodiment of the application, custom tags are additional category or attribute keywords that users define for target items based on their personal management habits, such as important, frequently used, tools, etc.

[0105] The system receives custom tag information input by the user through the interactive interface of the voice control screen. The user can speak the tag name or enter the tag text via touch on the control screen. After receiving the tag, the system binds the tag data to the association index created for the target item in step 1051. After binding, the custom tag becomes a searchable attribute of the association index and is linked to the item's audio, image, location, and other information. The system then stores the updated index information back into the database.

[0106] This application embodiment closely links the audio recordings, visual images, text keywords, and physical location of items, optimizes information presentation efficiency through intelligent summarization, and supports user-defined classifications, thereby achieving efficient, accurate, and flexible recording and retrieval of item location information, significantly improving the user experience and efficiency when searching for and managing items.

[0107] In some embodiments, step 106 includes: Step 1061: Receive the item query command input by the user via voice or touch.

[0108] In this embodiment, the voice control screen refers to an intelligent terminal device that integrates voice recognition, touch interaction, display, and computing capabilities, typically serving as the control hub of a smart home system. An item search command is a command issued by a user to the system to find a specific item, containing a clear search intent; its input form can be a voice command or a screen touch operation.

[0109] The system continuously listens for or monitors input signals from the user through its integrated microphone array or touch sensors. When the system detects a voice waveform that matches a preset wake-up pattern, or recognizes a specific touch gesture on the screen (such as clicking a query icon), it determines it as a valid command input event. The system then activates the command processing thread, sending the captured raw voice signal or touch coordinate data as the initial data packet for the item query command to the subsequent command parsing module for processing.

[0110] Step 1062: Identify the query keywords in the item query instruction and perform fuzzy matching retrieval in the associated index of the database.

[0111] In this embodiment, query keywords refer to words or phrases extracted from the user's query instructions using natural language processing technology that represent the core attributes of the item being searched, such as the item name, storage location, or user-defined tags. The database's association index is a pre-built data structure that systematically links audio files, image data, extracted keyword tags, and item location information, aiming to achieve rapid data cross-referencing. Fuzzy matching retrieval is an information retrieval technique that allows query keywords to have certain character differences, synonym substitutions, or semantic similarities with entries in the index, rather than requiring a completely precise literal match, thereby improving the retrieval's error tolerance and recall rate.

[0112] The system uses a speech recognition engine to translate received voice commands into text-based query statements; for touch commands, it directly retrieves a preset text query template. Subsequently, the natural language processing module performs word segmentation, part-of-speech tagging, and named entity recognition on the query text, extracting one or more query keywords. The system compares these keywords with the associated index in the database. The retrieval algorithm calculates the similarity between each query keyword and the keyword tags of each record in the index, setting a similarity threshold. All database records with similarity exceeding this threshold are considered preliminary matches, forming a candidate result set.

[0113] Step 1063: Retrieve the image data that successfully matches the query keywords and the storage location of the target item.

[0114] In this embodiment, image data refers to panoramic and partial images captured and stored by the system through a camera during the item placement recording stage, along with metadata (such as item category and bounding box coordinates) generated after image recognition and analysis. The storage location of the target item is a structured location description, which may originate from two-dimensional / three-dimensional coordinates derived from image analysis, area labels associated with a map (such as the second shelf of a bookshelf), or point coordinates manually confirmed by the user.

[0115] The system sorts the candidate results based on matching similarity and selects the record with the highest similarity as the final matching result. Then, based on the unique identifier of this record in the database, the system initiates a data retrieval request to the storage subsystem. The storage subsystem, using this identifier, reads the image data (which may be the image file path or already loaded image data) associated with the item placement record and the target item's storage location information (such as coordinate data or text description) from the corresponding data table. This data is encapsulated into a data object and returned to the core processing module for subsequent display.

[0116] Step 1064: Display a visual card containing the image data and location markers on the voice control screen, and play the relevant audio file or summary information through the audio output device.

[0117] In this embodiment, the visual card is an information aggregation view component rendered on the graphical user interface by the system. It embeds multimedia elements such as images, text labels, and location markers to intuitively present search results to the user. The audio output device refers to hardware such as speakers and headphones connected to the system for playing sound. The summary information is a condensed version of the key content generated after automatic speech recognition and text summarization processing of the original long audio file; it can be a short audio clip or a text summary.

[0118] The system invokes the graphics rendering engine to draw visual cards within the display area of ​​the voice control screen. The rendering engine uses the retrieved image data as the card's background or main view, and overlays a highlighted, flashing, or specially shaped graphic marker (such as an arrow or circle) on the corresponding coordinates of the image based on the target item's storage location information. Simultaneously, the card displays auxiliary information such as the item name and recording time in text form. During or after interface rendering, the system controls the audio output device via the audio driver interface to play the original audio file associated with the recording, or plays a speech version of the summary information synthesized by the text-to-speech engine. Audio playback and visual marker display are synchronized or sequential to provide multimodal cues.

[0119] The embodiments of this application ensure that the system can flexibly and accurately understand the user's search needs, efficiently locate and extract the most relevant historical data from the structured database, and transform abstract database records into intuitive and three-dimensional guidance information by integrating visual and auditory dual-channel feedback. Ultimately, this improves the efficiency of users in finding items, reduces operational complexity, and provides adaptive assistance for users with different perceptual impairments.

[0120] Reference Figure 2 When the central control screen is woken up, the central control camera first takes a panoramic view. When the user gives a voice command, the central control camera adjusts its angle and focus to take a partial image. By combining the panoramic view and the partial image, the area where the user can place items can be narrowed down. The system saves and parses the user's voice command audio, generates keyword tags, records them to the captured images, and stores the image information in association with the audio information.

[0121] Reference Figure 3 When a user needs to find an item, they can issue a voice query command to the central control screen (such as "Where did I put my keys?"). The system will extract keywords through voice recognition and perform a matching search in the database to find the most relevant audio record associated with the image.

[0122] Based on the foregoing embodiments, this application provides an item location retrieval device. The device includes various units and modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0123] Figure 4 This is a schematic diagram of the composition structure of an item location retrieval device provided in an embodiment of this application, as shown below. Figure 4 As shown, the item location retrieval device 20 includes: The recording module 201 is used to respond to the user's trigger operation on the central control screen, wake up the central control screen and receive the user's voice command containing the placement information of the item, convert the voice command into an audio file and associate it with a timestamp; Control the camera to capture image data of the scene containing the placement of objects; The audio file is parsed to extract keywords, and the image data is used for object recognition to obtain item category and location information; The keywords are matched with the item category and the location information using a multimodal association to determine and record the storage location of the target item. The audio file, the image data, the keywords, and the storage location of the target item are associated and stored in the database; The retrieval module 202 is used to receive an item query command input by the user on the central control screen, perform a retrieval in the database based on the item query command, and output the retrieval results associated with the storage location of the target item.

[0124] In some embodiments, the recording module 201 is further configured to: The system receives a wake-up signal from the user via voice command or physical action and activates the recording function of the voice control screen. The voice input device of the voice control screen collects the user's voice signal describing the name of the item and its location. The speech signal is converted into a digital audio file, and a timestamp representing the recording time is added to the audio file; If the voice signal does not contain preset key intent information, the user will be prompted to supplement the key intent information through the audio output device of the voice control screen.

[0125] In some embodiments, the recording module 201 is further configured to: When the voice control screen is activated, the camera is controlled to capture a first panoramic image with a first preset parameter; During the process of receiving the voice command, the angle and focus of the camera are dynamically adjusted according to the content of the voice command. A second partial image containing the area where the target item is placed is captured using the adjusted second preset parameters; The first panoramic image and the second partial image are saved together as the image data.

[0126] In some embodiments, the recording module 201 is further configured to: The audio file is subjected to speech recognition to convert the speech content into text information; Extract keywords representing item attributes and location attributes from the text information; The image data is uploaded to a preset recognition server, and multiple items in the image data are identified through an object detection model; For each identified item, generate location information including a category label and its coordinates in the image.

[0127] In some embodiments, the recording module 201 is further configured to: Match the item names in the keywords with the item categories based on similarity. If there is one item whose name and category are successfully matched in terms of similarity, the coordinates of that item are marked in the image data as the storage location of the target item. If multiple items are successfully matched between the item name and the item category, the location description in the keyword is matched with the location information. If there is an item whose location description in the keyword matches the location information, the coordinates of the item are marked in the image data as the storage location of the target item. If multiple items successfully match the location description in the keyword with the location information, multiple candidate locations are marked in the image data for the user to select and confirm. The storage location of the target item is determined based on the user's selection of multiple candidate locations.

[0128] In some embodiments, the recording module 201 is further configured to: Establish an association index between the audio file, the image data, the keywords, and the storage location of the target item; The associated index and corresponding file data are stored in a local database or a cloud database; The audio file is subjected to intelligent summarization processing to extract key information segments and generate summary text or summary audio; Receive custom tags added by the user for the target item, and bind and store the custom tags with the associated index.

[0129] In some embodiments, the retrieval module 202 is further configured to: Receive the item query command input by the user via voice or touch; Identify the query keywords in the item query command and perform fuzzy matching retrieval in the associated index of the database; Retrieve the image data that successfully matches the query keywords and the storage location of the target item; The voice control screen displays a visual card containing the image data and location markers, and plays the relevant audio file or summary information through an audio output device.

[0130] This application embodiment achieves accurate, intuitive, and efficient recording and retrieval of item placement information through response triggering, receiving voice commands, image acquisition, parsing and recognition, multimodal matching and association, structured storage, and intelligent retrieval output. It overcomes the shortcomings of single voice recording lacking spatial intuitiveness, utilizes images to provide visual anchors, and greatly improves the efficiency and accuracy of information management through the associated storage of voice and images and keyword-based retrieval. It provides users with a natural, convenient, and reliable multimodal item location management experience, and is particularly helpful in solving the problem of difficulty in finding forgotten items in daily life.

[0131] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this application can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0132] It should be noted that, in the embodiments of this application, if the above-mentioned method for retrieving the location of items is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0133] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0134] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.

[0135] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0136] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0137] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0138] It should be noted that, Figure 5 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application, such as... Figure 5 As shown, the hardware entity of the computer device 700 includes: one or more processors 701, a communication interface 702, and a memory 703, wherein: Processor 701 typically controls the overall operation of computer device 700.

[0139] Communication interface 702 enables computer devices to communicate with other terminals or servers over a network.

[0140] The memory 703 is configured to store instructions and applications executable by the processor 701, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 701 and various modules in the computer device 700. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 701, the communication interface 702, and the memory 703 can be performed via bus 704. Only one processor is shown in the figure; each processor 100 includes one or more cores.

[0141] It should be noted that the computer device may include multiple processors 701, and each processor 701 can interact with each other through aggregated communication methods such as all-to-all, all-gather, or all-reduce. The processors 701 may be central processing units (CPUs), graphics processing units (GPUs), embedded neural network processing units (NPUs), tensor processing units (TPUs), data processing units (DPUs), accelerated processing units (APUs), floating-point processing units (FPUs), or application-specific integrated circuits (ASICs). The processors may also be single-core or multi-core processors. The processor may consist of a CPU and hardware chips. The hardware chips may be ASICs, PLDs, or combinations thereof. The PLDs may be complex programmable logic devices (CPLDs), FPGAs, generic array logic (GALs), or any combination thereof. The processor can also be implemented using logic devices with built-in processing logic, such as FPGAs or digital signal processors (DSPs).

[0142] The communication interface 702 can be a wired interface or a wireless interface, used to communicate with other modules or devices. The wired interface can be an Ethernet interface, a local interconnect network (LIN), etc., and the wireless interface can be a cellular network interface or a wireless LAN interface, etc.

[0143] Memory 703 can be non-volatile memory, such as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Memory 703 can also be volatile memory, which can be random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synclink dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM), direct rambus DRAM (DRDRAM), and rambus DRAM.

[0144] The 704 bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc.

[0145] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0146] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0147] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0148] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0149] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0150] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0151] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0152] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for recording and retrieving the location of an item, characterized in that, The method includes: In response to the user's trigger operation on the central control screen, the central control screen is woken up and the user's voice command containing the placement information of the item is received. The voice command is converted into an audio file and associated with a timestamp. Control the camera to capture image data of the scene containing the placement of objects; The audio file is parsed to extract keywords, and the image data is used for object recognition to obtain item category and location information; The keywords are matched with the item category and the location information using a multimodal association to determine and record the storage location of the target item. The audio file, the image data, the keywords, and the storage location of the target item are associated and stored in the database; The central control screen receives item query commands input by the user, performs a search in the database based on the item query commands, and outputs search results associated with the storage location of the target item.

2. The method according to claim 1, characterized in that, The process of responding to a user's trigger operation on the central control screen, waking up the central control screen and receiving a voice command from the user containing information about item placement, converting the voice command into an audio file and associating it with a timestamp, includes: The system receives a wake-up signal from the user via voice command or physical action and activates the recording function of the voice control screen. The voice input device of the voice control screen collects the user's voice signal describing the name of the item and its location. The speech signal is converted into a digital audio file, and a timestamp representing the recording time is added to the audio file; If the voice signal does not contain preset key intent information, the user will be prompted to supplement the key intent information through the audio output device of the voice control screen.

3. The method according to claim 1, characterized in that, The control camera acquires image data containing the scene of object placement, including: When the voice control screen is activated, the camera is controlled to capture a first panoramic image with a first preset parameter; During the process of receiving the voice command, the angle and focus of the camera are dynamically adjusted according to the content of the voice command. A second partial image containing the area where the target item is placed is captured using the adjusted second preset parameters; The first panoramic image and the second partial image are saved together as the image data.

4. The method according to claim 1, characterized in that, The process of parsing the audio file to extract keywords and performing object recognition on the image data to obtain item category and location information includes: The audio file is subjected to speech recognition to convert the speech content into text information; Extract keywords representing item attributes and location attributes from the text information; The image data is uploaded to a preset recognition server, and multiple items in the image data are identified through an object detection model; For each identified item, generate location information including a category label and its coordinates in the image.

5. The method according to claim 1, characterized in that, The step of performing multimodal association matching between the keywords, the item category, and the location information to determine and record the storage location of the target item includes: Match the item names in the keywords with the item categories based on similarity. If there is one item whose name and category are successfully matched in terms of similarity, the coordinates of that item are marked in the image data as the storage location of the target item. If multiple items are successfully matched between the item name and the item category, the location description in the keyword is matched with the location information. If an item exists that successfully matches the location description in the keyword with the location information, the coordinates of the item are marked in the image data as the storage location of the target item.

6. The method according to claim 5, characterized in that, The method further includes: If multiple items successfully match the location description in the keyword with the location information, multiple candidate locations are marked in the image data for the user to select and confirm. The storage location of the target item is determined based on the user's selection of multiple candidate locations.

7. The method according to claim 1, characterized in that, The step of associating and storing the audio file, the image data, the keywords, and the storage location of the target item in the database includes: Establish an association index between the audio file, the image data, the keywords, and the storage location of the target item; The associated index and corresponding file data are stored in a local database or a cloud database; The audio file is subjected to intelligent summarization processing to extract key information segments and generate summary text or summary audio; Receive custom tags added by the user for the target item, and bind and store the custom tags with the associated index.

8. The method according to claim 1, characterized in that, The process of receiving a user-inputted item search command, performing a search in the database based on the command, and outputting search results associated with the storage location of the target item includes: Receive the item query command input by the user via voice or touch; Identify the query keywords in the item query command and perform fuzzy matching retrieval in the associated index of the database; Retrieve the image data that successfully matches the query keywords and the storage location of the target item; The voice control screen displays a visual card containing the image data and location markers, and plays the relevant audio file or summary information through an audio output device.

9. A device for retrieving the location of an item, characterized in that, The device includes: The recording module is used to respond to the user's trigger operation on the central control screen, wake up the central control screen and receive the user's voice command containing the placement information of the item, convert the voice command into an audio file and associate it with a timestamp; Control the camera to capture image data of the scene containing the placement of objects; The audio file is parsed to extract keywords, and the image data is used for object recognition to obtain item category and location information; The keywords are matched with the item category and the location information using a multimodal association to determine and record the storage location of the target item. The audio file, the image data, the keywords, and the storage location of the target item are associated and stored in the database; The retrieval module is used to receive item query instructions input by the user on the central control screen, perform a retrieval in the database based on the item query instructions, and output retrieval results associated with the storage location of the target item.

10. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for retrieving the location of an item as described in any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for retrieving the location of an item as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program, which, when read and executed by a computer, implements the steps of the method for retrieving the location of an item as described in any one of claims 1 to 8.