Converting wireless communication system-based positioning into location descriptions describing layout with respect to asset
By integrating a processor and memory into the ESL system, and utilizing wireless communication technology to acquire UE location information and convert it into an asset description layout, the problem of inaccurate asset location information in the ESL system is solved, improving the efficiency of users in finding assets and reducing the use of paper tags and environmental impact.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-08-06
- Publication Date
- 2026-05-01
AI Technical Summary
Existing electronic shelf label (ESL) systems fail to provide accurate information about asset locations, leading to inefficiency for users searching for specific assets in retail environments, and the replacement of paper labels increases maintenance costs and environmental burden.
By integrating a processor and memory into the ESL system, the location information of the user equipment (UE) is obtained using wireless communication technology, survey video frames are captured, features are extracted, and a location description is output based on the asset description layout identifier, thus realizing the conversion from wireless positioning to asset description layout.
It improves the efficiency of users in finding assets in the retail environment, reduces the frequency of paper label usage, lowers maintenance costs, and reduces environmental impact.
Smart Images

Figure CN121970077A_ABST
Abstract
Description
Transform location data based on wireless communication systems into location descriptions of asset layouts.
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Patent Application No. 18 / 485,111, filed October 11, 2023, entitled "CONVERTING WIRELESS COMMUNICATIONS SYSTEM-BASED POSITIONING TO POSITIONAL DESCRIPTIONS WITHRESPECT TO DESCRIPTIONAL LAYOUTS OF ASSETS", the entire contents of which are expressly incorporated herein by reference. Technical Field
[0003] This disclosure relates generally to wireless communication systems, and more specifically to wireless positioning. Some features enable and provide improved communication, including translating positioning based on the wireless communication system into a location description of the asset's layout, which can be utilized in an electronic shelf label (ESL) system. Background Technology
[0004] Retail stores typically use paper labels to display information about products displayed on shelves, such as price, discount rates, unit cost, and country of origin. Using such paper labels for price display has limitations. For example, when product information or positioning on the shelf changes, retailers must generate new paper labels and discard the old ones. This increases maintenance costs for both supply chain and employee labor. Furthermore, from an environmental perspective, replacing labels wastes raw materials such as paper, negatively impacting environmental protection. Moreover, human error is prone to occur, such as mislabeling shelves or products or forgetting to remove temporary price changes from certain shelves, which can lead to shopper frustration.
[0005] Electronic shelf label (ESL) devices are electronic devices used to display price information for items on retail store shelves, replacing paper labels. ESL devices are attached to the front edge of retail shelves and use display devices such as liquid crystal displays (LCDs) to show various pricing information. ESL devices can be programmed with new product information whenever information about a product or its positioning changes. Therefore, electronic shelf labels can be reused repeatedly.
[0006] ESL systems provide easily up-to-date information about the location of assets in retail locations (such as stores) or storage locations (such as warehouses). While ESL devices can display asset names, prices, or other information, this information may not provide users traversing the location with information about their own location or the location of other assets. Therefore, shoppers or stockers may spend significant time traversing the location to find a specific asset or the exact location of a stored asset. Although some ESL devices support location operations, these generate a position relative to other ESL devices and represent the user's position within the location's boundaries, such as a position in a Cartesian coordinate system. If a user possesses a map and their position relative to that map, they might be able to pinpoint the location of a specific asset; however, such a map may not exist or may not include the latest asset location, leaving the user searching for the specific asset without guidance. Summary of the Invention
[0007] The following summary outlines some aspects of this disclosure to provide a basic understanding of the techniques discussed. This summary is not an exhaustive overview of all the intended features of this disclosure, nor is it intended to identify key or essential elements of all aspects of this disclosure, nor to define the scope of any or all aspects of this disclosure. The sole purpose of this summary is to present, in a general form, some concepts of one or more aspects of this disclosure as a prelude to the more detailed description that follows.
[0008] In one aspect of this disclosure, an apparatus includes: at least one processor; and a memory coupled to the at least one processor. The at least one processor is configured to cause the apparatus to: acquire location information associated with a user equipment (UE). The at least one processor is further configured to cause the apparatus to: retrieve one or more frames of survey video of a location corresponding to an estimated location of the UE. The estimated location is based on the location information. The at least one processor is configured to cause the apparatus to: extract a set of features from the one or more frames. The at least one processor is further configured to cause the apparatus to: output a description of the estimated location of the descriptive layout based on one or more descriptions identified from the descriptive layout identifier of an asset associated with the location. The one or more descriptions are identified based on the set of features.
[0009] In an additional aspect of this disclosure, a method of communicating in an ESL system includes: obtaining location information associated with a UE. The method further includes: capturing one or more frames of a survey video of a location corresponding to an estimated location of the UE. The estimated location is based on the location information. The method includes: extracting a set of features from the one or more frames. The method further includes: outputting a description of the estimated location of the descriptive layout based on one or more descriptions identified from descriptive layout identifiers of assets associated with the location. The one or more descriptions are identified based on the set of features.
[0010] In an additional aspect of this disclosure, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations include: obtaining location information associated with a UE. The operations further include: capturing one or more frames of a survey video of a location corresponding to an estimated location of the UE. The estimated location is based on the location information. The operations include: extracting a set of features from the one or more frames. The operations further include: outputting a description of the estimated location of the descriptive layout based on one or more descriptions identified from the descriptive layout identifier of an asset associated with the location. The one or more descriptions are identified based on the set of features.
[0011] In an additional aspect of this disclosure, an Electronic Shelf Tag (ESL) system includes a server comprising: a memory; and at least one processor coupled to the memory and configured to perform operations. The operations include: obtaining location information associated with a UE. The operations further include: capturing one or more frames of survey video of a location corresponding to an estimated location of the UE. The estimated location is based on the location information. The operations include: extracting a set of features from the one or more frames. The operations further include: outputting a description of the estimated location of the described layout based on one or more descriptions identified from the described layout identifiers of assets associated with the location. The one or more descriptions are identified based on the set of features.
[0012] The features and technical advantages of the examples according to this disclosure have been summarized rather extensively above in order to better understand the detailed description below. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily utilized as the basis for modifying or designing other structures for achieving the same purpose of this disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein (both their organization and manner of operation) and the associated advantages will be better understood from the following description when considered in conjunction with the accompanying drawings. Each figure in the drawings is provided for illustrative and descriptive purposes and not as a limitation of the definitions in the claims.
[0013] Devices, networks, and systems can be configured to communicate via one or more portions of the electromagnetic spectrum. This disclosure refers to certain communication technologies, such as Bluetooth or Wi-Fi, to describe certain aspects. However, this description is not intended to be limited to any particular technology or application, and one or more aspects described with reference to one technology may be understood to be applicable to another technology. Furthermore, it should be understood that, in operation, wireless communication networks adapted according to the concepts herein may operate using any combination of licensed or unlicensed spectrum, depending on load and availability. Therefore, it will be apparent to those skilled in the art that the systems, apparatuses, and methods described herein can be applied to other communication systems and applications besides the specific examples provided.
[0014] For example, the specific implementation described can be implemented in any device, system, or network capable of transmitting and receiving RF signals according to any of the wireless communication standards, including any of the IEEE 802.11 standards, IEEE 802.15.1 Bluetooth, etc. ® Standards, Bluetooth Low Energy (BLE), Code Division Multiple Access (CDMA), Frequency Division Multiple Access (FDMA), Time Division Multiple Access (TDMA), Global System for Mobile Communications (GSM), GSM / General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), Terrestrial Trunking Radio (TETRA), Wideband CDMA (W-CDMA), Evolved Data Optimized (EV-DO), 1×EV-DO, EV-DO Revision A, EV-DO Revision B, High-Speed Packet Access (HSPA), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolved High-Speed Packet Access (HSPA+), Long Term Evolution (LTE), AMPS, 5G New Radio (5G NR), 6G, or other known signals used for communication within wireless networks, cellular networks, or Internet of Things (IoT) networks (such as systems utilizing 3G, 4G, 5G, or 6G technologies or further implementations thereof).
[0015] In various specific implementations, technologies and devices can be used in wireless communication networks such as Code Division Multiple Access (CDMA) networks, Time Division Multiple Access (TDMA) networks, Frequency Division Multiple Access (FDMA) networks, Orthogonal FDMA (OFDMA) networks, Single Carrier FDMA (SC-FDMA) networks, LTE networks, GSM networks, fifth-generation (5G) or new radio (NR) networks (sometimes referred to as "5G NR" networks, systems, or devices), and other communication networks. As described herein, the terms "network" and "system" are used interchangeably and can refer to a collection of devices capable of communicating with each other via one or more communication technologies.
[0016] While aspects and implementations are described herein by way of example, those skilled in the art will understand that additional implementations and use cases may arise in many different arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, or package arrangements. For example, implementations or uses may be achieved via integrated chip implementations or other devices based on non-modular components (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail or purchasing devices, medical devices, AI-enabled devices, etc.).
[0017] The scope of implementations can range from chip-level or modular components to non-modular, non-chip-level implementations, and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more of the described aspects. In some settings, devices incorporating the described aspects and features may also include additional components and features for implementing and practicing the claimed and described aspects. The innovations described herein are expected to be practiced in a wide variety of implementations of different sizes, shapes, or constructions, including both large and small devices, chip-level components, multi-component systems (e.g., radio frequency (RF) chains, communication interfaces, processors), distributed arrangements, end-user equipment, etc.
[0018] In the following description, numerous specific details (such as examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of this disclosure. As used herein, the term "coupled" means a direct connection or a connection via one or more intermediate components or circuits. Furthermore, specific terminology is set forth in the following description and for purposes of explanation to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that practicing the teachings disclosed herein may not require these specific details. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring the teachings of this disclosure.
[0019] Certain portions of the following detailed description are presented using other symbolic representations of procedures, logic blocks, processes, and data bit operations within computer memory. In this disclosure, procedures, logic blocks, processes, etc., are conceived as a self-consistent sequence of steps or instructions that produce a desired result. These steps are those that require physical operations on physical quantities. Although not strictly necessary, these physical quantities typically take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated within a computer system.
[0020] In the accompanying drawings, a single block can be described as performing one or more functions. The one or more functions performed by this block may be performed in a single component or across multiple components, and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps are described below in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure. Additionally, the example device may include components other than those shown, including well-known components such as processors, memory, etc.
[0021] Unless otherwise specifically stated, it will be apparent from the following discussion that, throughout this application, the use of terms such as “access,” “receive,” “transmit,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” “set,” and “generate” refers to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities in the registers and memories of the computer system into other data similarly represented as physical quantities in the registers, memories, or other such information storage, transmission, or display devices of the computer system.
[0022] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (such as a smartphone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more components that can implement at least some parts of this disclosure. Although the term "device" is used in the following description and examples to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. As used herein, an apparatus can include a device or part of a device for performing the described operations.
[0023] As used herein (including the claims), the term "or" in a list of two or more items means that any one of the listed items may be used alone, or any combination of two or more listed items may be used. For example, if a device is described as containing components A, B, or C, then the device may contain: A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C.
[0024] Furthermore, as used herein (including the claims), the word “or” in a list of items beginning with “at least one of” indicates a separate list such that a list such as “at least one of A, B or C” refers to A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items.
[0025] Additionally, as used herein, the term “substantially” is defined as being largely but not necessarily entirely what is specified (and includes what is specified; for example, substantially 90 degrees includes 90 degrees and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any specific implementation of the disclosure, the term “substantially” may be used in place of the specified content within “[percentage]”, where percentage includes 0.1%, 1%, 5%, or 10%.
[0026] Additionally, as used herein, relative terms, unless otherwise specified, can be understood as a quantity relative to a reference. For example, terms such as “higher” or “lower” or “more” or “less” can be understood as a threshold amount higher, lower, more, or less than a reference value. Attached Figure Description
[0027] A further understanding of the nature and advantages of this disclosure can be achieved by referring to the following figures. In the figures, similar components or features may have the same reference numerals. Furthermore, various components of the same type can be distinguished by adding a dash after the reference numerals and a second reference numeral for differentiation between similar components. If only the first reference numeral is used in the specification, the description applies to any one of the similar components having the same first reference numeral, regardless of the second reference numerals.
[0028] Figure 1A is a block diagram illustrating an example electronic shelf label (ESL) system according to some embodiments of the present disclosure.
[0029] Figure 1B is a diagram illustrating an example display of an ESL device.
[0030] Figure 2A is a perspective view of a display rack with ESL equipment according to some embodiments of this disclosure.
[0031] Figure 2B is a top view of a retail environment with user-accessible ESL equipment according to some embodiments of this disclosure.
[0032] Figure 3 is a timing diagram illustrating time-division multiplexing for communicating with multiple ESL devices according to some embodiments of the present disclosure.
[0033] Figure 4 is a block diagram illustrating an example ESL device according to some embodiments of the present disclosure.
[0034] Figure 5 is a block diagram illustrating an example wireless communication system that converts location based on a wireless communication system into a location description of an asset's layout, supported by one or more aspects of this disclosure.
[0035] Figure 6 is a block diagram illustrating an example of location estimation and description of a pipeline according to one or more aspects of this disclosure.
[0036] Figure 7 is a block diagram illustrating an example of location estimation and description of a pipeline according to one or more aspects of this disclosure.
[0037] Figure 8 is a perspective view of a portion of a retail location with ESL equipment and assets according to one or more aspects of this disclosure, for which location based on a wireless communication system is translated into a location description of the asset layout.
[0038] Figure 9 is a flowchart illustrating an example process of converting location based on a wireless communication system into a location description of the layout of an asset, supported by one or more aspects of this disclosure.
[0039] Figure 10 is a block diagram of an example ESL device that converts location based on a wireless communication system into a location description of the layout of an asset, supported by one or more aspects of this disclosure.
[0040] Figure 11 is a flowchart illustrating an example process of converting location based on a wireless communication system into a location description of the layout of an asset, supported by one or more aspects of this disclosure.
[0041] Figure 12 is a block diagram of an example user equipment (UE) that converts location based on a wireless communication system into a location description of the layout of an asset, supported by one or more aspects of this disclosure.
[0042] Similar reference numerals and names in the various figures indicate similar elements. Detailed Implementation
[0043] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to limit the scope of this disclosure. Rather, the detailed description includes specific details for providing a thorough understanding of the subject matter of the invention. It will be apparent to those skilled in the art that these specific details are not necessary in every situation, and in some cases, well-known structures and components are shown in block diagram form for clarity of presentation.
[0044] This disclosure provides systems, apparatus, methods, and computer-readable media that support the conversion of location based on a wireless communication system into location descriptions of asset layouts. The conversion of location based on a wireless communication system is performed by one or more frames of survey video, such as those derived from location information derived from location operations performed with respect to the wireless communication system, to estimate the location of a user equipment (UE) and identify locations corresponding to the estimated locations. Object detection can be performed on the one or more frames to identify assets, signs, or other such information in the frame, and a set of features representing the identified objects and / or signs can be extracted from the one or more frames. This set of features can be used to perform pattern matching on the asset layout descriptions (such as shelf diagrams or product information databases) to identify one or more descriptions corresponding to the estimated locations. In some embodiments, a Large Language Model (LLM) artificial intelligence (AI) model can be used to generate the description of this set of features for performing pattern matching with descriptions in the layout descriptions. The identified descriptions can be used to output a description of the estimated location of the UE describing the layout, such as a description of the estimated location relative to one or more assets (e.g., "the user is in front of the canned beans, and the canned beans are to the left of the canned peas") or in "retail semantics" (e.g., aisle labels, shelf labels, and rack labels). In some implementations, the descriptions of asset locations and estimated locations may be primarily or entirely based on markers located at various locations and identified in the one or more frames of the video survey.
[0045] Specific implementations of the subject matter described in this disclosure are possible to achieve one or more of the following potential advantages or benefits. In some aspects, this disclosure provides techniques for wireless communication systems that enable the conversion of wireless communication-based positioning into more asset-related positioning. This conversion allows devices to utilize wireless communication-based positioning operations and information for tracking assets (such as shelf maps or product information databases) to generate location descriptions that increase utility for users without requiring additional positioning operations using different coordinate systems or other reference frames. Such location descriptions are likely to be more useful to users than conventional location coordinates and reduce the time spent orienting themselves about assets stored in locations such as stores or warehouses. Furthermore, converting positioning information into more description-based positioning provides the information needed to support reverse lookups, which can be used to provide users with directions to specific assets. For example, a shopper in a store might query for directions to bread, and the ESL system could determine the direction from the user's current location to the bread and display such directions in retail semantics, such as "Move to aisle 3, look to the right, on shelf 7, the bread is on the third shelf (below the hamburger and hot dog buns)." Such orientation reduces the time users spend searching for specific assets compared to knowing their location coordinates within a store but not their placement. These advantages can be particularly beneficial in Electronic Shelf Labeling (ESL) systems, which provide an ESL-based positioning framework and, in some implementations, offer information associated with the location of ESL equipment and assets that can be used as a basis for or supplement to describing the layout.
[0046] Figure 1A is a block diagram illustrating an example Electronic Shelf Label (ESL) system according to some embodiments of the present disclosure. ESL system 100 may include a management server 122 integrated with or coupled to gateway node 120. Management server 122 may include at least one processor coupled to memory, wherein the at least one processor is configured to execute computer program code stored on a computer-readable medium to cause management server 122 to perform operations related to managing ESL devices 108A to 108D, APs 106A to 106B, gateway node 120, and / or other components within ESL system 100. For example, management server 122 may perform operations related to converting location based on a wireless communication system into a location description of the asset's layout. For example, management server 122 may perform operations described with reference to Figures 5, 6, and / or Figure 6.
[0047] Gateway node 120 can communicate with access points (APs) 106A and 106B. Although only two APs are shown in the example system, ESL system 100 may include fewer or more APs. APs 106A and 106B can communicate with gateway node 120 via a first communication network (wired or wireless communication network). APs 106A and 106B also communicate with electronic shelf label (ESL) tagging devices via a second communication network. For example, APs 106A and 106B can communicate with pairs of ESL devices in an assigned geographic area. In a first geographic assignment 110A, AP 106A can communicate with ESL devices 108A and 108B; in a second geographic assignment 110B, AP 106B can communicate with ESL devices 108C and 108D. The first and second communication networks can be different networks. In some implementations, the first communication network used for communication between AP 106A and gateway node 120 is a Wi-Fi network, and the second communication network used for communication between AP 106A and ESL device 108A is a Bluetooth network.
[0048] Bluetooth technology provides a secure way to connect and exchange information between electronic devices such as smartphones, other cellular phones, headsets, earphones, smartwatches, laptops, wearable devices, and / or shelf tags. Bluetooth communication may include establishing wireless personal area networks (PANs) (also known as “self-organizing” or “peer-to-peer” networks). These self-organizing networks are often referred to as “piconet”. Each device may belong to multiple piconet. Multiple interconnected piconet can be referred to as a distributed network. A distributed network is formed when members of a first piconet choose to participate in a second piconet. In the example of Figure 1, ESL device 108A may be located in a piconet with AP 106A.
[0049] Because many services provided via Bluetooth may expose private data or allow connected parties to control connected devices, Bluetooth networks may require devices to establish a "trust relationship" before being allowed to communicate private data with each other. This trust relationship can be established using a process called "pairing," in which a bond is formed between the two devices. This bond allows the devices to communicate with each other in the future without further authentication. ESL device 108A can be bonded to AP 106A in this way. The pairing process can be automatically triggered whenever a device is powered on or moves within a certain distance of another Bluetooth device. Pairing information associated with current and previously established pairings can be stored in a Paired Device List (PDL) in the memory of the Bluetooth devices (such as ESL device 108A and / or AP 106A). This pairing information may include a name field, an address field, a link key field, and other similar fields (such as "profile" type) used to authenticate the device or establish a Bluetooth communication link. When, for example, a power outage causes ESL system 100 to reset, the pairing information allows ESL device 108A to automatically reconnect to AP 106A.
[0050] A Bluetooth “profile” describes the general behavior of a Bluetooth-enabled device communicating with other Bluetooth devices. For example, the Hands-free Profile (HFP) describes how a Bluetooth device (such as a smartphone) can make and receive calls for another Bluetooth device, and the Advanced Audio Distribution Profile (A2DP) describes how stereo-quality audio can be streamed from a first Bluetooth device (such as a smartphone) to another Bluetooth device (such as an earphone). ESL devices 108A to 108D can be configured with an Electronic Shelf Label Profile (ESL) conforming to ESL Profile v1.0 dated March 28, 2023 (which is incorporated herein by reference). The ESL profile specifies how AP 106A can use one or more ESL services exposed by ESL device 108A.
[0051] The management server 122 can be implemented as a database (DB) server for storing and managing product information about products displayed in the distribution store. The management server 122 can store various information used during store operation, as well as product information. Furthermore, the management server 122 can write and manage command messages for performing various functions, such as synchronizing, updating, and changing the product information displayed on ESL devices 108A to 108D. The management server 122 can be configured with a database for ESL devices 108A to 108D and the product information displayed on them. That is, the management server 122 can be configured with a database that stores identification information associated with ESL devices 108A to 108D, which is linked to the product information displayed on a corresponding ESL device among ESL devices 108A to 108D.
[0052] Command messages (e.g., product information change messages or management information retrieval messages) created by management server 122 can be delivered to the gateway node using messages encapsulated in packets suitable for a communication scheme used with gateway node 120, and the configured packets can be delivered. Furthermore, management server 122 can receive reception acknowledgment messages transmitted from gateway node 120 via the communication scheme, convert the received messages into messages that management server 122 can receive, and deliver the converted messages. These messages may include notifications and / or instructions for changing ESL tags or notifications for providing a description of location or orientation to user equipment (UE) or other devices, as described with reference to Figures 5 through 7 and / or Figure 9.
[0053] Although only one gateway node 120 is shown in ESL system 100, several such gateway nodes may exist that communicate with management server 122. Each gateway node 120 analyzes the data received from management server 122 to confirm whether there is a message or data to be transmitted to ESL device 108A, and then transmits the confirmed message or data to the corresponding ESL device 108A. Gateway node 120 may configure messages to be transmitted to ESL device 108A into packets according to a communication scheme, and transmit the configured packets to ESL device 108A by commanding AP 106A to send the packets. In addition, gateway node 120 may forward reception acknowledgment messages received from ESL device 108A via AP 106A to management server 122.
[0054] ESL devices 108A to 108D may include multiple ESL devices 108A to 108D that display data about product information received from gateway node 120. ESL devices 108A to 108D that display product information associated with a product may be attached to a shelf. An example layout of the ESL system 100 is shown across multiple display shelves 112A to 112H. Each display shelf 112A to 112H may include one or more shelves to which ESL devices 108A to 108D are attached. ESL devices 108A to 108D may be configured, for example, as shown in FIG. 4, wherein a microcontroller is configured to perform the operations described with reference to FIGS. 5 to 6 and / or FIG. 9.
[0055] In some implementations, a video surveillance system may be included as part of ESL system 100 or used to enhance the capabilities of ESL system 100. For example, shelf cameras 104A to 104D may be positioned to have a field of view capturing one or more shelves of one or more display racks 112A to 112H. Shelf cameras 104A to 104D can be used to assist in tracking inventory levels and / or identifying items picked by users in the environment. As another example, over-the-top (OTT) cameras 102A to 102D may be positioned to have a large field of view capturing the environment of ESL system 100. An object recognition system may be applied to image frames received from cameras 102A to 102D or 104A to 104D to determine the presence or count of objects and people in the respective camera's field of view.
[0056] OTT cameras 102A to 102D can be used to support determining the location of ESL devices 108A to 108D, user mobile devices, or other devices in the environment. Bluetooth Low Energy (BLE)-enabled mobile devices (such as BLE device 124) can traverse the environment and communicate with ESL devices 108A to 108D, for example, to receive identification information from ESL devices 108A to 108D, wherein the location of ESL devices 108A to 108D is determined by identifying the location of BLE device 124 from camera image frames when BLE device 124 receives a signal and / or the strength of the signal received from ESL devices 108A to 108D.
[0057] ESL devices 108A to 108D can change pricing information or be activated or deactivated while communicating with gateway node 120. Store managers can send commands to management server 122 regarding the synchronization of products with ESL devices 108A and / or commands for correcting information about products assigned to ESL devices 108A. Figure 1B shows an example ESL device display for ESL device 108C, where such devices display information including product descriptions, product images, product prices, product barcodes, product ratings, stock keeping units (SKUs), and / or product links (e.g., URLs or QR codes).
[0058] As previously described, the environment may include ESL equipment organized on display shelves and racks. An example illustration of such an arrangement is shown in Figure 2A. Figure 2A is a perspective view of a display shelf with Electronic Shelf Labeling (ESL) equipment according to some embodiments of this disclosure. Display shelf 112A may include multiple racks 202A to 202C at different vertical levels from the floor. ESL equipment may be attached to racks 202A to 202C. For example, ESL equipment 108A may be attached to rack 202A to display information about products stored on rack 202A near ESL equipment 108A.
[0059] ESL devices can provide information to shoppers or store employees operating in the environment, such as providing information about products and / or assisting in product or user location determination. Figure 2B is a top view of a retail environment with user-accessible Electronic Shelf Label (ESL) devices according to some embodiments of this disclosure. A user pushing a shopping cart 212 across an aisle can use an ESL device to determine the location of a specific product. For example, a mobile device associated with the shopping cart 212 can guide the user to location 210 where the desired product is in stock.
[0060] Communication between APs and ESL devices within ESL system 100 can be performed according to a Time Division Multiple Access (TDMA) scheme, such as the one illustrated in Figure 3. Figure 3 is a timing diagram illustrating time division multiplexing for communication with multiple ESL devices according to some embodiments of this disclosure. An AP, such as AP 106A, can broadcast information received by all ESL devices, including ESL device 108A, during a first time period 302. ESL devices can communicate with the AP during subsequent time periods. For example, a first ESL device, such as ESL device 108A, can transmit during time period 304A, while other ESL devices transmit during time periods 304B to 304K. In an ESL system with a large number of ESL devices, the ESL devices can be configured to communicate in different groups. For example, ESL devices 1 to 11 can be configured to transmit to the AP during a first time period, and ESL devices 12 to 22 can be configured to transmit to the AP during a second time period. The first and second time periods can alternate during the operation of the wireless network. For example, after ESL devices 1 to 11 transmit during time periods 304A to 304K, the AP can transmit during the second time period 306, and ESL devices 12 to 22 can transmit during time periods 308A to 308K respectively.
[0061] ESL devices may include components configured together to provide some or all of the functionality described in this disclosure and / or to provide additional functionality. Figure 4 is a block diagram illustrating an example ESL device according to some embodiments of this disclosure. ESL device 108A may include a low-power microcontroller 410. Although the functionality of ESL device 108A may be configured by microcontroller 410 in embodiments of this disclosure, any single processor or combination of processors (e.g., at least one processor) may be used to perform the functionality described in embodiments of this disclosure.
[0062] Microcontroller 410 may include memory 416. Memory 416 may store computer program code that causes microprocessor 414 to perform operations implementing some or all of the functionalities described in the embodiments of this disclosure. Although shown as part of microcontroller 410, memory 416 may be located internally or externally to microcontroller 410. Microcontroller 410 may also include one or more wireless radio components 412. Wireless radio components 412 may include, for example, Bluetooth wireless radio components, including a front end coupled to antenna 408 for transmitting and receiving radio frequency (RF) signals at one or more frequencies in one or more frequency bands. In some embodiments, microcontroller 410 is a system-on-a-chip (SoC), wherein two or more components, including wireless radio component 412, microprocessor 414, and / or memory 416, are included in a single semiconductor package. In some embodiments, two or more components may be included on a single semiconductor die.
[0063] ESL device 108A may include I / O devices such as notification LED 402 and / or electronic display 404. Notification LED 402 may include one or more light-emitting diodes (LEDs), or other light sources configured to flash one or more colors. Notification LED 402 may be triggered to flash at a specific time and / or in a specific color based on a command received from gateway node 120. For example, notification LED 402 may flash to draw a user's attention to a specific location on a shelf. Electronic display 404 may be, for example, an electronic ink (e-Ink) display configured to output product information.
[0064] ESL device 108A may be coupled to battery 406 or other power source to power operations performed by ESL device 108A, such as operating wireless radio component 412, notification LED 402, electronic display 404, memory 416, and / or microprocessor 414. Battery 406 allows ESL device 108A to be placed in locations where a constant power supply is difficult to achieve. Therefore, to enable a single battery charge to provide a long usage period (e.g., lasting longer than several years), ESL device 108A may be configured to reduce power consumption during periods when frequent commands are not expected. For example, ESL device 108A may operate using a wake-up communication scheme. That is, ESL device 108A wakes up at predetermined time intervals to determine if data is waiting to be received. When no data is waiting, power to ESL device 108A is turned off until the next wake-up cycle to reduce power consumption. When data is to be received, ESL device 108A wakes up to perform communication operations.
[0065] Figure 5 is a block diagram illustrating an example wireless communication system 500 that converts location based on a wireless communication system into a location description of an asset's layout, supported by some aspects of this disclosure. As illustrated in the example of Figure 5, the wireless communication system 500 includes a UE 501 and a network entity 520. Although one network entity 520 and one UE 501 are illustrated in Figure 5, in other examples, the wireless communication system 500 may include multiple network entities 520 and / or multiple UEs 501. In some examples, the wireless communication system 500 may implement aspects of the ESL system 100. For example, the wireless communication system 500 may include an ESL network (such as an ESL infrastructure including one or more ESL devices or network entities) and one or more wireless devices interacting with the ESL infrastructure. In such examples, the network entity 520 may include or correspond to an ESL device, such as an ESL controller or an ESL server, and the wireless device may include a UE, such as UE 501. In some such implementations, the wireless communication system 500 may optionally include one or more ESL devices (hereinafter collectively referred to as "ESL devices 591"), one or more ESL tags (hereinafter collectively referred to as "ESL tags 593"), or a combination thereof.
[0066] Network entity 520 may include or correspond to entities in a wireless network, such as base stations, access points, servers, routers, switches, another network entity, or components or combinations thereof. Alternatively, network entity 520 may include or correspond to any of the ESL devices or infrastructure described herein, including gateway node 120, management server 122, ESL AP (e.g., AP 106A or 106B), ESL devices or controllers (e.g., ESL devices 108C or 108D), ESL tagging devices (e.g., IoT tags) as shown in Figure 1A, or ESL device 400 as shown in Figure 4. An ESL controller may include one or more ESL devices and wireless radio components and is referred to as an ESL track controller. Such an ESL controller may wirelessly communicate with and control one or more ESL devices, such as ESL device 591 including a display and configured to output information based on data received from the ESL controller, or ESL tag 593 configured to attach to an asset and provide location information associated with the asset. Alternatively, the ESL controller may be included in or integrated into an ESL AP, ESL hub, server, etc., configured to perform the operations described herein.
[0067] In some such ESL-based implementations, UE 501 (or another type of wireless device) may interact with ESL infrastructure or ESL devices. UE 501 may be part of or separate from the ESL infrastructure. For example, UE 501 may be associated with a worker or robot, or with a customer / shopper. It may interact with network entity 520 (which may be an ESL device), ESL device 591, ESL tag 593, or any other type of ESL device. As an illustrative, non-limiting example, network entity 520 may include or correspond to an ESL controller, ESL AP, server, etc., as described above. ESL device 591 may include or correspond to an ESL device of the same or different type as network entity 520. For example, network entity 520 may communicate with and control ESL device 591, which is coupled to a display associated with various assets or groups of assets. ESL tag 593 (e.g., an IoT tag) may include or correspond to a passive or battery-free radio component that can output signals or beacons based on received RF energy. ESL tag 593 may be coupled to or associated with one or more products or assets of an ESL system. In a particular implementation, ESL device 591 may include an actuator device configured to provide RF power to ESL tag 593 and / or trigger ESL tag 593 to broadcast a beacon for measurement.
[0068] Network entity 520 and UE 501 may be configured to communicate via one or more portions of the electromagnetic spectrum. For example, network entity 520, UE 501, or both may be configured to communicate via one or more portions of the electromagnetic spectrum associated with Bluetooth transmission, Wi-Fi transmission, Zigbee or Z-wave transmission, local area network (LAN) transmission, personal area network (PAN) transmission, or cellular transmission (including sub-6 GHz and 6 GHz).
[0069] Network entity 520 and UE 501 can be configured to communicate via one or more channels or component carriers (CCs), such as representative first channel 581, second channel 582, third channel 583, and fourth channel 584. Although four channels are illustrated, this is for illustrative purposes only, and more or fewer channels may be used. One or more channels may be used to communicate control channel transmission, data channel transmission, and / or sidelink channel transmission between network entity 520 and UE 501.
[0070] Each channel or CC may have a corresponding configuration, such as configuration parameters / settings. This configuration may include bandwidth, bandwidth portion, HARQ process, TCI status, RS, control channel resources, data channel resources, or a combination thereof. Additionally or alternatively, one or more channels or CCs may have or be assigned a cell ID or a bandwidth portion (BWP) ID. The cell ID may include a unique cell ID for the channel or CC, a virtual cell ID, or a specific cell ID for a particular channel or CC among multiple channels or CCs. Additionally or alternatively, one or more channels or CCs may have or be assigned a HARQ ID. Each channel or CC may also have corresponding management functionalities, such as beam management or BWP handover functionality. In some implementations, two or more channels or CCs are quasi-co-located, such that the channels or CCs have the same beam and / or the same symbol.
[0071] In some specific implementations, control information may be communicated via network entity 520 and UE 501. For example, control information may be communicated via Bluetooth, Zigbee or Z-wave, Wi-Fi, MAC-CE, RRC, DCI (downlink control information), UCI (uplink control information), SCI (sidelink control information), other types of transmission, or combinations thereof.
[0072] UE 501 may include various components (e.g., architecture, hardware components) for performing one or more of the functions described herein. These components may include, for example, a processor 502, a memory 504, a transmitter 510, a receiver 512, an encoder 513, a decoder 514, a location determiner 516, and antennas 511a to 511r. Processor 502 may be configured to execute instructions stored in memory 504 to perform the operations described herein. In some specific embodiments, processor 502 includes or corresponds to microcontroller 410 and / or microprocessor 414 of FIG. 4, and memory 504 includes or corresponds to memory 416 of FIG. 4. Memory 504 may also be configured to store location information 506. Location information 506 includes or corresponds to data associated with or corresponding to the location of UE 501. The location indicated by location information 506 may include or correspond to a signaling-based location that may be determined based on communications within wireless communication system 500. For example, location information 506 may include data for determining location (e.g., measurement data, ephemeris, fingerprint data, etc.), data indicating location (e.g., location coordinates or relative positioning data), data indicating a formula or method for calculating location, or a combination thereof. Additionally or alternatively, location information 506 may be information related to a graph of relevant elements derived from survey video 528, where each vertex of the graph corresponds to a feature and the edges between vertices define proximity and orientation information. In such examples, the orientation is relative to the specific video from which features are extracted, and this orientation is coordinated when a graph based on one video is combined with other graphs based on other videos to generate a combined graph that enables the determination of a consistent orientation of the camera with respect to all combined videos (e.g., from which “anchor points” indicating directions such as left, right, up, and down can be extracted). Location information 506 may enable the generation of notifications (e.g., notification information, indications, and / or instructions) or may enable the determination of an estimated location (e.g., such as estimation based on one or more signal measurements). Location information 506 may include the original location and / or a determined or updated location. In some specific implementations of the wireless communication system 500 that support the ESL system, the location information 506 is determined based on measurement information associated with the ESL system information and / or ESL wireless transmission.
[0073] Transmitter 510 is configured to transmit data to one or more other devices, and receiver 512 is configured to receive data from one or more other devices. For example, transmitter 510 may transmit data via a network (such as a wired network, a wireless network, or a combination thereof), and receiver 512 may receive data via that network. For example, UE 501 may be configured to transmit and / or receive data via: direct device-to-device connection, local area network (LAN), wide area network (WAN), modem-to-modem connection, Internet, intranet, extranet, cable transmission system, cellular communication network, any combination of the foregoing, or any other communication network now known or later developed that allows two or more electronic devices to communicate therein. In some implementations, transmitter 510 and receiver 512 may be replaced by transceivers. Additionally or alternatively, transmitter 510 or receiver 512 may include or correspond to one or more components of ESL device 108A as described with reference to FIG. 4.
[0074] Encoder 513 and decoder 514 can be configured to encode and decode data for transmission. Location determiner 516 can be configured to perform location determination and management operations. For example, location determiner 516 can be configured to determine location based on measurement information and signaling. For example, location determiner 516 can be configured to determine one or more locations of UE 501 based on measurements of signals from other network devices, such as network entity 520. Additionally or alternatively, location determiner 516 can be configured to determine one or more locations of UE 501 based on measurements of beacons and / or reference signals or responses to beacons and / or reference signals. Additionally, location determiner 516 can be configured to extract location or location information of other network or ESL devices from other networks or ESL transmissions. The location information 506 determined by location determiner 516 can be used to generate a textual description of the estimated location of UE 501, as further described herein.
[0075] Although the example in Figure 5 shows a single UE (i.e., UE 501), in other implementations, the network may include additional wireless devices that interact with the wireless communication system 500. These other wireless devices may include one or more elements similar to UE 501. In some implementations, UE 501 and the other wireless devices may be different types of UEs. For example, UE 501 may have higher quality or different operational constraints than other UEs, or vice versa. For illustration, one of the UEs (UE 501 or another UE) may have a larger form factor or may be a current-generating device, and therefore have more advanced capabilities and / or reduced battery constraints, higher processing constraints, etc. As another example, one UE may be associated with a human, while another UE may be associated with a robot or autonomous device.
[0076] Network entity 520 includes a processor 522, a memory 524, a transmitter 536, a receiver 538, an encoder 539, a decoder 540, a location mapper 542, an LLM manager 544, and antennas 537a to 537t. The processor 522 may be configured to execute instructions stored in the memory 524 to perform the operations described herein. In some specific implementations, the processor 522 includes or corresponds to a low-power microcontroller 410 and / or a microprocessor 414, and the memory 524 includes or corresponds to the memory 416 of FIG. 4. The memory 524 may be configured to store estimated locations 526, survey videos 528, feature groups 530 (e.g., a group of one or more features), descriptive layouts 532, a location information database 534, or combinations thereof, as further described herein.
[0077] Transmitter 536 is configured to transmit data to one or more other devices, and receiver 538 is configured to receive data from one or more other devices. For example, transmitter 536 may transmit data via a network (such as a wired network, a wireless network, or a combination thereof), and receiver 538 may receive data via that network. For example, UE and / or network entity 520 may be configured to transmit and / or receive data via: direct device-to-device connection, local area network (LAN), wide area network (WAN), modem-to-modem connection, Internet, intranet, extranet, cable transmission system, cellular communication network, any combination of the foregoing, or any other communication network now known or later developed that allows two or more electronic devices to communicate therein. In some embodiments, transmitter 536 and receiver 538 may be replaced by transceivers. Additionally or alternatively, transmitter 536 or receiver 538 may include or correspond to one or more components of ESL device 108A as described with reference to FIG. 4.
[0078] Encoder 539 and decoder 540 may include the same functionality as described with reference to encoder 513 and decoder 514, respectively. Location mapper 542 may be configured to map location information to estimated locations of devices within the network. For example, location mapper 542 may be configured to map location information associated with UE 501 (such as signal measurements of device signaling or ephemeris) to estimated locations of UE 501 (such as coordinates in a Cartesian coordinate system). LLM manager 544 may be configured to manage the operation of an artificial intelligence (AI) large language model (LLM) to generate descriptions for reference description layout 532 that convert location based on wireless communication systems into location descriptions. For example, LLM manager 544 may be configured to generate one or more prompts to be provided as input to the LLM to facilitate communication with and / or operation of the LLM, or both, thereby generating artificially created text output, such as words, monophthongs, sentences, paragraphs, or other documents, based on that input. In some specific implementations, such text output includes descriptions of features, descriptions of locations, descriptions of symbols, or combinations thereof, as further described herein.
[0079] As further described below, the wireless communication system 500 enables the use of survey video 528 to convert location based on the wireless communication system into a location description of the asset's layout. Survey video 528 may include multiple video frames (e.g., images) captured by the video capture device as a user navigates the location. For example, in a retail environment, store employees may use a mobile phone (e.g., UE) to move around the store to capture video of products on shelves in aisles. Similarly, robots may traverse warehouse aisles, recording video of items stored on various shelves, in storage areas, or other storage areas. Additionally or alternatively, the video capture device may be a fixed camera, such as one or more closed-circuit television (CCTV) or security cameras, or video captured by a mobile device may be supplemented by video captured by a fixed camera or other type of video capture device. As used herein, an asset refers to any product, item, or other element that is stored, displayed for sale, included for use, or otherwise positioned such that the asset's location relative to other assets is useful information enabling customers, employees, etc., to move efficiently to the asset. Each frame or subset of frames in the investigation video 528 is associated with location data, such as the Cartesian coordinates (e.g., x, y, z coordinates) of the corresponding video frame captured at that location, and optionally with the timestamp of the corresponding video frame being captured.
[0080] The measurement system problem can be defined as follows: for any t in the survey interval (time), define a function that maps time to location. The purpose of trajectory coordination is to define such a function (or multiple functions). Therefore, for a given trajectory... Any element T in a subset of the defined R:
[0081] Uncontrolled (e.g., arbitrary) video of a use case, such as survey video 528, can be used to estimate the structure and attitude (e.g., orientation and position) of the camera capturing the video. Video capture can be performed concurrently with capturing sensor data or performing radio-based positioning. Therefore, the device capturing the video or another device can collect observations from any type of signal or sensor, which can later be used in the location determination function. Determination of a function that coordinates time and position and is derived solely from the video allows for a mapping between observations and position (e.g., via timestamps or other time measurements). In this way, an ephemeris can be generated, which is auxiliary information used in determining or estimating the position. Furthermore, processing including <collected data, position> can be used to determine or estimate the position of a signal source (e.g., a WiFi access point (AP), ESL AP, etc.) or to attribute a signal's "fingerprint" (e.g., an RF signal or a magnetometer signal) to a specific location. The inferred information described above can then be collected in the ephemeris, which can subsequently be used to assist in the indoor positioning of entities such as customer-held mobile devices, robots performing warehouse inventory, etc. In such examples, indoor positioning (e.g., Locations are typically represented in a coordinate system consistent with the survey coordinate system, such as the coordinates of a store or other location. However, mapping calculated or estimated indoor locations from the survey coordinate system to the retail store coordinate system or other site coordinate system can be challenging and often requires a detailed and up-to-date floor plan that also includes asset location information. To address this, the wireless communication system 500 uses multimodal artificial intelligence (AI) to automatically extract semantic information from the survey video 528 to generate a description of the location in the retail store coordinate system (e.g., using "retail semantics") or other site coordinate system. This semantic information can provide navigational cues, such as "Tomato sauce is in aisle 4, to the left of shelf 4." If the user's location is estimated from the positioning system, this type of waypoint guidance can be provided to the user to navigate them to the product, such as providing the following in response to a query for tomato sauce: "Tomato sauce is in this aisle on the left, behind the pasta sauce." If the user's location information is unknown, navigation can be provided by using an image of the user's current location (such as an image captured via a mobile device or XR / VR glasses) to correlate the user's position relative to the semantic map.
[0082] During the operation of the wireless communication system 500, entities (e.g., customers, inventory handlers, warehouse employees, etc.) may carry UE 501 (e.g., mobile phones) in locations such as retail locations or warehouses. Network entity 520 may obtain location information 506 associated with UE 501 based on one or more location operations. As an example, network entity 520 (or one of ESL device 591 or ESL tag 593) may send a beacon, and ESL device 591, ESL tag 593, and UE 501 may send one or more beacon responses. For example, network entity 520 may be an ESL controller associated with multiple ESLs (e.g., ESL device 591), and each ESL device 591 may be associated with one or more ESL tags in ESL tag 593, which correspond to one or more assets and send a beacon response when network entity 520 sends a beacon. For example, network entity 520 may be an edge or cloud server associated with multiple ESL access points (APs) (e.g., ESL device 591), and each ESL device in ESL device 591 may send a beacon response when network entity 520 sends a beacon. Alternatively, network entity 520 may provide one or more reference signals to devices of wireless communication system 500 and / or may communicate with one or more other network devices. As a non-limiting example, network entity 520 may generate one or more measurements based on beacon responses or other signaling such as Received Signal Strength Indicator (RSSI), Reference Signal Received Power (RSRP), Angle of Arrival (AoA), or combinations thereof. Since network entity 520 may know the location of at least some of the ESL devices in ESL device 591, signal measurements based on their beacon responses and beacon responses from UE 501 can be used to determine location information 506. Additionally or alternatively, network entity 520 may perform one or more measurements to generate ephemeris data, which may be mapped to location data based on survey video 528. Alternatively, UE 501 may perform a positioning operation and send location information 506 to network entity 520. For example, UE 501 may include a location determiner 516 configured to perform positioning based on a wireless communication system, such as based on signal measurements of beacons and / or beacon responses, similar to that described with reference to network entity 520. In some specific implementations, location information 506 may include or be based on RF fingerprints (e.g., based on signal measurements described above) and sensor data from network entity 520 or UE 501. As a non-limiting example, UE 501 may include a GPS sensor or a GNSS sensor (or other types of sensors), and UE 501 may provide positioning coordinates determined based on the sensor data to network entity 520 as location information 506.
[0083] After obtaining location information 506, location mapper 542 can map location information 506 to estimated location 526 to generate an estimated location of UE 501 within that location. For example, location mapper 542 can map the relative location of UE 501 to ESL device 591 and / or ESL tag 593 (or other devices), or the location given by RF fingerprint or signal measurement, to estimated location coordinates (e.g., estimated location 526), such as the coordinate system of a reference location. If location information 506 includes other sensor data, location mapper 542 can perform other types of mapping. In some implementations, location mapper 542 also manages and controls positioning operations performed by network entity 520, such as beacon transmission, beacon response measurement, and determination of RSSI, RSRP, AOA, etc.
[0084] After determining the estimated location 526 of UE 501, network entity 520 may retrieve one or more frames of survey video 528 corresponding to the estimated location 526. For example, each frame of survey video 528 may be tagged or associated with the location of the captured video, and network entity 520 may retrieve one or more frames associated with the location coordinates (or other form of location data) of the matching estimated location 526. As described above, survey video 528 may be captured by an entity (such as a store employee) moving around the location where the video was captured, and the frames of survey video 528 may be tagged using location information generated by the video capture device at the location of the captured frame and / or at the location and orientation of the video capture device estimated using SFM on survey video 528. In some implementations, each frame of survey video 528 may also be associated with a timestamp, and network entity 520 may retrieve frames associated with the location of the matching estimated location 526 and one or more frames that appear before or after the matching frame based on the timestamp. For example, network entity 520 may be configured to retrieve a specific number of frames, and if not enough matching frames are identified, the remaining frames may be supplemented with frames from before or after the time of the matching frame. As another example, network entity 520 may be configured to extract a specific number of features for pattern matching, as further described below, and if the matching frame does not provide a sufficient number of features, additional frames from before or after the matching frame may be provided for extracting additional features.
[0085] After selecting one or more frames from the survey video 528, network entity 520 may extract feature set 530 from the selected frames. These features can be used to identify or associate with one or more elements in the selected frame (e.g., an image), such as products, logos, labels, colors, patterns, other elements, or combinations thereof. For example, feature set 530 may include products identified in the selected frame, text of logos or labels in the selected frame, directional features of one or more detected objects in the selected frame, geometric description of the selected frame, spatial description of the selected frame, textual description of the selected frame, other features, or combinations thereof. To extract feature set 530, network entity 520 may perform object detection on the selected frames to detect one or more objects within the selected frames. Object detection and object tracking image processing techniques may be used to detect these objects, and features representing the detected objects may be extracted as at least a portion of feature set 530. As a non-limiting example, network entity 520 may perform object detection on a selected image and detect a can of tomatoes, a box of noodles, a label on the box of noodles, and a logo above the can of tomatoes and the box of noodles. The extracted features may include text from the sign, text from the label, features indicating tomatoes, features indicating noodles, shape features associated with the can or box, spatial or directional features (e.g., the sign is above the tomato can and to the left and above the noodle box), and color features (e.g., the can is green and the box is red). Additional types of image processing operations, such as object tracking, text recognition, thresholding, and others, can also be performed. These features can be used to perform pattern matching or to generate descriptions for pattern matching, as further described herein.
[0086] After extracting feature set 530, network entity 520 can identify one or more descriptions from description layout 532 based on feature set 530. Description layout 532 may include asset names, asset descriptions, and asset locations cataloged and associated based on the arrangement of assets in a location. In other words, description layout 532 includes text representing visual descriptions of assets and their locations within a location. Therefore, description layout 532 may include or correspond to a store shelf diagram or a database of products and their corresponding locations. Description layout 532 may include descriptions of product appearance, product names, spatial relationships between at least some of the products, locations in "retail semantics" (or other contextual semantics), such as aisle numbers, display shelf numbers, and shelf numbers, other descriptions, or combinations thereof. As a non-limiting example, description layout 532 may include product entries, and each entry may include a product name, product container (e.g., box, jar, bag, etc.), product color, aisle number of the aisle containing the product, display shelf number of the display shelf containing the product, and shelf number of the shelf containing the product. This example is illustrative, and in other specific implementations, fewer, more, or different elements may be included in the description layout 532. The description layout 532 can be created and maintained by the entity that owns the location (such as an owner or manager), and can be changed based on changes in inventory, layout, and other circumstances. Because the description layout 532 does not include precise location information (e.g., coordinates of the location of assets within the location), maintaining and updating the description layout 532 can be easier and more efficient than maintaining a map of the location that includes asset location information.
[0087] To identify one or more descriptions representing assets located at the same location as UE 501, network entity 520 may perform pattern matching between feature group 530 and description layout 532 to identify one or more descriptions that most closely match feature group 530. For example, network entity 520 may compare feature group 530 or its textual description with asset names, asset locations, asset containers, shapes, or colors included in description layout 532 to identify descriptions (or descriptions derived from it) of assets most similar to feature group 530. In some implementations, at least a portion of the description in description layout 532 is identical to one of the features in feature group 530 to detect a match. Alternatively, features (or feature-based descriptions) may be compared with one or more descriptions in description layout 532 to generate corresponding similarity scores (or other metrics), and a match may be detected for feature-description pairs whose similarity scores meet a threshold. The threshold may be set based on matching accuracy and competition considerations to prevent false negatives.
[0088] Network entity 520 can utilize artificial intelligence (AI) models and logic to assist in the process of identifying descriptions in description layout 532 that match feature group 530. In some specific implementations, LLM manager 544 is configured to generate prompts based on feature group 530 and use such prompts as input data for AI large language models (LLMs) (such as ChatGPT, BERT, etc.) to generate text descriptions of feature group 530. These text descriptions can be used in pattern matching to identify matching descriptions from description layout 532. As a non-limiting example, if feature group 530 includes features indicating the shape of a can, a red label including the word "tomato," and directional features indicating the area below the blue box, then LLM manager 544 can combine these features with a prompt template or otherwise generate a prompt that causes the AI LLM output the description "Red tomato cans are located on the shelf below the blue box." In this example, a textual description is more likely to match a description from description layout 532 than a single feature itself. This description includes phrases like "Canned tomatoes are red cans labeled as tomatoes located below the boxes of pasta in aisle 4, display shelf 2, shelf 3." An LLM can be included or integrated into network entity 520 and managed by an LLM manager 544, or the LLM can be maintained at an external network connection, such as a corporate server or cloud server, and accessed by the LLM manager 544. Using an LLM in this way can generate textual descriptions of feature group 530, or more user-friendly and semantically relevant textual descriptions, for matching with description layout 532.
[0089] Additionally or alternatively, the LLM manager 544 may use an AI LLM to perform the pattern matching described above. For illustration, the LLM may be trained using information included in the description layout 532, enabling the LLM manager 544 to pose a question to the LLM, and the LLM to provide a description most similar to that question. In such a specific implementation, the LLM manager 544 may generate prompts based on feature set 530, such as by combining some of these features with a question or other prompt template, or otherwise generating a prompt to be provided to the LLM. Based on the received prompts as input data, the LLM may output text output including one or more descriptions (e.g., questions based on feature set 530) that best match (e.g., are most similar) to the information in the prompt. For example, the LLM manager 544 may generate a prompt including "Where is the canned tomato with the red label under the blue box?", and the LLM may output "The canned tomatoes are the red canned tomato labeled under the box of pasta in aisle 4, display shelf 2, shelf 3."
[0090] Although described as network entity 520 generating the description and / or questions of the LLM, in some other embodiments, UE 501 may include or access the LLM and perform at least some of the operations described above. In some such embodiments, network entity 520 may send feature set 530 to UE 501, and UE 501 may generate a description of feature set 530 by generating a prompt based on feature set 530 and providing that prompt as input to the LLM. Additionally or alternatively, UE 501 may generate questions for the LLM trained based on description layout 532. In these embodiments, the security of the description layout and the survey video is preserved by storing the description layout 532 and the survey video 528 at or accessible by network entity 520 and having network entity 520 provide feature set 530 to UE 501, while offloading AI interaction to UE 501. After the LLM generates the description or questions, UE 501 may send the description or questions to network entity 520 for performing pattern matching on description layout 532. Configuring UE 501 to interact with the LLM instead of network entity 520 allows less complex and computationally intensive devices to operate as network entity 520. Alternatively, if network entity 520 is a server or other device capable of performing such operations, assigning such operations to network entity 520 allows less complex UEs to participate in the location switching services described herein.
[0091] After identifying one or more descriptions from description layout 532 that match or are most similar to feature group 530, network entity 520 may output description 550 of the estimated location 526 of description layout 532. For example, network entity 520 may send description 550 to UE 501 for display to a user via touchscreen, audio output, or other forms of output. In some embodiments, description 550 may be relative to one or more assets, such as indicating the relative location of nearby assets. Additionally or alternatively, description 550 may be “retail semantics” and may therefore include aisle labels, shelf labels, rack labels, or combinations thereof. In other embodiments, description 550 may be semantics related to other contexts, such as warehouses, hangars, or other storage locations. In some embodiments, LLM manager 544 may generate a prompt based on the identified description, and LLM manager 544 may provide the prompt as input to LLM to generate description 550. Using LLM to generate location descriptions can result in more user-friendly, grammatically correct, and / or semantically relevant location descriptions than those captured directly from description layout 532. This article further describes additional details of a pipeline configured to perform at least some of the operations described above, with reference to Figure 6.
[0092] Now, referring to Figure 8, an illustrative example of the process described above for converting location based on a wireless communication system into a location description is described. Figure 8 is a perspective view of a portion of a retail location 800 having assets according to one or more aspects of this disclosure, for which location based on a wireless communication system is converted into a location description of the asset's layout. The perspective view shown in Figure 8 may correspond to one or more frames of survey video 528. In this example, the retail location 800 includes multiple display shelves 810A to 810B, each display shelf including multiple racks 812A to 812C and 812D to 812F located at different vertical levels above the ground. ESL devices may be attached to the shelves to display information about products stocked on various racks near the ESL devices. Various products are positioned on the racks of display shelves 810A to 810B, which in this example are located in the same aisle, such as aisle 8. For example, for display shelf 810A, peaches and pears are positioned on shelf 812A, apples are positioned on shelf 812B, and mixed fruits are positioned on shelf 812C. In this example, for display shelf 810B, cakes are positioned on shelf 812D, candies are positioned on shelf 812E, and muffins are positioned on shelf 812F.
[0093] As described above, network entity 520 can obtain location information 506 indicating the location of UE 501, and location information 506 can be used to determine estimated location 526. In this example, video frames displaying the view shown in FIG8 are associated with location data that are the same as (or close to before or after) the estimated location 526 in time. Based on the matching between the location of estimated location 526 and frames of survey video 528, network entity 520 can perform text recognition on the frame to detect text of product labels, such as text 802A (“peach”), text 802B (“pear”), text 802C (“mixed”), and text 802D (“cake”). Network entity 520 can also perform object detection on the frame to detect one or more objects, such as object 804A (apple), object 804B (candy), and object 806C (mug). Network entity 520 may extract feature group 530 from the frame, and feature group 530 may include text 802A to 802D, the shape of objects 804A to 804C, the geometry of objects 804A to 804C, the size of text 802A to 802D and objects 804A to 804C, the color of text 802A to 802D and objects 804A to 804C, spatial or directional features (e.g., object 804B above text 802C, text 802A to the left of text 802B, object 804B relatively close to object 804C, etc.), other features, or combinations thereof.
[0094] After extracting feature group 530, network entity 520 can use LLM manager 544 to generate and provide prompts to LLM to generate a description of the extracted features, which can then be compared with the description in description layout 532. For example, LLM can generate text output associated with peaches, which includes the text "A box of peaches is located to the left of a box of pears, and the box of peaches is located above a jar of applesauce." Based on this description, network entity 520 can identify the description "Peaches are on shelf 1 in aisle 8 (fruit and dessert aisle), on shelf 2, to the left of the pears" in description layout 532. This description can be a description stored in description layout 532, or it can be the output of LLM after providing prompts based on the description from description layout 532 as input, and this description can be provided to UE 501 as description 550. As another example, LLM can generate text output associated with muffins, which includes the text "Multiple rectangles with muffin images are located below a jar with candy images." Based on this description, network entity 520 can identify the description "Shelf 5 (Fruit and Dessert Aisle) of display shelf 2 in aisle 8 contains muffin and muffin powder" in description layout 532 (or use this description as the basis for input prompts in LLM to generate output). This description can be provided to UE 501 for display to the user. In this way, the location based on the wireless communication system can be converted into a location description referring to description layout 532, which can be more useful for users navigating in a store to search for one or more products.
[0095] Referring back to Figure 5, in some implementations, additional location determination can be performed using information from signs and locations in video frames. In such implementations, features associated with the sign can be extracted, and the description of the estimated location can be associated with the sign. For example, if the identified frames of investigation video 528 include signs, feature group 530 can include features based on the sign, such as the sign's text, the sign's description, the sign's shape, the sign's geometry, the sign's color, the text color, the sign's orientation or spatial characteristics (e.g., relative to other signs or assets), other sign-based features, or combinations thereof. Additionally or alternatively, description layout 532 can include fields indicating one or more entries for the nearest sign, which could be a label, aisle sign, shelf sign, or another type of sign displayed by ESL device 591 or ESL marker 593. Additionally, description 550 can refer to signs within the location. Similar to what has been described above, LLM manager 544 can use sign-based features, sign-related descriptions, or questions including sign-based information to generate prompts to be provided to the LLM to generate a description of the estimated location of the sign. In some specific implementations, before performing pattern matching on the description of description layout 532, network entity 520 may filter at least one description from description layout 532 based on the flag-based features. For example, if a flag for corridor 6 is detected within one or more frames and the corresponding features are extracted, network entity 520 may restrict pattern matching to descriptions in description layout 532 corresponding only to corridor 6. Additional details of the pipeline for at least some of the operations described above, configured to perform reference flag operations, are further described herein with reference to FIG7.
[0096] In some implementations, network entity 520 may be configured to determine location solely based on signs at various locations. In such implementations, network entity 520 may obtain an image or video frame (rather than location information 506) of the current location of UE 501 and perform object detection on the image or video frame to detect signs within the image or video frame. In such implementations, network entity 520 may extract a set of sign-based features from a portion of the image that includes the sign, such as the text of the sign, the color of the sign, the shape of the sign, etc., and network entity 520 may perform pattern matching between the sign-based features and a description of the sign layout. In such examples, network entity 520 may output a description of the current location of UE 501 regarding the signs described in the sign layout. For example, in such examples, description 550 may include "You are located in aisle 2, where fresh produce is located between aisle 1 (delicious food) and aisle 3 (beverages)." Additional details of the pipeline of at least some of the operations described above, configured to perform reference signs, are further described herein with reference to FIG7.
[0097] In some implementations, network entity 520 may use the ability to convert between location based on the wireless communication system and a textual description of the location of layout 532 to provide additional or advanced features, such as reverse lookup or description direction. In some implementations supporting reverse lookup functionality, network entity 520 is configured to maintain a location information database 534 to store information related to location data (e.g., estimated / calculated location coordinates) and a textual description of layout 532. In such implementations, after outputting description 550, network entity 520 may store description 550 and estimated location 526 in location information database 534 for later use. If a user of UE 501 or a user of wireless communication system 500 wishes to know the location coordinates of a selected product, network entity 520 may receive a reverse lookup request including a description of the location of the requested product. In response to receiving a reverse lookup request, network entity 520 may access location information database 534 to identify location data associated with the description of the requested location. The accessed location information can be output as location data by network entity 520 to the user of UE 501 or by the UE to be displayed to the user of UE 501.
[0098] In some specific implementations supporting directional functionality, network entity 520 is configured to retrieve a query indicating a description of a requested product for performing pattern matching with description layout 532, and the matching description from description layout 532 is used to generate the direction between UE 501 and the requested product. For example, network entity 520 may receive query 552 from UE 501. Query 552 may include or indicate the requested product (e.g., an asset). Network entity 520 may identify the matching description in description layout 532 by performing pattern matching based on query 552, and may compare description 550 and the matching description (e.g., to query 552) to determine a description 554 indicating the direction from estimated location 526 to the location of the requested item. For example, if description 550 indicates aisle 8, and the matching description of query 552 indicates that the requested product is located in aisle 4, on the left side of shelf 3 of display shelf 3, then description 554 may include the text “Walk down the four aisles to the aisle, and the item you requested is located in the middle of the right aisle, on the third shelf of display shelf 3.”
[0099] As described above with reference to Figure 5, the wireless communication system 500 enables the conversion of wireless communication-based positioning into a more asset-relative positioning, such as a description of the location relative to the descriptive layout 532. For example, network entity 520 can generate a description 550 based on location information 506, which provides a description of the current location of UE 501 in a way that is more meaningful than providing location coordinates within a store or other location, without referencing the asset and its location. Thus, the wireless communication system 500 utilizes wireless communication-based positioning operations (e.g., location information 506) and information for tracking assets (such as descriptive layout 532 (e.g., shelf diagrams or product information databases)) to generate a location description (e.g., description 550) that adds utility to the user, without requiring additional positioning operations using different coordinate systems or other reference frames. In other words, this positioning conversion utilizes a coordinate system derived from the survey video 528 (e.g., using the SFM method) and maps each frame (and each feature in that frame by transitivity) to a coordinate system associated with the asset. Features can be extracted and categorized by associating them with unique descriptors such as aisle, shelf, and shelf locations, and / or their relative locations to other objects (e.g., "to the left of the pasta sauce"). This allows locations to be expressed in terms associated with the extracted features (e.g., the product name at the location, the area name near the location, etc.), rather than referencing location coordinates. Such location descriptions are likely more useful to users than regular location coordinates and reduce the time spent orienting themselves about assets stored at locations such as stores or warehouses. Additionally, converting location information 506 into more descriptive positioning (e.g., description 550) provides information to support reverse lookups or searches, i.e., from features to coordinates. The conversion between positioning systems also enables the generation of organization-based, asset-based, or sign-based directions (e.g., description 554), which are more useful to users than coordinate-based directions. For example, a shopper in a store can query directions to bread, and network entity 520 can provide description 554, which includes the direction of the queried product in retail semantics (such as specific aisle and shelf locations) and the distance in the aisle from the current location. Compared to knowing an asset is about a store but not its location coordinates, this type of orientation reduces the time users spend searching for the asset. Similar descriptions can be provided based on signage within the location, offering users simple directions that can be easily verified by checking nearby signs or information displayed on ESL device 591 or ESL marker 593. The functionality described above can be particularly beneficial in ESL applications because it allows for the transformation of ESL-based positioning into a more descriptive positioning framework that utilizes the ESL system.
[0100] Figure 6 is a block diagram illustrating an example location estimation and description pipeline 600 according to one or more aspects of this disclosure. The location estimation and description pipeline 600 may be included or integrated into an ESL device (such as ESL device 108A of Figures 1A, 2B, and 4) or a network entity (such as network entity 520 of Figure 5). In the example shown in Figure 6, the location estimation and description pipeline 600 includes a location engine 610, a mapping engine 612 coupled to the location engine 610, a feature extractor 614 coupled to the mapping engine 612, and a pattern matcher 616 coupled to the feature extractor 614. Additionally, the location estimation and description pipeline 600 can access video surveys 618 and description layouts 620 of assets, similar to the survey video 528 and description layout 532 of Figure 5. In the aspects described herein, the location survey process is video-based, thereby generating ephemeris data that provides computed locations (e.g., camera poses) associated with video frames of the video survey 618. Additionally, similar to that described above, the description layout 620 provides a reference product organization. For example, the description layout 620 may include a shelf diagram or database that provides information about product organization using typical retail semantics such as aisles, display shelves, units, and racks.
[0101] During the operation of the location estimation and description pipeline 600, the location engine 610 receives location information and generates an estimated location based on that information. This location information may include, for example, an RF fingerprint (e.g., signal measurement), sensor readings, and other forms of wireless communication system-based or external system-based positioning, and the estimated location includes location coordinates, such as x, y, z coordinates in a coordinate system used to perform a video survey (e.g., a shop, warehouse, etc.). Such a coordinate system may also be referred to as a survey coordinate system. The mapping engine 612 may receive the estimated location and extract one or more frames from the video survey 618 based on that estimated location. For example, the mapping engine 612 may use the surveyed ephemeris (e.g., video survey 618) to extract an image associated with the current (or within a time window) location estimated by the location engine 610. For instance, the extracted frames may be associated with an RF fingerprint (or other wireless communication system-based positioning data) that matches the estimated location output by the location engine 610.
[0102] Feature extractor 614 can extract relevant features from ephemeris (e.g., the extracted frame) and reference product organization (e.g., description layout 620). In other words, feature extractor 614 can automatically derive vocabulary for products or other assets from the extracted video frames and can detect / identify objects (e.g., objects described in description layout 620) from the extracted frames and from the shelf diagram. In some implementations, the extracted features may be equipped with associated descriptors, such as geometric descriptions, like shape, spatial descriptions, text descriptions, size descriptions, color descriptions, etc. Additionally or alternatively, feature extractor 614 can extract features similar to those included in description layout 620 to achieve more accurate pattern matching, as shown in Figure 6 where feature extractor 614 receives asset descriptions from description layout 620. These asset descriptions can be used to identify the features used for extraction or to generate descriptions of the extracted features with vocabulary similar to the descriptions in description layout 620. Pattern matcher 616 can match the received features (and descriptions) with descriptions in description layout 620 to identify one or more descriptions to be output by location estimation and description pipeline 600 as location descriptions (e.g., descriptions of the estimated location of the UE). For illustration, pattern matcher 616 can identify and match patterns between extracted features (and descriptions) and reference product organization descriptions (such as assets, spaces, scaling invariants, and relative descriptions). Therefore, pattern matcher 616 can perform pattern matching between objects in a retail store image (e.g., based on extracted features) and objects in a shelf diagram (e.g., descriptions in description layout 620). This pattern matching may be at different levels. For example, depending on the complexity of description layout 620 and the processing capabilities of network entities, pattern matching may be performed at the unit level, at the shelf level, at the aisle level, etc. The description that matches or is most similar to the extracted features can be output as a location description, which can be retail semantic: as a non-limiting example, aisle-display shelf-unit-shelf-product. In some specific implementations, as described above with reference to Figure 5, a custom machine learning (ML) model or LLM can be used for feature extraction and description and / or for pattern matching with the description layout 620.
[0103] Figure 7 is a block diagram illustrating an example location estimation and description pipeline 700 according to one or more aspects of the present disclosure. The location estimation and description pipeline 700 may be included or integrated into an ESL device (such as ESL device 108A of Figures 1A, 2B, and 4) or a network entity (such as network entity 520 of Figure 5). In the example shown in Figure 7, the location estimation and description pipeline 700 includes a location engine 710, a mapping engine 712, a feature extractor 714, a pattern matcher 716, a video survey 718, and a description layout 720, each of which is similar to the location engine 610, mapping engine 612, feature extractor 614, pattern matcher 616, video survey 618, and description layout 620, respectively. The location estimation and description pipeline 700 also includes a flag extractor 722 coupled to the mapping engine 712, and a switch 724 coupled to the pattern matcher 716 and the flag extractor 722.
[0104] Retail stores typically include signs displaying information about which products are located in aisles, on shelves, etc., as well as signs indicating item categories such as milk, baked goods, produce, desserts, beverages, and chips. Compared to the ephemeris described with reference to Figure 6, the ephemeris generated via video survey 718 can be expanded by overlaying sign information (e.g., sign identification and product association) onto each surveyed location (e.g., each frame of video survey 718). To utilize this sign information, sign extractor 722 can scan, identify, and associate signs captured by a video capture device (e.g., a camera) during the survey and included in the extracted video frames, similar to extracting features related to other detected objects. These sign-based features may include the text of one or more signs in the one or more frames, the description of the one or more signs, or a combination thereof. In some implementations, sign extractor 722 may utilize a custom ML model or LLM to generate sign descriptions or otherwise perform sign extraction. The extracted signs (and corresponding descriptions) may be provided to pattern matcher 716 for performing pattern matching with the description layout 720, and the extracted signs may be provided to switch 724. Therefore, switch 724 provides two possibilities: first, switch 724 can output a description that matches the extracted sign and the extracted sign, similar to the description described above with reference to Figure 6; or second, switch 724 can output the extracted sign (and its corresponding description). Switch 724 can operate based on user input (e.g., the user's choice of asset and sign-based description or sign-specific description), based on one or more pre-programmed settings or parameters, or otherwise be configured to select between outputting a description taken for both the asset and the sign, or a description taken only for the sign. If the first option is selected, the sign information can be used to reduce the search space when performing pattern matching, thereby making the scheme optimal and robust. For example, if one or more descriptions in the description layout 720 are associated with a sign that does not match the extracted sign, those one or more descriptions can be filtered out. In this case, the estimated location is expressed in the retail semantics of both the product and the sign, such as aisle number, shelf number, unit number, shelf number, product type, or a combination thereof. If the second option is selected, the estimated location (e.g., x, y, z coordinates) is converted to a location relative to the sign representation. This eliminates the reliance on reference product organization (e.g., describing layout 720), but the location granularity depends on the density and granularity of signage at that location. For example, if only aisle and product type are indicated in the signage, such a description might only be aisle number and product type, rather than shelf number, rack number, or the orientation of other assets.
[0105] Referring again to Figure 8, an illustrative example is described. In the example shown in Figure 8, the video frames include signs 806A to 806B corresponding to display shelves 810A to 810B, respectively. Sign 806A indicates "canned fruit," and sign 806B indicates "dessert." Although signs associated with display shelves are shown in Figure 8, in other examples, signs may provide information related to aisles, shelves, or other units. Sign extractor 722 can identify signs 806A and 806B, similar to what has been described above for detected text 802A to 802D or objects 804A to 804C. The description of the estimated location may also include sign information or product type information. For example, based on pattern matching using the extracted features and the extracted signs, pattern matcher 716 can identify in description layout 720 the description "Peaches are on display shelf 1 in aisle 8 (fruit and dessert aisle), on shelf 2, to the left of pears." As another example, based on pattern matching using extracted features and extracted tags, pattern matcher 716 can identify in layout description 720 that "shelf 5 of display rack 2 in aisle 8 (fruit and dessert aisle) contains muffins and muffin powder, and other desserts are located below the shelf." These descriptions can provide more detailed information to users navigating the store to search for one or more products, thereby reducing the time they spend searching for products.
[0106] If the second option of switch 724 is selected, the location estimation and description pipeline 700 can output a description of the location based solely on signs (rather than products). For example, sign extractor 722 can receive image or video frames captured at the UE's current location as input, instead of extracting video frames from video survey 718. Sign extractor 722 can extract signs included in the image or frame to provide a direct estimate of the UE's location in retail semantics through sign detection / identification. However, such location granularity depends on the hierarchy and density of signs in the retail environment. For example, a location can be described by aisle signs, shelf signs, rack signs, individual item signs, product type signs, etc., depending on the granularity of the signs placed around the location. Such an implementation can eliminate the need to conduct surveys if the entity managing the location is willing to allow images or videos to be captured at that location and shared between devices.
[0107] Figure 9 is a flowchart illustrating an example process 900 that supports the conversion of location based on a wireless communication system into a location description of an asset's layout according to some embodiments of this disclosure. The operation of process 900 can be performed by a network entity (or a component thereof) as described herein or by an ESL device (or a component thereof). For example, process 900 can be performed by network entity 520 as described above with reference to Figure 5 or by network entity 1000 as described with reference to Figure 10.
[0108] At box 902, the network entity obtains location information associated with the UE. For example, this location information may include or correspond to location information 506 in Figure 5. At box 904, the network entity captures one or more frames of survey video of a location corresponding to the estimated location of the UE. The estimated location is based on the location information. For example, the estimated location may include or correspond to estimated location 526 in Figure 5, and the survey video may include or correspond to survey video 528 in Figure 5.
[0109] At box 906, the network entity extracts a set of features from the one or more frames. For example, this set of features may include or correspond to feature group 530 of FIG. 5. At box 908, the network entity outputs a description of the estimated location of the descriptive layout based on one or more descriptions identified from the descriptive layout identifier of the asset associated with the location. The one or more descriptions are identified based on the set of features. For example, the descriptive layout may include or correspond to descriptive layout 532 of FIG. 5, and the description of the estimated location may include or correspond to description 550 of FIG. 5. In some embodiments, the description of the estimated location is relative to one or more assets among the assets. Additionally or alternatively, the description of the estimated location may include aisle labels, shelf labels, shelf labels, or combinations thereof. Additionally or alternatively, the descriptive layout may include the name, description, and location of the product in retail semantics. Additionally or alternatively, the descriptive layout may represent a store shelf layout and / or include a description of the spatial relationships between at least some of the products among the products.
[0110] In some implementations, process 900 further includes mapping the location information to the estimated location of the UE within the location. For example, this mapping may be performed by the location mapper 542 of Figure 5. Additionally or alternatively, obtaining the location information may include: calculating one or more of the RSSI, RSRP, or AoA of one or more beacon responses from one or more assets in the assets and the beacon responses from the UE; and determining the location information based on the one or more beacon responses and the RSSI, RSRP, or AoA of the beacon responses. In some such implementations, the one or more assets in the assets may correspond to one or more radio frequency identification (RFID) tags for one or more products capable of transmitting beacon response messages within the ESL system. For example, the one or more RFID tags may include or correspond to the ESL tag 593 of Figure 5.
[0111] In some implementations, the survey video comprises multiple frames captured by a video capture device navigating various locations within the site. In some such implementations, each of these multiple frames is associated with the location coordinates, ephemeris data, or both of which were used to capture the corresponding frame, and the location information associated with the UE includes the location coordinates associated with the UE, the ephemeris data associated with the UE, or both. For example, a store employee may walk around the store capturing video of assets on shelves, or a CCTV camera may capture video of the store, and each frame of the video is labeled with the location coordinates of the video capture device as it is captured. Additionally or alternatively, an ephemeris associated with a specific time may be generated, and this ephemeris can be mapped to a location by mapping its time to the time of the captured video.
[0112] In some implementations, the device is an ESL controller associated with multiple ESLs, and each of these ESLs is associated with one or more tags corresponding to one or more assets in the asset set. For example, the device may correspond to network entity 520 of Figure 5 (e.g., an ESL controller) associated with ESL tag 593 of Figure 5. Alternatively, the device may be an edge or cloud server associated with multiple ESL access points (APs). For example, the device may correspond to network entity 520 of Figure 5 (e.g., an edge or cloud server) associated with ESL device 591 of Figure 5 (e.g., an ESL AP). Alternatively, the device may be a network entity, such as a base station or other network entity that is not part of the ESL system.
[0113] In some implementations, process 900 further includes performing object detection on the one or more frames to detect one or more objects within the one or more frames, and at least a portion of the set of features corresponds to the one or more objects. For example, network entity 520 of FIG5 may perform one or more object detection operations on the survey video 528 of FIG5 to detect objects, such as products or signs, associated with feature set 530. Additionally or alternatively, the set of features may include products identified in the one or more frames, text of signs in the one or more frames, directional features with respect to one or more detected objects in the one or more frames, geometric descriptions of the one or more frames, spatial descriptions of the one or more frames, textual descriptions of the one or more frames, or combinations thereof.
[0114] In some specific implementations, process 900 further includes providing one or more prompts to the AI LLM based on at least one feature from the set of features to generate one or more other features from the set of features. For example, the LLM manager 544 of Figure 5 may generate one or more prompts based on at least some of the extracted features, and the one or more prompts may be provided as input by the LLM manager 544 to the LLM to generate a text description corresponding to some of the features, or to generate a more natural, user-friendly text description. Additionally or alternatively, process 900 may include providing one or more prompts to the AI LLM based on at least one of the one or more descriptions to generate the description of the estimated location. For example, the LLM manager 544 of Figure 5 may generate one or more prompts based on a description from the description layout 532, and the one or more prompts may be provided as input by the LLM manager 544 to the LLM to generate a more natural, user-friendly text description as the output of description 550.
[0115] In some implementations, process 900 further includes performing pattern matching between the set of features and the description layout to identify the one or more descriptions that most closely match the set of features. For example, feature group 530 of FIG5 or its textual description may be compared with descriptions included in description layout 532 to identify which description is most similar to the feature / feature description, such as based on a similarity score or another metric. Additionally or alternatively, process 900 may include sending the description at the estimated location to the UE. For example, network entity 520 of FIG5 may send description 550 to UE 501.
[0116] In some implementations, this set of features also includes flag-based features, which include the text of one or more flags in the one or more frames, the description of the one or more flags, or a combination thereof. In some such implementations, process 900 may also include providing one or more prompts to the AI LLM based on at least one of the flag-based features to generate at least one description of the one or more flags. For example, the LLM manager 544 of FIG. 5 may generate one or more prompts based on at least some of the flag-based features, and the one or more prompts may be provided by the LLM manager 544 to the LLM to generate a text description of the information corresponding to the flag, or to generate a more natural, user-friendly text description of the flag. Additionally or alternatively, process 900 may also include filtering at least one description from the description layout based on the flag-based features before identifying the one or more descriptions. For example, when performing pattern matching, the network entity 520 of FIG. 5 may filter out descriptions from the description layout 532 that do not match the region associated with the flag detected in the frame of the investigation video 528. In some of these specific implementations, the description of the estimated location is referenced to signs within the site. For example, description 550 of Figure 5 may describe the location based on marked aisles, marked shelves, etc.
[0117] In some implementations, process 900 further includes storing the description of the estimated location and the estimated location in a location information database. The location information database may be configured to store descriptions of locations and associated location data. For example, the location information database may include or correspond to location information database 534 of Figure 5. In some such implementations, process 900 may also include receiving a reverse lookup request including a description of the requested location; accessing the location information database to identify location data associated with the description of the requested location; and outputting the identified location data. For example, network entity 520 of Figure 5 may access location information database 534 based on the location of the requested location described by it received from UE 501, in order to identify the location of the requested described location, such as in Cartesian coordinates of that location.
[0118] In some implementations, process 900 further includes: obtaining an image of the UE's current location; performing object detection on the image to detect signs within the image; extracting a second set of features from a portion of the image including the signs; and outputting a description of the current location of a plurality of signs described in the descriptive layout. This description of the current location is based on the description of one or more of the plurality of signs, which is identified from the descriptive layout based on the second set of features. For example, some implementations may perform sign-based localization, as further described above with reference to Figure 7.
[0119] In some specific implementations, process 900 further includes: receiving a query from the UE, the query including a requested item; determining a direction from the estimated location to the location of the requested item; determining a description of the direction based on the description layout; and outputting the description of the direction to the UE. For example, the query may include or correspond to query 552 of FIG. 5, and the description of the direction may include or correspond to description 554 of FIG. 5.
[0120] Figure 10 is a block diagram of an example network entity that converts location based on a wireless communication system into a location description of an asset's layout, supported by one or more aspects of this disclosure. Network entity 1000 can be configured to perform operations including the block diagram of process 900 described with reference to Figure 9. In some specific implementations, network entity 1000 includes the structures, hardware, and components shown and described with reference to network entity 520 of Figure 5. For example, network entity 1000 may include a controller 1040 that operates to execute logical or computer instructions stored in memory 1042, and components that control network entity 1000 and provide the characteristics and functionality of network entity 1000. Under the control of controller 1040, network entity 1000 transmits and receives signals via wireless radio components 1001a to 1001t and antennas 1034a to 1034t. Wireless radio components 1001a to 1001t include various components and hardware such as modulators and demodulators, multiple-input multiple-output (MIMO) detectors, receiver processors, transmitter processors, transmit (TX) MIMO processors, receive (RX) MIMO processors, Bluetooth receivers, Bluetooth transmitters, PAN receivers, PAN transmitters, other wireless radio components, or combinations thereof.
[0121] As shown in the figure, memory 1042 may include (or be configured to store) location logic 1002, feature extraction logic 1003, pattern matching logic 1004, location information 1005, survey video 1006, and description layout 1007. Location logic 1002 may be configured to perform a location operation based on a wireless communication system to generate location information 1005. Feature extraction logic 1003 may be configured to extract one or more features from frames of survey video 1006. Pattern matching logic 1004 may be configured to match descriptions of the extracted features with descriptions in description layout 1007. Location information 1005 may include or correspond to location information 506 of FIG. 5. Survey video 1006 may include or correspond to survey video 528 of FIG. 5. Description layout 1007 may include or correspond to description layout 532 of FIG. 5. Memory 1042 may also include (or be configured to store) communication logic configured to enable communication between network entity 1000 and one or more other devices. Network entity 1000 may receive signals from or send signals to one or more UEs (such as UE 501 in Figure 5 or UE 1200 as described with reference to Figure 12).
[0122] Figure 11 is a flowchart illustrating an example process 1100 that supports the conversion of location based on a wireless communication system into a location description of an asset's layout according to some embodiments of this disclosure. The operation of process 1100 can be performed by a UE or its components as described herein. For example, process 1100 can be performed by UE 501 as described above with reference to Figure 5 or UE 1200 as described with reference to Figure 12.
[0123] At box 1102, the UE generates location information. For example, this location information may include or correspond to location information 506 in Figure 5. At box 1104, the UE sends its estimated location to a network entity. This estimated location is based on the location information. For example, this estimated location may include or correspond to estimated location 526 in Figure 5.
[0124] At box 1106, the UE receives a set of features from the network entity. This set of features is extracted from one or more frames of a survey video corresponding to the estimated location. For example, this set of features may include or correspond to feature group 530 of Figure 5. At box 1108, the UE sends one or more descriptions to the network entity. These one or more descriptions are generated based on the set of features. For example, these one or more descriptions may resemble one or more descriptions generated by the LLM manager 544 of Figure 5. At box 1110, the UE receives from the network entity a description of the estimated location regarding the layout of assets associated with the location. For example, this description may include or correspond to description 550 of Figure 5. In some implementations, this description of the estimated location may be relative to one or more assets or expressed in “retail semantics,” such as including aisle labels, shelf labels, shelf labels, or combinations thereof.
[0125] In some implementations, the UE may map the location information (e.g., one or more beacons or beacon response RSSI, RSRP, or AOA) to the estimated location and send the estimate to a network entity. For example, similar to location mapper 542 in Figure 5, the UE may map location information based on the wireless communication system to an estimated location in a Cartesian coordinate system representing the location. Additionally or alternatively, based on this set of features, the UE provides one or more prompts to the LLM to generate one or more descriptions. For example, the LLM may include or correspond to the LLM managed by LLM manager 544 in Figure 5. In some such implementations, the set of features may include flag-based features, and based on these flag-based features, the UE may provide one or more prompts to the LLM to generate the one or more descriptions.
[0126] In some implementations, the UE may receive user input including a query for a requested item. The UE may send this query as a request for direction to the network entity. For example, the query may include or correspond to query 552 in Figure 5. In some such implementations, the UE may receive a description of the direction of the requested item, and the UE may output this description of direction to enable the user to navigate through the location to the requested asset. For example, this description of direction may include or correspond to description 554 in Figure 5.
[0127] Figure 12 is a block diagram of an example UE 1200 that converts location based on a wireless communication system into a location description of an asset's layout, supported by one or more aspects of this disclosure. The UE 1200 can be configured to perform operations including the block diagram of process 1100 described with reference to Figure 11. In some specific implementations, the UE 1200 includes the structures, hardware, and components shown and described with reference to UE 501 of Figure 5. For example, the UE 1200 includes a controller 1280 that operates to execute logical or computer instructions stored in memory 1282, and components that control the UE 1200 and provide the characteristics and functionality of the UE 1200. Under the control of the controller 1280, the UE 1200 transmits and receives signals via wireless radio components 1201a to 1201r and antennas 1252a to 1252r. Wireless radio components 1201a to 1201r may include a variety of components and hardware, such as modulators and demodulators, multiple inputs, MIMO detectors, receiver processors, transmitter processors, TX MIMO processors, RX MIMO processors, Bluetooth receivers, Bluetooth transmitters, PAN receivers, PAN transmitters, other wireless radio components, or combinations thereof.
[0128] As shown in the figure, memory 1282 may include (or be configured to store) location logic 1202, feature extraction logic 1203, pattern matching logic 1204, location information 1205, survey video 1206, and description layout 1207. Location logic 1202 may be configured to perform a location operation based on a wireless communication system to generate location information 1205. Feature extraction logic 1203 may be configured to extract one or more features from frames of survey video 1206. Pattern matching logic 1204 may be configured to match the description of the extracted features with the description in description layout 1207. Location information 1205 may include or correspond to location information 506 of FIG. 5. Survey video 1206 may include or correspond to survey video 528 of FIG. 5. Description layout 1207 may include or correspond to description layout 532 of FIG. 5. Memory 1282 may also include (or be configured to store) communication logic configured to enable communication between UE 1200 and one or more other devices. UE 1200 can receive signals from or send signals to one or more devices, such as network entity 520 in Figure 5 or network entity 1000 in Figure 10.
[0129] It should be noted that one or more boxes (or operations) described with reference to Figures 1 through 12 can be combined with one or more boxes (or operations) described in another figure in the reference figures. For example, one or more boxes (or operations) in Figure 5 can be combined with one or more boxes (or operations) in Figures 1 through 4, 6, 7, and 9 through 12. Similarly, one or more boxes associated with Figure 9 can be combined with one or more boxes associated with Figure 11. Furthermore, one or more boxes associated with Figure 10 or Figure 12 can be combined with one or more boxes associated with Figures 1 through 9 and Figure 11.
[0130] In one or more aspects, techniques for supporting the conversion of location based on a wireless communication system into a location description relating to a layout may include additional aspects, such as any single aspect or any combination of aspects described in connection with one or more other processes or devices described below or elsewhere herein. In some examples, the techniques of one or more aspects may be implemented in a method or process. In some other examples, the techniques of one or more aspects may be implemented in a wireless communication device, such as a network entity or a component of a network entity, an ESL device or a component of an ESL device, an AP or a component of an AP, a gateway node or a component of a gateway node, a server or a component of a server, a UE or a component of a UE, a base station, a component of a base station, a server, a component of a server, another network entity or a component of another network entity. In some examples, a wireless communication device may include at least one processor (which may include an application processor, a modem, or other component) and at least one memory device coupled to the processor. The processor may be configured to perform the operations described herein with respect to a wireless communication device. In some examples, the memory device includes a non-transitory computer-readable medium on which instructions or program code are stored, which, when executed by the processor, causes the wireless communication device to perform the operations described herein. Additionally or alternatively, a wireless communication device may include an interface (e.g., a wireless communication interface) comprising a transmitter, a receiver, or a combination thereof. Additionally or alternatively, a wireless communication device may include one or more components configured to perform the operations described herein.
[0131] Specific implementation examples are described in the following numbered clauses: Clause 1: An apparatus for wireless communication, the apparatus comprising: at least one processor; and a memory coupled to the at least one processor, the at least one processor being configured to cause the apparatus to: obtain location information associated with a user equipment (UE); retrieve one or more frames of survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information; extract a set of features from the one or more frames; and output a description of the estimated location of the described layout based on one or more descriptions identified from the described layout identifiers of assets associated with the location, the one or more descriptions being identified based on the set of features.
[0132] Clause 2: The device as described in Clause 1, wherein the description of the estimated location is relative to one or more of the assets.
[0133] Clause 3: The equipment as described in Clause 1, wherein the description of the estimated location includes aisle labels, display shelf labels, shelf labels, or combinations thereof.
[0134] Clause 4: The device according to Clause 1, wherein the at least one processor is configured to further cause the device to: map the location information to the estimated location of the UE within the location.
[0135] Clause 5: The device according to Clause 1, wherein, in order to obtain the location information, the at least one processor is configured to cause the device to: calculate one or more of a Received Signal Strength Indicator (RSSI), Reference Received Power (RSRP), or Angle of Arrival (AoA) of one or more beacon responses from one or more of the assets and beacon responses from the UE; and determine the location information based on the one or more beacon responses and one or more of the RSSI, RSRP, or AoA of the beacon responses.
[0136] Clause 6: The equipment described in Clause 5, wherein the one or more assets of the assets correspond to one or more radio frequency identification (RFID) tags for one or more products within an electronic shelf label (ESL) system.
[0137] Clause 7: The device as described in Clause 1, wherein the survey video comprises multiple frames captured by the video capture device while navigating at various locations.
[0138] Clause 8: The device according to Clause 7, wherein each of the plurality of frames is associated with location coordinates, ephemeris data or both of which captured the corresponding frame, and wherein the location information associated with the UE includes location coordinates associated with the UE, ephemeris data or both of which are associated with the UE.
[0139] Clause 9: The device described in Clause 1, wherein the device is a network entity.
[0140] Clause 10: The device as described in Clause 1, wherein the device is an edge or cloud server, and wherein the edge or cloud server is associated with a plurality of electronic shelf label (ESL) access points (APs) or a plurality of ESLs, each of the plurality of ESLs being associated with one or more tags corresponding to one or more assets in the assets.
[0141] Clause 11: A method for wireless communication, the method comprising: obtaining location information associated with a user equipment (UE); retrieving one or more frames of a survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information; extracting a set of features from the one or more frames; and outputting a description of the estimated location of the described layout based on one or more descriptions of the described layout identifiers of assets associated with the location, the one or more descriptions being identified based on the set of features.
[0142] Clause 12: The method according to Clause 11 further comprises: performing object detection on the one or more frames to detect one or more objects within the one or more frames, wherein at least a portion of the set of features corresponds to the one or more objects.
[0143] Clause 13: The method according to Clause 11, wherein the set of features includes a product identified in the one or more frames, text of a mark in the one or more frames, directional features with respect to one or more detected objects in the one or more frames, a geometric description of the one or more frames, a spatial description of the one or more frames, a textual description of the one or more frames, or a combination thereof.
[0144] Clause 14: The method according to Clause 11 further comprises: providing one or more cues to an artificial intelligence (AI) large language model (LLM) based on at least one of the set of features to generate one or more other features of the set of features.
[0145] Clause 15: The method according to Clause 11 further includes: performing pattern matching between the set of features and the description layout to identify the one or more descriptions that most closely match the set of features.
[0146] Clause 16: The method according to Clause 11 further comprises: providing one or more cues to an artificial intelligence (AI) large language model (LLM) based on at least one of the one or more descriptions to generate the description of the estimated location.
[0147] Clause 17: The method according to Clause 11, wherein the set of features further includes flag-based features, the flag-based features including text of one or more flags in the one or more frames, descriptions of the one or more flags, or combinations thereof.
[0148] Clause 18: The method according to Clause 17 further comprises: providing one or more cues to an artificial intelligence (AI) large language model (LLM) based on at least one of the cues based on the cues to generate at least one of the descriptions of the one or more cues.
[0149] Clause 19: The method according to Clause 17 further comprises: filtering at least one description from the description layout based on the flag-based features before identifying the one or more descriptions.
[0150] Clause 20: The method described in Clause 17, wherein the description of the estimated location refers to a sign within the location.
[0151] Clause 21: The method according to Clause 11 further includes: sending the description of the estimated location to the UE.
[0152] Clause 22: A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining location information associated with a user equipment (UE); retrieving one or more frames of a survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information; extracting a set of features from the one or more frames; and outputting a description of the estimated location of the descriptive layout based on one or more descriptions of descriptive layout identifiers of assets associated with the location, the one or more descriptions being identified based on the set of features.
[0153] Clause 23: The non-transitory computer-readable medium as described in Clause 22, wherein the operation further comprises: storing the description of the estimated location and the estimated location in a location information database, wherein the location information database is configured to store the description of the location and associated location data.
[0154] Clause 24: The non-transitory computer-readable medium as described in Clause 23, wherein said operation further comprises: receiving a reverse lookup request including a description of the requested location; accessing the location information database to identify location data associated with the description of the requested location; and outputting the identified location data.
[0155] Clause 25: A non-transitory computer-readable medium as described in Clause 22, wherein said operation further comprises: obtaining an image of the current location of the UE; performing object detection on the image to detect signs within the image; extracting a second set of features from a portion of the image including the signs; and outputting a description of the current location of a plurality of signs described in the description layout, the description of the current location being based on a description of one or more of the plurality of signs, the description of the one or more signs being identified from the description layout based on the second set of features.
[0156] Clause 26: A non-transitory computer-readable medium as described in Clause 22, wherein said operation further comprises: receiving a query from the UE, the query including a requested item; determining a direction from the estimated location to the location of the requested item; determining a description of the direction based on the description layout; and outputting the description of the direction to the UE.
[0157] Clause 27: An Electronic Shelf Label (ESL) system comprising: a server including: a memory; and at least one processor coupled to the memory and configured to perform operations including: obtaining location information associated with a user equipment (UE); retrieving one or more frames of survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information; extracting a set of features from the one or more frames; and outputting a description of the estimated location of the described layout based on one or more descriptions identified from the described layout identifiers of assets associated with the location, the one or more descriptions being identified based on the set of features.
[0158] Clause 28: The ESL system as described in Clause 27, wherein the descriptive layout includes the name, description, and location of the product in retail semantics.
[0159] Clause 29: The ESL system as described in Clause 28, wherein the described layout represents a store shelf diagram.
[0160] Clause 30: The ESL system as described in Clause 28, wherein the described layout includes a description of the spatial relationships between at least some of the products.
[0161] The components, functional blocks, and modules described herein, in relation to the accompanying figures, include processors, electronic devices, hardware devices, electronic components, logic circuits, memory, software code, firmware code, and so on, or any combination thereof. Software should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, and / or functions, whether referred to as software, firmware, middleware, microcode, hardware description languages, or other terms. Furthermore, the features discussed herein may be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.
[0162] Those skilled in the art will further understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein are merely examples, and that components, methods, or interactions of various aspects of this disclosure may be combined or performed in ways other than those illustrated and described herein.
[0163] The various exemplary logics, logic blocks, modules, circuits, and algorithmic processes described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. The interchangeability of hardware and software has been broadly described in terms of functionality and illustrated in the aforementioned exemplary components, blocks, modules, circuits, and processes. Whether such functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system.
[0164] Hardware and data processing means for implementing the various exemplary logic, logic blocks, modules, and circuits described herein can be implemented or executed using general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some embodiments, the processor may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be performed by circuitry specific to a given function.
[0165] In one or more aspects, the described functionality may be implemented in hardware, digital electronic circuits, computer software, firmware, including the structures disclosed in this specification and their structural equivalents or any combination thereof. Specific implementations of the subject matter described in this specification may also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus.
[0166] If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted through a computer-readable medium. The processes of the methods or algorithms disclosed herein can be implemented in a processor-executable software module that can reside on a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that can be implemented to transfer a computer program from one location to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible to a computer. Additionally, any connection may be appropriately referred to as a computer-readable medium. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media. In addition, the operation of a method or algorithm may reside as a set of code and instructions or any combination of code and instructions on a machine-readable medium and a computer-readable medium that may be incorporated into a computer program product.
[0167] Various modifications to the specific embodiments described in this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to some other specific embodiments without departing from the spirit or scope of this disclosure. Therefore, the claims are not intended to be limited to the implementations shown herein, but are to be granted the broadest scope consistent with this disclosure, the principles disclosed herein, and the novel features.
[0168] Certain features described in this specification in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even originally claimed in this way, one or more features from the claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0169] Similarly, although operations are depicted in a specific order in the figures, this should not be construed as requiring such operations to be performed in the indicated specific order or sequential order, or to perform all illustrated operations to achieve the desired result. Furthermore, the figures may schematically depict one or more example processes in the form of flowcharts. However, other operations not depicted may be combined with the schematically illustrated example processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any illustrated operation. In some contexts, multitasking and parallel processing are advantageous. Moreover, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result.
[0170] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A device for wireless communication, the device comprising: At least one processor; The device includes a memory coupled to the at least one processor, the at least one processor being configured to cause the device to: obtain location information associated with a user equipment (UE); and retrieve one or more frames of a survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information. Extract a set of features from the one or more frames; And output a description of the estimated location of the described layout based on one or more descriptions identified from the described layout identifiers of the assets associated with the location, the one or more descriptions being identified based on the set of features.
2. The device of claim 1, wherein the description of the estimated location is relative to one or more of the assets.
3. The device of claim 1, wherein the description of the estimated location includes aisle labels, display shelf labels, shelf labels, or combinations thereof.
4. The device of claim 1, wherein the at least one processor is configured to further cause the device to: map the location information to the estimated location of the UE within the location.
5. The device according to claim 1, wherein, To obtain the location information, the at least one processor is configured to cause the device to: calculate one or more of the received signal strength indicator (RSSI), reference received power (RSRP), or angle of arrival (AoA) of one or more beacon responses from one or more of the assets and beacon responses from the UE; and determine the location information based on the one or more beacon responses and one or more of the RSSI, RSRP, or AoA of the beacon responses.
6. The device of claim 5, wherein the one or more assets in the assets correspond to one or more radio frequency identification (RFID) tags for one or more products within an electronic shelf label (ESL) system.
7. The device of claim 1, wherein the survey video comprises a plurality of frames captured by the video capture device while navigating at various locations.
8. The device of claim 7, wherein each of the plurality of frames is associated with location coordinates, ephemeris data, or both of which captured the corresponding frame, and wherein the location information associated with the UE includes location coordinates associated with the UE, ephemeris data associated with the UE, or both.
9. The device according to claim 1, wherein the device is a network entity.
10. The device of claim 1, wherein the device is an edge or cloud server, and wherein the edge or cloud server is associated with a plurality of electronic shelf label (ESL) access points (APs) or a plurality of ESLs, each of the plurality of ESLs being associated with one or more tags corresponding to one or more assets among the assets.
11. A method for wireless communication, the method comprising: Obtain location information associated with user equipment (UE); Retrieve one or more frames of survey video corresponding to the estimated location of the UE, the estimated location being based on the location information; Extract a set of features from the one or more frames; And output a description of the estimated location of the described layout based on one or more descriptions identified from the described layout identifiers of the assets associated with the location, the one or more descriptions being identified based on the set of features.
12. The method according to claim 11, further comprising: Object detection is performed on the one or more frames to detect one or more objects within the one or more frames, wherein at least a portion of the set of features corresponds to the one or more objects.
13. The method of claim 11, wherein the set of features includes a product identified in the one or more frames, text of a mark in the one or more frames, directional features with respect to one or more detected objects in the one or more frames, a geometric description of the one or more frames, a spatial description of the one or more frames, a textual description of the one or more frames, or a combination thereof.
14. The method according to claim 11, further comprising: Based on at least one feature from the set of features, provide one or more prompts to an Artificial Intelligence (AI) Large Language Model (LLM) to generate one or more other features from the set of features.
15. The method according to claim 11, further comprising: Perform pattern matching between the set of features and the description layout to identify the one or more descriptions that most closely match the set of features.
16. The method according to claim 11, further comprising: One or more cues are provided to an artificial intelligence (AI) large language model (LLM) based on at least one of the one or more descriptions to generate the description of the estimated location.
17. The method of claim 11, wherein the set of features further comprises flag-based features, the flag-based features including text of one or more flags in the one or more frames, descriptions of the one or more flags, or combinations thereof.
18. The method according to claim 17, further comprising: Based on at least one of the flag-based features, one or more prompts are provided to an Artificial Intelligence (AI) Large Language Model (LLM) to generate at least one description of the descriptions of the one or more flags.
19. The method of claim 17, further comprising: Before identifying the one or more descriptions, at least one description is filtered from the description layout based on the flag-based features.
20. The method of claim 17, wherein the description of the estimated location refers to a marker within the location.
21. The method according to claim 11, further comprising: The description of the estimated location is sent to the UE.
22. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining location information associated with a user equipment (UE); retrieving one or more frames of a survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information; Extract a set of features from the one or more frames; And output a description of the estimated location of the described layout based on one or more descriptions identified from the described layout identifiers of the assets associated with the location, the one or more descriptions being identified based on the set of features.
23. The non-transitory computer-readable medium of claim 22, wherein the operation further comprises: The description of the estimated location and the estimated location are stored in a location information database, wherein the location information database is configured to store the description of the location and the associated location data.
24. The non-transitory computer-readable medium of claim 23, wherein the operation further comprises: Receive a reverse lookup request that includes a description of the requested location; Access the location information database to identify location data associated with the description of the requested location; And output the identified location data.
25. The non-transitory computer-readable medium of claim 22, wherein the operation further comprises: Obtain an image of the current location of the UE; Perform object detection on the image to detect landmarks within the image; Extract a second set of features from a portion of the image that includes the logo; And output a description of the current position of a plurality of flags described in the description layout, the description of the current position being based on the description of one or more of the plurality of flags, the description of the one or more flags being identified from the description layout based on the second set of features.
26. The non-transitory computer-readable medium of claim 22, wherein the operation further comprises: The UE receives a query, the query including the requested item; Determine the direction from the estimated location to the location of the requested item; The description of the direction is determined based on the described layout; And output the description of the direction to the UE.
27. An electronic shelf label (ESL) system, the electronic shelf label (ESL) system comprising: A server, comprising: a memory; and at least one processor coupled to the memory and configured to perform operations including: obtaining location information associated with a user equipment (UE); retrieving one or more frames of survey video of a location corresponding to an estimated location of the UE, the estimated location being based on the location information; extracting a set of features from the one or more frames; and outputting a description of the estimated location of the descriptive layout based on one or more descriptions of descriptive layout identifiers of assets associated with the location, the one or more descriptions being identified based on the set of features.
28. The ESL system of claim 27, wherein the description layout includes the name, description, and location of the product in retail semantics.
29. The ESL system of claim 28, wherein the described layout represents a store shelf diagram.
30. The ESL system of claim 28, wherein the described layout includes a description of the spatial relationships between at least some of the products.