System and methods for supporting smart spatial anchors
The WTRU in wireless communication systems identifies objects in video frames and overlays virtual content in real-time, addressing the lack of efficient methods for smart spatial anchors, thereby enhancing user interaction and experience.
Patent Information
- Application Number
- PCT/US2025/027979
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-09
- Filing Date
- 2025-05-06
- Publication Date
- 2025-11-13
AI Technical Summary
Existing wireless communication systems lack efficient methods for identifying and overlaying virtual content onto video streams in real-time, particularly for smart spatial anchors, which are crucial for enhancing user interaction and experience in mobile environments.
A wireless transmit/receive unit (WTRU) equipped with a processor receives messages to identify objects in video frames, applies a model to determine their position, and overlays virtual content, sending reports to a network node for further processing and display.
Enables real-time identification and overlay of virtual content on video streams, improving user interaction and experience by providing precise object recognition and augmented reality capabilities.
Smart Images

Figure US2025027979_13112025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHODS FOR SUPPORTING SMART SPATIAL ANCHORSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 644,763, filed May 9, 2024, the contents of which are hereby incorporated by reference herein.BACKGROUND
[0002] Mobile communications using wireless communication continue to evolve. A fifth generation may be referred to as 5G. A previous (legacy) generation of mobile communication for example, may be fourth generation (4G) long term evolution (LTE).SUMMARY
[0003] Systems, methods, and devices described herein are related to supporting smart spatial anchors (SSAs). A wireless transmit / receive unit (WTRU) may include a processor. The WTRU may receive a first message from a network node. The first message may indicate a request to identify an instance of an object. The WTRU may obtain a model associated with the object. The WTRU may identify the instance of the object in a video frame by applying the model to the video frame. The video frame may be associated with a video stream. The WTRU may send a second message to the network node when the instance of the object has been identified. The WTRU may receive a third message from the network node. The third message may indicate virtual content associated with the object. The WTRU may modify the video frame by overlaying the virtual content onto the video frame. The WTRU may display the modified video frame to a user.
[0004] One or more features may be included. For example, the WTRU may determine a position of the identified instance of the object in the video frame. The virtual content may be overlaid onto the video frame at the position of the identified instance of the object. The model may be obtained based on position information associated with the WTRU and position information associated with the instance of the object. The second message may include at least one of an identifier of the object, a location where the object is recognized in the video frame, or a location where the object is recognized in the video frame. The virtual content may include at least one of a bounding box or a class label identifying the bounding box. The WTRU may send a report to the network node. The report may include at least one of information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRUorientation information. The first message may (e.g., further) comprise at least one of WTRU location information, WTRU orientation information, or a filter value. The WTRU may be at least one of a smartphone, a tablet, a wearable device, a head mounted display, a connected vehicle, or a drone.
[0005] A network node may include a processor. The network node may determine an object associated with a WTRU. The network node may send a first message to the WTRU. The first message may indicate a request to identify an instance of the object in a video frame based on an application of a model to the video frame. The video frame may be associated with a video stream. The network node may receive a second message from the WTRU based on the instance of the object being identified by the WTRU. The network node may send a third message to the WTRU. The third message may indicate virtual content associated with the object. The virtual content may be overlaid onto the video frame.
[0006] One or more features may be included. For example, the second message may include at least one of an identifier of the object, a location where the object is recognized in the video frame, or a location where the object is recognized in the video frame. The network node may receive a report from the WTRU. The report may include at least one of: information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRU orientation information. The first message may (e.g., further) include at least one of: WTRU location information, WTRU orientation information, or a filter value. The network node may obtain a model associated with the determined object. The first message sent to the WTRU may (e.g., further) indicate that the model to be applied to the video frame is the model associated with the determined object.
[0007] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that, in operation, causes the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0008] In an example, a wireless transmit / receive unit may include a processor configured to at least: receive a first message from a network node, wherein the first message indicates a request to identify an instance of an object; determine a model associated with the object; identify the instance of the object in a video stream by applying the model to the video stream; send a second message to the network node when the instance of the object has been identified; receive a third message from the network node, wherein the third message indicates virtual content associated with the object; modify the video stream by overlaying the virtual content onto the video stream according to the identified object instance; and display the modified video stream to a user.
[0009] In an example, a WTRU may include one or more of the following features. The WTRU may include an indication based on an inference classification and a label, and wherein the indication may include at least one of an identifier of the object, a location where the object is recognized in a video frame of a video stream, or a location where the object is recognized in a video stream. The virtual content may include at least one of a bounding box or a class label identifying the bounding box. The WTRU may send a report to the network node, wherein the report may include at least one of information associated with the object, a WTRU identifier, WTRU location information, or WTRU orientation information. The first message may include at least one of WTRU location information, WTRU orientation information, a filter value, or an information element. The WTRU may be at least one of a smartphone, a tablet, a wearable device, a head mounted display, a connected vehicle, or a drone. Implementations of the described techniques may include hardware, a method or process, or a computer tangible medium.
[0010] In an example, a method may include receiving a first message from a network node, wherein the first message indicates a request to identify an instance of an object. The method may also include determining a model associated with an object. The method may include identifying an instance of the object in a video stream by applying the model to the video stream. The method may in addition include sending a second message to a network node when the instance of the object has been identified in the video stream. The method may include receiving a third message from the network node, wherein the third message indicates virtual content associated with the object. The method may also include modifying the video stream by overlaying the virtual content onto the video stream according to the identified object. The method may include displaying the modified video stream to a user. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0011] In an example, a computer program product may include receiving a first message from a network node, wherein the first message indicates a request to identify an instance of an object. The computer program product may include determining a model associated with an object. The computer program product may include identifying an instance of the object in a video stream by applying the model to the video stream. The computer program product may include sending a second message to a network node when the instance of the object has been identified in the video stream. The computer program product may include receiving a third message from the network node, wherein the third message indicates virtual content associated with the object. The computer program product may include modifying the video stream by overlaying the virtual content onto the video stream according to the identified object. The computer program product may include displaying the modified video stream to a user. Other embodimentsof this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] FIG. 1 A is a system diagram illustrating an example communications system in which one or more disclosed embodiments may be implemented.
[0013] FIG. 1 B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communications system illustrated in FIG. 1A according to an embodiment.
[0014] FIG. 1 C is a system diagram illustrating an example radio access network (RAN) and an example core network (ON) that may be used within the communications system illustrated in FIG. 1 A according to an embodiment.
[0015] FIG. 1 D is a system diagram illustrating a further example RAN and a further example ON that may be used within the communications system illustrated in FIG. 1A according to an embodiment.
[0016] FIG. 2 is a block diagram illustrating an example mobile metaverse architecture that may include localized spatial anchors.
[0017] FIG. 3 is a block diagram illustrating an example architecture for supporting smart spatial anchors in an example network.
[0018] FIG. 4 is a flowchart illustrating an example method for discovering and / or using smart spatial anchors.
[0019] FIG. 5 is a flowchart illustrating an example method for reporting smart spatial anchors.DETAILED DESCRIPTION
[0020] FIG. 1 A is a diagram illustrating an example communications system 100 in which one or more disclosed embodiments may be implemented. The communications system 100 may be a multiple access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communications system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communications systems 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), and the like.
[0021] As shown in FIG. 1A, the communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, though it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which may be referred to as a “station” and / or a “STA”, may be configured to transmit and / or receive wireless signals and may include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular telephone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (loT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and applications (e.g., remote surgery), an industrial device and applications (e.g., a robot and / or other wireless devices operating in an industrial and / or an automated processing chain contexts), a consumer electronics device, a device operating on commercial and / or industrial wireless networks, and the like. Any of the WTRUs 102a, 102b, 102c and 102d may be interchangeably referred to as a UE.
[0022] The communications systems 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as the CN 106 / 115, the I nternet 110, and / or the other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node-B, an eNode B, a Home Node B, a Home eNode B, a gNB, a NR NodeB, a site controller, an access point (AP), a wireless router, and the like. While the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0023] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or the base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a wireless service to a specific geographical area that may be relatively fixed or that may change over time. The cell may further be divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e. , one foreach sector of the cell. In an embodiment, the base station 114a may employ multiple-input multiple output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.
[0024] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0025] More specifically, as noted above, the communications system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a in the RAN 104 / 113 and the WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0026] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0027] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR Radio Access , which may establish the air interface 116 using New Radio (NR).
[0028] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together, for instance using dual connectivity (DC) principles. Thus, the air interface utilized by WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., an eNB and a gNB).
[0029] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System forMobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), and the like.
[0030] The base station 114b in FIG. 1 A may be a wireless router, Home Node B, Home eNode B, or access point, for example, and may utilize any suitable RAT for facilitating wireless connectivity in a localized area, such as a place of business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a roadway, and the like. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR etc.) to establish a picocell or femtocell. As shown in FIG. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not be required to access the Internet 110 via the CN 106 / 115.
[0031] The RAN 104 / 113 may be in communication with the CN 106 / 115, which may be any type of network configured to provide voice, data, applications, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have varying quality of service (QoS) requirements, such as differing throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. The CN 106 / 115 may provide call control, billing services, mobile location-based services, pre-paid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not shown in FIG. 1A, it will be appreciated that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may be utilizing a NR radio technology, the CN 106 / 115 may also be in communication with another RAN (not shown) employing a GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0032] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or the other networks 112. The PSTN 108 may include circuit- switched telephone networks that provide plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as the transmission control protocol (TCP), user datagram protocol (UDP) and / or the internet protocol (IP) in the TCP / IP internet protocol suite. The networks 112 may include wired and / or wireless communications networks owned and / or operated by other service providers. For example,the networks 112 may include another CN connected to one or more RANs, which may employ the same RAT as the RAN 104 / 113 or a different RAT.
[0033] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with the base station 114a, which may employ a cellular-based radio technology, and with the base station 114b, which may employ an IEEE 802 radio technology.
[0034] FIG. 1 B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1 B, the WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138, among others. It will be appreciated that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0035] The processor 118 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1 B depicts the processor 118 and the transceiver 120 as separate components, it will be appreciated that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0036] The transmit / receive element 122 may be configured to transmit signals to, or receive signals from, a base station (e.g., the base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be appreciated that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0037] Although the transmit / receive element 122 is depicted in FIG. 1 B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0038] The transceiver 120 may be configured to modulate the signals that are to be transmitted by the transmit / receive element 122 and to demodulate the signals that are received by the transmit / receive element 122. As noted above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as NR and I EEE 802.11 , for example.
[0039] The processor 118 of the WTRU 102 may be coupled to, and may receive user input data from, the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from, and store data in, any type of suitable memory, such as the non-removable memory 130 and / or the removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0040] The processor 118 may receive power from the power source 134, and may be configured to distribute and / or control the power to the other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, and the like.
[0041] The processor 118 may also be coupled to the GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or in lieu of, the information from the GPS chipset 136, the WTRU 102 may receive location information over the air interface 116 from a base station (e.g., base stations 114a, 114b) and / or determine its location based on the timing of the signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 may acquire location information by way of any suitable locationdetermination method while remaining consistent with an embodiment.
[0042] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, an activity tracker, and the like. The peripherals 138 may include one or more sensors, the sensors may be one or more of a gyroscope, an accelerometer, a hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor; a geolocation sensor; an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0043] The WTRU 102 may include a full duplex radio for which transmission and reception of some or all of the signals (e.g., associated with particular subframes for both the UL (e.g., for transmission) and downlink (e.g., for reception) may be concurrent and / or simultaneous. The full duplex radio may include an interference management unit to reduce and or substantially eliminate self-interference via either hardware (e.g., a choke) or signal processing via a processor (e.g., a separate processor (not shown) or via processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio for which transmission and reception of some or all of the signals (e.g., associated with particular subframes for either the UL (e.g., for transmission) or the downlink (e.g., for reception)).
[0044] FIG. 1 C is a system diagram illustrating the RAN 104 and the CN 106 according to an embodiment. As noted above, the RAN 104 may employ an E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 may also be in communication with the CN 106.
[0045] The RAN 104 may include eNode-Bs 160a, 160b, 160c, though it will be appreciated that the RAN 104 may include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, 160c may implement MIMO technology. Thus, the eNode-B 160a, for example, may use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a.
[0046] Each of the eNode-Bs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, and the like. As shown in FIG. 1 C, the eNode-Bs 160a, 160b, 160c may communicate with one another over an X2 interface.
[0047] The CN 106 shown in FIG. 1 C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the foregoing elements are depicted as part of the CN 106, it will be appreciated that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0048] The MME 162 may be connected to each of the eNode-Bs 160a, 160b, 160c in the RAN 104 via an S1 interface and may serve as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during an initial attach of the WTRUs 102a, 102b, 102c, and the like. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.
[0049] The SGW 164 may be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via the S1 interface. The SGW 164 may generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions, such as anchoring user planes during inter- eNode B handovers, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing contexts of the WTRUs 102a, 102b, 102c, and the like.
[0050] The SGW 164 may be connected to the PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0051] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional land-line communications devices. For example, the CN 106 may include, or may communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to the other networks 112, which may include other wired and / or wireless networks that are owned and / or operated by other service providers.
[0052] Although the WTRU is described in FIGS. 1 A-1 D as a wireless terminal, it is contemplated that in certain representative embodiments that such a terminal may use (e.g., temporarily or permanently) wired communication interfaces with the communication network.
[0053] In representative embodiments, the other network 112 may be a WLAN.
[0054] A WLAN in Infrastructure Basic Service Set (BSS) mode may have an Access Point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have an access or an interface to a Distribution System (DS) or another type of wired / wireless network that carries traffic in to and / or out ofthe BSS. Traffic to STAs that originates from outside the BSS may arrive through the AP and may be delivered to the STAs. Traffic originating from STAs to destinations outside the BSS may be sent to the AP to be delivered to respective destinations. Traffic between STAs within the BSS may be sent through the AP, for example, where the source STA may send traffic to the AP and the AP may deliver the traffic to the destination STA. The traffic between STAs within a BSS may be considered and / or referred to as peer-to- peer traffic. The peer-to-peer traffic may be sent between (e.g., directly between) the source and destination STAs with a direct link setup (DLS). In certain representative embodiments, the DLS may use an 802.11e DLS or an 802.11 z tunneled DLS (TDLS). A WLAN using an Independent BSS (I BSS) mode may not have an AP, and the STAs (e.g., all of the STAs) within or using the IBSS may communicate directly with each other. The IBSS mode of communication may sometimes be referred to herein as an “ad- hoc” mode of communication.
[0055] When using the 802.11 ac infrastructure mode of operation or a similar mode of operations, the AP may transmit a beacon on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., 20 MHz wide bandwidth) or a dynamically set width via signaling. The primary channel may be the operating channel of the BSS and may be used by the STAs to establish a connection with the AP. In certain representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) may be implemented, for example in in 802.11 systems. For CSMA / CA, the STAs (e.g., every STA), including the AP, may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit at any given time in a given BSS.
[0056] High Throughput (HT) STAs may use a 40 MHz wide channel for communication, for example, via a combination of the primary 20 MHz channel with an adjacent or nonadjacent 20 MHz channel to form a 40 MHz wide channel.
[0057] Very High Throughput (VHT) STAs may support 20MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. The 40 MHz, and / or 80 MHz, channels may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining 8 contiguous 20 MHz channels, or by combining two non-contiguous 80 MHz channels, which may be referred to as an 80+80 configuration. For the 80+80 configuration, the data, after channel encoding, may be passed through a segment parser that may divide the data into two streams. Inverse Fast Fourier Transform (IFFT) processing, and time domain processing, may be done on each stream separately. The streams may be mapped on to the two 80 MHz channels, and the data may be transmitted by a transmitting STA. At the receiver of the receiving STA, the above described operation for the 80+80 configuration may be reversed, and the combined data may be sent to the Medium Access Control (MAC).
[0058] Sub 1 GHz modes of operation are supported by 802.11af and 802.11 ah. The channel operating bandwidths, and carriers, are reduced in 802.11 af and 802.11 ah relative to those used in 802.11 n, and802.11 ac. 802.11 af supports 5 MHz, 10 MHz and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11 ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non- TVWS spectrum. According to a representative embodiment, 802.11 ah may support Meter Type Control / Machine-Type Communications, such as MTC devices in a macro coverage area. MTC devices may have certain capabilities, for example, limited capabilities including support for (e.g., only support for) certain and / or limited bandwidths. The MTC devices may include a battery with a battery life above a threshold (e.g., to maintain a very long battery life).
[0059] WLAN systems, which may support multiple channels, and channel bandwidths, such as802.11 n, 802.11 ac, 802.11 af, and 802.11 ah, include a channel which may be designated as the primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by a STA, from among all STAs in operating in a BSS, which supports the smallest bandwidth operating mode. In the example of 802.11 ah, the primary channel may be 1 MHz wide for STAs (e.g., MTC type devices) that support (e.g., only support) a 1 MHz mode, even if the AP, and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) settings may depend on the status of the primary channel. If the primary channel is busy, for example, due to a STA (which supports only a 1 MHz operating mode), transmitting to the AP, the entire available frequency bands may be considered busy even though a majority of the frequency bands remains idle and may be available.
[0060] In the United States, the available frequency bands, which may be used by 802.11 ah, are from 902 MHz to 928 MHz. In Korea, the available frequency bands are from 917.5 MHz to 923.5 MHz. In Japan, the available frequency bands are from 916.5 MHz to 927.5 MHz. The total bandwidth available for802.11 ah is 6 MHz to 26 MHz depending on the country code.
[0061] FIG. 1 D is a system diagram illustrating the RAN 113 and the CN 115 according to an embodiment. As noted above, the RAN 113 may employ an NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 may also be in communication with the CN 115.
[0062] The RAN 113 may include gNBs 180a, 180b, 180c, though it will be appreciated that the RAN 113 may include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, 180c may implement MIMOtechnology. For example, gNBs 180a, 108b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, 180c. Thus, the gNB 180a, for example, may use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a. In an embodiment, the gNBs 180a, 180b, 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum while the remaining component carriers may be on licensed spectrum. In an embodiment, the gNBs 180a, 180b, 180c may implement Coordinated Multi-Point (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0063] The WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using subframe or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing varying number of OFDM symbols and / or lasting varying lengths of absolute time).
[0064] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In the standalone configuration, WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c without also accessing other RANs (e.g., such as eNode-Bs 160a, 160b, 160c). In the standalone configuration, WTRUs 102a, 102b, 102c may utilize one or more of gNBs 180a, 180b, 180c as a mobility anchor point. In the standalone configuration, WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using signals in an unlicensed band. In a non-standalone configuration WTRUs 102a, 102b, 102c may communicate with / connect to gNBs 180a, 180b, 180c while also communicating with / connecting to another RAN such as eNode-Bs 160a, 160b, 160c. For example, WTRUs 102a, 102b, 102c may implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In the non-standalone configuration, eNode-Bs 160a, 160b, 160c may serve as a mobility anchor for WTRUs 102a, 102b, 102c and gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for servicing WTRUs 102a, 102b, 102c.
[0065] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support of network slicing, dual connectivity, interworking between NR and E- UTRA, routing of user plane data towards User Plane Function (UPF) 184a, 184b, routing of control planeinformation towards Access and Mobility Management Function (AMF) 182a, 182b and the like. As shown in FIG. 1 D, the gNBs 180a, 180b, 180c may communicate with one another over an Xn interface.
[0066] The CN 115 shown in FIG. 1 D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements are depicted as part of the CN 115, it will be appreciated that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0067] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may serve as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling of different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, management of the registration area, termination of NAS signaling, mobility management, and the like. Network slicing may be used by the AMF 182a, 182b in order to customize CN support for WTRUs 102a, 102b, 102c based on the types of services being utilized WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services for machine type communication (MTC) access, and / or the like. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.
[0068] The SMF 183a, 183b may be connected to an AMF 182a, 182b in the CN 115 via an N11 interface. The SMF 183a, 183b may also be connected to a UPF 184a, 184b in the CN 115 via an N4 interface. The SMF 183a, 183b may select and control the UPF 184a, 184b and configure the routing of traffic through the UPF 184a, 184b. The SMF 183a, 183b may perform other functions, such as managing and allocating UE IP address, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, and the like. A PDU session type may be IP-based, non-IP based, Ethernetbased, and the like.
[0069] The UPF 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet- switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPF 184, 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, and the like.
[0070] The CN 115 may facilitate communications with other networks. For example, the CN 115 may include, or may communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to the other networks 112, which may include other wired and / or wireless networks that are owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to a local Data Network (DN) 185a, 185b through the UPF 184a, 184b via the N3 interface to the UPF 184a, 184b and an N6 interface between the UPF 184a, 184b and the DN 185a, 185b.
[0071] In view of FIGS. 1 A-1 D, and the corresponding description of FIGS. 1 A-1 D, one or more, or all, of the functions described herein with regard to one or more of: WTRU 102a-d, Base Station 114a-b, eNode- B 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN 185a- b, and / or any other device(s) described herein, may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more, or all, of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functions.
[0072] The emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or in an operator network environment. For example, the one or more emulation devices may perform the one or more, or all, functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network in order to test other devices within the communication network. The one or more emulation devices may perform the one or more, or all, functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation device may be directly coupled to another device for purposes of testing and / or may perform testing using over-the-air wireless communications.
[0073] The one or more emulation devices may perform the one or more, including all, functions while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in a testing scenario in a testing laboratory and / or a non-deployed (e.g., testing) wired and / or wireless communication network in order to implement testing of one or more components. The one or more emulation devices may be testing equipment. Direct RF coupling and / or wireless communications via RF circuitry (e.g., which may include one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0074] Reference to a timer herein may refer to a time, a time period, a tracking of time, a tracking of a period of time, a combination thereof, and / or the like. Reference to a timer expiration herein may refer to determining that the time has occurred or that the period of time has expired.
[0075] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0076] In an example, a wireless transmit / receive unit (WTRU) may include a processor configured to at least: receive a first message from a network node, wherein the first message indicates a request to identify an instance of an object; determine a model associated with the object; identify the instance of the object in a video stream by applying the model to the video stream; send a second message to the network node when the instance of the object has been identified; receive a third message from the network node, wherein the third message indicates virtual content associated with the object; modify the video stream by overlaying the virtual content onto the video stream; and display the modified video stream to a user.
[0077] In an example, a WTRU may include one or more of the following features. The WTRU may include an indication based on an inference classification and a label, and wherein the indication may include at least one of an identifier of the object, a location where the object is recognized in a video frame of a video stream, or a location where the object is recognized in a video stream. The virtual content may include at least one of a bounding box or a class label identifying the bounding box. The WTRU may send a report to the network node, wherein the report may include at least one of information associated with the object, a WTRU identifier, WTRU location information, or WTRU orientation information. The first message may include at least one of WTRU location information, WTRU orientation information, a filter value, or an information element. The WTRU may be at least one of a smartphone, a tablet, a wearable device, a head mounted display, a connected vehicle, or a drone. Implementations of the described techniques may include hardware, a method or process, or a computer tangible medium.
[0078] In an example, a method may include receiving a first message from a network node, wherein the first message indicates a request to identify an instance of an object. The method may also include determining a model associated with an object. The method may include identifying an instance of the object in a video stream by applying the model to the video stream. The method may in addition include sending a second message to a network node when the instance of the object has been identified in the video stream. The method may include receiving a third message from the network node, wherein the third message indicates virtual content associated with the object. The method may also include modifying the video stream by overlaying the virtual content onto the video stream according to the identified object. The method may include displaying the modified video stream to a user. Other embodiments of this aspectinclude corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0079] In an example, a computer program product may include receiving a first message from a network node, wherein the first message indicates a request to identify an instance of an object. The computer program product may include determining a model associated with an object. The computer program product may include identifying an instance of the object in a video stream by applying the model to the video stream. The computer program product may include sending a second message to a network node when the instance of the object has been identified in the video stream. The computer program product may include receiving a third message from the network node, wherein the third message indicates virtual content associated with the object. The computer program product may include modifying the video stream by overlaying the virtual content onto the video stream according to the identified object. The computer program product may include displaying the modified video stream to a user. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0080] Systems, methods, and devices described herein are related to supporting smart spatial anchors (SSAs). A wireless transmit / receive unit (WTRU) may include a processor. The WTRU may receive a first message from a network node. The first message may indicate a request to identify an instance of an object. The WTRU may obtain a model associated with the object. The WTRU may identify the instance of the object in a video frame by applying the model to the video frame. The video frame may be associated with a video stream. The WTRU may send a second message to the network node when the instance of the object has been identified. The WTRU may receive a third message from the network node. The third message indicates virtual content associated with the object. The WTRU may modify the video frame by overlaying the virtual content onto the video frame. The WTRU may display the modified video frame to a user.
[0081] One or more features may be included. For example, the WTRU may determine a position of the identified instance of the object in the video frame. The virtual content may be overlaid onto the video frame at the position of the identified instance of the object. The model may be obtained based on position information associated with the WTRU and position information associated with the instance of the object. The second message may include at least one of an identifier of the object, a location where the object is recognized in the video frame, or a location where the object is recognized in the video frame. The virtual content may include at least one of a bounding box or a class label identifying the bounding box. The WTRU may send a report to the network node. The report may include at least one of information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRUorientation information. The first message may (e.g., further) comprise at least one of WTRU location information, WTRU orientation information, or a filter value. The WTRU may be at least one of a smartphone, a tablet, a wearable device, a head mounted display, a connected vehicle, or a drone.
[0082] A network node may include a processor. The network node may determine an object associated with a WTRU. The network node may send a first message to the WTRU. The first message may indicate a request to identify an instance of the object in a video frame based on an application of a model to the video frame. The video frame may be associated with a video stream. The network node may receive a second message from the WTRU based on the instance of the object being identified by the WTRU. The network node may send a third message to the WTRU. The third message may indicate virtual content associated with the object. The virtual content may be overlaid onto the video frame.
[0083] One or more features may be included. For example, the second message may include at least one of an identifier of the object, a location where the object is recognized in a video frame, or a location where the object is recognized in the video frame. The network node may receive a report from the WTRU. The report may include at least one of: information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRU orientation information. The first message may (e.g., further) include at least one of: WTRU location information, WTRU orientation information, or a filter value. The network node may obtain a model associated with the determined object. The first message sent to the WTRU may (e.g., further) indicate that the model to be applied to the video frame is the model associated with the determined object.
[0084] A Smart Spatial Anchor (SSA) including computing capabilities may be provided.
[0085] In examples, a mobile metaverse enabler client (MMEC) may be present on a WTRU. A MMEC may discover an SSA, determine the SSA validity for using an artificial intelligence machine learning (AIML) model, obtain an AIML model, establish connectivity to a metaverse media service, provide information about an AIML model and / or a metaverse media service to an SSA consumer, obtain information about object recognition, and / or send recognition reports to the network. The terms “model” and “AIML model” may be used interchangeably herein.
[0086] Disclosed herein is a system and / or methods for determining and / or providing an SSA. The system may include one or more of an architecture, usage, reporting, and / or functionality associated with an SSA. In examples, the system may include an information model for supporting SSA, procedures for SSA discovery, usage, and / or reporting associated with an SSA.
[0087] In examples, a WTRU may perform one or more actions and / or methods as described herein. For example, a WTRU may transmit a request to a network (e.g., MMES). A request may be triggered by, forexample, an SSA consumer on the WTRU, another system, and / or the like. A request may include filters to discover, among other things, SSA instances of interest.
[0088] A WTRU may receive a response from the network. A response may include, for example, SSA information as described herein (e.g., indicating a request to identify an instance of an object). A WTRU may determine the validity of an SSA and / or an applicability (e.g., a need) for an AIML model (e.g., for object recognition and / or the like). SSA validity and / or applicability (e.g., a need) for an AIML model (e.g., for object recognition and / or the like) may be based on, for example, the proximity of a WTRU from an SSA and / or the orientation of a WTRU.
[0089] In examples, a WTRU may transmit an AIML model request to the network (e.g., AIMLES). A WTRU may determine, for example, whether to transmit a request based on an objective of object recognition (e.g., determining an AIML model for object recognition needs). A request may include, among other things, an AIML model identifier from an SSA (e.g., SSA information, and / or the like).
[0090] A WTRU may configure connectivity and / or connect to a metaverse media service. In examples, SSA information may include connectivity and / or configuration information. For example, connectivity and / or configuration information may include information regarding establishing a PDU session, requesting a slice, and / or configuring the QoS. A WTRU may provide SSA information to, for example, an SSA consumer application. SSA information may include, among other things, AIML model information and / or metaverse media service information.
[0091] A WTRU may provide XR data (e.g., image, video frame, video stream, and / or the like) to the AIML model for object recognition. The terms video frame and video stream may be used interchangeably. An AIML model may, for example, indicate recognized objects, provide a label, and / or provide a bounding area for a recognized object (e.g., each recognized object). A WTRU may transmit a metaverse media request to the metaverse media service. A request may be triggered by, for example, obtaining object recognition information from an AIML model. A request may include label information to retrieve metaverse media associated with recognizing an object. In examples, a WTRU may receive a metaverse media response from the metaverse media service.
[0092] A response may include metaverse media (e.g., virtual content) and / or label information associated with recognizing an object. A WTRU may use the obtained recognition information and / or metaverse media in one or more ways. For example, a WTRU may use the recognized objects information (e.g., object label and / or a bounding box) to position virtual content associated with a (e.g., each) label according to an object position (e.g., bounding box) in an image, video frame, video stream, and / or the like.
[0093] In examples, a WTRU may use the recognized object’s information to overlay (e.g., extend, overwrite, override, superimpose, and / or the like) positioned virtual content over the image, video frame,video stream, and / or the like. In examples, a WTRU may use the recognized object’s information to report object recognition information (e.g., object labels, WTRU location, WTRU orientation, and / or the like) to a network.
[0094] A metaverse may be provided. A metaverse may be related to virtual technologies such as AR, MR, and / or VR (e.g., termed XR). The metaverse may be a loosely defined term referring to a persistent, shared, perceived set of interactive spaces.
[0095] A metaverse may also evoke, for example, possible new experiences, products, and / or services that may emerge once VR and / or AR become commonly available. Metaverse applications may be diverse and / or may eventually provide new experiences in work, leisure, and / or other activities.
[0096] As disclosed herein, metaverse services may be provided that may foster the adoption of metaverse technologies and / or provide mobile user experiences.
[0097] Localized spatial anchors (SAs) may be provided. Localized user experiences may occur in a user’s local environment. Such experiences may request (e.g., require that metaverse media displayed for XR) provided to a given user is adapted and / or integrated with a local physical world surrounding the user.
[0098] A localized spatial anchor may be information that is associated with a spatial location (e.g., a 3D location in the physical world). An association between a spatial location and / or a service (e.g., an associated service) may be a localized spatial anchor”. A service associated with a spatial anchor may be used to obtain metaverse media. The metaverse media may be displayed, for example, at a spatial location in an XR environment.
[0099] A localized mobile metaverse service may allow a user to provide localized spatial anchors. In examples, a localized mobile metaverse service may allow a user to discover localized spatial anchors (e.g., for the purpose of displaying metaverse media associated with localized spatial anchors).
[0100] A mobile metaverse architecture may be provided. FIG. 2 is an example of a mobile metaverse architecture including, for example, localized spatial anchors capabilities.
[0101] An architecture may be based on, for example, mobile metaverse services formed by an MMEC and / or mobile metaverse capabilities provided by a service enabler architecture layer (SEAL) client. An MMEC and / or SEAL client may reside on a WTRU. The terms MMEC and SEAL client may be used interchangeably. In examples, an architecture may be based on mobile metaverse services formed by a mobile metaverse enablement server (MMES) and / or mobile metaverse capabilities provided by a SEAL server. An MMES and / or SEAL server may reside in a network. The terms MMES and SEAL server may be used interchangeably.
[0102] Vertical application layer (VAL) clients (e.g., applications) and an MMEC may interact, for example, when a client requests (e.g., needs) to discover spatial anchors and / or when the MMEC has obtained information from the network about spatial anchors of interest to a VAL client. An MMEC and / or VAL client may interact with a SEAL client to access SEAL services (e.g., a SEAL location management system (LMS)).
[0103] VAL servers (e.g., application servers) may interact with an MMES. For example, VAL servers may interact with an MMES when a VAL server provisions (e.g., configures) a spatial anchor in the MMES and / or when the MMES has determined information regarding one or more WTRU(s). In examples, the determined information may be associated with one or more spatial anchors of interest to a VAL server. MMES and / or VAL servers may interact with a SEAL server to access, among other things, SEAL services (e.g., SEAL LMS). In examples, an MMES and / or a SEAL server may communicate with a core network.
[0104] An MMEC may interact with an MMES to discover, receive, and / or subscribe to spatial anchor information. In examples, spatial anchor information (e.g., location information, associated service information, and / or the like) may be provided to a VAL client (e.g., VAL server) such that the VAL client may obtain metaverse media from the associated service (e.g., based on the spatial anchor information) and / or display the obtained media in an XR environment.
[0105] Interconnections between functional elements of the architecture represent communication reference points (e.g., SEAL-C / S, MM-C / S / EE / UU, VAL-UU, and / or the like). A reference point may be a logical representation of one or more possible communication paths between one or more functional elements of the architecture.
[0106] A location precision feature may be provided. Latitude is a geographic coordinate that specifies the north-south position of a point on the surface of a planet. Latitude is given as an angle that ranges from -90 ° at the South Pole to 90° at the North Pole, with 0° at the Equator. In examples, the precision of latitude may be expressed as decimal degrees. For example, a latitude may be 40.123456. Increments of the fifth decimal (e.g., 5) may represent approximately 1.1 meters, whereas increments of the sixth decimal (e.g., 6) may represent approximately 0.11 meters. In examples, obtaining a six-decimal precision latitude measurement may be impractical as it may rely upon precise measurements / equipment.
[0107] Longitude is a geographic coordinate that specifies the east-west position of a point on the surface of a planet. Longitude is an angular measurement, usually expressed in degrees. Meridians are imaginary semicircular lines running from pole to pole that connect points with the same longitude. In examples, the precision of a longitude may be expressed as decimal degrees. Precision may vary depending on the latitude. For example, 1 degree of longitude (e.g., from 0 to 1 degree) at the equator is approximately 111 km, while 1 degree of longitude is approximately 85 km at a latitude of about 40 degrees(e . g . , from 40-41 degrees). For example, given a longitude of 40.123456, increments of the fifth decimal (e.g., 5) represent approximately 1.1 meters at the equator, whereas increments of the fifth decimal (e.g., 5) represent approximately 0.85 meters at a latitude of approximately 40 degrees. Increments of the sixth decimal (e.g., 6) represent approximately 0.11 meters at the equator and 0.085 meters at a latitude of approximately 40 degrees. Obtaining a six-decimal precision longitude measurement may be impractical as it may rely upon precise measurements / equipment.
[0108] For the purposes of localized spatial anchors, it may be impractical and / or challenging to measure high-precision spatial locations (e.g., 5 and / or 6 decimal degree precision for latitude and / or longitude) that are requested (e.g., required) to allow precision use cases (e.g., use cases that request a precision below 1 meter).
[0109] AIML object recognition may be provided. An AIML model for object recognition may be a class of computer vision AIML model(s).
[0110] Object recognition may refer to a collection of computer vision tasks that involve identifying objects in a digital image. A digital image may be, for example, a digital frame of a video stream.
[0111] Computer vision tasks involved in object recognition are described herein.
[0112] Image classification is a task that may involve predicting a class and / or classes of objects in a digital image and / or a video frame. In examples, an image classification algorithm may produce a list of object categories present in a digital image and / or video frame. For example, an input of an AIML model for image classification may be a digital image, a video frame, a video stream, and / or the likes. An output of an AIML model for image classification may be, among other things, one or more labels. One or more labels may identify one or more objects in the digital image, video frame, video stream, and / or the like.
[0113] The terms object recognition and object detection may be used interchangeably herein. In examples, object localization may be a task that involves identifying the location of one or more objects in a digital image and / or video frame. In examples, object localization may indicate the location of one or more objects (e.g., in the digital image, video frame, video stream, and / or the likes) via, for example, a bounding box. An object localization algorithm may produce an axis-aligned bounding box indicating the position and / or scale of an object’s instance present in a digital image and / or video frame. For example, an input of an AIML model for object localization may be a digital image, video frame, video stream, and / or the likes, including (e.g., containing) one or more objects. Output for an AIML model for object localization may be one or more bounding boxes (e.g., a point of origin, a height, a width, and / or the like). A point of origin may be a point located within a digital image, video frame, video stream, and / or the likes corresponding to at least one corner of a bounding box.
[0114] Object detection is a task that may combine image classification and / or object localization. Object detection may allow an AIML model to locate the presence of objects with a bounding box and / or classify (e.g., assign a class label and / or the like) to a (e.g., each) bounding box in a digital image, video frame, video stream, and / or the like. Object detection algorithms may produce, for example, a list of object categories present in an image, along with an axis-aligned bounding box indicating the position and / or scale of one or more (e.g., every) instances of one or more (e.g., each) object category (ies). For example, the input of an AIML model for object detection may be a digital image and / or video frame having one or more objects. An output of an AIML model for object detection may be one or more bounding boxes and / or a class label for a (e.g., each) bounding box. In examples, region-based convolutional neural network (R- CNN) and / or You Only Look Once (YOLO) may be examples of methods used for performing object recognition.
[0115] An SSA in a network may be provided. A localized SA may be an association between a spatial location and / or a service. For example, an associated service may be accessed by a WTRU that has discovered the SA. The WTRU may be configured to obtain metaverse media content that may be displayed in an XR environment at the WTRU.
[0116] As described herein, it may not be possible to measure high-precision spatial locations (e.g., latitude, longitude) that are requested (e.g., required) to allow precision use cases. Provisioning (e.g., configuring) an SA with an imprecise location may lead to a bad (e.g., degraded) user experience, for example, displaying metaverse media at an incorrect location.
[0117] For example, an SA may not be aware (e.g., may not include information associated with) of a local environment (e.g., the environment where metaverse media is intended to be displayed). A lack of awareness may lead to incorrectly displaying metaverse media (e.g., if an SA is associated with an object that is not present).
[0118] SSA(s), as disclosed herein, may be utilized to support spatial anchors for precision use cases and / or to avoid the display of metaverse media at unintended locations.
[0119] In examples, a spatial anchor may leverage its static location. For example, a spatial anchor may use the static location associated with the spatial anchor to display metaverse media content (e.g., only) for objects that are static (e.g., objects that will not move and / or will not change position).
[0120] Dynamic spatial anchor positioning may adapt to changing environments, enhancing interactivity and realism in the metaverse. While challenges in achieving ultra-precise location measurements exist, such as under 1 -meter. For example, spatial anchors may be used in scenarios where coarse location precision is sufficient. Such flexibility may allow spatial anchors to be effectively utilized in a broad range of applications, making them versatile tools in dynamic settings.
[0121] In examples, supporting spatial anchors for precision use cases may be provided. SSAs may be used for precision use cases. In examples, SSAs may adapt to a changing environment to deliver a feature-rich and / or improved (e.g., optimized) user experience based on spatial anchors.
[0122] SSAs may have the capability of adapting to an environment. An SSA may provide precise metaverse media positioning (e.g., at and / or less than 1 -meter precision).
[0123] An SSA may have characteristics similar to, and / or the same as, an SA (e.g., a regular SA) such that both an SA and / or an SSA may co-exist in a system. The system may offer one or more services. For example, a system may provision (e.g., configure) an SA and / or an SSA, maintain an SA and / or an SSA, or allow for discovery and / or use of an SA and / or an SSA (e.g., concurrently and / or the like). An SSA may be compatible with, for example, a mobile metaverse architecture (e.g., as described with reference to FIG. 2 and / or any architecture supporting an SA).
[0124] SSAs may be enabled by associating an AIML model for object recognition with an SA, such that the AIML model for object recognition may provide improved capabilities described herein.
[0125] SSA recognition principles may be provided. An AIML model may be associated with an SSA (e.g., interchangeably referred to herein as an AIML model for object recognition). An AIML model may provide capabilities for recognizing objects (e.g., things, animals, humans, and / or the like). An AIML model may recognize objects by, for example, performing an inference on an image, a video frame, a video stream, and / or the like.
[0126] An AIML model for object recognition may be trained (e.g., specifically) in accordance with an SSA objective. As an illustrative example, an AIML model for object recognition of an SSA may be trained to recognize cheese objects and / or the like as part of a cheese display in a grocery store.
[0127] An inference result (e.g., an output) of the AIML model for object recognition may include an indication that an object has been recognized, an identifier (e.g., label) of the recognized object, and / or the location (e.g., bounding box) where the recognized object was recognized in an image, a video frame, a video stream, and / or the like (e.g., used for XR).
[0128] SSA information principles may be provided. SSA information may include an AIML model for object recognition and / or an identifier of an AIML model for object recognition. For mobile wireless systems that support AIML model distribution and / or other systems as described herein, providing an AIML model identifier may be sufficient (e.g., to obtain an AIML model for object recognition). In examples, it may be more efficient for a consumer of an SSA (e.g., a WTRU) to obtain an AIML model for object recognition.
[0129] SSA information may include, for example, a range for coarse and / or precise positioning. A range for coarse and / or precise positioning may be used by the SSA consumer (e.g., a WTRU and / or the like) to determine whether an AIML model may be used for inferencing. An SSA consumer may (e.g., first)determine if an SSA is valid. An SSA consumer may (e.g., then) determine if a user (e.g., a WTRU) is within range for enabling the AIML model for object recognition.
[0130] SSA cardinality principles may be provided.
[0131] SSA cardinality with recognizable objects may be, for example, one-to-one, when one or more SSA instance(s) are associated with a (e.g., one) recognizable object and / or the AIML model for object recognition is trained to recognize an (e.g., only one) object.
[0132] In examples, SSA cardinality with recognizable objects may be one-to-many when one or more SSAs are associated with multiple recognizable objects, and / or the AIML model for object recognition is trained to recognize multiple objects.
[0133] SSA reporting principles may be provided. SSA recognition outputs may be monitored and / or reported to a network. An SSA recognition report may indicate object(s) that a WTRU has recognized, a WTRU location, and / or a WTRU orientation where the objects were recognized.
[0134] A network may provide (e.g., send) report information about an SSA to interested application function(s). For example, a network may send report information to an application function that provisioned an SSA (e.g., based on an application function’s interest in knowing if a WTRU has recognized certain objects).
[0135] An architecture for an SSA may be provided. FIG. 3 presents an architecture for supporting SSAs in a network (e.g., a mobile network and / or the like). Functional elements and / or reference points of an architecture (e.g., a system, a WTRU, and / or the like) are described respectively for a WTRU and / or a network hereafter.
[0136] An SSA consumer functional element may be an application executing on, for example, a WTRU. An SSA consumer may be an XR application that provides augmented capabilities by enhancing a real environment with virtual content (e.g., metaverse media). An XR application may consume a real-time video feed, augment the video feed by rendering metaverse media on top of the video feed (e.g., overlay metaverse media on the video feed and / or the like), and / or provide the augmented video to a WTRU. An SSA consumer may be an application executing on, for example, a WTRU (e.g., a connected device, an XR headset, and / or the like). An SSA consumer may be tethered to the WTRU (e.g., configured as part of and / or connected to a WTRU).
[0137] The MMEC functional element may be a service and / or application executing on a WTRU. An MMEC may provide metaverse services to applications executing on a WTRU and / or executing on a device tethered to the WTRU (e.g., configured as part of and / or connected to a WTRU). In examples, an MMEC may consume metaverse services provided by a network.
[0138] An AIML enabler client (AIMLEC) functional element may be a service and / or application executing on a WTRU. An AIMLEC may provide AIML services to applications executing on a WTRU and / or executing on a device tethered to the WTRU. In examples, an AIMLEC may consume AIML services from a network.
[0139] The AIML model local repository functional element may be a service and / or an application executing on a WTRU. An AIML model local repository may provide storage and / or management capabilities for AIML models. In examples, a WTRU may receive an AIML model from an AIML model local repository. An AIMLEC may manage the AIML models stored in the AIML model local repository.Applications present on a WTRU and / or a device tethered to the WTRU may access and / or utilize one or more AIML models stored in an AIML model local repository.
[0140] An SSA producer functional element may be an application function executing in a network. An SSA producer may be an application that may create, store, manage, and / or delete spatial anchors (e.g., SA and / or SSA) in an MMES.
[0141] A metaverse media service functional element may be an application function executing in a network. A metaverse media service may offer a service that may be associated with, for example, a spatial anchor (e.g., an SA and / or an SSA). The metaverse media service may provide metaverse media content to, for example, XR application(s). In examples, a metaverse media service may be discovered (e.g., identified) when an XR application discovers (e.g., identifies) one or more spatial anchors.
[0142] An MMES functional element may be an application server executing in a network. A MMES may provide metaverse services to WTRU(s) (e.g., MMEC), to application functions, and / or to application servers in a network.
[0143] An AIML enabler server (AIMLES) functional element may be an application server executing in a network. An AIMLES may provide AIML services to WTRU(s) (e.g., AIMLEC), to application functions and / or application servers in the network.
[0144] An AIML model repository functional element may be an application server executing in a network. An AIML model repository may provide storage and / or management capabilities for AIML models that may be used in a network and / or distributed to WTRU(s). An AIMLES may manage one or more AIML models stored in the AIML model repository. An AIMLES may distribute AIML models, which may be executed locally by a WTRU. An AIML model repository may store AIML models that may be accessed and / or executed by AIMLES, application functions (AF), and / or application servers.
[0145] As shown at 1 in FIG. 3, a system (e.g., a network node and / or a WTRU) may allow an AF to access services from an MMES. The AF may provision (e.g., configure) an SA and / or an SSA in theMMES. An MMES may manage (e.g., create, update, use, and / or delete) the provisioned SA and / or SSA, and / or may subscribe to receive notifications about events related to the provisioned SA and / or SSA.
[0146] At 2, a system (e.g., a network node and / or a WTRU) may allow an MMES to access services from an AIMLES. An MMES may, for example, provision AIML models in the AIMLES, manage (e.g., create, update, use, and / or delete) the provisioned AIML models, and / or subscribe to receive notifications about events related to the provisioned AIML models.
[0147] At 3, a system (e.g., a network node and / or a WTRU) may allow an AIMLES to access services from an AIML model repository. For example, an AIMLES may store AIML models in an AIML model repository. In examples, an AIMLES may manage (e.g., create, update, use, and / or delete) stored AIML models.
[0148] At 4, a system (e.g., a network node and / or a WTRU) may allow an SSA consumer, such as an XR application, to access services from, for example, a metaverse media service. An SSA consumer may obtain metaverse media content from the metaverse media service for the purpose of, among other things, rendering metaverse media in an XR environment.
[0149] At 5, a system (e.g., a network node and / or a WTRU) may allow an MMEC to access services from an MMES. The services offered by an MMES may allow the MMEC to discover an SA and / or an SSA that has been provisioned in the MMES. An MMEC may provide SSA recognition reports to the MMES.
[0150] At 6, a system (e.g., a network node and / or a WTRU) may allow an AIMLEC to consume services from an AIMLES. Services offered by an AIMLES may enable the AIMLEC to discover AIML models that are managed by the AIMLES. In examples, one or more services offered by an AIMLES may fetch AIML models (e.g., that are needed) for a WTRU.
[0151] At 7, a system (e.g., a network node and / or a WTRU) may allow an SSA consumer to access services from an MMEC. Services offered by an MMEC may allow an SSA consumer to discover an SA and / or an SSA that may be used by the SSA consumer. In examples, services offered by an MMEC may allow an SSA consumer to generate and / or transmit SSA recognition reports to a network.
[0152] At 8, a system (e.g., a network node and / or a WTRU) may allow an SSA consumer to access services from an AIML model local repository. Services offered by an AIML model local repository may, for example, allow the SSA consumer to use one or more stored AIML models.
[0153] At 9, a system (e.g., a network node and / or WTRU) may allow an MMEC to access services from an AIMLEC. Services offered by an AIMLEC may allow the MMEC to, for example, discover AIML models that may be available in a network (e.g., AIML models that may be needed by the WTRU).
[0154] At 10, a system (e.g., a network and / or WTRU) may allow the AIMLEC to access services from an AIML model local repository. An AIMLEC may store AIML models in an AIML model local repository. An AIMLEC may manage (e.g., create, update, use, delete, and / or the like) stored AIML models.
[0155] Information model(s) and / or information elements related to a smart spatial anchor may be provided. In examples, SSA information may be provided. An SSA instance may include, for example, information associating a spatial location and / or a service with an AIML model. An SSA instance may include policy information describing, among other things, conditions for using an AIML model. Described herein are examples of (e.g., specific) information elements included as part of an SSA instance.
[0156] In an example, an information element may include a spatial anchor recognition model identifier (SARMID). A SARMID may be an identifier of an AIML model for object recognition. A SARMID may indicate to an SA consumer (e.g., a WTRU) whether a spatial anchor may be an SSA. A SARMID may identify an AIML model for object recognition (e.g., an AMIL that may be associated with an SSA). An MMEC may use a SARMID to determine whether an AIML model for object recognition may be applicable (e.g., needed) to use an SSA. An MMEC may use a SARMID to determine whether an AIML model is already available locally (e.g., via an AIMLEC). An MMEC may use a SARMID to obtain and / or fetch an associated AIML model from a network.
[0157] In an example, an information element may include a spatial anchor recognition model (SARM). A SARM may be an AIML model for object recognition. A SARM may indicate to an SA consumer (e.g., WTRU) whether a spatial anchor is an SSA. A SARM may provide an AIML model for object recognition that may be associated with an SSA (e.g., when a system does not support AIML model distribution and / or the like). In examples, an MMEC may store a SARM in a local AIML repository and / or provide the SARM to an SSA consumer.
[0158] In an example, an information element may include a spatial anchor recognition model range (SARMR). A SARMR may be a distance and / or an area value. A SARMR may be used to determine the usage of an AIML model for object recognition. A SARMR may indicate to an SSA consumer (e.g., WTRU), whether an AIML model for object recognition may (e.g., only) be used when the SA consumer is within a distance range of a spatial anchor location and / or within an area. An MMEC may use an SARMR to determine (e.g., by calculating the distance between a WTRU and / or an SSA location), and / or evaluate the presence of a WTRU in an area. An MMEC may use an SARMR to determine whether an AIML model may be used for object recognition. An MMEC may indicate to the AR application whether an AIML model may be used. An MMEC may use an SARMR to obtain and / or fetch an AIML model. An MMEC may use a SARMR to generate and / or transmit an indication (e.g., to provide AIML model information) to an AR application and / or the like.
[0159] In an example, an information element may include spatial anchor recognition information (SARI). A SARI may indicate possible inference classifications and / or labels that may be produced by an AIML model for object recognition. A SARI may indicate (e.g., to an SSA consumer such as a WTRU) one or more labels (e.g., a list of labels) that an AIML model for object recognition may produce. An SSA consumer may use one or more labels for retrieving label-specific metaverse media from a media service.
[0160] Information elements (e.g., common) may be used by an SA and / or an SSA. Information elements (e.g., common information elements) may be described herein for highlighting SSA functionality (e.g., specificities). The term SA and / or SSA may be used interchangeably herein.
[0161] In an example, an information element may include a spatial anchor type (SAT). An SAT may identify (e.g., uniquely identify) a type of spatial anchor. An SAT may indicate whether a spatial anchor may be an SA and / or an SSA. For example, discovery filters may be provided to discover spatial anchors. A discovery filter may include an SAT (e.g., SA or SSA). In examples, a network (e.g., an MMES, and / or the like) may use an SAT to select one or more spatial anchors according to a discovery filter.
[0162] In an example, an information element may include a spatial anchor identifier (SAID). A SAID may identify (e.g., uniquely identify) an SA and / or an SSA instance. A SAID may be assigned when provisioning a spatial anchor. For example, one or more discovery filters may be provided to discover spatial anchors. A discovery filter may include one or more SAID(s). A network (e.g., an MMES and / or the like) may use a SAID to select one or more spatial anchors according to a discovery filter.
[0163] In an example, an information element may include a spatial anchor service identifier (SASID). A SASID may identify an application and / or service(s) associated with an SA and / or an SSA instance. An application and / or service(s) may be performed by a service (e.g., the metaverse media service). For example, an SASID may indicate a fully qualified domain name (FQDN), a unified resource identifier (URI), an IP address, an edge application server (EAS) identifier, a cloud application server (CAS) identifier, a combination thereof, and / or the like. An SASID may include information to retrieve (e.g., requested (e.g., required) to retrieve) service information, such as, for example, an edge enabler server (EES) identifier, an edge configuration server (ECS) identifier, an edge application server discovery function (EASDF) identifier, and / or the like. In examples, an SA and / or SSA consumer (e.g., WTRU) may use an SASID to retrieve metaverse media from a service. An SA and / or SSA consumer (e.g., WTRU) may display metaverse media (e.g., retrieved from a service) in an XR environment.
[0164] In examples, an information element may include a spatial anchor service connection information (SASCI). An SASCI may indicate information used (e.g., required) by a WTRU to establish a connection to a service associated with the SA and / or an SSA instance. For example, a SASCI may indicate a data network name (DNN), an access point name (APN), a single-network slice selection assistance information(S-NSSAI), and / or the like. An SA and / or SSA consumer (e.g., a WTRU) may use an SASCI to establish connectivity to a data network and / or to retrieve metaverse media from a service.
[0165] In examples, an information element may include a spatial anchor service characteristics (SASC). A SASC may identify one or more characteristics of an application and / or service(s) performed by a service associated with an SA and / or an SSA instance. For example, a SASC may indicate an application type (e.g., information service, media service, and / or the like), service version information (e.g., service version, API version, and / or the like), media content type delivered by a service, and / or the like.
[0166] In examples, an information element may include a spatial anchor traffic descriptor (SATD). An SATD may identify traffic that may be provided by an associated SA and / or SSA service. An SA and / or an SSA consumer may use an SATD to, for example, enable QoS policies for media content traffic (e.g., media content traffic that may be transmitted from a service). For example, a SATD may include a QoS reference (e.g., which may be used by a 5G system to determine QoS parameters, among other things).
[0167] In examples, an information element may include a service anchor producer identifier (SAPID). A SAPID may indicate an identity of an SA and / or SSA producer. A SAPID may be used, for example, for filtering SA and / or SSA instances. When an SA and / or an SSA producer is an application (e.g., a specific application), a SAPID may be used to filter SA and / or SSA instances provided for an application (e.g., an XR application, an AR application, and / or the like).
[0168] In examples, an information element may include a service anchor consumer identifier (SACID). An SACID may indicate the identity of an SA and / or an SSA consumer. An SACID may be static. A static SACID may, for example, represent an SA and / or an SSA consumer that is allowed to use an SA and / or an SSA instance. An SACID may be dynamic. A dynamic SACID may include, for example, a list of SA and / or SSA consumers that have been provided with SA and / or SSA instance information. An SACID may be used for filtering SA and / or SSA instances. An SACID may be used to authorize access to SA and / or SSA instances.
[0169] In examples, an information element may include a service anchor state. A service anchor state may indicate whether an SA and / or an SSA instance is enabled and / or disabled. An enabled SA and / or SSA instance may indicate whether an associated service is allowed to provide, for example, media content. In examples, a disabled SA and / or SSA instance may indicate whether an SA is not allowed to provide media content. In examples, a disabled SA and / or SSA instance may be discovered (e.g., indicating that there is an SA and / or SSA at a specific location). In examples, a disabled SA and / or SSA instance may indicate whether an associated media content is not available.
[0170] As described herein, SSA discovery and / or usage may be provided, along with actions performed by functional entities involved in the SSA discovery and / or usage.
[0171] FIG. 4 describes an example (e.g., a method) for discovering and / or using an SSA.
[0172] In examples (e.g., before the operations of FIG. 4 may be executed), one or more SSA instances may be provisioned in an MMES. SSA provisioning may be performed by interacting with an MMES via, for example, at 1 of FIG. 3 (e.g., SSA(s) provisioning may be initiated by an SSA producer).
[0173] In examples (e.g., before the method of FIG. 4 may be executed), AIML model(s) for object recognition may be provisioned (e.g., configured) in an AIMLES. AIML model provisioning may be performed by interacting with an AIMLES via, for example, at 2 of FIG. 3 (e.g., AIML model(s) provisioning may be initiated by an MMES if, for example, an API as part of an SSA producer and / or an MMES may use AIML model(s)). AIML model provisioning may be initiated by an SSA producer (e.g., directly) with an AIMLES via, for example, at 2 of FIG. 3 (e.g., if an SSA producer has the capability to initiate AIML model provisioning).
[0174] As shown in FIG. 4 at 1, an SSA consumer (e.g., present on a WTRU, such as, for example, an XR application) may transmit a request (e.g., 1a) to an MMEC (e.g., of a WTRU). At 1 a, a request may be sent to discover spatial anchors, including an SA and / or an SSA. An SSA consumer may provide instructions (e.g., requirements) for discovering spatial anchors. The instructions (e.g., parameters or requirements) may contain one or more information elements as described herein.
[0175] At 1 b, an MMEC may send an SA discovery request to an MMES. The request (e.g., 1 b) may include an indicated instruction (e.g., requirement) as spatial anchor discovery filters. The request (e.g., 1 b) may include contextual information related to a WTRU (e.g., WTRU identifiers, application identifier, WTRU location, WTRU orientation, and / or the like). An MMES may use information received in an SA discovery request to determine an SA and / or an SSA. For example, an MMES may use WTRU location and / or orientation information to determine an SA and / or an SSA in proximity to a WTRU. An MMES may use discovery filter values to identify an SA and / or an SSA corresponding to one or more SSA consumer instructions (e.g., consumer requirements).
[0176] At 1c, an MMES may send an SA discovery response to an MMEC. An SA discovery response (e.g., 1c) may include determined SA(s) and / or SSA(s) information.
[0177] At 2, a WTRU may determine valid SA and / or SSA instances from multiple discovered SA and / or SSA instances. A validity determination may be based on WTRU sensor readings (e.g., WTRU location and / or orientation information).
[0178] An MMEC may determine whether a discovered SSA is valid. If an MMEC determines that an SSA is valid, the MMEC may further make one or more additional determinations for SSA usage.
[0179] An MMEC may make a first determination for SSA usage (e.g., using SSA information), whether an SSA has an associated AIML model for object recognition. The first determination may be based on thespatial anchor type (SAT) indicating an SSA type, on the presence of a SARM, and / or based on a SARMID.
[0180] An MMEC may make a second determination for SSA usage (e.g., using SSA information), whether a WTRU is in range of an SSA for using an AIML model for object recognition. The second determination may be based on an SARMR, WTRU positioning, and / or an SSA spatial location.
[0181] In examples, if an MMEC’s first determination for SSA usage has concluded that a SARM is provided with an SSA, the MMEC may proceed to 4. At 4, an MMEC may request that an AIMLEC store a SARM in an AIML repository for local models at 9 in FIG. 3.
[0182] Referring to FIG. 4, if an MMEC first determination for SSA usage has concluded that a SARM is not provided with an SSA, that an SARMID is present, and / or that a second determination for SSA usage has concluded that a WTRU is within range for using an AIML model for object recognition, the MMEC may proceed to 3 to retrieve the AIML model for object recognition.
[0183] At 3a in FIG. 4, an MMEC may request an AIML service (e.g., AIMLEC) of a WTRU to obtain an AIML model. The request may include AIML model instructions (e.g., requirements). AIML model instructions (e.g., requirements) may include an SARMID and / or the like.
[0184] At 3b in FIG. 4, an AIMLEC may send an AIML model request to an AIML enablement server (AIMLES). The request may include indicated AIML model instructions (e.g., requirements) as AIML model filters. The request may include contextual information related to a WTRU (e.g., WTRU identifiers, an application identifier, WTRU location, and / or the like). An AIMLES may use information received in an AIML model request to determine, among other things, AIML models. For example, an AIMLES may use a provided SARMID to identify a model with a provided identifier. An AIMLES may obtain an AIML model from an AIML model repository (e.g., at 3 in FIG. 3).
[0185] At 3c in FIG. 4, an AIMLES may send an AIML model response to an AIMLEC. A request may include a determined AIML model and / or information for fetching the AIML model from an AIML model repository. An AIMLEC may fetch and / or store an AIML model in an AIML model local repository (e.g., via at 10 in FIG. 3).
[0186] At 3d in FIG. 4, an AIMLEC may provide an indication to an MMEC that an AIML model for object recognition was retrieved. In examples, an indication may include a success indication, a SARMID, and / or a retrieved model.
[0187] At 4 in FIG. 4, a WTRU may configure the connectivity (e.g., to a metaverse media service) according to a determined valid SA and / or SSA instances. For example, a WTRU may use an SASID, an SASCI, and / or an SATD to configure connectivity to a metaverse media service associated with a valid SA and / or a valid SSA.
[0188] At 5 in FIG. 4, an MMEC may provide SSA information to an SSA consumer. SSA information may include SSA information as defined herein, for example, information obtained at 1c, 3d, etc. Upon receiving SSA information, an SSA consumer may begin an SSA usage phase described herein with reference to at least 6-8 in FIG. 4.
[0189] At 6a in FIG. 4, an SSA consumer may provide XR video data (e.g., an image, a video frame, a video stream, and / or the like) to an AIML model for object recognition. An AIML model for object recognition may perform an inference using XR video data (e.g., as provided by an SSA consumer) as input. XR video data may include an image of an XR video flow, video frames of an XR video flow, an XR video flow and / or the like. An AIML model for object recognition may perform inferencing on XR video data (e.g., an image, a video frame, a video stream, and / or the like), to identify, for example, objects based on recognition capabilities of the AIML model for object recognition.
[0190] At 6b in FIG. 4, an AIML model for object recognition may return recognition information to an SSA consumer. SSA information may include SARI, which may indicate one or more recognized object labels, an identifier for a (e.g., each) recognized object, and / or object recognition area information (e.g., a bounding box information where recognized objects were identified).
[0191] At 7a in FIG. 4, an SSA consumer may send a metaverse media request to a metaverse media service. The metaverse media service may be associated with an SSA. A request may include, for example, SARI information (e.g., as obtained at 6a and / or 6b). The request may include labels for one or more (e.g., a) identified object(s). The request may include range information indicating whether the WTRU is within a certain range of identified object(s). A metaverse media service may use SARI information, range information, and label information to determine virtual media content for one or more (e.g., each) recognized object(s) (e.g., one or more labels).
[0192] At 7b in FIG. 4, a metaverse media service may send a response to an SSA consumer. A response may include, for example, media content for one or more (e.g., a) recognized objects (e.g., labels).
[0193] At 8 in FIG. 4, an SSA consumer may use an object label and / or object recognition area information of a SARI (e.g., obtained at 6a and / or 6b) to position media content (e.g., obtained at 7a and / or 7b). An SSA consumer may render precisely positioned metaverse media content in a live XR video stream.
[0194] A spatial anchor discovery (e.g., as described at 1 in FIG. 4) may occur via a subscribe-notify model. An MMEC may subscribe with an MMES (e.g., to receive a notification when an SA and / or an SSA of interest is in range of a WTRU). When an MMES detects that an SA and / or an SSA of interest is in range of a WTRU, the MMES may send an SA discovery notification to the MMEC to indicate the SA and / or SSA.The processing resulting from a discovery notification may be similar to and / or the same as the discovery response at 1c and / or the like.
[0195] An MMEC may determine (e.g., at 2 in FIG. 4) if an AIML model for object recognition associated with an SSA may be used. An MMEC may re-evaluate whether a WTRU is within range of an SSA for using the AIML model for object recognition. An MMEC may notify an SSA consumer if the MMEC may stop and / or start using an AIML model for object recognition based on a re-evaluation.
[0196] An MMEC may determine (e.g., at 2 in FIG. 4) whether an AIML model for object recognition associated with an SSA may be used. An MMEC may determine if an SSA may be valid based on a WTRU position and / or orientation (e.g., prior to a determination whether an AIML model for object recognition associated with an SSA may be used). For example, if a WTRU is positioned such that the WTRU may not see an SSA, an MMEC may determine that it may be unnecessary to use an AIML model.
[0197] An example for recovering an AIML model for object recognition (e.g., at 3 in FIG. 4) may differ based on, for example, different system capabilities and / or implementations. For example, an operation (e.g., the operation shown at 3 in FIG. 4) may not occur if an AIML model for object recognition is included in an SSA. An operation (e.g., shown at 3 in FIG. 4) may not occur if the AIMLEC determines whether the AIML model for object recognition of the SSA is already present on a WTRU. An operation (e.g., shown at 3 in FIG. 4) may occur via a subscribe-notify model. In a subscribe-notify model, an AIML model may be transmitted to a WTRU via an AIML model notification. Processing resulting from an AIML model notification may be similar to, for example, the AIML model response shown at 3c and / or the like.
[0198] Establishing connectivity (e.g., shown at 4 in FIG. 4) may occur in another operation, such as at 2. Connectivity may be established once SSA validity is determined. If an MMEC determines whether connectivity is established, the operation at 4 may not be performed.
[0199] SSA information provided to an SSA consumer (e.g., at 5 in FIG. 4) may differ based on, for example, system capabilities, and / or one or more implementations. System capabilities and / or one or more implementations may influence the usage of an AIML model for object recognition (e.g., as shown at 6 in FIG. 4).
[0200] An AIML model may be provided to an SSA consumer by an MMEC (e.g., at 5 in FIG. 4). Execution of an AIML model may occur in an SSA consumer context and / or interactions (e.g., at 6 in FIG. 4) may occur as an intra-process interaction.
[0201] An AIML model may be stored in an AIML model local repository. An MMEC may provide an AIML model identifier to an SSA consumer. An SSA consumer may retrieve an AIML model from a repository. Execution of an AIML model may occur in an SSA consumer context and / or interactions (e.g., at 6 in FIG. 4) may occur as an intra-process interaction.
[0202] An AIML model may execute (e.g., on a WTRU) as a service in the AIML model’s own context. An SSA consumer may be informed in the SSA information on how to access the service. Interactions (e.g., at 6 FIG. 4) may include inter-process communication.
[0203] SSA recognition reporting may be provided. SSA reporting and / or the actions performed by functional entities involved in SSA reporting are described hereafter.
[0204] FIG. 5 describes, among other things, an example of reporting SSA recognition information.
[0205] At 0, an SSA producer may provision one or more SSA instances with MMES. At 0, one or more operations may be the same as and / or similar to one or more examples described herein, such as some of the examples described with reference to FIG. 4. An SSA producer may be interested in obtaining recognition information observed at an SSA consumer (e.g., to evaluate the performance of an AIML model associated with an SSA), and / or to detect anomalies (e.g., an object may not be detected anymore).
[0206] An SSA producer may (e.g., need to) subscribe with an MMES to obtain, for example, SSA recognition reporting notifications.
[0207] At 1 in FIG. 5, a SA consumer may discover an SSA. Discovering an SSA may be the same as and / or similar to one or more examples as described with reference to FIG. 4. For example, an SSA consumer may perform recognition inferencing on XR video data (e.g., an image, a video frame, a video stream, and / or the like), such as described at 6a and / or 6b of FIG. 4.
[0208] At 2a in FIG. 5, an SSA consumer may provide SSA recognition information (e.g., SARI) to an MMEC. The SSA recognition information may include information such as described at 6a and / or 6b of FIG. 4.
[0209] At 2b in FIG. 5, an SSA recognition report is sent to an MMES. An SSA recognition report may include SARI information, SSA information, a WTRU identifier, a WTRU location, a WTRU orientation, and / or the like. An MMES may store the received recognition report information as analytics related to an SSA. In examples, an MMES may store the received recognition report information as analytics related to an AIML model for object recognition.
[0210] An MMEC may identify whether an SSA producer has subscribed to be notified of SSA recognition reports. Based on a determination of whether an SSA producer has subscribed, an MMEC may provide the SSA recognition report information to an SSA producer.
[0211] At 3 in FIG. 5, an SSA producer may take recognition-based actions. In an example recognitionbased action, an SSA producer for a retail store may restock a product based on a report indicating that a product is not recognized anymore. For example, an SSA producer may re-train a model if the recognition failure rate is too high.
[0212] At 4 in FIG. 5, an SSA consumer may perform one or more actions as described at 7a, 7b, and / or 8 in FIG. 4.
[0213] Although features and elements described above are described in particular combinations, each feature or element may be used alone without the other features and elements of the preferred embodiments, or in various combinations with or without other features and elements.
[0214] Although the implementations described herein may consider 3GPP specific protocols, it is understood that the implementations described herein are not restricted to this scenario and may be applicable to other wireless systems. For example, although the solutions described herein consider LTE, LTE-A, New Radio (NR) or 5G specific protocols, it is understood that the solutions described herein are not restricted to this scenario and are applicable to other wireless systems as well.
[0215] The processes described above may be implemented in a computer program, software, and / or firmware incorporated in a computer-readable medium for execution by a computer and / or processor. Examples of computer-readable media include, but are not limited to, electronic signals (transmitted over wired and / or wireless connections) and / or computer-readable storage media. Examples of computer- readable storage media include, but are not limited to, a read only memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as, but not limited to, internal hard disks and removable disks, magneto-optical media, and / or optical media such as compact disc (CD)-ROM disks, and / or digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, terminal, base station, RNC, and / or any host computer.
Claims
CLAIMSWhat is Claimed:
1. A wireless transmit / receive unit (WTRU) comprising: a processor configured to at least: receive a first message from a network node, wherein the first message indicates a request to identify an instance of an object; obtain a model associated with the object; identify the instance of the object in a video frame by applying the model to the video frame; send a second message to the network node when the instance of the object has been identified; receive a third message from the network node, wherein the third message indicates virtual content associated with the object; modify the video frame by overlaying the virtual content onto the video frame; and display the modified video frame to a user.
2. The WTRU of claim 1 , wherein the processor is further configured to determine a position of the identified instance of the object in the video frame, and wherein the virtual content is overlaid onto the video frame at the position of the identified instance of the object.
3. The WTRU of claim 1 , wherein the model is obtained based on position information associated with the WTRU and position information associated with the instance of the object.
4. The WTRU of claim 1, wherein the second message comprises at least one of an identifier of the object, a location where the object is recognized in the video frame, or a location where the object is recognized in the video frame.
5. The WTRU of claim 1 , wherein the virtual content comprises at least one of a bounding box or a class label identifying the bounding box.
6. The WTRU of claim 1, wherein the processor is further configured to: send a report to the network node, wherein the report comprises at least one of information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRU orientation information.
7. The WTRU of claim 1, wherein the first message further comprises at least one of WTRU location information, WTRU orientation information, or a filter value.
8. The WTRU of claim 1 , wherein the WTRU is at least one of a smartphone, a tablet, a wearable device, a head mounted display, a connected vehicle, or a drone.
9. A method comprising: receiving a first message from a network node, wherein the first message indicates a request to identify an instance of an object; obtaining a model associated with the object; identifying the instance of the object in a video frame by applying the model to the video frame; sending a second message to the network node when the instance of the object has been identified; receiving a third message from the network node, wherein the third message indicates virtual content associated with the object; modifying the video frame by overlaying the virtual content onto the video frame; and displaying the modified video frame to a user.
10. The method of claim 9, wherein the method further comprises determining a position of the identified instance of the object in the video frame, and wherein the virtual content is overlaid onto the video frame at the position of the identified instance of the object.
11. The method of claim 9, wherein the model is obtained based on position information associated with a wireless transmit / receive unit (WTRU) and position information associated with the instance of the object.
12. The method of claim 9, wherein the second message comprises at least one of an identifier of the object, a location where the object is recognized in the video frame, or a location where the object is recognized in the video frame.
13. The method of claim 9, wherein the virtual content comprises at least one of a bounding box or a class label identifying the bounding box.
14. The method of claim 9, further comprising: sending a report to the network node, wherein the report comprises at least one of information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRU orientation information.
15. The method of claim 9, wherein the first message further comprises at least one of WTRU location information, WTRU orientation information, or a filter value.
16. A network node comprising: a processor configured to at least: determine an object associated with a wireless transmit / receive unit (WTRU); send a first message to the WTRU, wherein the first message indicates a request to identify an instance of the object in a video frame based on an application of a model to the video frame; receive a second message from the WTRU based on the instance of the object being identified by the WTRU; and send a third message to the WTRU, wherein the third message indicates virtual content associated with the object, wherein the virtual content is to be overlaid onto the video frame.
17. The network node of claim 16, wherein the second message comprises at least one of an identifier of the object, a location where the object is recognized in the video frame, or a location where the object is recognized in the video frame.
18. The network node of claim 16, wherein the processor is further configured to: receive a report from the WTRU, wherein the report comprises at least one of: information associated with the instance of the object, a WTRU identifier, WTRU location information, or WTRU orientation information.
19. The network node of claim 16, wherein the first message further comprises at least one of: WTRU location information, WTRU orientation information, or a filter value.
20. The network node of claim 16, wherein the processor is further configured to: obtain a model associated with the determined object, wherein the first message sent to the WTRU further indicates that the model to be applied to the video frame is the model associated with the determined object.
Citation Information
Patent Citations
Augmented reality data dissemination method, system and terminal and storage medium
US20200372687A1
Connecting spatial anchors for augmented reality
US20210350612A1
Methods and systems of combining video content with one or more augmentations to produce augmented video
US20220327830A1
Methods, architectures, apparatuses and systems for local mobile metaverse wireless / transmit receive unit service enablers
WO2024030359A1
Enabling sensing and sensing fusion for a metaverse service in a wireless communication system
WO2024088584A1