Apparatus, methods and computer programs
By transmitting messages between user equipment and using video codecs and virtual avatar codecs, the implementation difficulties of virtual avatar call sessions in the prior art are solved, enabling efficient virtual avatar call establishment and modification, and improving the quality of virtual avatar calls between user equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-10-16
- Publication Date
- 2026-05-26
Smart Images

Figure CN122095618A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This patent application claims the priority benefit of Indian Patent Application No. 202341073361, filed on October 27, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure relates to an apparatus, method, and computer program for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between a first user equipment and a second user equipment. Background Technology
[0003] A communication network (i.e., a communication system) can be viewed as a facility that enables a communication session between two or more entities (such as communication devices, base stations, and / or other nodes) by providing carriers between various entities involved in the communication path.
[0004] The communication network can be a wireless communication network. Examples of wireless systems include Public Land Mobile Networks (PLMNs) operating based on radio standards provided by 3GPP, satellite-based communication networks, and various wireless local networks such as Wireless Local Area Networks (WLANs). Wireless systems can typically be divided into cells and are therefore often referred to as cellular systems.
[0005] The communication network and associated equipment typically operate according to a given standard or specification that defines what the various entities associated with the system are allowed to do and how they should be implemented. Communication protocols and / or parameters applied to the connection are also typically defined. An example of such a standard is the so-called 5G standard. Summary of the Invention
[0006] According to one aspect, a method is provided for a first user equipment, comprising: transmitting to a second user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; generating a first video media including a video of a user; converting the first video media including the video of the user into a second video media including a video of a avatar representation of the user, or into avatar media including a avatar representation of the user and the user's pose within the video of the user; and transmitting the second video media or the avatar media to the second user equipment.
[0007] Converting the first video media into the second video media may include: decoding the first video media using a video codec to obtain the user's video; replacing the user in the user's video with a virtual avatar representation of the user to obtain a video of the user's virtual avatar representation; and encoding the video of the user's virtual avatar representation using the video codec to obtain the second video media.
[0008] Converting the first video media into virtual avatar media may include: decoding the first video media using a video codec to obtain the user's video; determining the user's posture within the user's video; and encoding the user's posture and the user's virtual avatar representation using a virtual avatar codec to obtain the virtual avatar media.
[0009] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a codec identifier that identifies a video codec or an avatar codec; and the method may include: receiving from the second user equipment a message indicating acceptance of the codec identifier; and converting the first video media into the second video media or the avatar media using the video codec or the avatar codec identified by the codec identifier.
[0010] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a avatar identifier that identifies the avatar representation of the user; and the method may include: receiving from the second user equipment a message indicating acceptance of the avatar identifier; and converting the first video media into the second video media or the avatar media using the avatar representation identified by the avatar identifier.
[0011] The virtual avatar identifier may include integers, strings, universally unique identifiers, uniform resource names, or uniform resource locators.
[0012] The method may include: retrieving from the memory of the first user device a virtual avatar representation of the user identified by the virtual avatar identifier.
[0013] The method may include: receiving a virtual avatar representation of the user identified by the virtual avatar identifier from a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem (IPMS) network.
[0014] The virtual avatar representation of the user, identified by the virtual avatar identifier, can be received using the Internet Protocol Multimedia Subsystem data channel.
[0015] The messages used to establish or modify an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment may be transmitted using an IPS data channel or using Hypertext Transfer Protocol.
[0016] The messages used to establish or modify an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment can use the session. The second video media or the virtual avatar media can be transmitted to the second user device using a real-time protocol.
[0017] According to one aspect, an apparatus for a first user equipment is provided, comprising: means for transmitting to a second user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; means for generating a first video media including a video of a user; means for converting the first video media including the video of the user into a second video media including a video representation of the user, or into avatar media including a avatar representation of the user and the user's pose within the video of the user; and means for transmitting the second video media or the avatar media to the second user equipment.
[0018] Converting the first video media into the second video media may include: decoding the first video media using a video codec to obtain the user's video; replacing the user in the user's video with a virtual avatar representation of the user to obtain a video of the user's virtual avatar representation; and encoding the video of the user's virtual avatar representation using the video codec to obtain the second video media.
[0019] Converting the first video media into virtual avatar media may include: decoding the first video media using a video codec to obtain the user's video; determining the user's posture within the user's video; and encoding the user's posture and the user's virtual avatar representation using a virtual avatar codec to obtain the virtual avatar media.
[0020] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a codec identifier that identifies a video codec or an avatar codec; and the apparatus may include: components for receiving from the second user equipment a message indicating acceptance of the codec identifier; and components for converting the first video media into the second video media or the avatar media using the video codec or the avatar codec identified by the codec identifier.
[0021] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a avatar identifier that identifies the avatar representation of the user; and the apparatus may include: components for receiving from the second user equipment a message indicating acceptance of the avatar identifier; and components for converting the first video media into the second video media or the avatar media using the avatar representation identified by the avatar identifier.
[0022] The virtual avatar identifier may include integers, strings, universally unique identifiers, uniform resource names, or uniform resource locators.
[0023] The apparatus may include a component for retrieving from the memory of the first user device a virtual avatar representation of the user identified by the virtual avatar identifier.
[0024] The apparatus may include a component for receiving a virtual avatar representation of the user identified by the virtual avatar identifier, either from a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem (IPMS) network.
[0025] The virtual avatar representation of the user, identified by the virtual avatar identifier, can be received using the Internet Protocol Multimedia Subsystem data channel.
[0026] The messages used to establish or modify an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment may be transmitted using an IPS data channel or using Hypertext Transfer Protocol.
[0027] Messages used to establish or modify an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment can be transmitted using the Session Description Protocol.
[0028] The second video media or the virtual avatar media can be transmitted to the second user device using a real-time protocol.
[0029] According to one aspect, an apparatus for a first user equipment is provided, the apparatus comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the apparatus to at least: transmit to a second user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; generate a first video media including a video of a user; convert the first video media including the video of the user into a second video media including a video of a avatar representation of the user, or into avatar media including a avatar representation of the user and the user's pose within the video of the user; and transmit the second video media or the avatar media to the second user equipment.
[0030] According to one aspect, an apparatus for a first user equipment is provided, the apparatus comprising circuitry configured to perform the following operations: transmitting to a second user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; generating a first video media including a video of a user; converting the first video media including the video of the user into a second video media including a video of a avatar representation of the user, or into avatar media including a avatar representation of the user and the user's pose within the video of the user; and transmitting the second video media or the avatar media to the second user equipment.
[0031] According to one aspect, a computer program for a first user equipment is provided, the computer program including computer-executable code, the computer-executable code being configured, when run on at least one processor, to perform: transmitting to a second user equipment messages for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; generating a first video media including a video of a user; converting the first video media including the video of the user into a second video media including a video of a avatar representation of the user, or into avatar media including a avatar representation of the user and the user's pose within the video of the user; and transmitting the second video media or the avatar media to the second user equipment.
[0032] According to one aspect, a method for an Internet Protocol Multimedia Subsystem (IPS) is provided, the method comprising: receiving from a first user equipment a first video medium including a video of a user or a first virtual avatar medium including the user's pose within the video of the user; converting the first video medium including the user's video or the first virtual avatar medium including the user's pose within the video of the user into a second video medium including a video of a virtual avatar representation of the user, or converting it into a second virtual avatar medium including a virtual avatar representation of the user and the user's pose within the video of the user; and transmitting the second video medium or the second virtual avatar medium to a second user equipment.
[0033] The method may include receiving a virtual avatar representation of the user from the first user equipment, a network function belonging to the Internet Protocol Multimedia Subsystem (IPMS) network, or a network function not belonging to the IMS network.
[0034] The virtual avatar can be received using the Internet Protocol Multimedia Subsystem Data Channel or the Hypertext Transfer Protocol.
[0035] Converting the first video media into the second video media may include: decoding the first video media using a video codec to obtain the user's video; replacing the user with a virtual avatar representation of the user in the user's video to obtain a video of the user's virtual avatar representation; and encoding the video of the user's virtual avatar representation using the video codec to obtain the second video media.
[0036] Converting the first virtual avatar media into the second video media may include: decoding the first virtual avatar media, which includes the user's pose within a video, using a virtual avatar codec; animate the user's virtual avatar representation using the user's pose within the video to obtain a video of the user's virtual avatar representation; and encoding the video of the user's virtual avatar representation using the video codec to obtain the second video media.
[0037] Converting the first video media into the second virtual avatar media may include: decoding the first video media using a video codec to obtain the user's video; determining the user's posture within the user's video; and encoding the user's posture and the user's virtual avatar representation using a virtual avatar codec to obtain the second virtual avatar media.
[0038] Converting the first virtual avatar media into the second virtual avatar media may include: using a virtual avatar codec to decode the first virtual avatar media containing the user's pose within a video; and using the virtual avatar codec to encode the user's pose and the user's virtual avatar representation to obtain the second virtual avatar media.
[0039] The method may include: receiving a message indicating a codec identifier for a video codec or a avatar codec from a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem Network (IPMS); and using the video codec or the avatar codec identified by the codec identifier to convert the first video media or the first avatar media into the second video media or the second avatar media.
[0040] The method may include: receiving a virtual avatar representation of the user from the first user equipment, a network function belonging to the Internet Protocol Multimedia Subsystem Network (IPMS) network, or a network function not belonging to the IMS network; and using the virtual avatar representation to convert the first video media or the first virtual avatar media into the second video media or the second virtual avatar media.
[0041] The method may include: transmitting the second video media or the second virtual avatar media to the first user equipment.
[0042] The second video media or the second virtual avatar media may be transmitted to at least one of the first user equipment or the second user equipment using a real-time protocol.
[0043] The user's posture and virtual avatar representation within the user's video can be transmitted via the Internet Protocol Multimedia Subsystem data channel.
[0044] The method can be executed by media functions or media resource functions.
[0045] According to one aspect, an apparatus for network functions within an Internet Protocol Multimedia Subsystem is provided, the apparatus comprising: components for receiving from a first user equipment a first video medium including a video of a user or a first virtual avatar medium including the user's pose within the video of the user; components for converting the first video medium including the user's video or the first virtual avatar medium including the user's pose within the video of the user into a second video medium including a virtual avatar representation of the user or into a second virtual avatar medium including both the user's virtual avatar representation and the user's pose within the video of the user; and components for transmitting the second video medium or the second virtual avatar medium to a second user equipment.
[0046] The apparatus may include a component for receiving a virtual avatar representation of the user from the first user equipment, a network function belonging to the Internet Protocol Multimedia Subsystem (IPMS) network, or a network function not belonging to the IMS network.
[0047] The virtual avatar can be received using the Internet Protocol Multimedia Subsystem Data Channel or the Hypertext Transfer Protocol.
[0048] Converting the first video media into the second video media may include: decoding the first video media using a video codec to obtain the user's video; replacing the user with a virtual avatar representation of the user in the user's video to obtain a video of the user's virtual avatar representation; and encoding the video of the user's virtual avatar representation using the video codec to obtain the second video media.
[0049] Converting the first virtual avatar media into the second video media may include: decoding the first virtual avatar media, which includes the user's pose within a video, using a virtual avatar codec; animate the user's virtual avatar representation using the user's pose within the video to obtain a video of the user's virtual avatar representation; and encoding the video of the user's virtual avatar representation using the video codec to obtain the second video media.
[0050] Converting the first video media into the second virtual avatar media may include: decoding the first video media using a video codec to obtain the user's video; determining the user's posture within the user's video; and encoding the user's posture and the user's virtual avatar representation using a virtual avatar codec to obtain the second virtual avatar media.
[0051] Converting the first virtual avatar media into the second virtual avatar media may include: using a virtual avatar codec to decode the first virtual avatar media containing the user's pose within a video; and using the virtual avatar codec to encode the user's pose and the user's virtual avatar representation to obtain the second virtual avatar media.
[0052] The apparatus may include: components for receiving a message indicating a codec identifier of a video codec or a avatar codec from a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem Network (IPMS); and components for converting the first video media or the first avatar media into the second video media or the second avatar media using the video codec or the avatar codec identified by the codec identifier.
[0053] The apparatus may include: components for receiving a virtual avatar representation of the user from the first user equipment, a network function belonging to the Internet Protocol Multimedia Subsystem (IPS) network, or a network function not belonging to the IPS network; and components for converting the first video media or the first virtual avatar media into the second video media or the second virtual avatar media using the virtual avatar representation.
[0054] The device may include a component for transmitting the second video media or the second virtual avatar media to the first user equipment.
[0055] The second video media or the second virtual avatar media may be transmitted to at least one of the first user equipment or the second user equipment using a real-time protocol.
[0056] The user's posture and virtual avatar representation within the user's video can be transmitted via the Internet Protocol Multimedia Subsystem data channel.
[0057] The device may include media functions or media resource functions.
[0058] According to one aspect, an apparatus for network functions within an Internet Protocol Multimedia Subsystem is provided, the apparatus comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the apparatus to at least: receive from a first user equipment a first video medium comprising a video of a user or a first virtual avatar medium comprising the user's pose within the video of the user; convert the first video medium comprising the user's video or the first virtual avatar medium comprising the user's pose within the video of the user into a second video medium comprising a video comprising a virtual avatar representation of the user, or converting it into a second virtual avatar medium comprising a virtual avatar representation of the user and the user's pose within the video of the user; and transmit the second video medium or the second virtual avatar medium to a second user equipment.
[0059] According to one aspect, an apparatus for network functions within an Internet Protocol Multimedia Subsystem is provided, comprising circuitry configured to perform the following operations: receiving from a first user equipment a first video medium including a video of a user or a first virtual avatar medium including the user's pose within the video of the user; converting the first video medium including the user's video or the first virtual avatar medium including the user's pose within the video of the user into a second video medium including a video of a virtual avatar representation of the user or a second virtual avatar medium including both the user's virtual avatar representation and the user's pose within the video of the user; and transmitting the second video medium or the second virtual avatar medium to a second user equipment.
[0060] According to one aspect, a computer program is provided for network functions within an Internet Protocol Multimedia Subsystem, comprising computer-executable code configured, when executed on at least one processor, to perform the following operations: receiving from a first user equipment a first video medium including a video of a user or a first virtual avatar medium including the user's pose within the video; converting the first video medium including the user's video or the first virtual avatar medium including the user's pose within the video into a second video medium including a video representing the user's virtual avatar or a second virtual avatar medium including the user's virtual avatar representation and the user's pose within the video; and transmitting the second video medium or the second virtual avatar medium to a second user equipment.
[0061] According to one aspect, a method for a second user equipment is provided, comprising: receiving from a first user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; receiving from the first user equipment avatar media including the user's pose within a video of the user; decoding the avatar media to obtain the user's pose within the video of the user; animate the avatar representation using the user's pose within the video of the user to obtain a video of the user's avatar representation; and displaying the video of the user's avatar representation.
[0062] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a codec identifier that identifies the avatar codec; and the method may include: transmitting to the first user equipment a message indicating acceptance of the codec identifier; and decoding the avatar media using the avatar codec identified by the codec identifier to obtain the user's pose within the user's video.
[0063] The message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment may include a avatar identifier that identifies the avatar representation of the user; and the method may include: transmitting a message to the first user equipment indicating acceptance of the avatar identifier; transmitting a message including the avatar identifier that identifies the avatar representation of the user to a network function belonging to or not belonging to the IPS network; and receiving the avatar representation of the user from a network function belonging to or not belonging to the IPS network.
[0064] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a avatar identifier that identifies the avatar representation of the user; and the method may include: transmitting a message to the first user equipment indicating acceptance of the avatar identifier; and retrieving the avatar representation of the user identified by the avatar identifier from the memory of the second user equipment.
[0065] The messages used to establish or modify an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment can be transmitted using an IPS data channel.
[0066] The messages used to establish or modify an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment may be transmitted using a session description protocol.
[0067] The virtual avatar media can be received from the first user device using a real-time protocol.
[0068] The virtual avatar media, including the user's posture within the user's video, can be received via the Internet Protocol Multimedia Subsystem data channel.
[0069] According to one aspect, an apparatus for a second user equipment is provided, comprising: means for receiving from a first user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; means for receiving from the first user equipment avatar media including the user's pose within a video of the user; means for decoding the avatar media to obtain the user's pose within the video of the user; means for animateing the avatar representation using the user's pose within the video of the user to obtain a video of the user's avatar representation; and means for displaying the video of the user's avatar representation.
[0070] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a codec identifier that identifies the avatar codec; and the apparatus may include: components for transmitting to the first user equipment a message indicating acceptance of the codec identifier; and components for decoding the avatar media using the avatar codec identified by the codec identifier to obtain the user's pose in the user's video.
[0071] The message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment may include a avatar identifier that identifies the avatar representation of the user; and the apparatus may include: components for transmitting a message to the first user equipment indicating acceptance of the avatar identifier; components for transmitting a message including the avatar identifier that identifies the avatar representation of the user to a network function belonging to or not belonging to the IPS network; and components for receiving the avatar representation of the user from a network function belonging to or not belonging to the IPS network.
[0072] The message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment may include a avatar identifier that identifies the avatar representation of the user; and the means may include: transmitting a message to the first user equipment indicating acceptance of the avatar identifier; and retrieving the avatar representation of the user identified by the avatar identifier from the memory of the second user equipment.
[0073] The messages used to establish or modify an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment can be transmitted using an IPS data channel.
[0074] The messages used to establish or modify an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment can be transmitted using a session description protocol.
[0075] The virtual avatar media can be received from the first user device using a real-time protocol.
[0076] The virtual avatar media, including the user's posture within the user's video, can be received via the Internet Protocol Multimedia Subsystem data channel.
[0077] According to one aspect, an apparatus for a second user equipment is provided, comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the apparatus to perform at least the following operations: receiving from a first user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; receiving from the first user equipment avatar media, the avatar media including the user's pose within a video of the user; decoding the avatar media to obtain the user's pose within the video of the user; animate the avatar representation using the user's pose within the video of the user to obtain a video of the user's avatar representation; and displaying the video of the user's avatar representation.
[0078] According to one aspect, an apparatus for a second user equipment is provided, comprising circuitry configured to perform the following operations: receiving from a first user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; receiving from the first user equipment avatar media including the user's pose within a video of the user; decoding the avatar media to obtain the user's pose within the video of the user; animate the avatar representation using the user's pose within the video of the user to obtain a video of the user's avatar representation; and displaying the video of the user's avatar representation.
[0079] According to one aspect, a computer program for a second user equipment is provided, comprising computer-executable code configured, when executed on at least one processor, to perform the following operations: receiving from a first user equipment a message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for avatar calling between the first user equipment and the second user equipment; receiving from the first user equipment avatar media including the user's pose within a video of the user; decoding the avatar media to obtain the user's pose within the video of the user; animateing the avatar representation using the user's pose within the video of the user to obtain a video of the user's avatar representation; and displaying the video of the user's avatar representation.
[0080] According to one aspect, a computer-readable medium is provided, including program instructions stored thereon for performing at least one of the methods described above.
[0081] According to one aspect, a non-transitory computer-readable medium is provided, comprising program instructions stored thereon for performing at least one of the methods described above.
[0082] According to one aspect, a non-volatile tangible storage medium is provided, including program instructions stored thereon for performing at least one of the methods described above.
[0083] Many different aspects have been described above. It should be understood that other aspects can be provided through any combination of two or more of the aspects mentioned above.
[0084] Various other aspects are also described in the following detailed description and the appended claims.
[0085] List of abbreviations AF: Application Functions AMF: Access and Mobility Management Functions AR: Augmented Reality BS: Base Station CU: Centralized Unit DAC: Digital Asset Container DL: Downlink DTLS: Datagram Transport Layer Security DU: Distributed Unit gLTF: Graphics Language Transmission Format gNB: gNodeB HSS: Home Subscriber Server IBCF: Interconnection Border Control Function I-CSCF: Query Interrogating Call Session Control Function IMS: Internet Protocol Multimedia Subsystem IMS AGW: Internet Protocol Multimedia Subsystem Application Level Gateway IMS HSS: Internet Protocol Multimedia Subsystem Home Subscriber Server LMF: Location Management Function LPP: Location Positioning Protocol LRDD: Link-based Resource Descriptor Document LTE: Long Term Evolution MAC: Medium Access Control ML: Machine Learning MS: Mobile Station MTC: Machine Type Communication NEF: Network Exposure Function NF: Network Function NR: New Radio NRF: Network Repository Function NWDAF: Network Data Analytics Function P-CSCF: Proxy Call Session Control Function PDU: Protocol Data Unit RAM: Random Access Memory (R)AN: (Radio) Access Network ROM: Read Only Memory RTP: Real-Time Protocol S-CSCF: Serving Call Session Control Function SDP: Session Description Protocol SIP: Session Initiation Protocol SMF: Session Management Function TrGW: Transition Gateway TS: Technical Specification UE: User Equipment URI: Uniform Resource Identifier URL: Uniform Resource Locator URN: Uniform Resource Name UUID: Universal Unique Identifier XR: Extended Reality 3GPP: 3rd Generation Partnership Project 5G: 5th Generation 5GC: 5G Core network 5GS: 5G System. Attached Figure Description
[0086] The embodiments will now be described by way of example only, with reference to the accompanying drawings, wherein: Figure 1A schematic diagram illustrating communication between an originating user equipment and a terminating user equipment via a communication network according to an example embodiment is shown, the communication network including two Internet Protocol Multimedia Subsystem networks; Figure 2 A schematic diagram of an apparatus for implementing one or more network functions of an Internet Protocol Multimedia Subsystem (IPS) network is shown. Figure 3 A schematic diagram of the user equipment is shown; Figure 4 A schematic diagram of a rendering based on the originating user equipment (OUE) is shown for a virtual avatar call between an OUE and a terminating user equipment (TTU). Figure 5 This diagram illustrates another rendering based on the originating user equipment (OUE) for a virtual avatar call between the originating and terminating OUEs. Figure 6 A schematic diagram of a rendering based on the terminating user equipment is shown for a virtual avatar call between the originating user equipment and the terminating user equipment. Figure 7 The process of web-based rendering for virtual avatar calls between originating and terminating user equipment is illustrated. Figure 8 Another process for web-based rendering of generic virtual avatar calls between originating and terminating user equipment is shown. Figure 9 The process of rendering based on the originating user equipment (OUE) for performing a virtual avatar call between the originating user equipment (OUE) and the terminating user equipment (TERM) is illustrated. Figure 10 Another process for rendering based on the originating user equipment (OUE) is shown for virtual avatar calls between the originating and terminating user equipment (TTU). Figure 11 A block diagram showing the rendering based on the originating user equipment (OUE) for virtual avatar calls between the originating and terminating OUEs is shown. Figure 12 A block diagram showing network-based rendering for virtual avatar calls between originating and terminating user equipment is presented. Figure 13 A block diagram showing a rendering based on the terminating user equipment (TUE) for virtual avatar calls between originating and terminating TUEs is illustrated; and Figure 14 A schematic diagram is shown of a non-volatile storage medium for storing instructions, which, when executed by a processor, allow the processor to perform... Figures 11 to 13 One or more steps of the method. Detailed Implementation
[0087] Figure 1 A schematic diagram illustrating communication between a first user equipment (UE-A) and a second user equipment is shown, including a communication network comprising two IMS networks. The communication network may include a Public Land Mobile Network (PLMN). The communication network may include a (Radio) Access Network ((R)AN) (not shown), a core network (5GC) (not shown), and the two Internet Protocol Multimedia Subsystem (IMS) networks shown. The IMS network serving the first user equipment (e.g., UE-A) may be referred to as the originating IMS network, while the IMS network serving the second user equipment (e.g., UE-B) may be referred to as the terminating IMS network.
[0088] In this disclosure, the terms "Internet Protocol Multimedia Subsystem" and "Internet Protocol Multimedia Subsystem Network" are used interchangeably.
[0089] The (R)AN (not shown) may include one or more base stations. The base station may be an evolved NodeB (eNB) or a gNodeB (gNB). The gNB may include distributed units connected to one or more centralized units. The base station may be configured to provide a wireless connection between the first user equipment and the base station, and to provide wired and / or wireless connections between the base station and the IMS network, and between the second user equipment and the 5GC, via the 5GC. The 5GC may include Access and Mobility Management Function (AMF), Session Management Function (SMF), Authentication Server Function (AUSF), User Data Management (UDM), Network Openness Function (NEF), Unified Data Repository (UDR), Network Repository Function (NRF), and / or Network Data Analysis Function (NWDAF), as well as a user plane including User Plane Function (UPF). The 5GC may be connected to an IMS network (e.g., the originating IMS network or the terminating IMS network). The 5GC may be used to manage services and routing from the UE to the IMS network via the (R)AN (e.g., from the first UE to the originating IMS network or the second UE to the terminating IMS network).
[0090] An IMS network (e.g., the originating IMS network or the terminating IMS network) may include a proxy call session control function (P-CSCF), a service call session control function (S-CSCF), an inquiry call session control function (I-CSCF), an IMS application layer gateway (IMS-AGW), an IMS home subscriber server (IMS HSS), an IMS application server (IMS AS), a media function (MF) or a media resource function (MRF), a transition gateway (TrGW), an interconnection boundary control function (IBCF), a terminating IMS core, a data channel signaling function (DCSF), an XR application server (AS), a digital asset container (DAC), or a network server (not shown).
[0091] The DAC can be separate from other network functions of the originating IMS network, or it can be integrated with other network functions of the originating IMS network. For example, the DAC can be integrated into the IMS HSS or the XR AS.
[0092] Despite Figure 1 The DAC described herein is shown as part of the originating IMS network, but it should be understood that the DAC is not necessarily part of the originating IMS network, let alone the communication network (e.g., PLMN). The DAC may be part of the 5GC. The DAC may be separate from or integrated with other network functions of the 5GC. For example, the DAC may be integrated within the UDM. Alternatively, the DAC may not be part of the communication network (e.g., PLMN). For example, the DAC may be part of a network server or a cloud provider's computing system connected to the communication network (e.g., PLMN).
[0093] Figure 2 It shows the use of Figure 1An example of the IMS device 200 shown is illustrated. The device 200 may include at least one random access memory (RAM) 211a, at least one read-only memory (ROM) 211b, at least one processor 212, 213, and an input / output interface 214. The at least one processor 212, 213 may be coupled to the RAM 211a and the ROM 211b. The at least one processor 212, 213 may be configured to execute appropriate software code 215. The software code 215 may, for example, be software code for the network functions of the IMS, which, when executed by the at least one processor 212, 213, causes the device 200 to perform one or more operations of the network functions, including those operations described in further detail below. The software code 215 may be stored in the ROM 211b. The device 200 may interconnect with another device 200 that implements or includes another network function of the IMS network.
[0094] Figure 3 An example of UE 300 is shown, for example Figure 1 The UE 300 is shown as UE-A or UE-B. The UE 300 is a communication device capable of (or configured to) transmit and receive radio signals. Non-limiting examples include mobile stations (MS) or mobile devices, such as mobile phones or so-called "smartphones," computers equipped with wireless interface cards or other wireless interface facilities (e.g., USB dongles), personal data assistants (PDAs) or tablets equipped with wireless communication capabilities, machine-type communication (MTC) devices, cellular Internet of Things (CIoT) devices, or any combination of these devices. The UE 300 may, for example, provide data communication for carrying communication. The communication may be one or more of voice, email, text messages, multimedia, data, machine data, etc.
[0095] The UE 300 can receive signals via air or radio interface 307 through appropriate means for receiving, and can transmit signals via appropriate means for transmitting radio signals. Figure 3 In the diagram, the transceiver device is schematically represented by block 306. The transceiver device 306 may be provided, for example, via a radio section and an associated antenna arrangement. The antenna arrangement may be located inside or outside the mobile device.
[0096] The UE 300 may be equipped with at least one processor 301, at least one memory ROM 302a, at least one RAM 302b, and other possible components 303 for performing tasks designed to be performed with software and hardware assistance, including controlling access to and communication with access systems and other communication devices. The at least one processor 301 is coupled to the RAM 302b and the ROM 302a. The at least one processor 301 may be configured to execute appropriate software code 308. The software code 308 may, for example, allow the execution of one or more aspects of this disclosure. The software code 308 may be stored in the ROM 302a.
[0097] The processor, memory, and other related control devices can be housed on a suitable circuit board and / or in a chipset. This feature is indicated by reference numeral 304. The device may optionally have a user interface, such as a keyboard 305, a touch-sensitive screen or touchpad, or a combination thereof. Optionally, depending on the type of device, one or more of a display, speaker, and microphone may be provided.
[0098] One or more aspects of this disclosure can establish or modify an IMS session for virtual avatar calls between two communication devices (e.g., a first communication device denoted as UE-A and a second communication device denoted as UE-B).
[0099] In this disclosure, UE-A may be referred to as the first UE or the originating UE, and UE-B may be referred to as the second UE or the terminating UE.
[0100] The virtual avatar call between UE-A and UE-B can be a call in which a user of UE-A is represented by a virtual avatar in the user interface of UE-B, and another user of UE-B is represented by a virtual avatar in the user interface of UE-A.
[0101] In this disclosure, the terms "user" and "subscriber" are used interchangeably.
[0102] In this disclosure, the terms "virtual avatar representation" and "virtual avatar" are used interchangeably.
[0103] The user's virtual avatar representation may include a two-dimensional or three-dimensional representation of the user's face, part of the body, or the entire body. The user's virtual avatar representation may be stored in the DAC.
[0104] A user's avatar representation can be identified by an avatar identifier. The avatar identifier can include a reference or link to the avatar representation. The avatar identifier can include an integer, a string, a universally unique identifier (UUID), or a uniform resource identifier (URI). The URL can include a uniform resource name (URN) or a uniform resource locator (URL). When the avatar identifier is a URI, it can be a dereferenceable URI so that the avatar representation can be easily retrieved.
[0105] Various techniques can be used to retrieve a user's virtual avatar representation (e.g., by the UE-A, the UE-B, or the originating IMS network), as described in further detail below. For example, webfinger, host-meta, or link-based resource descriptor discovery (LRDD) can be used to retrieve the virtual avatar representation.
[0106] A user's virtual avatar representation can be associated with the user or the user's user identity. A user can have multiple user identities. For example, a user can have a personal user identity and a professional user identity. That is, a user can be associated with multiple user identities and multiple virtual avatar representations of the user identified by multiple virtual avatar identifiers.
[0107] A user's virtual avatar representation can be associated with metadata. This metadata may include format, resolution, version, access or usage permissions (e.g., number of uses), the creation date of the virtual avatar representation, or a license to use the virtual avatar representation. The format may include Graphical Language Transfer Format (gLTF) or Virtual Reality Modeling (VRM).
[0108] A user's virtual avatar representation can be associated with a human-readable description or name. This description or name allows a user of UE-A or a user of UE-B to recognize the virtual avatar representation of UE-A or UE-B.
[0109] One or more aspects of this disclosure enable a first UE (e.g., UE-A) or UE-B to initiate a virtual avatar call with a second UE (e.g., UE-B or UE-A).
[0110] One or more aspects of this disclosure enable a first UE (e.g., UE-A), the originating IMS network, or the second UE (e.g., UE-B) to retrieve a virtual avatar identifier that identifies a user's virtual avatar representation.
[0111] One or more aspects of this disclosure enable the first UE (e.g., UE-A), the originating IMS network, or the second UE (e.g., UE-B) to provide a virtual avatar representation of the user.
[0112] One or more aspects of this disclosure enable the first UE (e.g., UE-A) and the second UE (e.g., UE-B) to negotiate, together with the originating and terminating IMS networks, a virtual avatar representation for the user for the virtual avatar call.
[0113] One or more aspects of this disclosure enable the first UE (e.g., UE-A) and the second UE (e.g., UE-B) to negotiate codecs for virtual avatar calls together with the originating and terminating IMS networks.
[0114] One or more aspects of this disclosure enable the first UE (e.g., UE-A), the originating IMS network, the terminating IMS network, or the second UE (e.g., UE-B) to perform rendering.
[0115] In this disclosure, the term "rendering" may refer to a video that provides a virtual representation of a user or to providing the user's pose within the video and the virtual representation of the user.
[0116] One or more aspects of this disclosure enable the first and second UEs to switch from virtual avatar calls to audio and / or video calls or vice versa.
[0117] The following describes various implementations of a virtual avatar representation and codec for negotiating a virtual avatar representation and codec for virtual avatar calls between a first UE (e.g., UE-A) and UE-B via a communication network including originating and terminating IMS networks.
[0118] Initially, the UE-A may or may not store one or more virtual avatar representations identifying the user in its memory (e.g., local cache). The UE-A may or may not store the user's one or more virtual avatar representations in its memory.
[0119] If the UE-A stores in its memory one or more virtual avatar representations that identify the user, the UE-A may select a virtual avatar identifier.
[0120] If UE-A does not store one or more virtual avatar representations identifying the user in its memory, UE-A may receive virtual avatar representations identifying the user from the originating IMS network (e.g., IMS AS).
[0121] UE-A can send a message to the originating IMS network to register UE-A with the originating IMS network.
[0122] UE-A can send messages to UE-B via the originating IMS network (e.g., IMS AS) and the terminating IMS network to establish or modify an IMS session between UE-A and UE-B for virtual avatar calls between UE-A and UE-B.
[0123] An IMS session can refer to an IMS data channel, which is established between two communication devices via an originating and terminating IMS network and can be used for audio calls, video calls, and / or avatar calls.
[0124] The message may include a virtual avatar identifier that identifies a user's virtual avatar representation. The virtual avatar identifier may indicate the intent of UE-A to establish or modify an IMS session between UE-A and UE-B for virtual avatar calls. The virtual avatar identifier may indicate a suggestion to use the user's virtual avatar representation (by UE-A and / or UE-B).
[0125] The message may include a codec identifier that identifies a video codec or a avatar codec. The codec identifier may indicate a suggestion to use a video codec or a avatar codec for decoding video media or avatar media (by UE-A and / or UE-B) before displaying video representing the user's avatar. The video codec or avatar codec identified by the codec identifier may be based on UE-A capabilities, UE-B capabilities, or XR application preferences or requirements originating from and / or terminating IMS networks (e.g., the originating and / or terminating IMS network prefers or requires the video codec and / or avatar codec).
[0126] The message can be sent using an IMS data channel, SIP, or SDP. The message may include a SIP message or SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the avatar identifier may be sent as a SIP parameter or an SDP parameter.
[0127] The IMS data channel may include one or more IMS application data channels. For example, the IMS data channel may include a first IMS application channel between UE-A and IMS (e.g., IMS AS or XR AS) and a second IMS application channel between UE-B and IMS (e.g., IMS AS or XR AS).
[0128] In this disclosure, the terms "IMS application data channel", "IMS peering to application data channel" or "IMS application to peer data channel" are used interchangeably.
[0129] UE-A can receive a message from UE-B indicating acceptance of the virtual avatar identifier and the codec identifier via the originating IMS network (e.g., IMS AS) and the terminating IMS network. The message can be transmitted using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters.
[0130] I - Rendering Implementation Based on the Originating User Interface (UE) In a rendering implementation based on the originating UE, the originating UE (e.g., UE-A) can perform rendering.
[0131] If the originating UE (e.g., UE-A) stores in its memory a virtual avatar representation of the user of the originating UE (e.g., UE-A) (selected by the user of the originating UE (e.g., UE-A), as described in detail above), then UE-A can retrieve the virtual avatar representation of the user.
[0132] If the originating UE (e.g., UE-A) does not store the user's virtual avatar representation in its memory, the originating UE (e.g., UE-A) can retrieve the user's virtual avatar representation from the network functions of the originating IMS network (e.g., XR AS of the originating IMS network).
[0133] The originating UE (e.g., UE-A) can capture a user's video or receive the user's video from a camera connected to the originating UE (e.g., UE-A). The originating UE (e.g., UE-A) can use the video codec to encode the user's video to generate video media that includes the user's video.
[0134] Ia - Video Media to Video Media Transcoding If the codec identifier identifies a video codec, the originating UE (e.g., UE-A) can use the video codec to convert (e.g., transcode) video media containing the user's video into video media containing a representation of the user's virtual avatar.
[0135] The originating UE (e.g., UE-A) can use the video codec to decode the video media to obtain the user's video. The originating UE (e.g., UE-A) can replace the user with the user's virtual avatar representation to obtain the video of the user's virtual avatar representation.
[0136] The originating UE (e.g., UE-A) may display a video (e.g., a thumbnail view) representing the user's virtual avatar.
[0137] The originating UE (e.g., UE-A) can use the video codec to encode the video of the user's virtual avatar representation to obtain video media that includes the video of the user's virtual avatar representation.
[0138] The originating UE (e.g., UE-A) can send video media, including a video representation of the user's virtual avatar, to UE-B via the originating and terminating IMS networks. The video media can be sent using Real-Time Protocol (RTP) or Datagram Transport Layer Security (DTLS).
[0139] The terminating UE (e.g., UE-B) can use the video codec to decode video media including a virtual avatar representation of the user of the originating UE (e.g., UE-A). UE-B can display the video representing the virtual avatar of the user of the originating UE (e.g., UE-A).
[0140] Ib – Transcoding from video media to virtual avatar media If the codec identifier identifies a virtual avatar video codec, the originating UE (e.g., UE-A) can use the virtual avatar codec to convert (e.g., transcode) video media including the user's video into virtual avatar media including the user's pose within the user's video and the virtual avatar representation.
[0141] The originating UE (e.g., UE-A) can use the video codec to decode the video media to obtain the user's video. The originating UE (e.g., UE-A) can determine the user's posture within the user's video. For example, the originating UE (e.g., UE-A) can use artificial intelligence or machine learning to perform detection or recognition techniques and determine the user's posture (e.g., actions, gestures, or facial expressions) within the user's video.
[0142] The originating UE (e.g., UE-A) can animate the user's virtual avatar representation using the user's pose within the user's video to obtain a video of the user's virtual avatar representation. The originating UE (e.g., UE-A) can display the video of the user's virtual avatar representation (e.g., a thumbnail view).
[0143] The originating UE (e.g., UE-A) can use the virtual avatar codec to encode the user's pose within the user's video and the virtual avatar representation to obtain virtual avatar media that includes the user's pose within the user's video.
[0144] The originating UE (e.g., UE-A) can transmit virtual avatar media, including the user's pose within the user's video and the user's virtual avatar representation, to the terminating UE (e.g., UE-B) via the originating and terminating IMS networks. The virtual avatar media can be transmitted using RTP or DTLS.
[0145] The terminating UE (e.g., UE-B) can use the virtual avatar codec to decode virtual avatar media, including the user's pose within the user's video and the user's virtual avatar representation. UE-B can animate the user's virtual avatar representation using the user's pose within the user's video to obtain a video of the user's virtual avatar representation. UE-B can then display the video of the user's virtual avatar representation.
[0146] II – Rendering Implementation Based on IMS Network In an IMS-based rendering implementation, the originating and / or terminating IMS network (e.g., MF or MRF or another network function of the originating and / or terminating IMS network) can perform rendering.
[0147] If the MF or MRF of the originating IMS network stores the user's virtual avatar representation in its memory, the MF or MRF of the originating IMS network can retrieve the user's virtual avatar representation from its memory.
[0148] If the MF or MRF of the originating IMS network does not store the user's virtual avatar representation in its memory, the MF or MRF of the originating IMS network can receive the user's virtual avatar representation from another network function of the originating IMS network (e.g., XR AS).
[0149] The originating UE (e.g., UE-A) can capture or receive video of the user, captured by at least one camera connected to the originating UE (e.g., UE-A). UE-A can use a video codec to generate video media including the user's video. The originating UE (e.g., UE-A) can transmit the video media including the user's video to the originating IMS network. The video media can be transmitted using RTP or DTLS.
[0150] Alternatively, the originating UE (e.g., UE-A) can determine the user's pose within the user's video. For example, UE-A can process the video to determine the pose performed by the user in the video. The originating UE (e.g., UE-A) can process the video using an artificial intelligence or machine learning model configured to perform pose detection or pose recognition to identify the user's pose within the user's video. The originating UE (e.g., UE-A) can use a virtual avatar codec to generate virtual avatar media that includes the user's pose within the user's video. The originating UE (e.g., UE-A) can transmit the virtual avatar media, which includes the user's pose within the user's video, to the originating IMS network. The virtual avatar media can be transmitted using RTP.
[0151] The MF or MRF of the originating IMS network can receive video media including the user's video. Alternatively, the MF or MRF of the originating IMS network can receive virtual avatar media including the user's posture within the video.
[0152] II.a - Transcoding from video media or virtual avatar media to video media If the codec identifier identifies a video codec, the MF or MRF of the originating IMS network can use the video codec to convert (e.g., transcode) video media containing the user's video into video media containing a representation of the user's virtual avatar.
[0153] The originating IMS network's MF or MRF can use a video codec to decode the video media to obtain the user's video. The originating IMS network's MF or MRF can replace the user with a virtual avatar representation to obtain a video of the user's virtual avatar representation. The originating IMS network's MF or MRF can use the video codec to encode the video of the user's virtual avatar representation to obtain video media including the video of the user's virtual avatar representation.
[0154] The MF or MRF of the originating IMS network can transmit video media, including a virtual avatar representation of the user of the originating UE (e.g., UE-A), to UE-B via the terminating IMS network. The video media can be transmitted using RTP or DTLS.
[0155] The terminating UE (e.g., UE-B) can use the video codec to decode video media including a video of the virtual avatar representation of the user of the originating UE (e.g., UE-A) to obtain the video of the virtual avatar representation of the user of the originating UE (e.g., UE-A). The terminating UE (e.g., UE-B) can display the video of the virtual avatar representation of the user of the originating UE (e.g., UE-A).
[0156] The MF or MRF of the originating IMS network can send video media to the originating UE (e.g., UE-A) including a video of a virtual avatar representation of the user of the originating UE (e.g., UE-A). The video media can be sent using RTP or DTLS.
[0157] UE-A can use the video codec to decode video media including a virtual avatar representation of the user of the originating UE (e.g., UE-A) to obtain a video of the virtual avatar representation of the user of the originating UE (e.g., UE-A). The originating UE (e.g., UE-A) can display the video of the virtual avatar representation of the user of the originating UE (e.g., UE-A) (e.g., a thumbnail view).
[0158] Alternatively, the MF or MRF of the originating IMS network may use the video codec to convert (e.g., transcode) the virtual avatar media of the user's pose within the video including the user of the originating UE (e.g., UE-A) into video media of the video including the virtual avatar representation of the user of the originating UE (e.g., UE-A).
[0159] The MF or MRF of the originating IMS network can use a virtual avatar codec to decode the virtual avatar media to obtain the user's pose within the video of the user of the originating UE (e.g., UE-A). The MF or MRF of the originating IMS network can animate the user's virtual avatar representation within the video of the user of the originating UE (e.g., UE-A) to obtain a video of the virtual avatar representation of the user of the originating UE (e.g., UE-A). The MF or MRF of the originating IMS network can use the video codec to encode the video of the virtual avatar representation of the user of the originating UE (e.g., UE-A) to obtain video media including the video of the virtual avatar representation of the user of the originating UE (e.g., UE-A).
[0160] The MF or MRF of the originating IMS network can transmit video media, including a virtual avatar representation of the user of the originating UE (e.g., UE-A), to the terminating UE (e.g., UE-B) via the terminating IMS network.
[0161] The terminating UE (e.g., UE-B) can use the video codec to decode video media including the virtual avatar representation of the user of the originating UE (e.g., UE-A) to obtain the video including the virtual avatar representation of the user of the originating UE (e.g., UE-A). The terminating UE (e.g., UE-B) can display the video including the virtual avatar representation of the user of UE-A.
[0162] The MF or MRF of the originating IMS network can send video media to the originating UE (e.g., UE-A) including a video of a virtual avatar representation of the user of the originating UE (e.g., UE-A). The video media can be sent using RTP or DTLS.
[0163] The originating UE (e.g., UE-A) can use the video codec to decode video media including a video representing the user's virtual avatar to obtain the video representing the user's virtual avatar. The originating UE (e.g., UE-A) can display the video representing the user's virtual avatar (e.g., a thumbnail view).
[0164] II-b - Transcoding from video media or avatar media to avatar media If the codec identifier identifies a virtual avatar video codec, the MF or MRF of the originating IMS network can use the virtual avatar codec to convert (e.g., transcode) the video media including the user's video into virtual avatar media including the user's pose and virtual avatar representation within the video.
[0165] The MF or MRF of the originating IMS network can use a video codec to decode the video media to obtain the user's video. The MF or MRF of the originating IMS network can determine the user's pose within the video. For example, the MF or MRF of the originating IMS network can use artificial intelligence or machine learning to perform detection or recognition techniques and determine the user's pose within the video. The MF or MRF of the originating IMS network can use the avatar codec to encode the user's pose within the video and the avatar representation to obtain avatar media including the user's pose within the video.
[0166] The MF or MRF of the originating IMS network can transmit virtual avatar media, including the user's posture within the video of the user of the originating UE (e.g., UE-A) and the virtual avatar representation of the user, to the terminating UE (e.g., UE-B) via the terminating IMS network. The virtual avatar media can be transmitted using RTP or DTLS.
[0167] The terminating UE (e.g., UE-B) can use the virtual avatar codec to decode virtual avatar media, including the user's pose within the user's video and the user's virtual avatar representation, to obtain the user's pose within the user's video and the user's virtual avatar representation. The terminating UE (e.g., UE-B) can animate the user's virtual avatar representation using the user's pose within the user's video to obtain a video of the user's virtual avatar representation. The terminating UE (e.g., UE-B) can display the video of the user's virtual avatar representation.
[0168] The MF or MRF of the originating IMS network can send virtual avatar media to the originating UE (e.g., UE-A), including the user's posture within the video of the user of the originating UE (e.g., UE-A) and the virtual avatar representation of the user. The virtual avatar media can be sent using RTP or DTLS.
[0169] The originating UE (e.g., UE-A) can use the virtual avatar codec to decode virtual avatar media, including the user's pose within the user's video and the user's virtual avatar representation, to obtain the user's pose within the user's video and the user's virtual avatar representation. The originating UE (e.g., UE-A) can animate the user's virtual avatar representation using the user's pose within the user's video to obtain a video of the user's virtual avatar representation. The originating UE (e.g., UE-A) can display the video of the user's virtual avatar representation (e.g., a thumbnail view).
[0170] Alternatively, the MF or MRF of the originating IMS network may use the avatar codec to convert (e.g., transcode) the avatar media, which includes the user's pose within the video, into avatar media, which includes the user's pose and avatar representation within the video.
[0171] The MF or MRF of the originating IMS network can use the virtual avatar codec to decode the virtual avatar media to obtain the user's pose within the user's video. The MF or MRF of the originating IMS network can use the virtual avatar codec to encode the user's pose within the user's video and the user's virtual avatar representation to obtain virtual avatar media including the user's pose within the user's video and the user's virtual avatar representation.
[0172] The MF or MRF of the originating IMS network can transmit virtual avatar media, including the user's posture within the user's video and the user's virtual avatar representation, to the terminating UE (e.g., UE-B) via the terminating IMS network. The virtual avatar media can be transmitted using RTP or DTLS.
[0173] The terminating UE (e.g., UE-B) can use the virtual avatar codec to decode virtual avatar media, including the user's poses within the user's video and the user's virtual avatar representation. The terminating UE (e.g., UE-B) can animate the user's virtual avatar representation using the user's poses within the user's video to obtain a video of the user's virtual avatar representation. The terminating UE (e.g., UE-B) can display the video of the user's virtual avatar representation.
[0174] The MF or MRF of the originating IMS network can send virtual avatar media, including the user's posture within the user's video and the user's virtual avatar representation, to the originating UE (e.g., UE-A). The virtual avatar media can be encrypted using RTP or DTLS.
[0175] The originating UE (e.g., UE-A) can use the virtual avatar codec to decode virtual avatar media, including the user's pose within the user's video and the user's virtual avatar representation. The originating UE (e.g., UE-A) can animate the user's virtual avatar representation using the user's pose within the user's video to obtain a video of the user's virtual avatar representation. The originating UE (e.g., UE-A) can display the video of the user's virtual avatar representation (e.g., a thumbnail view).
[0176] III – Rendering Implementation Based on UE Termination In a rendering implementation based on a terminated UE, the terminated UE (e.g., UE-B) can perform the rendering.
[0177] If the terminating UE (e.g., UE-B) stores the user's virtual avatar representation in memory, then the terminating UE (e.g., UE-B) can retrieve the user's virtual avatar representation from memory.
[0178] If the terminating UE (e.g., UE-B) does not store the user's virtual avatar representation in its memory, the terminating UE (e.g., UE-B) can receive the user's virtual avatar representation from the originating IMS network (e.g., XR AS) via the terminating IMS network.
[0179] The originating UE (e.g., UE-A) can capture video of the user or receive video of the user from a camera connected to the originating UE (e.g., UE-A).
[0180] If the codec identifier identifies a virtual avatar codec, the originating UE (e.g., UE-A) can determine the user's pose within the user's video. For example, the originating UE (e.g., UE-A) can use artificial intelligence or machine learning to perform detection or recognition techniques and determine the user's pose within the user's video.
[0181] The originating UE (e.g., UE-A) can animate the user's virtual avatar representation using the user's posture within the user's video to obtain a video of the user's virtual avatar representation. The originating UE (e.g., UE-A) can then display the video of the user's virtual avatar representation.
[0182] The originating UE (e.g., UE-A) can use a virtual avatar codec to encode the user's pose within the user's video to obtain virtual avatar media that includes the user's pose within the user's video.
[0183] The originating UE (e.g., UE-A) can send virtual avatar media, including the user's posture within the user's video, to the terminating UE (e.g., UE-B) via the originating and terminating IMS networks. The virtual avatar media can be transmitted using RTP or DTLS.
[0184] The terminating UE (e.g., UE-B) can use a virtual avatar codec to decode virtual avatar media including the user's pose within the user's video to obtain the user's pose within the user's video. The terminating UE (e.g., UE-B) can mix (combine) the user's pose within the user's video with the user's virtual avatar representation.
[0185] The terminating UE (e.g., UE-B) can animate the user's virtual avatar representation using the user's posture within the user's video to obtain a video of the user's virtual avatar representation. The terminating UE (e.g., UE-B) can display the video of the user's virtual avatar representation. The virtual avatar media can be transmitted using RTP or DTLS.
[0186] It should be understood that in the rendering implementation based on the originating UE, the rendering implementation based on the IMS network, or the rendering implementation based on the terminating UE, the roles of the originating UE (e.g., UE-A) and the terminating UE (e.g., UE-B) can be interchanged.
[0187] It should be understood that in the rendering implementation based on the originating UE, the rendering implementation based on the IMS network, or the rendering implementation based on the terminating UE, the originating UE (e.g., UE-A), the originating IMS network, the terminating IMS network, or the terminating UE (e.g., UE-B) can modify the IMS session to switch from a virtual avatar call to an audio / video call at any time, for example, using a re-invitation or update message. The originating UE (e.g., UE-A), the originating IMS network, the terminating IMS network, or the terminating UE (e.g., UE-B) can switch from a virtual avatar call to an audio / video call at any time based on a request from the user, from the originating IMS network (e.g., XR AS), or from the terminating IMS network.
[0188] Figure 4 A schematic diagram of another example of the process for an implementation of "originating UE" (e.g., video media to video media transcoding) is shown.
[0189] In step 1, the originating UE (e.g., UE-A) can receive the user's virtual avatar representation from the originating IMS network (e.g., XR AS).
[0190] In step 2, the originating UE (e.g., UE-A) can perform rendering. The originating UE (e.g., UE-A) can convert (e.g., transcode) video media including user video into video media including video representing the user's virtual avatar.
[0191] In step 3, the originating UE (e.g., UE-A) can transmit video media, including a video representing the user's virtual avatar, to the terminating UE (e.g., UE-B) via the originating IMS network and the terminating IMS network.
[0192] Figure 5 A schematic diagram illustrating an example of the implementation of the "based on originating UE" (i.e., Ib - video media to virtual avatar media transcoding).
[0193] In step 1, the originating UE (e.g., UE-A) can receive the user's virtual avatar representation from the originating IMS network (e.g., XR AS).
[0194] In step 2, the originating UE (e.g., UE-A) can perform rendering. The originating UE (e.g., UE-A) can convert (e.g., transcode) video media including user video into virtual avatar media including user poses and user avatar representations within the user video.
[0195] In step 3, the originating UE (e.g., UE-A) can transmit virtual avatar media, including the user's posture within the user's video and the user's virtual avatar representation, to the terminating UE (e.g., UE-B) via the originating IMS network and the terminating IMS network.
[0196] Figure 6 A schematic diagram of an example process for an implementation of "UE-based termination" (e.g., avatar media to avatar media transcoding) is shown.
[0197] In step 1, the terminating UE (e.g., UE-B) can receive the user's virtual avatar representation from the originating IMS network (e.g., XR AS) via the terminating IMS network.
[0198] In step 2, the originating UE (e.g., UE-A) can (e.g., transcode) video media including user video into virtual avatar media including user gestures within the user video.
[0199] In step 3, the originating UE (e.g., UE-A) can send virtual avatar media, including the user's posture within the user video, to the terminating UE (e.g., UE-B) via the originating IMS network and the terminating IMS network.
[0200] In step 4, the terminating UE (e.g., UE-B) can decode virtual avatar media including user gestures within the user's video. The terminating UE (e.g., UE-B) can combine user gestures within the user's video with the user's virtual avatar representation.
[0201] Figure 7 The process for implementing network-based rendering is shown.
[0202] In step 1, an IMS session can be established between the originating UE (e.g., UE-A) and the terminating UE (e.g., UE-B). A bootstrap data channel can be established between the originating UE (e.g., UE-A) and the terminating UE (e.g., UE-B).
[0203] In step 2, the originating UE (e.g., UE-A) may decide to request network-based rendering (i.e., network-based rendering implementation) based on the state of the originating UE (e.g., UE-A) (e.g., power, signal strength, computing power, internal storage, etc.).
[0204] In step 3, the originating UE (e.g., UE-A) and the originating IMS network (e.g., XR AS) can perform rendering negotiation.
[0205] In step 4, if the rendering negotiation result described in step 3 is successful, the originating UE (e.g., UE-A) can initiate a peering-to-application data channel, which can be used for XR data transmission between the originating UE (e.g., UE-A) and the originating IMS network. During the peering-to-application data channel establishment process, the originating IMS network's DCSF can instruct the originating IMS network's MF or MRF on how to establish the peering-to-application data channel and the corresponding media processing specifications via the originating IMS network's IMS AS.
[0206] In step 5, the originating UE (e.g., UE-A) and the IMS AS of the originating IMS network can perform media renegotiation to connect the originating UE's audio and video media streams to the MF or MRF of the originating IMS network. The originating UE (e.g., UE-A) can send a message including a avatar identifier identifying the user's avatar representation and a codec identifier identifying the video codec or avatar codec. The media renegotiation request can be transmitted using a peer-to-peer application data channel.
[0207] In step 6, the terminating UE (e.g., UE-B) and the IMS AS of the originating IMS network can perform media renegotiation to connect the audio and video media streams of the terminating UE (e.g., UE-B) to the MF or MRF of the originating IMS network.
[0208] The originating IMS network (e.g., IMS AS) can send a message including a avatar identifier and a codec identifier to the terminating UE (e.g., UE-B) via the terminating IMS network, using a peer-to-peer application data channel. The terminating UE (e.g., UE-B) can accept or reject the audio or video session. If the terminating UE (e.g., UE-B) rejects the audio or video session, it can terminate the audio or video session. The terminating UE (e.g., UE-B) can accept or reject the avatar identifier or the codec identifier. If the terminating UE (e.g., UE-B) rejects the avatar identifier or the codec identifier, it can terminate the audio or video session.
[0209] In step 7, the XR AS of the originating IMS network can send a message including the virtual avatar identifier to the DAC. The message may include a UE identifier that identifies the originating UE (e.g., UE-A).
[0210] In step 8, the XR AS originating from the IMS network can receive a message including the user's virtual avatar representation from the DAC. The user's virtual avatar representation can be identified based on the virtual avatar identifier. The user's virtual avatar representation can be identified based on the user identifier.
[0211] In step 9, the XR AS of the originating IMS network can start rendering.
[0212] In step 10, the XR AS of the originating IMS network can send a message including the virtual avatar representation of the user to the MF or MRF of the originating IMS network.
[0213] In step 11, the originating UE (e.g., UE-A) may send video media including the user's video or virtual avatar media including the user's posture within the video to the MF or MRF of the originating IMS network.
[0214] In step 12, the MF or MRF of the originating IMS network can perform rendering.
[0215] The MF or MRF originating from the IMS network can use a video codec identified by the codec identifier to transform video media containing the user's video into video media containing a representation of the user's virtual avatar.
[0216] Alternatively, the MF or MRF originating from the IMS network can use a virtual avatar codec identified by the codec identifier to transform video media including the user's video into virtual avatar media including the user's pose and the user's virtual avatar representation within the video.
[0217] Alternatively, the MF or MRF originating from the IMS network can use a virtual avatar codec identified by the codec identifier to transform virtual avatar media including the user's pose within the video of the user into virtual avatar media including the user's pose within the video of the user and the virtual avatar representation of the user.
[0218] Alternatively, the MF or MRF originating from the IMS network can use a video codec identified by the codec identifier to transform the virtual avatar media of the user's pose within the video, which includes the user, into video media that includes the video of the user's virtual avatar representation.
[0219] In step 13, the MF or MRF of the originating IMS network can transmit video media, including a video representation of the user's virtual avatar, to the terminating UE (e.g., UE-B) via the terminating IMS network. The video media can be transmitted using RTP.
[0220] Alternatively, the MF or MRF of the originating IMS network can send virtual avatar media, including the user's posture within the user's video and the user's virtual avatar representation, to the terminating UE (e.g., UE-B). The virtual avatar media can be sent using RTP.
[0221] In step 14, the MF or MRF of the originating IMS network can send video media, including a video representing the user's virtual avatar, to the originating UE (e.g., UE-A). The video media can be sent using RTP.
[0222] Alternatively, the MF or MRF of the originating IMS network can send virtual avatar media, including the user's posture within the user's video and the user's virtual avatar representation, to the originating UE (e.g., UE-A). The virtual avatar media can be sent using RTP.
[0223] Figure 8 The process for implementing network-based rendering is shown.
[0224] In step 1, the originating UE (e.g., UE-A) may decide to request network-based rendering based on the state or capabilities (e.g., power, signal strength, computing power, internal storage, etc.) of the originating UE (e.g., UE-A).
[0225] In step 2, the originating UE (e.g., UE-A) may send a message to the IMS AS of the originating IMS network via the P-CSCF, I-CSCF, S-CSF, or IMS AGW of the originating IMS network. The message includes a virtual avatar identifier identifying the user's virtual avatar representation and a codec identifier identifying the video codec or virtual avatar codec. The message may include metadata associated with the virtual avatar representation. The message may include a description or name associated with the virtual avatar representation. The message may be sent using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the virtual avatar identifier may be sent as a SIP parameter or an SDP parameter. The codec identifier or the virtual avatar identifier may be transmitted via m- or a- lines.
[0226] In step 3, the IMS AS of the originating IMS network can send a message to the XR AS of the originating IMS network, including a virtual avatar identifier that identifies the user's virtual avatar representation and a codec identifier that identifies the video codec or virtual avatar codec. The message can be sent using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the virtual avatar identifier can be sent as a SIP parameter or an SDP parameter. The message enables the XR AS of the originating IMS network to retrieve the virtual avatar representation from the DAC.
[0227] In step 4, the IMS AS of the originating IMS network can send a message to the terminating UE (e.g., UE-B) including a virtual avatar identifier that identifies the user's virtual avatar representation and a codec identifier that identifies the video codec or virtual avatar codec. The message can be sent using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the virtual avatar identifier can be sent as a SIP parameter or an SDP parameter. The terminating UE (e.g., UE-B) can accept or reject the virtual avatar identifier or the codec identifier. If the terminating UE (e.g., UE-B) rejects the virtual avatar identifier or the codec identifier, the terminating UE (e.g., UE-B) can terminate the session.
[0228] In step 5, the XR AS originating from the IMS network can send a message to the DAC for receiving the user's virtual avatar representation. The message may include a virtual avatar identifier identifying the user's virtual avatar representation. The message may also include the user's user identifier.
[0229] In step 6, the DAC can send a message including the user's virtual avatar representation to the XR AS of the originating IMS network. The user's virtual avatar representation can be identified based on the virtual avatar identifier. The user's virtual avatar representation can be identified based on the user identifier.
[0230] In step 7, the terminating UE (e.g., UE-B) may send a message to the IMS AS of the originating IMS network that includes an indication to accept the virtual avatar identifier and the codec identifier.
[0231] In step 8, the IMS AS of the originating IMS network may send a message to the originating UE (e.g., UE-A) including an indication to accept the virtual avatar identifier and the codec identifier.
[0232] In step 9, the IMS AS originating from the IMS network can send a message including the virtual avatar identifier and the codec identifier to the XR AS originating from the IMS network. This message can cause the XR AS originating from the IMS network to initiate rendering.
[0233] In step 10, the XR AS of the originating IMS network can start rendering.
[0234] In step 11, the XR AS of the originating IMS network can send a message including the virtual avatar representation of the user to the MF or MRF of the originating IMS network.
[0235] In step 12, the originating UE (e.g., UE-A) may send video media including the user's video or virtual avatar media including the user's posture within the video to the MF or MRF of the originating IMS network.
[0236] In step 13, the MF or MRF of the originating IMS network can perform rendering.
[0237] The MF or MRF originating from the IMS network can use a video codec identified by the codec identifier to transform video media containing the user's video into video media containing a representation of the user's virtual avatar.
[0238] Alternatively, the MF or MRF originating from the IMS network can use a virtual avatar codec identified by the codec identifier to transform video media including the user's video into virtual avatar media including the user's pose and the user's virtual avatar representation within the video.
[0239] Alternatively, the MF or MRF originating from the IMS network can convert virtual avatar media, which includes user poses within the user's video, into virtual avatar media, which includes user poses within the user's video and a virtual avatar representation of the user using a virtual avatar codec identified by a codec identifier.
[0240] Alternatively, the MF or MRF originating from the IMS network can convert virtual avatar media, which includes user poses within the user's video, into video media that includes a representation of the user's virtual avatar using a video codec identified by a codec identifier.
[0241] In step 14, the MF or MRF of the originating IMS network can transmit video media, including a video representation of the user's virtual avatar, to the terminating UE (e.g., UE-B) via the terminating IMS network. The video media can be transmitted using RTP.
[0242] Alternatively, the MF or MRF of the originating IMS network can send virtual avatar media, including user gestures and a representation of the user's virtual avatar within the user's video, to the terminating UE (e.g., UE-B). This virtual avatar media can be sent using RTP.
[0243] In step 15, the MF or MRF of the originating IMS network can send video media, including a video representation of the user's virtual avatar, to the originating UE (e.g., UE-A). The video media can be sent using RTP.
[0244] Alternatively, the MF or MRF of the originating IMS network can transmit virtual avatar media, including the user's pose within the user's video and the user's virtual avatar representation, to the originating UE (e.g., UE-A) via IMS. This virtual avatar media can be transmitted using RTP.
[0245] Figure 9 The process for rendering based on the originating UE is shown.
[0246] In step 1, an IMS session can be established between the originating UE (e.g., UE-A) and the terminating UE (e.g., UE-B). A bootstrap data channel can be established for both the originating UE (e.g., UE-A) and the terminating UE (e.g., UE-B).
[0247] In step 2, the originating UE (e.g., UE-A) can upgrade the IMS session using the XR experience.
[0248] The originating UE (e.g., UE-A) can initiate a re-INVITE to add media descriptors (e.g., RTP and / or application data channels) required for end-to-end establishment by XR applications. The application data channel can be anchored in the MF or MRF of the originating IMS network.
[0249] The originating UE (e.g., UE-A) may send a message to the XR AS of the originating IMS network, including a virtual avatar identifier that identifies the user. The message may include a codec identifier that identifies the video codec or the virtual avatar codec. The message may include a user identifier that identifies the user. The message may be sent using an application data channel.
[0250] In step 3, the XR AS originating from the IMS network can send a message to the DAC to request a virtual avatar representation of the user. The message may include a virtual avatar identifier that identifies the user's virtual avatar representation. The message may also include a user identifier that identifies the user.
[0251] In step 4, the DAC can send a message including the user's virtual avatar representation to the XR application server originating the IMS network.
[0252] In step 5, the XR AS of the originating IMS network can send a message including the virtual avatar representation of the user to the originating UE (e.g., UE-A). The message may include a virtual avatar identifier identifying the user's virtual avatar representation. The message may also include a user identifier identifying the user. The message can be sent using an application data channel, using Hypertext Transfer Protocol (HTTP), using WebSocket, or using SIP.
[0253] In step 6, the originating UE (e.g., UE-A) may send a message indicating success or failure to the XR AS of the originating IMS network.
[0254] In step 7, the originating UE (e.g., UE-A) can convert video media including the user's video into video media including the user's virtual avatar representation using a video codec identified by a codec identifier.
[0255] Alternatively, the originating UE (e.g., UE-A) may convert video media including the user's video into virtual avatar media including the user's pose within the user's video and the user's virtual avatar representation using a virtual avatar codec identified by a codec identifier.
[0256] In step 8, the originating UE (e.g., UE-A) may transmit video media including a video of the user's virtual avatar representation or virtual avatar media including the user's pose within the video and the user's virtual avatar representation to the terminating UE (e.g., UE-B) via the originating and terminating IMS networks. The video media can be received using RTP.
[0257] In step 9, the originating UE (e.g., UE-A) can receive video media, including a virtual avatar representation of another user, from the terminating UE (e.g., UE-B) via the originating and terminating IMS networks. The video media can be received using RTP.
[0258] Alternatively, the originating UE (e.g., UE-A) can receive virtual avatar media from the terminating UE (e.g., UE-B), which includes the gestures of another user within a video stream and a virtual avatar representation of that other user. This virtual avatar media can be received using RTP.
[0259] In step 10, the originating UE (e.g., UE-A) can use a video codec to decode the video media including the video representing the virtual avatar of the other user. The originating UE (e.g., UE-A) can then display the video representing the virtual avatar of the other user.
[0260] Alternatively, the originating UE (e.g., UE-A) can decode virtual avatar media that includes the other user's pose and virtual avatar representation within the other user's video. The originating UE (e.g., UE-A) can animate the other user's virtual avatar representation using the other user's pose to obtain a video of the other user's virtual avatar representation. The originating UE (e.g., UE-A) can display the video of the other user's virtual avatar representation.
[0261] Figure 10 The process for rendering based on the originating UE is shown.
[0262] In step 1, the originating UE (e.g., UE-A) may decide to request rendering based on the originating UE (e.g., UE-A) based on the state or capabilities of the originating UE (e.g., UE-A) (e.g., power, signal strength, computing power, internal storage, etc.).
[0263] In step 2, the originating UE (e.g., UE-A) may send a message to the originating IMS network (e.g., an IMS AS via the originating IMS network's P-CSCF, I-CSCF, S-CSF, or IMS AGW) including a virtual avatar identifier that identifies the user's virtual avatar representation and a codec identifier that identifies the video codec or virtual avatar codec. The message may include metadata associated with the virtual avatar representation. The message may include a description or name associated with the virtual avatar representation. The message may be sent using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or the message may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the virtual avatar identifier may be sent as a SIP parameter or an SDP parameter.
[0264] In step 3, the IMS AS originating from the IMS network can send a message to the XR AS originating from the IMS network, including a virtual avatar identifier that identifies the user's virtual avatar representation and a codec identifier that identifies the video codec or virtual avatar codec. The message can be sent using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the virtual avatar identifier can be sent as a SIP parameter or an SDP parameter. The message enables the XR AS originating from the IMS network to retrieve the virtual avatar representation from the DAC.
[0265] In step 4, the originating IMS network (e.g., IMS AS) may send a message to the terminating UE (e.g., UE-B) including a virtual avatar identifier that identifies the user's virtual avatar representation and a codec identifier that identifies the video codec or virtual avatar codec. The message may be sent using an IMS data channel, using SIP, or using SDP. The message may include a SIP message or may include SDP content (e.g., an SDP proposal). The SIP message may include one or more SIP parameters, or the SDP content may include one or more SDP parameters. The codec identifier or the virtual avatar identifier may be sent as a SIP parameter or an SDP parameter. The terminating UE (e.g., UE-B) may accept or reject the virtual avatar identifier or the codec identifier. If the terminating UE (e.g., UE-B) rejects the virtual avatar identifier or the codec identifier, the terminating UE (e.g., UE-B) may terminate the session.
[0266] It will be understood that the terminating UE (e.g., UE-B) can reject a avatar call by rejecting the avatar identifier or the codec identifier. Alternatively, the terminating UE (e.g., UE-B) can reject a avatar call by sending a SIP message (e.g., re-invite) to the IMSAS of the originating IMS network without a avatar identifier representing the user's avatar.
[0267] In step 5, the XR AS originating from the IMS network can send a message to the DAC to receive the user's virtual avatar representation. The message may include a virtual avatar identifier identifying the user's virtual avatar representation. The message may also include the user's user identifier.
[0268] In step 6, the DAC can send a message including the user's virtual avatar representation to the XR AS of the originating IMS network. The user's virtual avatar representation can be identified based on the virtual avatar identifier. The user's virtual avatar representation can be identified based on the user identifier.
[0269] In step 7, the terminating UE (e.g., UE-B) may send a message to the IMS AS of the originating IMS network that includes an indication to accept the virtual avatar identifier and the codec identifier.
[0270] In step 8, the IMS AS of the originating IMS network may send a message to the originating UE (e.g., UE-A) including an indication to accept the virtual avatar identifier and the codec identifier.
[0271] In step 9, the originating UE (e.g., UE-A) may send a message including a virtual avatar identifier and a codec identifier to the XR AS of the originating IMS network. This message may cause the XR AS of the originating IMS network to initiate rendering.
[0272] In step 10, the XR AS of the originating IMS network can start rendering.
[0273] In step 11, the XR AS of the originating IMS network can send a message including the virtual avatar representation of the user to the originating UE (e.g., UE-A).
[0274] In step 12, the originating UE (e.g., UE-A) can perform rendering.
[0275] The originating UE (e.g., UE-A) can use a video codec identified by the codec identifier to convert video media containing the user's video into video media containing a representation of the user's virtual avatar.
[0276] Alternatively, the originating UE (e.g., UE-A) may convert video media including the user's video into virtual avatar media including the user's pose within the video and the user's virtual avatar representation using a virtual avatar codec identified by the codec identifier.
[0277] In step 13, the originating UE (e.g., UE-A) may transmit video media including a video of the user's virtual avatar representation or virtual avatar media including the user's posture within a video and the user's virtual avatar representation to the terminating UE (e.g., UE-B) via the originating and terminating IMS network. The video media or the virtual avatar media may be transmitted using RTP.
[0278] In step 14, the originating UE (e.g., UE-A) can receive video media, including a virtual avatar representation of another user, from the terminating UE (e.g., UE-B) via the originating and terminating IMS networks. The video media can be received using RTP.
[0279] Alternatively, the originating UE (e.g., UE-A) may receive virtual avatar media, including the other user's posture within a video and the other user's virtual avatar representation, from the terminating UE (e.g., UE-B) via the originating and terminating IMS networks. The virtual avatar media may be received using RTP.
[0280] The originating UE (e.g., UE-A) can use a video codec to decode video media including a video representation of the other user's virtual avatar. The originating UE (e.g., UE-A) can display the video representing the other user's virtual avatar.
[0281] Alternatively, the originating UE (e.g., UE-A) may decode the other user's pose and the virtual avatar representation of the other user within the video. The originating UE (e.g., UE-A) may animate the other user's virtual avatar representation using the other user's pose to obtain a video of the other user's virtual avatar representation. The originating UE (e.g., UE-A) may display the video of the other user's virtual avatar representation.
[0282] Figure 11 A block diagram is shown illustrating a rendering-based virtual avatar call between an originating user equipment (e.g., UE-A) and a terminating user equipment (e.g., UE-B).
[0283] In step 1100, the originating UE (e.g., UE-A) may transmit to the terminating UE (e.g., UE-B) a message for establishing or modifying an IMS session for virtual avatar calling between the originating UE (e.g., UE-A) and the terminating UE (e.g., UE-B).
[0284] In step 1002, the originating UE (e.g., UE-A) may generate a first video media that includes the user's video.
[0285] In step 1004, the originating UE (e.g., UE-A) can convert a first video media including the user's video into a second video media including the user's virtual avatar representation, or convert it into virtual avatar media including the user's virtual avatar representation and the user's posture within the user's video.
[0286] In step 1006, the originating UE (e.g., UE-A) may transmit the second video media or the virtual avatar media to the terminating UE (e.g., UE-B).
[0287] Figure 12 A block diagram is shown for a network-based rendering-based virtual avatar call between an originating user equipment (e.g., UE-A) and a terminating user equipment (e.g., UE-B).
[0288] In step 1200, the MF or MRF can receive from the originating UE (e.g., UE-A) a first video media including the user's video or a first virtual avatar media including the user's pose within the video.
[0289] In step 1202, the MF or MRF can convert a first video media including the user's video or a first virtual avatar media including the user's pose within the video into a second video media including the user's virtual avatar representation, or into a second virtual avatar media including the user's virtual avatar representation and the user's pose within the video.
[0290] In step 1204, the MF or MRF can transmit the second video media or the second virtual avatar media to the terminating UE (e.g., UE-B).
[0291] Figure 13 A block diagram is shown of a rendering-based virtual avatar call between a terminating user equipment (UE-A) and a terminating user equipment (UE-B).
[0292] In step 1300, the terminating UE (e.g., UE-B) can receive from the originating user equipment (e.g., UE-A) a message for establishing or modifying an IMS session for virtual avatar calling between the originating user equipment (e.g., UE-A) and the terminating UE (e.g., UE-B). In step 1302, the terminating UE (e.g., UE-B) can receive virtual avatar media including the user's posture within the video of the originating user equipment (e.g., UE-A).
[0293] In step 1304, the terminating UE (e.g., UE-B) can decode the virtual avatar media to obtain the user's posture within the user's video.
[0294] In step 1306, the terminating UE (e.g., UE-B) can animate the virtual avatar representation using the user's posture within the user's video to obtain the video of the user's virtual avatar representation.
[0295] In step 1306, the terminating UE (e.g., UE-B) may display a video representing the user's virtual avatar.
[0296] Figure 14 A schematic representation of a non-volatile memory medium 1400 is shown, which stores instructions that, when executed by a processor, cause the processor to perform... Figure 11 , Figure 12 and Figure 13 Any of the methods described in [the document / document].
[0297] It should be noted that although exemplary embodiments have been described above, several variations and modifications can be made to the disclosed solutions without departing from the scope of the invention.
[0298] It should be understood that although the above concepts have been discussed in the context of 5G, one or more of these concepts can be applied to other generations.
[0299] Therefore, the embodiments described herein can vary within the scope of the appended claims. Generally, some embodiments can be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while others can be implemented using firmware or software executable by a controller, microprocessor, or other computing device, but the embodiments are not limited thereto. While various embodiments may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be well understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0300] The embodiments described herein can be implemented by computer software stored in memory and executable by at least one data processor of the entity concerned, or by hardware, or by a combination of software and hardware. It should also be noted in this regard that any process, such as Figure 11 , Figure 12 and Figure 13 The process in the software can represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on memory blocks such as memory chips or implemented within a processor, magnetic media such as hard disks or floppy disks, and physical media such as optical media such as DVDs and their data variants, CDs.
[0301] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include one or more of general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures, as non-limiting examples.
[0302] Alternatively or additionally, some embodiments may be implemented using circuitry. This circuitry may be configured to perform one or more of the previously described functional and / or method steps. The circuitry may be provided in a base station and / or user equipment.
[0303] As used in this application, the term "circuit" may refer to all or one of the following: (a) Pure hardware circuits (such as analog and / or digital circuits only); (b) A combination of hardware circuits and software, such as: (i) A combination of analog and / or digital hardware circuitry with software / firmware, and (ii) Any part of the hardware processor works in conjunction with software (including digital signal processors), software, and memory(s) to enable the apparatus (such as a communication device or base station) to perform the various functions previously described; and (c) Hardware circuitry and / or processors, such as microprocessors or parts thereof, which require software (e.g., firmware) to operate, but may be absent when the software is not required to operate.
[0304] This definition of "circuit" applies to all uses of the term "component" in this application, including in any claim. As another example, as used herein, the term "circuit" also encompasses only hardware circuitry or a processor (or multiple processors) or a portion of hardware circuitry or a processor and its accompanying software and / or firmware implementation. The term "circuit" also encompasses, for example, integrated devices.
[0305] The foregoing description has provided a complete and informative description of some embodiments by way of exemplary and non-limiting examples. However, in view of the foregoing description, various modifications and adjustments will be apparent to those skilled in the art when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar teaching modifications will still fall within the scope defined in the appended claims.
Claims
1. A method for a first user equipment, the method comprising: Transmits messages to the second user equipment for establishing or modifying an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment; Generate the first video media that includes the user's video; The first video media, which includes the video of the user, is converted into a second video media, which includes a video representing the user's virtual avatar, or converted into a virtual avatar media, which includes the user's virtual avatar representation and the user's posture within the video. as well as Transmit the second video media or the virtual avatar media to the second user equipment.
2. The method according to claim 1, wherein converting the first video media into the second video media comprises: The first video media is decoded using a video codec to obtain the user's video; Replace the user in the user's video with the user's virtual avatar representation to obtain the video of the user's virtual avatar representation; as well as The video represented by the user's virtual avatar is encoded using the video codec to obtain the second video media.
3. The method according to claim 1, wherein converting the first video media into virtual avatar media comprises: The first video media is decoded using a video codec to obtain the user's video; Determine the user's pose within the video; as well as The user's pose and the user's virtual avatar representation are encoded using a virtual avatar codec to obtain the virtual avatar media.
4. The method according to any one of claims 1 to 3, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment includes a codec identifier that identifies a video codec or an avatar codec; and The method includes: Receive a message from the second user equipment indicating acceptance of the codec identifier; as well as The first video media is converted into the second video media or the virtual avatar media using the video codec or the virtual avatar codec identified by the codec identifier.
5. The method according to any one of claims 1 to 4, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment includes a virtual avatar identifier that identifies the virtual avatar representation of the user; and The method includes: Receive a message from the second user equipment indicating acceptance of the virtual avatar identifier; as well as The first video media is converted into the second video media or the virtual avatar media using a virtual avatar representation identified by the virtual avatar identifier.
6. The method according to claim 5, wherein the virtual avatar identifier includes an integer, a string, a universally unique identifier, a uniform resource name, or a uniform resource locator.
7. The method according to any one of claims 1 to 6, comprising: Retrieve from the memory of the first user equipment the virtual avatar representation of the user identified by the virtual avatar identifier.
8. The method according to any one of claims 1 to 7, comprising: The user's virtual avatar representation, identified by the virtual avatar identifier, is received by a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem (IPMS) network.
9. The method of claim 8, wherein the virtual avatar representation of the user, identified by the virtual avatar identifier, is received using an Internet Protocol Multimedia Subsystem data channel.
10. The method according to any one of claims 1 to 9, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment is transmitted using an IPS data channel or using Hypertext Transfer Protocol.
11. The method according to any one of claims 1 to 10, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment is transmitted using a session description protocol.
12. The method according to any one of claims 1 to 11, wherein the second video media or the virtual avatar media is transmitted to the second user equipment using a real-time protocol.
13. A method for an Internet Protocol Multimedia Subsystem (IPS) network, the method comprising: Receive a first video medium including a video of the user from a first user device, or a first virtual avatar medium including the user's posture within the video of the user; Convert the first video media including the user's video or the first virtual avatar media including the user's posture within the video into a second video media including the user's virtual avatar representation, or convert it into a second virtual avatar media including the user's virtual avatar representation and the user's posture within the video; and Transmit the second video media or the second virtual avatar media to the second user equipment.
14. The method of claim 13, comprising: The user's virtual avatar representation is received from the first user equipment, a network function belonging to the Internet Protocol Multimedia Subsystem (IPMS) network, or a network function not belonging to the IMPMS network.
15. The method of claim 13 or claim 14, wherein the virtual avatar is received using an Internet Protocol Multimedia Subsystem data channel or Hypertext Transfer Protocol.
16. The method according to any one of claims 13 to 15, wherein converting the first video media into the second video media comprises: The first video media is decoded using a video codec to obtain the user's video; In the user's video, the user is replaced with a virtual avatar representation of the user to obtain the video of the user's virtual avatar representation; as well as The video represented by the user's virtual avatar is encoded using the video codec to obtain the second video media.
17. The method according to any one of claims 13 to 15, wherein converting the first virtual avatar media into the second video media comprises: The first virtual avatar media, including the user's pose within the video, is decoded using a virtual avatar codec. Using the user's posture within the user's video, the user's virtual avatar representation is animated to obtain the video of the user's virtual avatar representation; as well as The video represented by the user's virtual avatar is encoded using the video codec to obtain the second video media.
18. The method according to any one of claims 13 to 15, wherein converting the first video media into the second virtual avatar media comprises: The first video media is decoded using a video codec to obtain the user's video; Determine the user's pose within the video; as well as The user's posture and the user's virtual avatar representation are encoded using a virtual avatar codec to obtain the second virtual avatar media.
19. The method according to any one of claims 13 to 15, wherein converting the first virtual avatar media into the second virtual avatar media comprises: The first virtual avatar media, which includes the user's pose within the video, is decoded using a virtual avatar codec. as well as The user's pose and the user's virtual avatar representation are encoded using the virtual avatar codec to obtain the second virtual avatar media.
20. The method according to any one of claims 13 to 19, comprising: Receive a message indicating a codec identifier from a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem Network (IPMS) network function, wherein the codec identifier identifies a video codec or a virtual avatar codec. as well as The first video media or the first virtual avatar media is converted into the second video media or the second virtual avatar media using the video codec or the virtual avatar codec identified by the codec identifier.
21. The method according to any one of claims 13 to 20, comprising: The virtual avatar representation of the user is received from the first user equipment, a network function belonging to the Internet Protocol Multimedia Subsystem Network (IPMS), or a network function not belonging to the Internet Protocol Multimedia Subsystem Network (IPMS). as well as Using the virtual avatar represents converting the first video media or the first virtual avatar media into the second video media or the second virtual avatar media.
22. The method according to any one of claims 13 to 21, comprising: Transmit the second video media or the second virtual avatar media to the first user equipment.
23. The method according to any one of claims 13 to 22, wherein the second video media or the second virtual avatar media is transmitted to at least one of the first user equipment or the second user equipment using a real-time protocol.
24. The method according to any one of claims 13 to 23, wherein the virtual avatar representation of the user and the posture of the user in the video of the user are transmitted via an Internet Protocol Multimedia Subsystem data channel.
25. The method according to any one of claims 13 to 24, wherein the method is performed by a media function or a media resource function.
26. A method for a second user equipment, the method comprising: Receive messages from the first user equipment for establishing or modifying an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment; Receive virtual avatar media containing the user's posture within a video stream from the first user equipment; Decode the virtual avatar media to obtain the user's posture within the user's video; Using the user's posture within the user's video, the virtual avatar representation is animated to obtain a video of the user's virtual avatar representation; as well as The video is displayed as a representation of the user's virtual avatar.
27. The message of claim 26, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment includes a codec identifier that identifies the avatar codec. as well as The method includes: Transmit a message to the first user equipment indicating acceptance of the codec identifier; as well as The virtual avatar media is decoded using a virtual avatar codec identified by the codec identifier to obtain the user's pose within the user's video.
28. The method of claim 26 or claim 27, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment includes a avatar identifier that identifies the avatar representation of the user; as well as The method includes: Transmit a message to the first user equipment indicating acceptance of the virtual avatar identifier; Transmitting a message including a virtual avatar identifier that identifies the user's virtual avatar representation to a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem (IPMS) network; and The user's virtual avatar representation is received from a network function belonging to or not belonging to the Internet Protocol Multimedia Subsystem (IPMS) network.
29. The method of claim 26 or claim 27, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for avatar calling between the first user equipment and the second user equipment includes a avatar identifier that identifies the avatar representation of the user; as well as The method includes: Transmit a message to the first user equipment indicating acceptance of the virtual avatar identifier; as well as Retrieve from the memory of the second user equipment the virtual avatar representation of the user identified by the virtual avatar identifier.
30. The method according to any one of claims 26 to 29, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem (IPS) session for virtual avatar calling between the first user equipment and the second user equipment is transmitted using an IPS data channel.
31. The method according to any one of claims 26 to 30, wherein the message for establishing or modifying an Internet Protocol Multimedia Subsystem session for virtual avatar calling between the first user equipment and the second user equipment is transmitted using a session description protocol.
32. The method according to any one of claims 26 to 31, wherein the virtual avatar media is received from the first user equipment using a real-time protocol.
33. The method according to any one of claims 26 to 32, wherein the virtual avatar media of the user's posture within the user's video is received via an Internet Protocol Multimedia Subsystem data channel.
34. A first user equipment, comprising: At least one processor; as well as At least one memory stores instructions that, when executed by the at least one processor, cause the first user equipment to perform the method according to any one of claims 1 to 12.
35. An Internet Protocol Multimedia Subsystem (IPS) network, comprising: At least one processor; as well as At least one memory stores instructions that, when executed by the at least one processor, cause the Internet Protocol Multimedia Subsystem network to perform the method according to any one of claims 13 to 25.
36. A second user equipment, comprising: At least one processor; as well as At least one memory stores instructions that, when executed by the at least one processor, cause the second user equipment to perform the method according to any one of claims 26 to 33.
37. A computer program comprising computer-executable code, which, when run on at least one processor of a user equipment, causes the user equipment to perform the method according to any one of claims 1 to 12 or 26 to 33.
38. A computer program comprising computer-executable code, which, when run on at least one processor of an Internet Protocol Multimedia Subsystem (IPMS) network, causes the IMS network to perform the method according to any one of claims 13 to 25.
39. A computer-readable medium comprising instructions that, when executed by a user equipment, cause the user equipment to perform the method according to any one of claims 1 to 12 or 26 to 33.
40. A computer-readable medium comprising instructions that, when executed by at least one processor of an Internet Protocol Multimedia Subsystem (IPMS) network, cause the IMS network to perform the method according to any one of claims 13 to 25.