Method and apparatus for avatar communication based on real time communication (RTC) in next-generation mobile communication system

The method addresses the challenge of format incompatibility in avatar calls by specifying avatar data formats and converting between them, ensuring seamless communication across different terminals.

WO2025198386A1PCT designated stage Publication Date: 2025-09-25SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/095027
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-20
Filing Date
2025-03-20
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing technologies face challenges in providing compatibility and format negotiation for avatar calls between different terminals and solution implementations, as they lack methods to specify avatar data formats and convert between incompatible formats during data sessions.

Method used

A method is introduced to specify avatar calls by using Session Description Protocol (SDP) offers and addresses to establish connections, and includes a conversion process for avatar data formats such as facial landmarks and blend shapes, ensuring compatibility and seamless communication.

Benefits of technology

Enables compatibility of avatar calls between different terminals by identifying format differences, converting between incompatible formats, and establishing data channels for real-time avatar communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025095027_25092025_PF_FP_ABST
    Figure KR2025095027_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a 5G or 6G communication system for supporting higher data transmission rates. According to an embodiment of the present disclosure, it is possible to provide compatibility of avatar calls between different terminals and solution implementations.
Need to check novelty before this filing date? Find Prior Art

Description

AVATAR communication method and device based on real-time communication (RTC) in next-generation mobile communication system

[0001] The present technology relates to the operation of a device and server for performing avatar communication based on real time communication (RTC) in a mobile communication system.

[0002] 5G mobile communication technology defines a wide frequency band to enable fast transmission speeds and new services, and can be implemented not only in the sub-6GHz frequency band such as 3.5 gigahertz (3.5GHz), but also in the ultra-high frequency band called millimeter wave (mmWave) such as 28GHz and 39GHz ('Above 6GHz'). In addition, for 6G mobile communication technology, which is called the system after 5G communication (Beyond 5G), implementation in the terahertz band (for example, the 3 terahertz (3THz) band at 95GHz) is being considered to achieve a transmission speed that is 50 times faster than 5G mobile communication technology and an ultra-low latency time that is reduced to one-tenth.

[0003] In the early stages of 5G mobile communication technology, the goal is to support services and satisfy performance requirements for enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communications (URLLC), and massive Machine-Type Communications (mMTC). These include beamforming and massive MIMO to mitigate path loss of radio waves in ultra-high frequency bands and increase the transmission distance of radio waves, support for various numerologies (such as operation of multiple subcarrier intervals) and dynamic operation of slot formats for efficient use of ultra-high frequency resources, initial access technology to support multi-beam transmission and wideband, definition and operation of BWP (Bidth Part), new channel coding methods such as LDPC (Low Density Parity Check) codes for large-capacity data transmission and Polar Code for reliable transmission of control information, and L2 pre-processing (L2). Standardization has been made for network slicing, which provides dedicated networks specialized for specific services, and pre-processing.

[0004] Currently, discussions are underway to improve and enhance the initial 5G mobile communication technology in consideration of the services that 5G mobile communication technology was intended to support, and physical layer standardization is in progress for technologies such as V2X (Vehicle-to-Everything) to help autonomous vehicles make driving decisions and increase user convenience based on their own location and status information transmitted by vehicles, NR-U (New Radio Unlicensed) for the purpose of system operation that complies with various regulatory requirements in unlicensed bands, NR terminal low power consumption technology (UE Power Saving), Non-Terrestrial Network (NTN), which is direct terminal-satellite communication to secure coverage in areas where communication with terrestrial networks is impossible, and Positioning.

[0005] In addition, standardization of wireless interface architecture / protocols is in progress for technologies such as intelligent factories (Industrial Internet of Things, IIoT) to support new services through linkage and convergence with other industries, Integrated Access and Backhaul (IAB) that provides nodes for expanding network service areas by integrating wireless backhaul links and access links, Mobility Enhancement technology including Conditional Handover and Dual Active Protocol Stack (DAPS) handover, and 2-step random access (2-step RACH for NR) that simplifies random access procedures. Standardization is also in progress for system architecture / services such as 5G baseline architecture (e.g., Service-based Architecture, Service-based Interface) for grafting Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies, and Mobile Edge Computing (MEC) that provides services based on the location of the terminal.

[0006] Once these 5G mobile communication systems are commercialized, an explosive increase in connected devices will be connected to the communication network, necessitating enhanced functionality and performance of 5G mobile communication systems and integrated operation of these connected devices. To this end, new research will be conducted on improving 5G performance and reducing complexity, supporting AI services, supporting metaverse services, and drone communications by utilizing eXtended Reality (XR), Artificial Intelligence (AI), and Machine Learning (ML) to efficiently support Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR).

[0007] In addition, the development of these 5G mobile communication systems includes new waveforms to ensure coverage in the terahertz band of 6G mobile communication technology, multi-antenna transmission technologies such as Full Dimensional MIMO (FD-MIMO), Array Antenna, and Large Scale Antenna, metamaterial-based lenses and antennas to improve the coverage of terahertz band signals, high-dimensional spatial multiplexing technology using Orbital Angular Momentum (OAM), Reconfigurable Intelligent Surface (RIS) technology, as well as full duplex technology to improve the frequency efficiency and system network of 6G mobile communication technology, satellite, AI (Artificial Intelligence) from the design stage and AI-based communication technology that realizes system optimization by internalizing end-to-end AI support functions, and ultra-high-performance communication and computing resources to provide services with complexity that exceeds the limits of terminal computing capabilities. It can serve as a basis for the development of next-generation distributed computing technologies that can be realized by utilizing them.

[0008] Meanwhile, the need for a method and device that provides compatibility of avatar calls between different terminals and solution implementations has arisen.

[0009] An object of the present invention is to provide a method and device for providing compatibility of avatar calls between different terminals and solution implementations.

[0010] In a wireless communication system according to one embodiment of the present disclosure, a method performed by a server for performing an avatar call may include the steps of: receiving a registration message indicating a matching specification for the avatar call from a first terminal; receiving a specific message for the avatar call from the first terminal after receiving the registration message; determining a second terminal for performing the avatar call based on at least one piece of information included in the specific message; receiving a connection message from the first terminal, the connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal, when a connection step is performed between the first terminal and the second terminal; and transmitting the connection message to the second terminal, and when an acceptance message including an SDP response is received from the second terminal, transmitting the acceptance message to the first terminal so that communication for the avatar call is initiated.

[0011] In another embodiment of the present disclosure, a method performed by a first terminal for performing an avatar call in a wireless communication system may include the steps of: transmitting a registration message indicating a matching specification for the avatar call to a server; transmitting a specific message for the avatar call to the server after transmitting the registration message; transmitting a connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal to the server when a connection step is performed with a second terminal determined based on the specific message; and initiating communication for the avatar call when an acceptance message including an SDP response transmitted by the second terminal is transmitted from the server.

[0012] In another embodiment of the present disclosure, a server for performing an avatar call in a wireless communication system includes a transceiver; And a control unit that receives a registration message indicating a matching standard for the avatar call from the first terminal through the transceiver, and after receiving the registration message, receives a specific message for the avatar call from the first terminal through the transceiver, and determines a second terminal for performing the avatar call based on at least one piece of information included in the specific message, and when a connection step is performed between the first terminal and the second terminal, receives a connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal from the first terminal through the transceiver, transmits the connection message to the second terminal, and when an acceptance message including an SDP response is received from the second terminal, transmits the acceptance message to the first terminal through the transceiver so that communication for the avatar call is initiated.

[0013] In another embodiment of the present disclosure, a first terminal for performing an avatar call in a wireless communication system may include a transceiver; and a control unit configured to transmit a registration message indicating a matching specification for the avatar call to a server through the transceiver, and to transmit a specific message for the avatar call to the server through the transceiver after transmitting the registration message, and when a connection step is performed with a second terminal determined based on the specific message, transmit a connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal to the server through the transceiver, and when an acceptance message including an SDP response transmitted by the second terminal is transmitted from the server, control to initiate communication for the avatar call.

[0014] According to an embodiment of the present invention, compatibility of avatar calls between different terminals and solution implementations can be provided.

[0015] Figure 1 is a diagram showing the workflow flow of a typical avatar call.

[0016] FIG. 2 is a drawing for explaining animation data according to one embodiment of the present disclosure.

[0017] FIG. 3 is a drawing for explaining animation data according to one embodiment of the present disclosure.

[0018] FIG. 4 is a sequence diagram specifically illustrating an embodiment of a method for initiating data communication for avatar calling according to one embodiment of the present disclosure.

[0019] FIG. 5a is a sequence diagram specifically explaining the above-described conversion steps according to one embodiment of the present disclosure.

[0020] FIG. 5b is a sequence diagram specifically explaining the above-described conversion steps according to one embodiment of the present disclosure.

[0021] FIG. 6 is a block diagram illustrating components of a terminal according to an embodiment of the present disclosure.

[0022] FIG. 7 is a block diagram illustrating components of a server according to one embodiment of the present disclosure.

[0023] The operating principles of the present invention will be described in detail below with reference to the attached drawings. In the following description of the present invention, detailed descriptions of known functions or components will be omitted if they are deemed to unnecessarily obscure the gist of the invention. Furthermore, the terms described below are defined based on their functions in the present invention and may vary depending on the intentions or practices of the user or operator. Therefore, their definitions should be based on the overall content of this specification.

[0024] The terms used in the following description to identify connection nodes, terms referring to network entities, terms referring to messages, terms referring to interfaces between network entities, and terms referring to various identification information are provided for convenience of explanation. Therefore, the present invention is not limited to the terms described below, and other terms referring to objects with equivalent technical meanings may be used.

[0025] Hereinafter, the base station is an entity that performs resource allocation of the terminal, and may be at least one of a gNode B, an eNode B, a Node B, a BS (Base Station), a wireless access unit, a base station controller, or a node on a network. The terminal may include a UE (User Equipment), an MS (Mobile Station), a cellular phone, a smartphone, a computer, or a multimedia system capable of performing a communication function. In the present disclosure, downlink (DL) refers to a wireless transmission path of a signal transmitted from a base station to a terminal, and uplink (UL) refers to a wireless transmission path of a signal transmitted from a terminal to a base station. In addition, although the LTE or LTE-A system may be described below as an example, the embodiments of the present disclosure may also be applied to other communication systems having similar technical backgrounds or channel types. For example, the 5th generation mobile communication technology (5G, new radio, NR) developed after LTE-A may be included in a system to which the embodiments of the present disclosure may be applied, and 5G below may also be a concept that includes existing LTE, LTE-A, and other similar services. Furthermore, the present disclosure may be applied to other communication systems with some modifications, as determined by a person skilled in the art, without significantly departing from the scope of the present disclosure. It will be appreciated that each block of the processing flow diagrams and combinations of the flow diagrams can be executed by computer program instructions.

[0026] These computer program instructions may be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for performing the functions described in the flowchart block(s). These computer program instructions may also be stored in a computer-available or computer-readable memory that can be directed to a computer or other programmable data processing apparatus to implement functions in a particular manner, so that the instructions stored in the computer-available or computer-readable memory can produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions may also be installed on a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable data processing apparatus to create a computer-implemented process, so that the instructions executing on the computer or other programmable data processing apparatus can provide steps for performing the functions described in the flowchart block(s).

[0027] Additionally, each block may represent a module, segment, or portion of code that contains one or more executable instructions for executing a specific logical function(s). It should also be noted that in some alternative implementation examples, the functions mentioned in the blocks may occur out of order. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order depending on the corresponding function. In this case, the term '~unit' used in the present embodiment means software or a hardware component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the '~unit' may perform certain roles. However, the '~unit' is not limited to software or hardware. The '~unit' may be configured to be on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the '~ unit' includes components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and '~ units' may be combined into a smaller number of components and '~ units' or further separated into additional components and '~ units'. In addition, the components and '~ units' may be implemented to reproduce one or more CPUs within a device or a secure multimedia card. Also, in an embodiment, the '~ unit' may include one or more processors.

[0028] For convenience of explanation, the present invention uses terms and names defined in the 5GS and NR standards, which are standards defined by the 3rd Generation Partnership Project (3GPP) among the existing communication standards. However, the present invention is not limited to the above terms and names and can be equally applied to wireless communication networks that follow other standards. For example, the present invention can be applied to the 3GPP 5GS / NR (5th generation mobile communication standard).

[0029] Below, the invention proposed in this disclosure is explained based on the premise that it applies technologies such as WebRTC data channel and SWAP.

[0030] Typically, next-generation interactive services can be launched between users. For example, voice and video calls over the Internet have become increasingly popular, and many users are using Over-the-Top (OTT) services like Skype, WhatsApp, and Zoom for their communication needs. These services can provide users with a convenient and cost-effective way to make voice and video calls over the Internet. However, these OTT services use proprietary formats for data exchange between applications, which can lead to interoperability issues when communicating with applications for third-party OTT services. This is because each service is designed to differentiate itself from other services. Therefore, there is generally no need to exchange information between apps operating within other services or exchange data.

[0031] A Mobile Network Operator (MNO) provides a standardized method for multiple network-connected device / app vendors to implement their products on the network's connectivity services, ensuring seamless interoperability between devices from different manufacturers. For example, MNOs provide connectivity between endpoints (e.g., mobile devices or other network entities) and can therefore provide and guarantee the Quality of Service (QoS) of inter-endpoint communications. QoS is one of the most important quality metrics for real-time communication use cases, such as video conferencing, voice calls, or online gaming, where communication quality is crucial to the success of the application.

[0032] When providing these services on standardized mobile networks like 3GPP, preferred formats can be selected depending on the terminal or carrier. This necessitates supporting more than one format for a single service. Therefore, when the formats supported by a first terminal from a first manufacturer and a second terminal from a second manufacturer differ, it is necessary to understand the differences in desired formats between the terminals and to address compatibility issues arising from these format differences.

[0033] For audio and video formats, the Session Description Protocol (SDP) can be used as an example of a means for negotiating content codecs and profiles between parties (endpoints) seeking to communicate. SDP is a protocol that describes the parameters of a communication session, and codecs can be negotiated using SDP offers and answers.

[0034] The endpoint initiating the call generates an SDP Offer, which may include available codecs, media types (audio, video, etc.), port information, and more. For example, when an SDP Offer is transmitted to the target endpoint, the target endpoint can select one or more media and codecs from among the codecs and media types supported by the terminal's hardware and operating system design, and return them in an SDP answer. This allows the call originating endpoint and the target endpoint to decide which media and codecs to use, such as after negotiation, before transmission can begin. At this time, data can be transmitted using the port and IP information specified in the SDP.

[0035] Typically, the endpoint wishing to initiate a call doesn't know the IP address of the target endpoint. Therefore, it can use an identifier, such as a phone number or the ID of a service to which the target user subscribes, to send an SDP offer to the call broker. The call broker then forwards the SDP offer to the endpoint matching the identifier (such as a phone number or ID), receives an SDP answer in return, and forwards it to the call initiating endpoint.

[0036] If the media's codecs do not match, the broker can communicate with the Media Resource Function (MRF) to request conversion from the first codec to the second codec and vice versa. Furthermore, the broker can authorize the MRF to convert the first codec transmitted from the call originating endpoint, and then transmit it to the call destination endpoint as the second codec.

[0037] Below, we describe the SDP and MRF problems described above.

[0038] Video and audio have limited codec types. Furthermore, because they require hardware that supports the corresponding codec, negotiation is essential before establishing a call session. Data transfer sessions, on the other hand, are used to exchange data such as text, files, and signals between clients like browsers. Therefore, there is no prior negotiation regarding the data specifications to be transmitted. WebRTC data channels, an example of a data transfer session being considered by 3GPP, use the aforementioned SDP to establish an end-to-end session, but do not involve data negotiation.

[0039] Accordingly, 3GPP defined the Simple WebRTC Application Protocol (SWAP) to complement WebRTC. The SWAP protocol is a protocol for negotiating between endpoints or servers in WebRTC data communication. When an endpoint transmits a SWAP message to WebRTC, the WebRTC signaling server acts as a SWAP server and can perform actions in response to the requested message.

[0040] As an example, a SWAP message can include messages such as Register, Response, Connect, Accept, Update, Reject, Close, and Application, each of which is described below.

[0041] - The Register message is for the endpoint to register itself with the SWAP server and can provide Matching Criteria that can be used to match incoming connection requests to the SWAP server.

[0042] A match condition may include at least one of an address, a service identifier, a user identifier, an application identifier, a location identifier, a QoS identifier, and a processing profile identifier.

[0043] - The Response message is a response that the SWAP server must send to every request it receives, and can indicate whether the message was acknowledged or contained an error.

[0044] - The Connect message can be used to establish a connection from a source endpoint (the entity to be connected, such as a terminal) to another endpoint (the target entity, such as a terminal or server). The Connect message must include an SDP offer and a match condition, which is a parameter that identifies the target endpoint.

[0045] - The Accept message is a message sent when the endpoint accepts the source and may include an SDP answer.

[0046] - The Update message may contain an updated SDP for adding / updating / removing media streams. The endpoint responds to an Update request from the source with an Accept message.

[0047] - A Reject message is sent when an endpoint does not accept a request from a source, and may include an explanation of why the message was rejected.

[0048] - The Close message can be triggered by either the source or the endpoint. Upon receiving the Close message, the opposite endpoint must respond with an Accept message, after which the WebRTC session can be released and any associated resources released.

[0049] - Application messages are messages defined by the application and can be exchanged between endpoints. The application message must include type information (e.g., a URN) that uniquely identifies the type of application message, and if the message type is not supported, the endpoint may reject it.

[0050] The operation of the SWAP protocol using at least one of the above SWAP messages is as follows.

[0051] First, a registration step is performed. An endpoint can register itself with the SWAP server. Other endpoints can also register themselves with the SWAP server. During registration, matching conditions can be submitted to describe the endpoint's capabilities and address.

[0052] - The App-specific configuration phase is performed. In this phase, a matching condition is transmitted to specify the opposing endpoint. When the first endpoint sends an application message to the SWAP server, transmitting the configuration information for the application and the matching condition, the SWAP server can transmit the configuration information of the first endpoint to the second endpoint that satisfies the matching condition. Once the configuration is completed at the second endpoint, the SWAP server transmits the connection information of the second endpoint to the first endpoint, thereby completing the application-specific negotiation phase.

[0053] The connection phase is implemented. The first endpoint can send a Connect message with the address of the second endpoint and an SDP offer. The second endpoint can then send an Accept message and an SDP answer. This completes the negotiation and establishes a data channel.

[0054] Recently, 3GPP has been exploring various next-generation calling technologies to enhance existing voice and video calling experiences. One of these technologies is real-time calling using avatars. Avatars are characters that can express a user's facial expressions and posture in mixed virtual reality environments like AR / VR. Avatars have already been introduced in games and messaging services. 3GPP aims to enhance users' immersive experiences in both real and virtual spaces by providing avatar calls that convey the user's facial expressions and posture in real time.

[0055] Avatars must transmit control signals and related information about facial expressions and postures through data channels, rather than through traditional media such as audio or video.

[0056] When providing avatar calling services on mobile networks, multiple terminal manufacturers and avatar solution providers may offer different avatar format preferences and hardware-dependent product implementations. The aforementioned data transmission session can use a WebRTC data channel, and the negotiation phase can be performed using SWAP.

[0057] SWAP is designed for pre-negotiation of data channels, but the following issues need to be addressed:

[0058] - SWAP does not have a way to specify information for the avatar currencies supported by the endpoint in the "Match Conditions".

[0059] - Even if a method for describing information for avatar calls is provided, the "application-specific negotiation step" does not provide a solution for cases where, for example, addresses match but the formats supported by each endpoint do not match.

[0060] Formats can be either mutually exclusive or inclusive, but there's no way to determine this. For example, if they're exclusive, formats A and B may be incompatible. If they're inclusive, A may be a subset of B, making them mutually supportive.

[0061] In order to provide compatibility of avatar calls between different terminals and solution implementations, the present invention provides a method for indicating that a data session to be established is an avatar call, and a method for indicating that the data to be transmitted is one of the avatar call data and which format it corresponds to.

[0062] In order to provide compatibility of avatar calls between different terminals and solution implementations, the present invention provides a method for identifying differences in formats between terminals, a method comprising a conversion process for converting between formats, and a method for establishing an avatar call data channel including the conversion process between formats.

[0063] Below we describe the avatar workflow.

[0064] Figure 1 is a diagram illustrating the workflow of a typical avatar call. According to one embodiment, an avatar call may consist of a concatenation of different types of avatar data and avatar processes that generate one avatar data type from another. A terminal supporting 3GPP avatar calls can understand and support at least one avatar data and one avatar process.

[0065] First, a user can be captured. Data (110) referred to as "Captured Data" can be generated through the generation of captured data (100) about the user. Capturing can be performed using a camera or sensor device. Examples of captured data include the user's facial expressions and / or posture. Alternatively, any input used to represent the user as an avatar, such as voice or text (e.g., user-generated commands), can be captured.

[0066] Users can create their own avatars (300) in real time or in advance through avatar calls. Avatar data that can represent a user is called a Base Avatar (310). The Base Avatar can be generated from captured data by a process called Base Avatar Generation (310) and stored or retrieved by a process called Avatar Storage (400).

[0067] Meanwhile, the captured data can be converted (200) into animation data (210), which is a command that can move the base avatar. The animation data is a command that can move the base avatar and can be converted from the captured data. For example, the captured data can be a video image taken of the user, and the animation data can be converted into a command that pulls the right eyebrow downward by about 10%. The conversion from the captured data to the animation data can be performed by the animation data generation process.

[0068] The above animation data is specifically compared based on Figures 2 and 3.

[0069] The formats of Animation Data that can be considered are as follows. First, Facial landmarks can be considered. According to one embodiment, as illustrated in FIG. 2, a Facial landmark is a set of facial feature points, such as the location of the nose, the start and end points of the eyes, or the location of the pupils. Depending on the type of solution that extracts and tracks Facial landmarks from a facial image, there may be solutions that extract 87, 300, or 478 Facial landmarks from the same image, for example. Since formats are continuously proposed by solution developers, the present invention is not limited to the solutions listed as examples, but for convenience, the description is based on the similarities and differences between the three solutions. Since each solution stores Facial landmarks in a different format, it can be said that there are at least three different formats in terms of format. However, from a semantic perspective of information, the three formats can be analyzed as solutions that are mutually inclusive and include each other, such as a solution that extracts 300 including 87 from the solution with the smallest number of 87 extractions, and a solution that extracts 478 including 300. In other words, the format conversion process provided by the present invention can extract 300 facial landmarks or 87 facial landmarks from a format storing 478 facial landmarks and convert them into each format, or can infer 300 and 478 facial landmarks from 87 based on an algorithm that understands the human muscle structure and convert them into each format.

[0070] Blend shapes are another format for animation data. Blend shapes represent reserved words and the degree of transformation for human body parts that interact with each other based on the structure of human facial muscles, such as raising an eyebrow by about 10%. For example, some people can move only their eyebrows independently, but raising one eyebrow can cause the pupil to dilate, and the other eyebrow to rise proportionally. Therefore, multiple reserved words are typically used to describe a single expression. Depending on the solution, the total number of reserved words included in Blend shapes, how they are displayed, and how the degree of transformation is expressed may vary. However, from the perspective of describing the effect of facial muscles on various facial organs, like Facial Landmarks, Blend shape formats can be interchanged based on the common information they contain.

[0071] Facial landmarks and blend shapes are different representations of the same information, and can be generated and inferred based on the human facial muscle structure, so a mutual transformation process can also be established between facial landmarks and blend shapes.

[0072] A user's posture, along with facial expressions, can be described by the position, orientation, and rotation of the skeleton and its connecting parts, as illustrated in Figure 3. As illustrated in Figure 3, there may be various formats, each varying depending on the name of the skeleton, the method of describing the motion, and the format in which the information is stored.

[0073] Hereafter, the following steps will be described based on Fig. 1 again. The 3GPP Avatar Animation process (500) can generate an Animated Avatar (510) using a Base Avatar (310) and Animation Data (210). For example, an avatar that mimics a user can be generated from information describing the user's status. If the Base Avatar is human body information stored in 3D graphics, changing each graphic element of the Base Avatar according to facial expression and posture can produce a result that reflects the user's facial expression and posture. In one embodiment, the Base Avatar and the Animated Avatar may be information in the same format but with different variable values. In another example, a process that can directly generate a 3D Animated Avatar from the aforementioned Captured Data, for example, if there is an artificial intelligence model, can be a solution that can handle data that is compatible from the perspectives of the Captured Data and the Animated Avatar without going through the Animation Data generation step.

[0074] The generated Animated Avatar can be placed in space by the Scene Management process (600). The Scene (610) is composed of a space and objects placed in the space, and the Animated Avatar can be considered as one object among the objects. The size of the space and the placement of objects in the space can vary depending on the users participating in the avatar call session. For example, for the first user, the placement can be such that the other party faces each other within a limited area considering the interior of an indoor space, whereas for the second user in the same session, the other party can be placed to the right of the viewing direction within a relatively wide area considering movement outdoors. The positions of the conversation participants themselves and the other party within the Scene can be determined by the participants, the other party, or the Scene Management process can determine them at its own discretion.

[0075] A Scene can be rendered by a SceneRenderer (700) existing on a terminal or server and converted into a format that can be displayed on the terminal's display. The converted content can be referred to as a RenderedScene (710).

[0076] Below we describe in detail how to describe information for avatar currencies supported by an endpoint.

[0077] As described above, the workflow of an avatar currency can consist of various avatar data types, various formats within a data type, and processes that perform conversions between data types.

[0078] An endpoint wishing to initiate an avatar call must establish a session with the other endpoint to transfer avatar data. To obtain the address of the other endpoint, the endpoint must register itself with the server provided by the avatar call service provider. For this purpose, the endpoint can use the aforementioned SWAP server.

[0079] When registering with a call service provider that also offers voice and video calls, an endpoint must specify its ability to make avatar calls in the matching criteria that other endpoints search for. Once this matching criteria is specified, other endpoints can identify voice, video, and avatar communication capabilities when searching for a call with the endpoint and select one of the available call methods.

[0080] The following information elements are suggested as methods for specifying avatar currency, but are not limited to:

[0081] The matching conditions for SWAP's Register include:

[0082] - ipv4: The IPv4 address of the target endpoint

[0083] -ipv6: The IPv6 address of the target endpoint

[0084] - fqdn: The FQDN of the target endpoint

[0085] - service: An identifier of a service or an application

[0086] - user: An identifier of the user such as a SIP address, a GPSI, or an MSISDN

[0087] - eas: An EAS identifier

[0088] - app: application-specific matching criteria that is compared using binary or string comparison

[0089] - location: one or more identifiers of a geographic location or area

[0090] - qos: a description of the QoS that is supported by the connection to the endpoint

[0091] - processing: a profile description of the processing capabilities of the endpoint.

[0092] Since the details of each item are not defined, the present invention provides a method for describing the matching conditions for avatar call performance of an endpoint as a subitem of the app item for the purpose of negotiation between endpoints for avatar calls. For convenience, the items are described based on the app, but it is understood that each service provider's usage is not limited to this and can be defined and used, for example, under processing.

[0093] Below we describe how to indicate that a data session is an avatar call.

[0094] The display method of avatar communication according to the present invention can describe urn, callType, and callSubType as sub-items under app.

[0095] - urn is used to distinguish this application from other applications, and can indicate that it is registered with 3GPP in URN format, that it is XR communication of SA4 in detailed categories, that it is Avatar communication, and that it is based on Scene Description. Accordingly, it can be expressed as urn:3GPP:SA4:NewCall:Avatar:SD.

[0096] - callType and callSubType can be described with the URN. callType can describe one or more of the various communication methods supported by the endpoint, such as Voice, Video, Avatar, ScreenContent, Haptics, VR, and AR. callSubType can describe information to further explain the callType, and can describe Spatial, 2D, 3D, Scene description, and Implicit Neural Representation. For example, callType:Avatar, callSubType:SceneDescription can indicate an avatar call based on Scene description. callType:Voice, callSubType:Spatial can indicate an audio call in spatial audio format.

[0097] Table 1 is a table that specifically explains the above urn, callType, and callSubType.

[0098] FieldDescription urn Indicates the various communication methods and sub-divisions supported by the endpoint itself. For example, urn:3GPP:SA4:NewCall:Avatar:2024 can be used for avatar calls, which were defined by the 3GPP SA4 working group in 2024 for a new calling experience. callType Lists all communication methods supported by the endpoint itself. Voice, Video, Avatar, ScreenContent, Haptics, VR, AR, etc. can be considered. callSubType Lists all information that provides additional explanation for each communication method supported by the endpoint itself. Additional explanation information can include Spatial, 2D, 3D, Scene description, and Implicit Neural Representation.

[0099] Another sub-item under app is to describe the avatar data to be transmitted for avatar calls and the format of the avatar data.

[0100] - An endpoint can specify one or more avatar data that it can provide to other endpoints, the format of the data, one or more avatar data that it can receive and play from other endpoints, and the format of the one or more avatar data.

[0101] - If the data that can be transmitted and the data that can be received do not match, additional directions for providing / receiving, such as sendOnly, receiveOnly, or sendAndReceive, can be provided as Direction.

[0102] - Avatar data is as shown in Fig. 1 and the above, and the avatar data format may have different formats for each avatar data. In the case of captured data, sub-items that can be considered include 2D / Stereoscopic / 3D Video, Audio, Text, Sensor data, etc. For example, the CMAF streaming format can be used as a video format, HEVC can be used as a video compression codec, and Main10 can be used as a compression profile of a compression codec, so these need to be described.

[0103] - As a sub-item of the App item mentioned above, you can display avatar data as DataType, avatar data format as DataFormat, and codec and profile as DataCodec and DataProfile, respectively.

[0104] Table 2 is a table that specifically describes the sub-items of the above app item.

[0105] FieldDescriptiondataType Specifies all avatar data in the avatar currency workflow that the endpoint itself supports. CapturedData, AnimationData, AnimatedAvatar, Scene, RenderedScene, etc. can be considered.dataFormat Specifies all transmission formats of data that the endpoint itself supports. It can be a list of urns registered with the service provider / organization or server, such as urn:mpeg:dash:schema:mpd:2011 (DASH), urn:mpeg:dash:profile:cmaf:2019 (CMAF), urn:3gpp:ar-mtsi:v1:sd (Scene description), etc.dataProfile Specifies all compression methods of data that the endpoint itself supports. It can be written as a list of urns registered with service providers / organizations or servers, such as urn:3GPP:26119:18:HEVC-FullHD-Dec. The direction endpoint indicates whether it can only provide (sendOnly), only receive (receiveOnly), or provide and receive (sendAndReceive) for the DataProfile of the DataFormat of the above DataType.

[0106] Another sub-item under app could be the avatar process that handles avatar calls.

[0107] - An endpoint can specify one or more avatar processes that it can handle, and the input and output data of those avatar processes and their formats.

[0108] If the endpoint is intended for processing, such as a virtual machine on a cloud computing server, performance metrics such as processing time can be specified in addition to simple input / output formats. Since processing time can be correlated with energy consumption, energy consumption metrics can also be specified. Since energy consumption can lead to a virtual or actual temperature increase in the server machine, temperature metrics related to energy consumption can also be specified.

[0109] Table 3 is a table that specifically describes the avatar process that can be described as a sub-item of the above app item.

[0110] FieldDescriptionprocess The avatar process that the endpoint itself has stated that it supports, such as UserCaptureDataGeneration, AnimationDataGeneration, BaseAvatarGeneration, AvatarStorage, AvatarAnimation, SceneManagement, and SceneRendering. Additionally, transcode can be considered, and transcode means conversion between the data format specified in the input and the data format specified in the output.performanceIndex It means the performance of the avatar process. It can contain time, energy, temperature, memory, input, and output.time It means the time that the avatar process processes for a unit of daily avatar data. For example, when the input of the AvatarAnimation process is AnimationData, the user's daily facial expression and posture can be a unit of AnimationData, and the time it takes the AvatarAnimation process to move the avatar with the corresponding facial expression and posture is the processing time.energy It describes the energy consumed by the avatar process for a unit of time when it is in continuous operation, in watts. The end point of the job that the avatar process will use can be multiplied by the usage time to calculate the total amount of energy the avatar process will consume in units of KHW. temperature Describes the range of temperature change that the avatar process can rise in unit time while it is running continuously. memory Describes the minimum amount of memory that the avatar process requires while running. input Describes the input data and data format for each avatar process. It can include dataType, dataFormat, dataProfile, etc. output Describes the output data and data format for each avatar process.It can include dataType, dataFormat, dataProfile, etc.

[0111] When an endpoint sends a Register message to the SWAP server, including its matching conditions, the SWAP server can store the endpoint and its matching conditions and send a Response message to the endpoint.

[0112] If the response does not contain any errors, the endpoint can consider itself to have registered successfully.

[0113] Below, an example is specifically described in which the address matches and the format matches.

[0114] According to one embodiment of the present disclosure, an endpoint may request a SWAP server to initiate an avatar call by sending an App-specific message specifying the address or identifier of the target endpoint and specifying that the intended call type is an avatar call. For example, a first endpoint may transmit an App-specific message to the SWAP server specifying the address of a second endpoint and an avatar call within a matching condition.

[0115] The SWAP server can find an endpoint that matches the matching criteria, forward an app-specific message to that endpoint, and receive a response and forward it to the original endpoint. For example, the SWAP server can forward the app-specific message from the first endpoint to a second endpoint. Once the second endpoint is ready for the avatar call, it can reply with a data transfer port for the call and forward it to the first endpoint.

[0116] FIG. 4 is a sequence diagram specifically illustrating an embodiment of a method for initiating data communication for avatar calling according to one embodiment of the present disclosure.

[0117] In step 401, the registration step of endpoints can be performed.

[0118] In step 402, the first endpoint can register itself with the SWAP server. Matching conditions may include avatar currency.

[0119] In step 403, the second endpoint can register itself with the SWAP server. Matching conditions may include avatar currency.

[0120] At step 404, an App-specific message exchange step may be performed.

[0121] In step 405, the first endpoint can transmit an app-specific message to the SWAP server for an avatar call with the second endpoint. The matching condition created to request an avatar call may include an identifier for the second endpoint, an indicator (e.g., urn, callType) indicating that the call is an avatar call, and avatar data that the first endpoint can generate and transmit, along with information about the format of the avatar data.

[0122] In step 406, the SWAP server can find an endpoint that matches all other matching conditions, including the identifier.

[0123] In step 407, the configuration message of the first endpoint can be delivered to the second endpoint as a result of the search.

[0124] At step 408, the SWAP server may notify the first endpoint that the message has been delivered to the second endpoint.

[0125] At step 409, the second endpoint may prepare to initiate the requested avatar call. For example, actions may be taken, such as launching an avatar call app, opening a port for data communication, or suspending or backgrounding other currently running top-level apps.

[0126] At step 410, the second endpoint is ready to initiate an avatar call and may pass the address and port information for accessing the second endpoint to the SWAP server.

[0127] At step 411, the SWAP server can forward the message from the second endpoint to the first endpoint.

[0128] At step 412, the SWAP server may notify the second endpoint that the message has been delivered to the first endpoint.

[0129] At step 413, a connection step may be performed.

[0130] In step 414, the first endpoint may transmit a CONNECT message containing an SDP offer to the SWAP server. The SWAP server may then transmit the CONNECT message containing the SDP offer to the second endpoint. The CONNECT message may include the address of the second endpoint, and the SDP offer may include a data channel request. Additional media, such as audio or video, may be included in the SDP offer.

[0131] At step 415, the SWAP server may notify the first endpoint that the message has been delivered to the second endpoint.

[0132] In step 416, the second endpoint may forward an ACCEPT message containing an SDP answer to the SWAP server. The SWAP server may forward the ACCEPT message containing the SDP answer to the first endpoint. The SDP answer may include whether the second endpoint accepts additional media, such as voice or video.

[0133] At step 417, the SWAP server may notify the second endpoint that the message has been delivered to the first endpoint.

[0134] Data communication for avatar calls may be initiated at step 418. Media communication for voice and video may also be initiated at the same time.

[0135] FIG. 4 is characterized by including a step 404 for determining whether avatar calls between the endpoints are possible before establishing a data channel for data transmission and reception for avatar calls between the first and second endpoints. In FIG. 4, it is assumed that the first and second endpoints can support all types of avatar data and data formats between the endpoints if avatar calls are possible.

[0136] If the format of the avatar data provided by the first endpoint is not supported by the second endpoint, avatar calls between the two endpoints will not be possible if the data channel is established without proper pre-negotiation procedures. Since the data transmitted after the CONNECT connection does not include the negotiation phase, both endpoints may become unable to process data in an unfamiliar format.

[0137] To prevent this, the types of avatar data proposed in the present invention, as well as the format, codec, and profile information for each data type, need to be described. Furthermore, the SWAP server according to the present invention compares the types, formats, codecs, and profiles of avatar data provided by the first endpoint with the types, formats, codecs, and profiles that the second endpoint can receive. If at least one of the lists of each endpoint matches, the first endpoint and the second endpoint are determined to be capable of conducting an avatar call, and the second endpoint can be selected as the matching endpoint.

[0138] If the avatar data provided by the first endpoint is not included in the list of avatar data that can be received by the second endpoint, the SWAP server according to the present invention must report to the first endpoint that there is no matching endpoint.

[0139] To this end, the SWAP server according to the present invention may reply to the first endpoint with a Response message, indicate the type of the Response message as an error, and describe the content as disclosed in Table 4 below as a description.

[0140] Error messageDescriptionnotSupportedCallTypeThe requested call format is not supported. If an avatar call is requested, notSupportedCallType means that no endpoint supports avatar calls.notSupportedDataTypeThe requested call format is supported by at least one endpoint, but the data format is not supported by that endpoint.notSupportedDataFormatThe requested call format and data format are supported by at least one endpoint, but there is no matching profile among the supported data formats.notSupportedDataProfileThe requested call format and data format are supported by at least one endpoint, but there is no matching profile among the supported data formats.

[0141] The endpoint that requested the avatar call receives a response message from the SWAP server, and if the type is error, it receives one of the error messages mentioned above in the description, which allows it to determine why the avatar call is not possible.

[0142] Meanwhile, the following explains how to resolve the issue when the addresses match but the formats do not.

[0143] According to one embodiment of the present disclosure, the first endpoint and the SWAP server can make the following judgment when there is no counterpart endpoint matching the specified matching condition.

[0144] If the address or identifier matches, it can be determined that the corresponding endpoint is the counterparty to the conversation, but there is a compatibility issue in the form of the currency it provides.

[0145] Based on this initial judgment, you can initiate a call other than an avatar call, such as a voice or video call. For example, if you receive a notSupportedCallType, you can determine that the other endpoint is unlikely to accept an avatar call in any format.

[0146] Another second consideration is that format conversion may be required to facilitate avatar calls. For example, if the call type is not "notSupportedCallType," avatar calls are supported, but the data format, format, or profile are outside the supported range. Therefore, conversion may be considered to bring the call within the supported range.

[0147] Conversion between formats can be selected and determined by the first endpoint requesting the conversation, the second endpoint receiving the conversation request, or the SWAP server connecting the conversation. When requested by an endpoint, instructions such as those described in Table 5 below may be included within a match condition or app-specific configuration message.

[0148] FieldDescriptionConditionalTranscoding: Requests that a transformation be performed if a transformation is required. The SWAP server or the other endpoint can perform the transformation using an endpoint within or outside the endpoint for conditionalTranscoding. ForceMFProcessing: Forces at least one network processing operation on the received data. For example, in the case of avatar calls, for a single data type such as CapturedData, the associated processing is GenerateAnimationData, so the MF process that generates AnimationData is performed, and the endpoint for performing this operation is searched for or created before connection is made.

[0149] In one embodiment, upon receipt of conditionalTranscoding or forceMFProcessing, the SWAP server may determine the connection method that most closely matches the matching conditions specified by the first and second endpoints. For example, if the callType matches, the callSubType matches, and the dataType matches but the dataFormat does not match, the SWAP server may determine whether to perform format conversion to resolve the data format mismatch between the two endpoints. For example, if the dataFormat matches but the dataProfile does not match, the SWAP server may determine whether to perform profile conversion to resolve the profile mismatch.

[0150] In another embodiment, the first endpoint does not request either conditionalTranscoding or forceMFProcessing, but simply requests a connection with the second endpoint via an app-specific configuration message, but the SWAP server may determine, at its discretion, whether to perform a conversion to one of the data types, data formats, or data profiles.

[0151] Accordingly, the first and second endpoints may be judged to be partially matched even if they present matching conditions that do not completely match, and may be supplemented to match according to the conversion process of the separate endpoints.

[0152] FIGS. 5A and 5B are sequence diagrams specifically illustrating the aforementioned conversion steps according to one embodiment of the present disclosure. The detailed steps are as follows.

[0153] In step 501, the registration step of endpoints and Media Functions can be performed.

[0154] In step 502, the first endpoint can register itself with the SWAP server. Matching conditions may include avatar currency, animationData data type, and data format #1.

[0155] In step 503, the second endpoint can register itself with the SWAP server. Matching conditions may include avatar currency, animationData data type, and data format #2.

[0156] In step 504, the Media Function (MF) can register itself with the SWAP server. Matching conditions may include avatar currency and conversion process, and input and output formats #1 and #2, respectively.

[0157] At step 505, an App-specific message exchange step may be performed.

[0158] At step 506, the first endpoint may transmit an app-specific message to the SWAP server for an avatar call with the second endpoint. The matching condition created to request an avatar call may include an identifier for the second endpoint, an indicator (e.g., urn, callType) indicating that the call is an avatar call, and avatar data that the first endpoint can generate and transmit, along with information about the format of the avatar data.

[0159] In step 507, the SWAP server may find an endpoint that matches all other matching conditions, including the identifier. In one embodiment, the SWAP server may find a second endpoint that matches some conditions (identifier, avatar currency support, avatar data), but does not match the conditions regarding the avatar data format.

[0160] Transcoding configuration by the SWAP server can be performed at step 508.

[0161] At step 509, the SWAP server may determine that the second endpoint matches the request from the first endpoint. However, it may determine that this requires conversion from the first format to the second format.

[0162] In step 510, the SWAP server can find an MF that has registered a conversion process that takes the first format as input and the second format as output and request a conversion that takes the data of the first endpoint as input.

[0163] At step 511, the Media function (MF) can prepare for conversion. It can open ports to prepare the process for execution, receive input data, and send output data.

[0164] After the conversion preparation is completed in MF at step 512, the preparation is completed and the input / output information at step 511 can be transmitted.

[0165] Avatar call configuration can be performed by the SWAP server at step 513.

[0166] At step 514, the SWAP server may decide to match the MF with the second endpoint.

[0167] At step 515, the SWAP server may forward the App-specific message received from the MF to the second endpoint. The format information forwarded to the second endpoint is based on the second format that the MF converts from the first format and that the second endpoint can receive. The caller address of the second endpoint may be notified as being the MF.

[0168] At step 516, the SWAP server may notify the first endpoint that an App-specific message regarding the avatar call request has been delivered.

[0169] At step 517, the second endpoint may prepare to initiate the requested avatar call. For example, actions may be taken, such as launching an avatar call app, opening a port for data communication, or suspending or backgrounding other currently running top-level apps.

[0170] At step 518, the second endpoint is ready to initiate an avatar call and may pass the address and port information for accessing the second endpoint to the SWAP server.

[0171] At step 519, the SWAP server can forward the message from the second endpoint to the MF.

[0172] At step 520, the MF may notify the SWAP server that the conversion is ready.

[0173] At step 521, the SWAP server may forward the second endpoint message (call preparation complete) of step 18 to the first endpoint. The first endpoint may be notified of the MF with the address of the target endpoint.

[0174] At step 522, the SWAP server may notify the second endpoint that the second endpoint message of step 18 (call preparation complete) has been delivered.

[0175] A connection step may be performed at step 523.

[0176] In step 524, the first endpoint may transmit a CONNECT message containing an SDP offer targeting the MF to the SWAP server. The SWAP server may then transmit a CONNECT message containing the SDP offer to the MF. The CONNECT message may include the MF's address, and the SDP offer may include a data channel request for transmitting the first format. Additionally, additional media, such as voice or video, may be included in the SDP offer.

[0177] At step 525, the SWAP server may send a CONNECT message to the MF.

[0178] At step 526, the MF may forward a CONNECT message containing an SDP offer targeting the second endpoint to the second endpoint. The CONNECT message may include the address of the second endpoint. The SDP offer may also include a data channel request for transmitting the second format. Additional media, such as voice or video, may be included in the SDP offer.

[0179] In step 527, the second endpoint may send an ACCEPT message containing an SDP answer to the SWAP server. The SDP answer may include whether the second endpoint accepts additional media, such as voice or video.

[0180] At step 528, the SWAP server may forward an ACCEPT message to the MF.

[0181] At step 529, the MF may forward an ACCEPT message containing an SDP answer to the first endpoint via the SWAP server. If the content for the second format is included, it may be replaced with the content for the first format.

[0182] At step 530, data communication for avatar calls may be initiated. Data from the first endpoint may be transmitted to the MF. Media communication for voice and video may also be initiated at the same time.

[0183] Data communication for avatar calls may be initiated at step 531. The MF may convert the first format into a second format and transmit it to the second endpoint.

[0184] In order to determine whether the formats of avatar data can be mutually converted, such as conversion between formats as described above, or processing between data, etc., the formats of avatar data can be provided.

[0185] For example, a second format may be capable of conversion to and from a first format. In another example, a third format may not be capable of conversion to and from a first or second format.

[0186] If the SWAP server supports receiving and configuring app-specific messages for avatar calls, it can have 'information on whether conversion is possible' between the above formats and instruct the MF to perform conversion processing based on the above information.

[0187] According to one embodiment, the first MF may be registered with the SWAP server as capable of mutual conversion from the first format to the second format during the pre-registration stage. Furthermore, if the supported formats between the first and second endpoints do not match, the SWAP server may select the first MF that supports the conversion among the pre-registered MFs and establish a connection as illustrated in FIG. 5.

[0188] In another embodiment, if the supported formats between the third endpoint and the fourth endpoint do not match, the SWAP server may check the 'information on whether conversion is possible' described above to see if conversion between the third format of the third endpoint and the fourth format of the fourth endpoint is possible, and if conversion is determined to be possible, instruct the creation of an MF that supports the conversion, and wait until the MF is created and registered before finalizing the connection between the third endpoint and the fourth endpoint.

[0189] During the waiting time until the MF providing the reciprocal conversion from the 3rd format to the 4th format is created and registered, the SWAP server can respond with a message such as "waiting for response from endpoint" as the description of the Response message so that the 3rd endpoint can wait for the call connection in an appropriate waiting state.

[0190] Information on whether conversion is possible may be as disclosed in Table 6 below.

[0191] FieldTypeDescriptionformatIdentifierstring The identifier of the format. dataTypestring The data type to which the format belongs (e.g., one of the avatar data types). transcodableFormatsformatIdentifier The identifiers of formats that can be converted to each other.

[0192] Meanwhile, FIG. 6 is a block diagram illustrating components of a terminal according to an embodiment of the present disclosure.

[0193] Each of the terminals described in the present disclosure (e.g., the first endpoint and the second endpoint) may correspond to either the first endpoint or the second endpoint, such as the caller or callee described in FIG. 4 or FIG. 5a / 5b.

[0194] Referring to FIG. 6, the terminal may include a transceiver unit (605), a control unit (e.g., at least one processor) (610), and a storage unit (615).

[0195] The transmitter and receiver (605) can transmit and receive signals, information, data, etc. to and from the server.

[0196] A transceiver (605) according to one embodiment of the present disclosure can communicate with a swap server.

[0197] Meanwhile, the control unit (610) is a component for overall control of the terminal. The control unit (610) can control the overall operation of the terminal according to various embodiments of the present disclosure. According to one embodiment of the present disclosure, the control unit (610) may include at least one processor.

[0198] A control unit (610) according to one embodiment of the present disclosure can control to transmit a setup message, a message regarding specifications, etc. to a swap server.

[0199] Meanwhile, the terminal may further include a storage unit (615) and may store data such as basic programs, application programs, and setting information for the operation of the terminal. In addition, the storage unit may include at least one storage medium among a Flash Memory Type, a Hard Disk Type, a Multimedia Card Micro Type, a memory of a card type (e.g., an SD or XD memory, etc.), a magnetic memory, a magnetic disk, an optical disk, a Random Access Memory (RAM), a Static Random Access Memory (SRAM), a Read-Only Memory (ROM), a Programmable Read-Only Memory (PROM), and an Electrically Erasable Programmable Read-Only Memory (EEPROM). In addition, the control unit (610) may perform various operations using various programs, contents, data, etc. stored in the storage unit (615).

[0200] Meanwhile, FIG. 7 is a block diagram illustrating components of a server according to an embodiment of the present disclosure. According to one embodiment, the server of FIG. 7 may be a swap server according to the aforementioned embodiment.

[0201] Referring to FIG. 7, the server may include a transceiver unit (705), a control unit (e.g., at least one processor) (710), and a storage unit (715).

[0202] The transmitter / receiver (705) can transmit and receive signals, information, data, etc. to and from a terminal or media function.

[0203] Meanwhile, the control unit (710) is a component for overall control of the server. The control unit (710) can control the overall operation of the server according to various embodiments of the present disclosure. According to one embodiment of the present disclosure, the control unit (720) may include at least one processor.

[0204] A control unit (710) according to one embodiment of the present disclosure can receive a setting message, a message regarding a standard, etc. from a terminal and control to determine a matched endpoint.

[0205] Meanwhile, the server may further include a storage unit (715) and may store data such as basic programs, application programs, and setting information for the operation of the server. In addition, the storage unit may include at least one storage medium among a Flash Memory Type, a Hard Disk Type, a Multimedia Card Micro Type, a memory of a card type (e.g., an SD or XD memory, etc.), a magnetic memory, a magnetic disk, an optical disk, a Random Access Memory (RAM), a Static Random Access Memory (SRAM), a Read-Only Memory (ROM), a Programmable Read-Only Memory (PROM), and an Electrically Erasable Programmable Read-Only Memory (EEPROM). In addition, the control unit (710) may perform various operations using various programs, contents, data, etc. stored in the storage unit (715).

[0206] In the specific embodiments of the present disclosure described above, components included in the disclosure are expressed in the singular or plural form, depending on the specific embodiment presented. However, the singular or plural expressions are selected to suit the presented situation for convenience of explanation, and the present disclosure is not limited to singular or plural components. Components expressed in the plural form may be composed of singular elements, or components expressed in the singular form may be composed of plural elements.

[0207] While the detailed description of this disclosure has described specific embodiments, it should be understood that various modifications are possible without departing from the scope of this disclosure. Therefore, the scope of this disclosure should not be limited to the described embodiments, but should be defined not only by the scope of the claims described below, but also by equivalents thereof.

[0208] The various embodiments of the present disclosure and the terminology used therein are not intended to limit the technology described in the present disclosure to a specific embodiment, but should be understood to include various modifications, equivalents, and / or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar components. The singular expression may include plural expressions unless the context clearly indicates otherwise. In the present disclosure, expressions such as “A or B,” “at least one of A and / or B,” “A, B, or C,” or “at least one of A, B, and / or C” may include all possible combinations of the items listed together. Expressions such as “first,” “second,” “first,” or “second” may modify the corresponding components regardless of order or importance, and are only used to distinguish one component from another, but do not limit the corresponding components. When it is said that a component (e.g., a first component) is “(functionally or communicatively) connected” or “connected” to another component (e.g., a second component), the component may be directly connected to the other component, or may be connected through another component (e.g., a third component).

[0209] The term "module" as used herein includes a unit composed of hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integrally formed component, or a minimum unit or portion thereof that performs one or more functions. For example, a module may be composed of an application-specific integrated circuit (ASIC).

[0210] Various embodiments of the present disclosure may be implemented as software (e.g., a program) including instructions stored in a machine-readable storage medium (e.g., an internal memory or an external memory) that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include a terminal (e.g., a first terminal (210), a second terminal (220)) according to various embodiments of the present disclosure. When an instruction is executed by a processor (e.g., a processor (620) of FIG. 6 or a processor (720) of FIG. 10), the processor may directly, or under the control of the processor, perform a function corresponding to the instruction by using other components. The instruction may include code generated or executed by a compiler or an interpreter.

[0211] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium does not contain signals and is tangible, but does not distinguish between whether data is stored semi-permanently or temporarily on the storage medium.

[0212] The methods according to various embodiments disclosed in the present disclosure may be provided as a computer program product. The computer program product may be traded as a commodity between sellers and buyers. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0213] Each component (e.g., a module or a program) according to various embodiments may be composed of one or more entities, and some of the aforementioned sub-components may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., a module or a program) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration. Operations performed by a module, program, or other component according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

Claims

1. A method performed by a server for performing an avatar call in a wireless communication system, A step of receiving a registration message indicating a matching specification for the avatar call from the first terminal; A step of receiving a specific message for the avatar call from the first terminal after receiving the registration message; A step of determining a second terminal to perform the avatar call based on at least one piece of information included in the specific message; When a connection step is performed between the first terminal and the second terminal, a step of receiving a connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal from the first terminal; and A method comprising: transmitting the connection message to the second terminal, and when an acceptance message including an SDP response is received from the second terminal, transmitting the acceptance message to the first terminal so that communication for the avatar call is initiated; 2. In paragraph 1, The above specific message is, A method characterized in that it further includes at least one of an identifier of the second terminal, an indicator for the avatar call, avatar data associated with the first terminal, and format information corresponding to the avatar data.

3. In paragraph 1, A step of identifying the second terminal that meets any condition for the avatar call but has different format information based on the at least one piece of information included in the specific message; A step of determining that a conversion is performed from a first format corresponding to the first terminal to a second format corresponding to the second terminal; A step of identifying a media function (MF) entity associated with the first format and the second format; and A method characterized by further comprising: a step of requesting conversion of data of the first terminal into input using the MF entity; 4. In paragraph 3, The above connection message is, Includes an SDP offer targeting the above MF entity, The above connection message further includes information about the address of the MF entity, The above SDP offer further includes a channel request for transmitting data corresponding to the first format, The above received connection message is forwarded to the MF entity, A method characterized in that when communication for the avatar call is initiated, data corresponding to the first format transmitted by the first terminal is converted into the second format by the MF entity and transmitted to the second terminal.

5. A method performed by a first terminal for performing an avatar call in a wireless communication system, A step of transmitting a registration message indicating matching specifications for the above avatar call to the server; A step of transmitting a specific message for the avatar call to the server after transmitting the registration message; When a connection step is performed with a second terminal determined based on the specific message, a step of transmitting a connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal to the server; and A method comprising: a step of initiating communication for the avatar call when an acceptance message including an SDP response transmitted by the second terminal is transmitted from the server; 6. In paragraph 5, The above specific message is, A method characterized in that it further includes at least one of an identifier of the second terminal, an indicator for the avatar call, avatar data associated with the first terminal, and format information corresponding to the avatar data.

7. In paragraph 5, The above connection message is, When the second terminal is identified by the server based on at least one piece of information included in the specific message and it is determined that conversion is performed from a first format corresponding to the first terminal to a second format corresponding to the second terminal, an SDP offer targeting a media function (MF) entity is included, A method characterized in that when communication for the avatar call is initiated, data corresponding to the first format transmitted by the first terminal is converted into the second format by the MF entity and transmitted to the second terminal.

8. In a server for performing an avatar call in a wireless communication system, Transmitter and receiver; and Receive a registration message indicating a matching specification for the avatar call from the first terminal through the transceiver, After receiving the above registration message, a specific message for the avatar call is received from the first terminal through the transceiver, Based on at least one piece of information included in the specific message, control is provided to determine a second terminal to perform the avatar call; When a connection step is performed between the first terminal and the second terminal, a connection message including at least one of a session description protocol (SDP) offer and an address of the second terminal is received from the first terminal through the transceiver, A server comprising a control unit that transmits the connection message to the second terminal and, when an acceptance message including an SDP response is received from the second terminal, controls transmission of the acceptance message to the first terminal through the transceiver so that communication for the avatar call is initiated.

9. In paragraph 8, The above specific message is, A server characterized in that it further includes at least one of an identifier of the second terminal, an indicator for the avatar call, avatar data associated with the first terminal, and format information corresponding to the avatar data.

10. In paragraph 8, The above control unit, Based on the at least one piece of information included in the specific message, identifying the second terminal that meets any condition for the avatar call but has different format information; It is determined that a conversion is performed from a first format corresponding to the first terminal to a second format corresponding to the second terminal, Identifying a media function (MF) entity associated with the first format and the second format, A server characterized in that it controls the above MF entity to request conversion of data of the first terminal into input.

11. In paragraph 10, The above connection message is, Includes an SDP offer targeting the above MF entity, The above connection message further includes information about the address of the MF entity, The above SDP offer further includes a channel request for transmitting data corresponding to the first format, A server characterized in that the received connection message is transmitted to the MF entity.

12. In paragraph 10, A server characterized in that when communication for the avatar call is initiated, data corresponding to the first format transmitted by the first terminal is converted into the second format by the MF entity and transmitted to the second terminal.

13. In a first terminal for performing an avatar call in a wireless communication system, Transmitter and receiver; and A registration message indicating the matching specifications for the above avatar call is transmitted to the server through the above transceiver, After transmitting the above registration message, a specific message for the above avatar call is transmitted to the above server through the above transmitter / receiver. When a connection step is performed with a second terminal determined based on the specific message, a connection message including at least one session description protocol (SDP) offer and an address of the second terminal is transmitted to the server through the transceiver, A first terminal including a control unit that controls to initiate communication for the avatar call when an acceptance message including an SDP response transmitted by the second terminal is transmitted from the server.

14. In paragraph 13, The above specific message is, A first terminal characterized in that it further includes at least one of an identifier of the second terminal, an indicator for the avatar call, avatar data associated with the first terminal, and format information corresponding to the avatar data.

15. In paragraph 13, The above connection message is, When the second terminal is identified by the server based on at least one piece of information included in the specific message and it is determined that conversion is performed from a first format corresponding to the first terminal to a second format corresponding to the second terminal, an SDP offer targeting a media function (MF) entity is included, A first terminal, characterized in that when communication for the avatar call is initiated, data corresponding to the first format transmitted by the first terminal is converted into the second format by the MF entity and transmitted to the second terminal.

Citation Information

Patent Citations

  • Intelligent blockchain point-to-point instant messaging system

    CN117714409A

  • Air Cleaner For Desktop

    KR1020230064670A

  • Avatar call platform

    US20230199147A1