Methods and systems for content creation in media streaming
By leveraging GenAI to generate and transmit semantic information for XR media streaming, the solution addresses bandwidth and latency issues, enhancing media quality and user experience in XR applications.
Patent Information
- Application Number
- PCT/US2024/032629
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-11
AI Technical Summary
Existing media streaming technologies struggle to provide high-quality, immersive extended reality (XR) experiences over wireless networks due to bandwidth and latency constraints, particularly when users move quickly, as they require high-resolution content rendering with low latency.
Utilizing generative artificial intelligence (GenAI) models to generate semantic information from lower quality media frames, which is then transmitted over the network, allowing application servers to enhance and reconstruct media frames with higher resolution and anticipate user movements, offloading resource-intensive processing from XR devices.
This approach minimizes network latency and bandwidth consumption while maintaining high-quality media streaming experiences by transmitting smaller semantic data, enabling efficient resource utilization and improved XR application performance.
Smart Images

Figure US2024032629_11122025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR CONTENT CREATION IN MEDIA STREAMINGTECHNICAL FIELD
[0001] Embodiments of the invention relate to the field of networking; and more specifically, to content creation in media streaming.BACKGROUND ART
[0002] With readily available cameras of a variety of types and explosion of extended reality (XR) applications, the sheer volume of media content to be transmitted in such applications is staggering. These extended reality applications include immersive ones in Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). The convergence of the virtual and augmented physical and digital experience in a collective space may be referred to as the metaverse. A critical obstacle in these XR applications is to stream high-resolution content to users for providing interactive immersive experience (e.g., in the metaverse), in both real-time and video on demand use cases.
[0003] This problem is perhaps most apparent when a user quickly moves their head, as an XR application needs to render a potentially novel scene at high resolution within a very short latency window. State-of-the-art codecs solve this problem by buffering medium-to-high resolution content around the viewport of an XR device and selectively upgrading the resolution as the user changes position. A viewport refers to a user's field of view within the virtual environment (e.g., the metaverse) displayed to the user (e.g., by the XR device), and it represents the portion of the virtual environment that the user can see at any given moment, e.g., through the VR device, which may include a head-mounted display (HMD) headset and / or XR glasses.
[0004] In a wireless network, these XR applications may utilize available bandwidth to upgrade resolution around the viewport and prevent low-quality rendering of scenes with head movement. Yet unlike traditional video transmission, XR videos demand substantially more data due to their panoramic nature. An XR device needs to maintain high refresh rates to avoid user discomfort. For a genuinely immersive experience, XR video frames may range between 30 and 100 frames per second. Moreover, studies have shown that to guarantee optimal XR quality, the total latency should be kept below 20 milliseconds (ms). Due to XR's interactive essence, such latency requirements are stringent, and current networks, with network delays exceeding 90 milliseconds, fall short. Thus, providing satisfactory XR experience through a wireless network is challenging with existing solutions.SUMMARY
[0005] Embodiments include methods, electronic device, storage medium, and computer program for editing a media content stream. In one embodiment, a method comprises: generating semantic information based on a first set of media frames, the semantic information to identify objects within the first set of media frames, wherein the first set of media frames are from a media stream of a media streaming application to be experienced by a user; transmitting the semantic information generated based on the first set of media frames to a server through a wireless network; receiving updated semantic information from the server through the wireless network, wherein the updated semantic information is generated to anticipate movement of the user and is generated based on the semantic information; and generating a second set of media frames based on the updated semantic information using a generative artificial intelligence model, the second set of media frames to be included in the media stream.
[0006] Embodiments include electronic devices for editing a media content stream. In one embodiment, an electronic device is disclosed to comprise a processor and non-transitory machine-readable storage medium that provides instructions that, when executed by the processor, are capable of causing the processor to perform: generating semantic information based on a first set of media frames, the semantic information to identify objects within the first set of media frames, wherein the first set of media frames are from a media stream of a media streaming application to be experienced by a user; transmitting the semantic information generated based on the first set of media frames to a server through a wireless network; receiving updated semantic information from the server through the wireless network, wherein the updated semantic information is generated to anticipate movement of the user and is generated based on the semantic information; and generating a second set of media frames based on the updated semantic information using a generative artificial intelligence model, the second set of media frames to be included in the media stream.
[0007] Embodiments include machine-readable storage media for editing a media content stream. In one embodiment, a machine-readable storage medium is disclosed, and it provides instructions that, when executed by a processor, are capable of causing the processor to perform: generating semantic information based on a first set of media frames, the semantic information to identify objects within the first set of media frames, wherein the first set of media frames are from a media stream of a media streaming application to be experienced by a user; transmitting the semantic information generated based on the first set of media frames to a server through a wireless network; receiving updated semantic information from the server through the wireless network, wherein the updated semantic information is generated to anticipate movement of the user and is generated based on the semantic information; and generating a second set of mediaframes based on the updated semantic information using a generative artificial intelligence model, the second set of media frames to be included in the media stream.
[0008] By implementing embodiments as described, a media streaming application may provide high quality media frames generated based on semantic information extracted from lower quality media frames as provided. Through transmitting semantic information instead of media frames themselves through a communication network, the embodiments additionally save resources of the communication network to allow the communication network to support interactive streaming applications with network delay that satisfy user experience.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The invention may best be understood by referring to the following description and accompanying drawings that are used to illustrate embodiments of the invention. In the drawings:
[0010] Figure 1 illustrates operations of semantic coding for media content transmission per some embodiments.
[0011] Figure 2 illustrates operations at a transmitting subsystem to generate semantic information per some embodiments.
[0012] Figure 3 illustrates operations at a receiving subsystem to utilize received semantic information per some embodiments.
[0013] Figure 4 is a flow diagram illustrating operations for content creation in media streaming per some embodiments.
[0014] Figure 5 illustrates an electronic device for content creation in media streaming per some embodiments.
[0015] Figure 6 illustrates an example of a communication system per some embodiments.
[0016] Figure 7 illustrates a user equipment (UE) per some embodiments.
[0017] Figure 8 illustrates a network node per some embodiments.
[0018] Figure 9 is a block diagram illustrating a virtualization environment in which functions implemented by some embodiments may be virtualized.DETAILED DESCRIPTION
[0019] Generally, all terms used herein are to be interpreted according to their ordinary meaning in the relevant technical field, unless a different meaning is clearly given and / or is implied from the context in which it is used. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. Thesteps of any methods disclosed herein do not have to be performed in the exact order disclosed, unless a step is explicitly described as following or preceding another step and / or where it is implicit that a step must follow or precede another step. Any feature of any of the embodiments disclosed herein may be applied to any other embodiment, wherever appropriate. Likewise, any advantage of any of the embodiments may apply to any other embodiments, and vice versa. Other objectives, features, and advantages of the enclosed embodiments will be apparent from the following description.
[0020] As discussed herein above, providing satisfactory XR experience through a wireless network is challenging with existing solutions. Yet generative Al (GenAI) has recently attracted attention due to its capacity to mimic human-like comprehension to follow instructions and generation. Leveraging the self-attention mechanisms inherent in transformers and enriched by extensive training datasets, GenAI models can effectively recognize and replicate statistical patterns, enabling accurate predictions and data synthesis. Note that GenAI models are a type of machine learning model designed to generate new data based on the data they were trained on. The GenAI models may thus be used to generate updated semantic information and / or corresponding media frames.
[0021] Embodiments described in this disclosure utilize GenAI to streamline fewer essential elements of XR media, aiming for enhanced media transmission with minimized latency and bandwidth consumption in a wireless network. Through generatively upgrading the lower resolution content on the client end (e.g., a XR device) and / or the server end, the lower resolution encodings received from the transmitting subsystem may be taken as input to generate a higher resolution version of the content within a larger buffer, where the higher resolution content may be stored / displayed at the client end for the client viewport and / or the server end for other clients’ viewports. To generate the higher resolution version, the lower resolution encodings are supplemented with semantic information. Through the encoding at the transmitting subsystem and decoding at the receiving subsystem, only essential information (e.g., the lower resolution content and semantic information) is required to be transmitted through the wireless network, which may then accommodate the media transmission with minimized latency and bandwidth consumption.Semantic Coding
[0022] Figure 1 illustrates operations of semantic coding for media content transmission per some embodiments. System 100 includes a set of extended reality (XR) devices 150 to 154, a set of application servers 132 to provide XR applications, and a communication network 190 that couples the set of XR devices and application servers and provide communication in between.The XR application may be a metaverse application or other extended virtual reality application through which a human user interacts with media content.
[0023] Each XR device is an electronic device, and it may comprise a user equipment (UE) (e.g., one of UEs 612A to 612D of Figure 6) and include an XR headset, glasses, display, speaker or other media delivery device in some embodiments. The set of application servers 132 are electronic devices, and they provide media streaming applications for XR devices 150 to 154. The application servers 132 may be implemented in one of host 616 or a network node within telecommunication network 602, and communication network 190 may comprise telecommunication network 602.
[0024] While the interaction between XR device 150 and application servers 132 are discussed in detail herein, multiple XR devices may interact with each other and with application servers 132 in some embodiments. For example, a media streaming application may be multi-user paintball games where each user is equipped with an XR headset, and users need to identify each other’s positions and hit each other to score in a virtual environment (e.g., metaverse). The media streaming application needs to determine (1) the movement of all users and (2) where a user is looking at a given moment to provide high resolution video / audio to an area in the virtual environment, including the any other users’ positions at the moment in the area, based on the user’s focus so that users in the application have good quality of experience.
[0025] To do that, an XR device of a user may include or couple with a set of motion sensors and / or eye tracking sensors to detect the movement of the user and / or the gaze of the user. The motion sensors may include an accelerometer to measure acceleration forces to detect change in movement and orientation; a vibration sensor to detect vibrations or changes in motion through a piezoelectric element; a photoelectric sensor to use a light beam (visible or infrared) and detect motion when the beam is interrupted; a microwave sensor to emit microwave pulses and detect the frequency shift of the reflected waves due to the Doppler effect; a passive infrared (PIR) sensor to detect infrared radiation (heat) emitted by living beings; an ultrasonic sensor to emit ultrasonic waves and measures the time it takes for the waves to reflect back from an object; or another type of sensor to detect user motion.
[0026] The eye tracking sensors may include a video-based eye tracker that utilizes cameras to capture images of the eyes and advanced image processing algorithms to determine the direction of gaze; an infrared (IR) eye tracker to emit infrared light to illuminate the eyes and use cameras to capture the reflection patterns from the cornea and retina, allowing precise calculation of the gaze direction; an electrooculography (EOS) device to measure the electrical potential around the eyes using electrodes placed on the skin around the eyes to determine the eye moment corresponding to measured changes in potential; a magnetic tracking device to use smallmagnetic sensors placed on the eyes or eyelids and external magnetic fields to track the position and movement of the eyes; a photoelectric sensor to detect eye movements by measuring the reflection of light from the eyes; an electromagnetic field sensor in close proximity to the eyes to detect eye movements; or another type of sensor to detect user’s gaze.
[0027] A media streaming application provides one or more media streams to users of XR devices 150 to 154. Each media stream may include one or more video / audio streams (also referred to as a video / audio flows), which may be defined as a set of packets whose headers match a given pattern of bits. A media stream may be identified by a set of attributes embedded to one or more packets of the traffic flow. The media stream may include a sequence of media frames, e.g., a sequence of image frames that comprises a video segment and that may be encoded differently such as intra-coded I-frames, predictive P-frames, and bidirectional B- frames but sharing the same destination for an application. A media stream and corresponding semantic information may be transmitted through communication network 190 in packets. The term packet, as used herein, is intended to be broadly construed to include a frame, a datagram, a packet, or a cell; a fragment of a frame, a fragment of a datagram, a fragment of a packet, or a fragment of a cell; or another type, arrangement, or packaging of data. While video frames are used as examples to discuss some embodiments, audio frames and other frames may be implemented in these and other embodiments as well.
[0028] References numbers 102 to 1 12 illustrate operations performed according to some embodiments of the disclosure. At operation 102, XR device 150 captures media frames in a media stream for a media streaming application. The XR device 150 may capture the media frames based on motion data such as ones to indicate user’ s movement and / or user’ s gaze, and the capturing is a continuing process as the user continues to use the application. The media frames, as captured and as adjusted by application servers 132 as discussed herein, form the media stream displayed (e.g., as video) and / or broadcasted (e.g., as audio) to the user to experience the application. In some embodiments, the media frames as captured may be in a lower resolution than desirable for a satisfactory experience of the user because XR device 150 may not have sufficient resources (e.g., bandwidth resources, storage resources, and processing resources) to provide the sufficient resolution for the user, particularly when the user makes quick moves or changes gaze rapidly.
[0029] At operation 104, XR device 150 generates semantic information from the captured media frames. The semantic information may be generated through one or more machine learning models. The machine learning models may encode the viewpoint of the user of XR device 150 and / or the virtual environment at the moment to allow the media frames to be reconstructed (in another device such as application servers 132 or another XR device or later inthe same XR device). The semantic information includes the high-level, meaningful content or concepts that are identified and interpreted from the captured media frames.
[0030] Contrary to media information (e.g., pixel values in a video image), the semantic information generated from the media frames may be used to understand the scene, objects, action, and context presented in the media frames. For example, the semantic information may include the description of a detected object / act included in the media frames. In some embodiments, the semantic information includes motion data such as one collected from the motion sensors and / or eye tracking sensors so that the current and future user motion / gaze can be determined.
[0031] Since the semantic information provides a higher-level content than the media frames themselves, generating the corresponding semantic information is sometimes referred to as encoding the media frames. The encoding is performed by the transmitting subsystem 182 as a part of transmission pipeline and more detail about the encoding will be discussed herein below (e.g., relating to Figure 2).
[0032] At operation 106, the generated semantic information is transferred to application servers 132 through communication network 190. The semantic information is typically much smaller in size comparing to the media frames themselves (e.g., in the order of megabytes for the former vs. gigabyte for the latter), and transmitting the semantic information thus takes much less resources of communication network 190, including bandwidth resources, storage resources, and processing resources of network nodes within communication network 190 (e.g., ones within access network 604 and core network 606 of Figure 6). In some embodiments, a subset of the media frames (e.g., raw image data of one or more media frames) may be transmitted to application servers 132 as well (e.g., to help with application servers 132 to understand / reconstruct the media frames quicker and / or more accurately).
[0033] At operation 108, application servers 132 update semantic information and optionally reconstruct the media streams captured earlier. The update may enhance the semantic information received from XR device 150 in a variety of ways. Application servers 132 may update the received semantic information to provide higher resolution media frames in some embodiments, as application servers 132 have more resources to enhance the semantic information, e.g., through generative artificial intelligence (GenAI) models. For example, the semantic information extracted from a media frame captured XR device 150 (see Circle 1) indicates that the media frame includes a player of a paintball game application wearing a red hat. Yet the resolution of the media frame is low (e.g., 640 x 480 pixels Standard Definition (SD)), and through a GenAI model, the semantic information may be updated to generate a frame including the player wearing the red hat with a higher resolution (e.g., 1920 x 1080 pixelsFull High Definition (FHD)). The higher resolution with additional pixels may be generated using generic media content for the object (a red hat), which may be generated through machine learning, and it may be generated prior to the media processing or at run time based on the certain existing content of the media stream as provided. Such resolution enhancement through application servers 132 is achieved without sending the raw low resolution media frame from XR device 150 to application servers 132.
[0034] Additionally, the deficiency of media frames from one user may be enhanced from media frames from other users. For example, the viewpoint of XR device 150 in a multi-user paintball game may have limited visibility of a virtual environment at a moment due to the visual limitations of XR device 150. Yet another XR device included in the same multi-user paintball game, XR device 152, may have better visibility of the virtual environment at the moment. Application servers 132 receives semantic information from other XR devices such as XR device 152 and may update the semantic information from a particular user using semantic information received from XR device 152. Furthermore, users in a gaming application are often in motion, and based on the motion data from the semantic information, application servers 132 may anticipate where the user will be in the virtual environment and generate updated semantic information and / or corresponding media frames accordingly.
[0035] In some embodiments, the updated semantic information and / or corresponding media frames are generated through one or more Gen Al models, and the GenAI model generates the new content (the updated semantic information and / or corresponding media frames) based on the received semantic information. The new content enhances the semantic information received from XR device 150 as discussed herein above. For example, from semantic information indicating a video frame with a lower resolution, the GenAI model may generate updated semantic information to construct a corresponding video frame with a higher resolution.
[0036] After application servers 132 generate updated semantic information and corresponding media frames, and the update semantic information and / or corresponding media frames are then transmitted back to XR device 150 at operation 110. In some embodiments, one or more media frames constructed at application servers 132 (e.g., based on semantic information received from another XR device in the same streaming application) may be sent along with the updated semantic information. For example, when the updated semantic information indicates objects not seen by XR device 150 and it will be challenging for the XR device 150 to generate media frames to include the objects.
[0037] At operation 112, XR device 150 generates media frames based on the updated semantic information, and the media frames are then included in the media stream to be experienced by the user of XR device 150 as shown. Using the same example at operation 108,the updated semantic information from application servers 132 may allow XR device 150 to generate a frame including the player wearing the red hat with FHD. The operation is referred to as decoding, which is performed by the receiving subsystem 184 as a part of reception pipeline and more detail about the encoding will be discussed herein below (e.g., relating to Figure 3).
[0038] The generated media frames through these operations may provide better experience to the user because (1) the generated media frames may provide better resolution than what was captured earlier through Gen Al models, (2) the generated media frames may incorporate semantic information from other users that the original media frames captured by XR device 150 do not have, and / or (3) the generated media frames may anticipate user’s motion better than XR device 150 could by itself because application servers 132 may have more resources to generate media frames with finer details. It is common that XR device 150 has more limited resources due to its mobile requirement (an XR device needs to be agile enough to be coupled closed to a user), while application servers 132 can have far more sources as they may be shared by many applications / XR devices and implemented in a remote / economic location.
[0039] In some embodiments, the semantic information (at reference 106) or updated semantic information (at reference 110) may be transmitted (e.g., through communication network 190 between XR device 150 and application servers 132) along with additional information as metadata for identification, feature enhancement, and authentication. In alternative embodiments, the metadata for identification, feature enhancement, and authentication are treated as a part of semantic information transmitted between XR device and application server.
[0040] The metadata for identification may identify (1) XR device 150 using a Universal Unique Identifier (UUID), an International Mobile Equipment Identity (IMEI), a Media Access Control (MAC) address, an IP address, a serial number, a Bluetooth address, or another ID that uniquely identifies XR device 150 within system 100, (2) an application server within application servers 132 using one or more similar or different types of IDs, and (3) an electronic device or components of communication network 190 (e.g., the ID of the next hop in the network to transmit the packet containing the semantic information). The metadata for identification may be used by the receiving subsystem to determine the source of the semantic information so the receiving subsystem may return updated semantic information to the correct transmitting subsystem.
[0041] In some embodiment, the identification information may be included through markers (also referred to as watermarking), which embeds imperceptible or semi-perceptible identifying information within the generated semantic information for identifying the corresponding one or more objects, noting ownership attribution, providing content verification, and / or enforcingprotection against unauthorized use or manipulation. The markers may also outline the nested shapes inside of an identified object in some embodiments.
[0042] The metadata for feature enhancement may include (1) timestamps of object segmentation indicating when objects within a media frame are segmented, (2) hashes applied to machine learning models and / or IDs of the machine learning models used on the transmitting subsystem, (3) hyperparameters configured to the machine learning models, and other information to notify the receiving subsystem how encoding was done in the transmitting subsystem.
[0043] The metadata for authentication may include signature signed on the semantic information to prevent the semantic information from being eavesdropped on during transmission, e.g., being signed by the private key of the transmitting subsystem prior to transmission. On receipt of the information, the receiving subsystem (e.g., an application server) will verify the digital signature and check the authenticity of the message by using the public key of the transmitting subsystem. The receiving subsystem can also check if the other additional information satisfies the policy of the receiving subsystem. For instance, if the receiving subsystem finds that the model hash used at the transmitting subsystem is different than what the receiving subsystem expects, the receiving subsystem can choose to reject the media frame generated by the corresponding model.
[0044] The signing and verification of the authentication information may be performed repeatedly on the information received at the receiving subsystem (e.g., on received packets from the transmitting subsystem that each include semantic information as pay loads). Signing every packet would provide maximum assurance but have the most verification overhead. Selectively signing some information (e.g., packets including shape / position deformation updates reflecting motions) would be a beneficial optimization in some embodiments.
[0045] The semantic information may be further secured through crypto processors or other secure hardware such as Trusted Platform Module (TPM), Software Guard Extension (SGX), Hardware Security Module (TSM), Secure Encrypted Virtualization (SEV), when available to prove the trustworthiness of the software and hardware state of the transmitting subsystem (e.g. Operating System (OS), firmware state) along with the other metadata. For instance, with TPMs, an application server as the transmitting subsystem can produce a TPM quote with metadata along with a hash of media frame content. The information may be transmitted along with semantic information (e.g., in packet payloads) so that the receiving subsystem may verify the software and hardware state of the transmitting subsystem to determine if the receiving subsystem can trust the information being sent over.
[0046] The transmission of subset of the media frames along with the semantic information in some embodiments allows the receiving subsystem to utilize the semantic information more efficiently. For example, in an XR application, rarely seen (or never seen before) objects and shapes may be created at the transmitting subsystem, and it is challenging to train machine learning models at the receiving subsystem to reliably label these objects, and transmitting the media frames with these objects (along with labels of these object) will allow the receiving subsystem to recognize these objects and further enhance them easier and more reliably.
[0047] Through embodiments of the disclosure, semantic information is transmitted through a communication network between an XR device and an application server for a media streaming application, and since semantic information is typically much smaller in size comparing to the media frames themselves, it utilizes less resources in the communication network and can be transmitted more efficiently. Additionally, since the semantic information is then used to reconstruct and enhance the original media frames as captured, the computation is offload to the application server from the XR device, the former of which may provide much more resources for media frame processing. The offload allows system 100 to provide a better overall quality of experience to the user.Transmitting Subsystem
[0048] Figure 2 illustrates operations at a transmitting subsystem to generate semantic information per some embodiments. The transmitting subsystem 290 provides semantic information for the corresponding receiving subsystem to update semantic information and / or reconstruct / enhance media frames obtained at the transmitting subsystem. When the transmitting subsystem 290 is implemented at an XR device, it may perform operation 104 of Figure 1 in some embodiments.
[0049] At reference 201, media frames are received at the transmitting subsystem 290 in some embodiments. The media frames may be captured by a camera or other media capture devices within or coupled to an XR device. The media frames may be received with deficiency due to resource limitations of the XR device. For example, an XR device may only be able to provide frames at a lower resolution less than desirable for the user in some embodiments. Additionally, media frames may not be provided at a constant rate due to frame drop or degradation.
[0050] At reference 202, the transmitting subsystem 290 initiates the initial transmission setup. At reference 204, the transmitting subsystem 290 segments all disconnected top level objects of one or more media frames, and the segmentation utilizes a machine learning model for segmentation 252. The segmentation is performed through object detection and utilizes classification parameters for object retrieval, deformation, and generation in some embodiments.
[0051] Machine learning model for segmentation 252 may include one or more discriminative and / or machine learning models. For example, a discriminative machine learning model may be used to classify or differentiate data into predefined categories based on input data. XR device 150 and / or application servers 132 may provide an object category and corresponding metadata to a discriminative machine learning model as input data, so the discriminative machine learning model may search repository to identify one or more objects that match ones in the object category and corresponding metadata. The discriminative machine learning models in machine learning models for segmentation 252 may include one or more of logistic regression, Support Vector Machines (SVM), neural networks, decision trees, random forests, Naive Bayes, and K- Nearest Neighbors (K-NN).
[0052] The machine learning model for segmentation 252 identifies objects with the one or more frames and may provide additional hierarchical information or refer to existing hierarchical semantic library / method for extracting information to be used in a generative pipeline of a GenAI model. A set of process hyperparameters that determines the structure and behavior of machine learning may be provided to perform machine learning using the machine learning model for segmentation 252.
[0053] The hyperparameters are not directly learned from the data but are instead chosen for the machine learning model based on prior knowledge, experimentation, or optimization techniques. For example, the hyperparameters may include (1) the degree of replacement desired, e.g., how much new content is generated by the model versus how much is borrowed or modified from the original content (e.g., the captured media frames), (2) the buffer size to store the objects of interest, and (3) the degrees of abstract desired (e.g., level of abstraction in the hierarchical structure of nested objects). In some embodiments, the process hyperparameters may be set prior to the training of the machine learning model.
[0054] The hyperparameters may be selected based on the characteristics of the media frames and / or the user. For example, the details of another user within the game are more important than those of the background road for a paintball game, while the details of the road for a car chasing game is more important than those of another user, and the degrees of abstract may be assigned accordingly. Additionally, each user may have their own preference, and hyperparameters may be selected based on the characteristics of the user, e.g., a female gamer may pay more attention to attire of an opponent then a male gamer in the same paintball game, and the latter may be more receptive to replacing the opponent’s attire with generic content thus tolerating a higher degree of replacement than the former.
[0055] Through the machine learning model for segmentation 252, objects of interest are segmented. The objects of interest may be determined based on motion data captured by sensors.For example, based on the user’s gaze, the machine learning model may determine to which region with an image frame that the user is or will pay attention, so that the objects within that region are made to as objects of interest and they are marked as the ones to be provided with a higher resolution.
[0056] In some embodiments, the objects are segmented with a deformation buffer for rectangular bounding box retrieval. The objects and classification metadata are added to a bounding box with a segmented object in some embodiments. In some embodiments, boxed artifacts are extracted from object segmentation with object tags and classification metadata.
[0057] At reference 206, additional objects / shapes are identified within the top-level objects, and the corresponding frames are prepared to be represented by the objects / shapes (e.g., with the original size of the frames). A hierarchical structure of nested objects may be created based on the top-level objects. For example, for a top-level object of a human face identified from an image frame, additional objects / shapes may be identified, including eyes, lips, nostrils, eyebrows, ears, wrinkles, dimples, of the face. The level of the hierarchical structure may be determined based on the hyperparameters configured earlier. Some features of the objects within a media frame may be generated independently while others are generated with each other.
[0058] At reference 208, the positions and hierarchy of the nested objects inside the top-level objects are estimated. These estimated positions and hierarchy of the nested objects may relate to one another, and they preserve motion characteristics of the object of a media frame relating to other media frames within a sequence.
[0059] In some embodiments, objects of interest may be identified and grouped according to one or more policies, including the nearest neighbor grouping. Additionally, a tracking matrix may be generated to track the movement between objects within the sequence of frames. The object grouping information may be stored in a frame-to-frame buffer, and a set of values may be generated to store in the tracking matrix across various object groupings within the sequence of frames.
[0060] The tracking matrix, the positions and hierarchy of the nested objects, and other information are provided as semantic information to be compressed at reference 212.
[0061] In some embodiments, for the next n frames in the sequence, where n corresponds to the potential buffered information in the media stream, the orientation and deformation of all objects in the previous frame may be encoded at reference 210 using a generative Al model for encoding shape deformation at reference 254. The orientation and deformation are indicated through the shapes and relative positions of the objects. The data reflecting the orientation and deformation may be stored in a deformation buffer. This data indicates the manipulation or alteration of objects in generated content, such as audios, images, videos. This manipulation caninvolve changing the sound, shape, appearance, or position of objects within the generated content while maintaining the overall coherence and realism of the scene.
[0062] The generative Al model for encoding shape deformation 254 generates new data that is similar to the data it was trained on, where the produced novel, synthetic data samples share characteristics with the original dataset. The generative Al model for encoding shape deformation 254 may be trained with various object categories, corresponding deformation data. The trained models are then used to provide enhanced semantic information based on the provided object category and corresponding deformation data. The GenAI models may include one or more of Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Recurrent Neural Networks (RNNs), transformers (e.g., Generative Pre-trained Transformer), and Long Short-Term Memory (LSTM) Networks.
[0063] Note that in some embodiments, the GenAI model may also generate entries within the mapping data structure to identify object categories of a detected object when no applicable object categories are identified in the mapping data structure for the object.
[0064] The generative Al model for encoding shape deformation 254 may be run iteratively to upgrade the texture and mesh information of objects identified in the bounding box.Additionally, the hierarchical structure of nested objects may be used in the iterative upgrade, and the tracing matrix may be used to map positions of objects to adapt rendering of upgraded objects in each subsequent frame.
[0065] The encoded orientation and deformation of the objects at reference 210 are then compressed for transmission at reference 212, along with the estimated position and the hierarchy of nested objects. The compression may be performed along with markers and / or authentication information (e.g., signature signed on the semantic information). In some embodiments, one or more media frames may be transmitted along with the semantic information, and they may be compressed with different compression techniques as the semantic information, e.g., the former may be compressed with lossy compression while the latter with lossless compression since the semantic information, with the higher-level content, contains more information per bit than the raw media frames.
[0066] By encoding the objects within media frames and transmitting the resulting semantic information, the information necessary to reconstruct the media frames may be transmitted with a lower bandwidth and / or less storage / processing resources in a wireless network comparing to transmitting the media frames themselves.Receiving Subsystem
[0067] Figure 3 illustrates operations at a receiving subsystem to utilize received semantic information per some embodiments. The receiving subsystem 390 receives semantic informationfrom a corresponding transmitting subsystem so that the receiving subsystem 390 may reconstruct and update received semantic information and / or media frames. In such setup, the transmitting subsystem performs encoding while the receiving subsystem performs decoding. When the receiving subsystem 390 is implemented at an XR device, it may perform operation 112 of Figure 1 in some embodiments.
[0068] At reference 301, the receiving subsystem 390 receives data from a corresponding transmitting subsystem. At reference 302, the receiving subsystem 390 may initialize the setup for starting video / audio once the corresponding media frames are reconstructed.
[0069] At reference 304, the receiving subsystem 390 decodes and generates objects for one or more frames based on the received semantic information using a generative Al model for decoding shape deformation 354, which performs operations opposite to that of generative Al model for encoding shape deformation 254 to extract the deformation information within the transmitted data. The nested objects and the deformation information, including the orientations, shapes, and positions of the objects are then decoded at the receiving subsystem 390 (e.g., within a buffer within or coupled with the receiving subsystem 390) for generating the corresponding one or more media frames. The corresponding one or more media frames generated through the generative Al model for decoding shape deformation 354 can be an enhancement of the media frames from which the earlier semantic information is generated.
[0070] The receiving subsystem may run another GenAI model, the trained adversarial model for generating original objects 356 to generate an estimation of the original media frames from the generated media frames. The estimation of the original media frames may then be compared with the original media frames to evaluate the object accuracy at reference 306.
[0071] For example, the received semantic information of a media frame indicates that the media frame includes a player of a paintball game application wearing a red hat at HFD resolution, which is an enhancement of the originally transmitted semantic information from the corresponding transmitting subsystem 290 indicating the player at SD resolution, based on (1) an original frame including the player wearing the red hat at SD resolution. The GenAI model 354 generates (2) an enhanced frame with the player wearing the red hat at HFD resolution. The trained adversarial model for generating original objects 356 may then generate (3) a frame with the player wearing the red hat at SD resolution, based on (2) the enhanced frame with the player wearing the red hat at HFD resolution. Then (1) the original frame earlier provided at the transmitting subsystem 290 may then be compared with (3) the frame with the player wearing the red hat at SD resolution as reconstructed by the adversarial model 356. The comparison may be done pixel by pixel with an accuracy score generated for the full frame to determine if the GenAI model 354 generates an accurate enhancement of the original frame.
[0072] At reference 308, the objects with imperfect deformation encoding are identified and corrected. The deformation imperfect may occur when the motion data is used improperly by the Gen Al model for decoding shape deformation 354; when the imperfection is identified, the Gen Al model for decoding shape deformation 354 is to be retrained, and the objects may be corrected. At reference 310, the receiving subsystem 390 determines whether the generated objects and frames are accurate. If so at reference 312, one or more media enhanced frames are generated and provided to be experienced by the user. If not at reference 314, the receiving subsystem 390 may request the corresponding transmitting subsystem to retransmit semantic information; and if enough inaccurate reconstruction is encountered, the Gen Al model retraining will be triggered.
[0073] Note that media frames may not be provided at a constant rate due to frame drop or degradation, and the receiving subsystem 390 may use the GenAI model to compensate for the frame drop or degradation by generating deformed object from a buffer of latest set of received frames of objects. When the delayed frame arrives, the relevant details are updated in the next frame.
[0074] Application servers 132 examines semantic information received from XR device 150, and extracts information useful for the update, and the process may be referred to as decoding the semantic information and application servers 132 is then the receiving subsystem to decode information encoded by a corresponding transmitting subsystem. In this process, application servers 132 may be viewed as the transmitting subsystem that encodes the media frames into the updated semantic information, and XR device 150 is a corresponding receiving subsystem that decodes the updated semantic information.
[0075] Note while the operations at the transmitting subsystem 290 and receiving subsystem 390 are explained as performed at an XR device, an application server includes corresponding receiving subsystem and transmitting receiving subsystems as well. For example, the transmitting subsystem 290 may be implemented at an application server such as one of application servers 132 to generate and transmit semantic information, e.g., the application server may get media frames for which semantic information may be generated. The receiving subsystem 390 may be implemented at an application server to receive semantic information and construct media frames (e.g., for the usage by another XR device within a same media streaming application.
[0076] In these embodiments, high resolution media frames may be generated for user’s experience at an XR device without burdening the XR device with excessive resources; instead, the resources at the servers are utilized to offload the necessary process, particularly to accommodate user’s movement (e.g., user’s body movement and / or user’s gaze change).Additionally, the comparison score may be used to calibrate the accuracy of the Gen Al models so that the Gen Al models may remain reliable over time.Operations per some embodiments
[0077] Figure 4 is a flow diagram illustrating operations for content creation in media streaming per some embodiments. The operations of method 400 may be implemented in an electronic device (e.g., one of the XR devices 150 to 154).
[0078] At reference 402, semantic information is generated based on a first set of media frames, the semantic information to identify objects within the first set of media frames, where the first set of media frames are from a media stream of a media streaming application to be experienced by a user.
[0079] At reference 404, the semantic information generated based on the first set of media frames is transmitted to a server through a wireless network.
[0080] At reference 406, updated semantic information is received from the server through the wireless network, where the updated semantic information is generated to anticipate movement of the user and is generated based on the semantic information.
[0081] At reference 408, a second set of media frames is generated based on the updated semantic information using a generative artificial intelligence model, the second set of media frames to be included in the media stream.
[0082] In some embodiments, the semantic information generated based on the first set of media frames includes one or more of viewpoint information tracking viewpoint of the user and deformation information of the objects.
[0083] In some embodiments, generating the semantic information based on the first set of media frames is performed based on a machine learning model using a set of hyperparameters, and wherein the set of hyperparameters are selected based on characteristics of (1) the first set of media frames and (2) the user.
[0084] In some embodiments, generating the semantic information based on the first set of media frames comprises creating a hierarchical structure of nested objects based on object identification. The hierarchical structure of nested objects may provide details of a top object as discussed relating to Figure 2.
[0085] In some embodiments, generating the semantic information based on the first set of media frames comprises identifying and grouping a set of objects of interest, where relative movement between objects groups or objects are tracked. The tracking may be performed through a tracking matrix discussed herein.
[0086] In some embodiments, the semantic information generated based on the first set of media frames includes a marker or markers indicating the objects as identified. The marker ormarkers may be transmitting in the transfer or semantic information or updated semantic information in operations 106 and / or 110. The transfers may be performed with the indication of source and destination identifiers. The transfers may also be encrypted.
[0087] In some embodiments, transmitting the semantic information comprises adding metadata to the semantic information to identify from where the media stream is transmitted. In some embodiments, the metadata is used by the server to authenticate the semantic information.
[0088] In some embodiments, transmitting the semantic information comprises encrypting the semantic information to prevent eavesdropping during transmission.
[0089] In some embodiments, generating the second set of media frames based on the updated semantic information comprises performing error correction on an object identified based on the updated semantic information. The error correction is discussed herein relating to reference 308.
[0090] In some embodiments, at least a subset of the first set of media frames is transmitted along with the semantic information to the server.
[0091] In some embodiments, the media streaming application is to be experienced by another user, and wherein the updated semantic information is generated further based on semantic information from media frames for the another user.
[0092] These features of the embodiments offer unique advantages over prior approaches, so that the embodiments may be used to provide high quality media frames generated based on semantic information extracted from lower quality media frames. Through transmitting semantic information instead of media frames themselves through a communication network, the embodiments additionally save resources of the communication network to allow the communication network to support interactive streaming applications with network delay that satisfy user experience.Devices and Environments for Implementing Embodiments of the Invention
[0093] Figure 5 illustrates an electronic device for content creation in media streaming per some embodiments. The electronic device 502 may be a host in a cloud system, or a network node or UE in a wireless / wireline network, and the operating environment and further embodiments the host and the network node are discussed in more details discussed relating to Figures 6 to 9. The electronic device 502 may be implemented using custom applicationspecific integrated-circuits (ASICs) as processors and a special-purpose operating system (OS), or common off-the-shelf (COTS) processors and a standard OS. In some embodiments, electronic device 502 implements the operations discussed herein relating to Figures 1 to 4.
[0094] The electronic device 502 includes hardware 540 comprising a set of one or more processors 542 (which are typically COTS processors or processor cores including Central Processing Units (CPUs), Graphics Processing Units (GPUs), Tensor Processing Units (TPUs),microcontrollers (MCUs), ASICs, Field-Programmable Gate Arrays (FPGAs) and physical NIs 546, as well as non- transitory machine-readable storage media 549 having software 550 stored therein. During operation, the one or more processors 542 may execute the software 550 to instantiate one or more sets of one or more applications 564A-R. While one embodiment does not implement virtualization, alternative embodiments may use different forms of virtualization. For example, in one such alternative embodiment, the virtualization layer 554 represents the kernel of an operating system (or a shim executing on a base operating system) that allows for the creation of multiple instances 562A-R called software containers that may each be used to execute one (or more) of the sets of applications 564A-R. The multiple software containers (also called virtualization engines, virtual private servers, or jails) are user spaces (typically a virtual memory space) that are separate from each other and separate from the kernel space in which the operating system is run. The set of applications running in a given user space, unless explicitly allowed, cannot access the memory of the other processes. In another such alternative embodiment, the virtualization layer 554 represents a hypervisor (sometimes referred to as a virtual machine monitor (VMM)) or a hypervisor executing on top of a host operating system, and each of the sets of applications 564A-R run on top of a guest operating system within an instance 562A-R called a virtual machine (which may in some cases be considered a tightly isolated form of software container) that run on top of the hypervisor - the guest operating system and application may not know that they are running on a virtual machine as opposed to running on a “bare metal” host electronic device, or through para- virtualization the operating system and / or application may be aware of the presence of virtualization for optimization purposes. In yet other alternative embodiments, one, some, or all of the applications are implemented as unikemel(s), which can be generated by compiling directly with an application only a limited sei of libraries (e.g., from a library operating system (LibOS) including drivers / libraries of OS services) that provide the particular OS services needed by the application. As a unikernel can be implemented to run directly on hardware 540, directly on a hypervisor (in which case the unikernel is sometimes described as running within a LibOS virtual machine), or in a software container, embodiments can be implemented fully with unikernels running directly on a hypervisor represented by virtualization layer 554, unikemels running within software containers represented by instances 562A-R, or as a combination of unikernels and the above-described techniques (e.g., unikernels and virtual machines both run directly on a hypervisor, unikemels, and sets of applications that are run in different software containers).
[0095] The software 550 includes a content creator module 555 that performs operations described with reference to operations as discussed relating to Figures 1 to 4. The contentreplacement management module 555 may perform operations relating to an XR device (e.g., XR device 150), or alternatively, ones relating to an application server (e.g., one of application server 132. The content replacement management module 555 be instantiated within the applications 564A-R. The instantiation of the one or more sets of one or more applications 564A-R, as well as virtualization if implemented, are collectively referred to as software instance(s) 552. Each set of applications 564A-R, corresponding virtualization construct (e.g., instance 562A-R) if implemented, and that part of the hardware 540 that executes them (be it hardware dedicated to that execution and / or time slices of hardware temporally shared), forms a separate virtual electronic device 560A-R.
[0096] A network interface (NI) may be physical or virtual. In the context of Internet Protocol (IP), an interface address is an IP address assigned to an NI, be it a physical NI or virtual NI. A virtual NI may be associated with a physical NI, with another virtual interface, or stand on its own (e.g., a loopback interface, a point-to-point protocol interface). A NI (physical or virtual) may be numbered (a NI with an IP address) or unnumbered (a NI without an IP address). The NI is shown as network interface card (NIC) 544. The physical network interface 546 may include one or more antenna of the electronic device 502. An antenna port may or may not correspond to a physical antenna. The antenna comprises one or more radio interfaces.A Wireless Network per Some Embodiments
[0097] Figure 6 illustrates an example of a communication system 600 per some embodiments. The communication system 600 provides additional details regarding implementation of system 100 of Figure 1. In the example, the communication system 600 includes a telecommunication network 602 that includes an access network 604, such as a radio access network (RAN), and a core network 606, which includes one or more core network nodes 608. The access network 604 includes one or more access network nodes, such as network nodes 610a and 610b (one or more of which may be generally referred to as network nodes 610), or any other similar 3rdGeneration Partnership Project (3GPP) access nodes or non-3GPP access points. Moreover, as will be appreciated by those of skill in the art, a network node is not necessarily limited to an implementation in which a radio portion and a baseband portion are supplied and integrated by a single vendor. Thus, it will be understood that network nodes include disaggregated implementations or portions thereof. For example, in some embodiments, the telecommunication network 602 includes one or more Open-RAN (ORAN) network nodes. An ORAN network node is a node in the telecommunication network 602 that supports an ORAN specification (e.g., a specification published by the O-RAN Alliance, or any similar organization) and may operate alone or together with other nodes to implement one or more functionalities of any node in the telecommunication network 602, including one or morenetwork nodes 610 and / or core network nodes 608. A network can correspond to a 3GPP network (4G / 5G / 6G), Wi-Fi, or any other standard-based or proprietary network.
[0098] Examples of an ORAN network node include an open radio unit (O-RU), an open distributed unit (O-DU), an open central unit (O-CU), including an O-CU control plane (O-CU- CP) or an O-CU user plane (O-CU-UP), a RAN intelligent controller (near-real time or non-real time) hosting software or software plug-ins, such as a near-real time control application (e.g., xApp) or a non-real time control application (e.g., rApp), or any combination thereof (the adjective “open” designating support of an ORAN specification). The network node may support a specification by, for example, supporting an interface defined by the ORAN specification, such as an Al, Fl, Wl, El, E2, X2, Xn interface, an open fronthaul user plane interface, or an open fronthaul management plane interface. Moreover, an ORAN access node may be a logical node in a physical node. Furthermore, an ORAN network node may be implemented in a virtualization environment (described further below) in which one or more network functions are virtualized. For example, the virtualization environment may include an O-Cloud computing platform orchestrated by a Service Management and Orchestration Framework via an 0-2 interface defined by the O-RAN Alliance or comparable technologies. The network nodes 610 facilitate direct or indirect connection of user equipment (UE), such as by connecting UEs 612A, 612B, 612C, and 612D (one or more of which may be generally referred to as UEs 612) to the core network 606 over one or more wireless connections.
[0099] Example wireless communications over a wireless connection include transmitting and / or receiving wireless signals using electromagnetic waves, radio waves, infrared waves, and / or other types of signals suitable for conveying information without the use of wires, cables, or other material conductors. Moreover, in different embodiments, the communication system 600 may include any number of wired or wireless networks, network nodes, UEs, and / or any other components or systems that may facilitate or participate in the communication of data and / or signals whether via wired or wireless connections. The communication system 600 may include and / or interface with any type of communication, telecommunication, data, cellular, radio network, and / or other similar type of system.
[0100] The UEs 612 may be any of a wide variety of communication devices, including wireless devices arranged, configured, and / or operable to communicate wirelessly with the network nodes 610 and other communication devices. Similarly, the network nodes 610 are arranged, capable, configured, and / or operable to communicate directly or indirectly with the UEs 612 and / or with other network nodes or equipment in the telecommunication network 602 to enable and / or provide network access, such as wireless network access, and / or to perform other functions, such as administration in the telecommunication network 602.
[0101] In the depicted example, the core network 606 connects the network nodes 610 to one or more hosts, such as host 616. These connections may be direct or indirect via one or more intermediary networks or devices. In other examples, network nodes may be directly coupled to hosts. The core network 606 includes one more core network nodes (e.g., core network node 608) that are structured with hardware and software components. Features of these components may be substantially similar to those described with respect to the UEs, network nodes, and / or hosts, such that the descriptions thereof are generally applicable to the corresponding components of the core network node 608. Example core network nodes include functions of one or more of a Mobile Switching Center (MSC), Mobility Management Entity (MME), Home Subscriber Server (HSS), Access and Mobility Management Function (AMF), Session Management Function (SMF), Authentication Server Function (AUSF), Subscription Identifier De-concealing function (SIDE), Unified Data Management (UDM), Security Edge Protection Proxy (SEPP), Network Exposure Function (NEF), and / or a User Plane Function (UPF).
[0102] The host 616 may be under the ownership or control of a service provider other than an operator or provider of the access network 604 and / or the telecommunication network 602, and may be operated by the service provider or on behalf of the service provider. The host 616 provides an embodiment of application server(s) 132 of FIG. 1, and may host a variety of applications to provide one or more services. Examples of such applications include live and pre-recorded audio / video content, data collection services such as retrieving and compiling data on various ambient conditions detected by a plurality of UEs, analytics functionality, social media, functions for controlling or otherwise interacting with remote devices, functions for an alarm and surveillance center, or any other such function performed by a server.
[0103] As a whole, the communication system 600 of Figure 6 enables connectivity between the UEs, network nodes, and hosts. In that sense, the communication system may be configured to operate according to predefined rules or procedures, such as specific standards that include, but are not limited to: Global System for Mobile Communications (GSM); Universal Mobile Telecommunications System (UMTS); Long Term Evolution (LTE), and / or other suitable 2G, 3G, 4G, 5G standards, or any applicable future generation standard (e.g., 6G); wireless local area network (WLAN) standards, such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards (WiFi); and / or any other appropriate wireless communication standard, such as the Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave, Near Field Communication (NFC) ZigBee, LiFi, and / or any low-power wide-area network (LPWAN) standards such as LoRa and Sigfox.
[0104] In some examples, the telecommunication network 602 is a cellular network that implements 3GPP standardized features. Accordingly, the telecommunication network 602 maysupport network slicing to provide different logical networks to different devices that are connected to the telecommunication network 602. For example, the telecommunication network 602 may provide Ultra Reliable Low Latency Communication (URLLC) services to some UEs, while providing Enhanced Mobile Broadband (eMBB) services to other UEs, and / or Massive Machine Type Communication (mMTC) / Massive loT services to yet further UEs.
[0105] In some examples, the UEs 612 are configured to transmit and / or receive information without direct human interaction. For instance, a UE may be designed to transmit information to the access network 604 on a predetermined schedule, when triggered by an internal or external event, or in response to requests from the access network 604. Additionally, a UE may be configured for operating in single- or multiple radio access technology (multi-RAT) or multistandard mode. For example, a UE may operate with any one or combination of Wi-Fi, NR (New Radio) and LTE, i.e., being configured for multi-radio dual connectivity (MR-DC), such as E-UTRAN (Evolved-UMTS Terrestrial Radio Access Network) New Radio - Dual Connectivity (EN-DC).
[0106] In the example, the hub 614 communicates with the access network 604 to facilitate indirect communication between one or more UEs (e.g., UE 612C and / or 612D) and network nodes (e.g., network node 610B). In some examples, the hub 614 may be a controller, router, content source and analytics, or any of the other communication devices described herein regarding UEs. For example, the hub 614 may be a broadband router enabling access to the core network 606 for the UEs. As another example, the hub 614 may be a controller that sends commands or instructions to one or more actuators in the UEs. Commands or instructions may be received from the UEs, network nodes 610, or by executable code, script, process, or other instructions in the hub 614. As another example, the hub 614 may be a data collector that acts as temporary storage for UE data and, in some embodiments, may perform analysis or other processing of the data. As another example, the hub 614 may be a content source. For example, for a UE that is an extended reality (XR) headset, display, loudspeaker or other media delivery device, the hub 614 may retrieve VR assets, video, audio, or other media or data related to sensory information via a network node, which the hub 614 then provides to the UE either directly, after performing local processing, and / or after adding additional local content. In still another example, the hub 614 acts as a proxy server or orchestrator for the UEs, in particular if one or more of the UEs are low energy loT devices.
[0107] The hub 614 may have a constan t / persistent or intermittent connection to the network node 610b. The hub 614 may also allow for a different communication scheme and / or schedule between the hub 614 and UEs (e.g., UE 612C and / or 612D), and between the hub 614 and the core network 606. In other examples, the hub 614 is connected to the core network 606 and / orone or more UEs via a wired connection. Moreover, the hub 614 may be configured to connect to a machine-to-machine (M2M) service provider over the access network 604 and / or to another UE over a direct connection. In some scenarios, UEs may establish a wireless connection with the network nodes 610 while still connected via the hub 614 via a wired or wireless connection. In some embodiments, the hub 614 may be a dedicated hub - that is, a hub whose primary function is to route communications to / from the UEs from / to the network node 610B. In other embodiments, the hub 614 may be a non-dedicated hub - that is, a device which is capable of operating to route communications between the UEs and network node 61 OB, but which is additionally capable of operating as a communication start and / or end point for certain data channels. In some embodiments, electronic device 502 that implements content creator module 555 may be, comprise, or be coupled to one of UE 612A-112D, network nodes 608, 610A-B, or host 617.UE per Some Embodiments
[0108] Figure 7 illustrates a UE 700 per some embodiments. As used herein, a UE refers to a device capable, configured, arranged and / or operable to communicate wirelessly with network nodes and / or other UEs. Examples of a UE include, but are not limited to, a smart phone, mobile phone, cell phone, voice over IP (VoIP) phone, wireless local loop phone, desktop computer, personal digital assistant (PDA), wireless cameras, gaming console or device, music storage device, playback appliance, wearable terminal device, wireless endpoint, mobile station, tablet, laptop, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), smart device, wireless customer-premise equipment (CPE), vehicle, vehicle-mounted or vehicle embedded / integrated wireless device, etc. Other examples include any UE identified by the 3rd Generation Partnership Project (3GPP), including a narrow band internet of things (NB-IoT) UE, a machine type communication (MTC) UE, and / or an enhanced MTC (eMTC) UE. Electronic device 502 and XR devices 150 to 154 that implement content creator module 555 may be, comprise, or be coupled to UE 700.
[0109] A UE may support device-to-device (D2D) communication, for example by implementing a 3GPP standard for sidelink communication, Dedicated Short-Range Communication (DSRC), vehicle-to- vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle- to-everything (V2X). In other examples, a UE may not necessarily have a user in the sense of a human user who owns and / or operates the relevant device. Instead, a UE may represent a device that is intended for sale to, or operation by, a human user but which may not, or which may not initially, be associated with a specific human user (e.g., a smart sprinkler controller). Alternatively, a UE may represent a device that is not intended for sale to, or operation by, auser but which may be associated with or operated for the benefit of a user (e.g., a smart power meter).
[0110] The UE 700 includes processing circuitry 702 that is operatively coupled via a bus 704 to an input / output interface 706, a power source 708, a memory 710, a communication interface 712, and / or any other component, or any combination thereof. Certain UEs may utilize all or a subset of the components shown in Figure 7. The level of integration between the components may vary from one UE to another UE. Further, certain UEs may contain multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0111] The processing circuitry 702 is configured to process instructions and data and may be configured to implement any sequential state machine operative to execute instructions stored as machine-readable computer programs in the memory 710. The processing circuitry 702 may be implemented as one or more hardware- implemented state machines (e.g., in discrete logic, field- programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.); programmable logic together with appropriate firmware; one or more stored computer programs, general-purpose processors, such as a microprocessor or digital signal processor (DSP), together with appropriate software; or any combination of the above. For example, the processing circuitry 702 may include multiple central processing units (CPUs) or graphics processing units (GPUs).
[0112] In the example, the input / output interface 706 may be configured to provide an interface or interfaces to an input device, output device, or one or more input and / or output devices. Examples of an output device include a speaker, a sound card, a video card, a display, a monitor, a printer, an actuator, an emitter, a smartcard, another output device, or any combination thereof. An input device may allow a user to capture information into the UE 700. Examples of an input device include a touch-sensitive or presence-sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, and the like. The presence-sensitive display may include a capacitive or resistive touch sensor to sense input from a user. A sensor may be, for instance, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. An output device may use the same type of interface port as an input device. For example, a Universal Serial Bus (USB) port may be used to provide an input device and an output device.
[0113] In some embodiments, the power source 708 is structured as a battery or battery pack. Other types of power sources, such as an external power source (e.g., an electricity outlet), photovoltaic device, or power cell, may be used. The power source 708 may further includepower circuitry for delivering power from the power source 708 itself, and / or an external power source, to the various parts of the UE 700 via input circuitry or an interface such as an electrical power cable. Delivering power may be, for example, for charging of the power source 708. Power circuitry may perform any formatting, converting, or other modification to the power from the power source 708 to make the power suitable for the respective components of the UE 700 to which power is supplied.
[0114] The memory 710 may be or be configured to include memory such as random-access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable readonly memory (EEPROM), magnetic disks, optical disks, hard disks, removable cartridges, flash drives, and so forth. In one example, the memory 710 includes one or more application programs 714, such as an operating system, web browser application, a widget, gadget engine, or other application, and corresponding data 716. The memory 710 may store, for use by the UE 700, any of a variety of various operating systems or combinations of operating systems.
[0115] The memory 710 may be configured to include a number of physical drive units, such as redundant array of independent disks (RAID), flash memory, USB flash drive, external hard disk drive, thumb drive, pen drive, key drive, high-density digital versatile disc (HD-DVD) optical disc drive, internal hard disk drive, Blu-Ray optical disc drive, holographic digital data storage (HDDS) optical disc drive, external mini-dual in-line memory module (DIMM), synchronous dynamic random access memory (SDRAM), external micro-DIMM SDRAM, smartcard memory such as tamper resistant module in the form of a universal integrated circuit card (UICC) including one or more subscriber identity modules (SIMs), such as a Universal Subscriber Identity Module (USIM) and / or IP Multimedia Services Identity Module (ISIM), other memory, or any combination thereof. The UICC may for example be an embedded UICC (eUICC), integrated UICC (iUICC) or a removable UICC commonly known as ‘SIM card.’ The memory 710 may allow the UE 700 to access instructions, application programs and the like, stored on transitory or non-transitory memory media, to off-load data, or to upload data. An article of manufacture, such as one utilizing a communication system may be tangibly embodied as or in the memory 710, which may be or comprise a device-readable storage medium.
[0116] The processing circuitry 702 may be configured to communicate with an access network or other network using the communication interface 712. The communication interface 712 may comprise one or more communication subsystems and may include or be communicatively coupled to an antenna 722. The communication interface 712 may include one or more transceivers used to communicate, such as by communicating with one or more remote transceivers of another device capable of wireless communication (e.g., another UE or a networknode in an access network). Each transceiver may include a transmitter 718 and / or a receiver 720 appropriate to provide network communications (e.g., optical, electrical, frequency allocations, and so forth). Moreover, the transmitter 718 and receiver 720 may be coupled to one or more antennas (e.g., antenna 722) and may share circuit components, software or firmware, or alternatively be implemented separately.
[0117] In the illustrated embodiment, communication functions of the communication interface 712 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short- range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof. Communications may be implemented in according to one or more communication protocols and / or standards, such as IEEE 802.11, Code Division Multiplexing Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, transmission control protocol / intemet protocol (TCP / IP), synchronous optical networking (SONET), Asynchronous Transfer Mode (ATM), Quick UDP Internet Connections (QUIC), Hypertext Transfer Protocol (HTTP), WebRTC (Web Real-Time Communication), and so forth.
[0118] Regardless of the type of sensor, a UE may provide an output of data captured by its sensors, through its communication interface 712, via a wireless connection to a network node. Data captured by sensors of a UE can be communicated through a wireless connection to a network node via another UE. The output may be periodic (e.g., once every 15 minutes if it reports the sensed temperature), random (e.g., to even out the load from reporting from several sensors), in response to a triggering event (e.g., when moisture is detected an alert is sent), in response to a request (e.g., a user initiated request), or a continuous stream (e.g., a live video feed of a patient).
[0119] As another example, a UE comprises an actuator, a motor, or a switch, related to a communication interface configured to receive wireless input from a network node via a wireless connection. In response to the received wireless input the states of the actuator, the motor, or the switch may change. For example, the UE may comprise a motor that adjusts the control surfaces or rotors of a drone in flight according to the received input or to a robotic arm performing a medical procedure according to the received input.
[0120] A UE, when in the form of an Internet of Things (loT) device, may be a device for use in one or more application domains, these domains comprising, but not limited to, city wearable technology, extended industrial application and healthcare. Non-limiting examples of such an loT device are a device which is or which is embedded in: a connected refrigerator or freezer, aTV, a connected lighting device, an electricity meter, a robot vacuum cleaner, a voice controlled smart speaker, a home security camera, a motion detector, a thermostat, a smoke detector, a door / window sensor, a flood / moisture sensor, an electrical door lock, a connected doorbell, an air conditioning system like a heat pump, an autonomous vehicle, a surveillance system, a weather monitoring device, a vehicle parking monitoring device, an electric vehicle charging station, a smart watch, a fitness tracker, a head-mounted display, glasses, or speaker for Augmented Reality (AR) or Virtual Reality (VR), a wearable for tactile augmentation or sensory enhancement, a water sprinkler, an animal- or item-tracking device, a sensor for monitoring a plant or animal, an industrial robot, an Unmanned Aerial Vehicle (UAV), and any kind of medical device, like a heart rate monitor or a remote controlled surgical robot. A UE in the form of an loT device comprises circuitry and / or software in dependence of the intended application of the loT device in addition to other components as described in relation to the UE 700 shown in Figure 7.
[0121] As yet another specific example, in an loT scenario, a UE may represent a machine or other device that performs monitoring and / or measurements, and transmits the results of such monitoring and / or measurements to another UE and / or a network node. The UE may in this case be an M2M device, which may in a 3 GPP context be referred to as an MTC device. As one particular example, the UE may implement the 3GPP NB-IoT standard. In other scenarios, a UE may represent a vehicle, such as a car, a bus, a truck, a ship and an airplane, or other equipment that is capable of monitoring and / or reporting on its operational status or other functions associated with its operation.
[0122] In practice, any number of UEs may be used together with respect to a single use case. For example, a first UE might be or be integrated in a drone and provide the drone’s speed information (obtained through a speed sensor) to a second UE that is a remote controller operating the drone. When the user makes changes from the remote controller, the first UE may adjust the throttle on the drone (e.g., by controlling an actuator) to increase or decrease the drone’ s speed. The first and / or the second UE can also include more than one of the functionalities described above. For example, a UE might comprise the sensor and the actuator, and handle communication of data for both the speed sensor and the actuators.Network Node per Some Embodiments
[0123] Figure 8 illustrates a network node 800 per some embodiments. As used herein, network node refers to equipment capable, configured, arranged and / or operable to communicate directly or indirectly with a UE and / or with other network nodes or equipment, in a telecommunication network. Examples of network nodes include, but are not limited to, access points (APs) (e.g., radio access points), base stations (BSs) (e.g., radio base stations, Node Bs,evolved Node Bs (eNBs) and NR NodeBs (gNBs)), O-RAN nodes or components of an O-RAN node (e.g., radio base stations, Node Bs, evolved Node Bs (eNBs) and NR NodeBs (gNBs)).
[0124] Base stations may be categorized based on the amount of coverage they provide (or, stated differently, their transmit power level) and so, depending on the provided amount of coverage, may be referred to as femto base stations, pico base stations, micro base stations, or macro base stations. A base station may be a relay node or a relay donor node controlling a relay. A network node may also include one or more (or all) parts of a distributed radio base station such as centralized digital units, distributed units (e.g., in an O-RAN access node) and / or remote radio units (RRUs), sometimes referred to as Remote Radio Heads (RRHs). Such remote radio units may or may not be integrated with an antenna as an antenna integrated radio. Parts of a distributed radio base station may also be referred to as nodes in a distributed antenna system (DAS).
[0125] Other examples of network nodes include multiple transmission point (multi-TRP) 5G access nodes, multi-standard radio (MSR) equipment such as MSR BSs, network controllers such as radio network controllers (RNCs) or base station controllers (BSCs), base transceiver stations (BTSs), transmission points, transmission nodes, multi-cell / multicast coordination entities (MCEs), Operation and Maintenance (O&M) nodes, Operations Support System (OSS) nodes, Self-Organizing Network (SON) nodes, positioning nodes (e.g., Evolved Serving Mobile Location Centers (E-SMLCs)), and / or Minimization of Drive Tests (MDTs).
[0126] The network node 800 includes a processing circuitry 802, a memory 804, a communication interface 806, and a power source 808. The network node 800 may be composed of multiple physically separate components (e.g., a NodeB component and a RNC component, or a BTS component and a BSC component, etc.), which may each have their own respective components. In certain scenarios in which the network node 800 comprises multiple separate components (e.g., BTS and BSC components), one or more of the separate components may be shared among several network nodes. For example, a single RNC may control multiple NodeBs. In such a scenario, each unique NodeB and RNC pair may in some instances be considered a single separate network node. In some embodiments, the network node 800 may be configured to support multiple radio access technologies (RATs). In such embodiments, some components may be duplicated (e.g., separate memory 804 for different RATs) and some components may be reused (e.g., a same antenna 810 may be shared by different RATs). The network node 800 may also include multiple sets of the various illustrated components for different wireless technologies integrated into network node 800, for example GSM, WCDMA, LTE, NR, WiFi, Zigbee, Z-wave, LoRaWAN, Radio Frequency Identification (RFID) or Bluetooth wirelesstechnologies. These wireless technologies may be integrated into the same or different chip or set of chips and other components within network node 800.
[0127] The processing circuitry 802 may comprise a combination of one or more of a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application- specific integrated circuit, field programmable gate array, or any other suitable computing device, resource, or combination of hardware, software and / or encoded logic operable to provide, either alone or in conjunction with other network node 800 components, such as the memory 804, to provide network node 800 functionality.
[0128] In some embodiments, the processing circuitry 802 includes a system on a chip (SOC). In some embodiments, the processing circuitry 802 includes one or more of radio frequency (RF) transceiver circuitry 812 and baseband processing circuitry 814. In some embodiments, the radio frequency (RF) transceiver circuitry 812 and the baseband processing circuitry 814 may be on separate chips (or sets of chips), boards, or units, such as radio units and digital units. In alternative embodiments, part or all of RF transceiver circuitry 812 and baseband processing circuitry 814 may be on the same chip or set of chips, boards, or units.
[0129] The memory 804 may comprise any form of volatile or non-volatile computer-readable memory including, without limitation, persistent storage, solid-state memory, remotely mounted memory, magnetic media, optical media, random access memory (RAM), read-only memory (ROM), mass storage media (for example, a hard disk), removable storage media (for example, a flash drive, a Compact Disk (CD) or a Digital Video Disk (DVD)), and / or any other volatile or non-volatile, non-transitory device-readable and / or computer-executable memory devices that store information, data, and / or instructions that may be used by the processing circuitry 802. The memory 804 may store any suitable instructions, data, or information, including a computer program, software, an application including one or more of logic, rules, code, tables, and / or other instructions capable of being executed by the processing circuitry 802 and utilized by the network node 800. The memory 804 may be used to store any calculations made by the processing circuitry 802 and / or any data received via the communication interface 806. In some embodiments, the processing circuitry 802 and memory 804 are integrated.
[0130] The communication interface 806 is used in wired or wireless communication of signaling and / or data between a network node, access network, and / or UE. As illustrated, the communication interface 806 comprises port(s) / terminal(s) 816 to send and receive data, for example to and from a network over a wired connection. The communication interface 806 also includes radio front-end circuitry 818 that may be coupled to, or in certain embodiments a part of, the antenna 810. Radio front-end circuitry 818 comprises filters 820 and amplifiers 822. The radio front-end circuitry 818 may be connected to an antenna 810 and processing circuitry 802.The radio front-end circuitry may be configured to condition signals communicated between antenna 810 and processing circuitry 802. The radio front-end circuitry 818 may receive digital data that is to be sent out to other network nodes or UEs via a wireless connection. The radio front-end circuitry 818 may convert the digital data into a radio signal having the appropriate channel and bandwidth parameters using a combination of filters 820 and / or amplifiers 822. The radio signal may then be transmitted via the antenna 810. Similarly, when receiving data, the antenna 810 may collect radio signals which are then converted into digital data by the radio front-end circuitry 818. The digital data may be passed to the processing circuitry 802. In other embodiments, the communication interface may comprise different components and / or different combinations of components.
[0131] In certain alternative embodiments, the network node 800 does not include separate radio front-end circuitry 818, instead, the processing circuitry 802 includes radio front-end circuitry and is connected to the antenna 810. Similarly, in some embodiments, all or some of the RF transceiver circuitry 812 is part of the communication interface 806. In still other embodiments, the communication interface 806 includes one or more ports or terminals 816, the radio front-end circuitry 818, and the RF transceiver circuitry 812, as part of a radio unit (not shown), and the communication interface 806 communicates with the baseband processing circuitry 814, which is part of a digital unit (not shown).
[0132] The antenna 810 may include one or more antennas, or antenna arrays, configured to send and / or receive wireless signals. The antenna 810 may be coupled to the radio front-end circuitry 818 and may be any type of antenna capable of transmitting and receiving data and / or signals wirelessly. In certain embodiments, the antenna 810 is separate from the network node 800 and connectable to the network node 800 through an interface or port.
[0133] The antenna 810, communication interface 806, and / or the processing circuitry 802 may be configured to perform any receiving operations and / or certain obtaining operations described herein as being performed by the network node. Any information, data and / or signals may be received from a UE, another network node and / or any other network equipment.Similarly, the antenna 810, the communication interface 806, and / or the processing circuitry 802 may be configured to perform any transmitting operations described herein as being performed by the network node. Any information, data and / or signals may be transmitted to a UE, another network node and / or any other network equipment.
[0134] The power source 808 provides power to the various components of network node 800 in a form suitable for the respective components (e.g., at a voltage and current level needed for each respective component). The power source 808 may further comprise, or be coupled to, power management circuitry to supply the components of the network node 800 with power forperforming the functionality described herein. For example, the network node 800 may be connectable to an external power source (e.g., the power grid, an electricity outlet) via an input circuitry or interface such as an electrical cable, whereby the external power source supplies power to power circuitry of the power source 808. As a further example, the power source 808 may comprise a source of power in the form of a battery or battery pack which is connected to, or integrated in, power circuitry. The battery may provide backup power should the external power source fail.
[0135] Embodiments of the network node 800 may include additional components beyond those shown in Figure 8 for providing certain aspects of the network node’s functionality, including any of the functionality described herein and / or any functionality necessary to support the subject matter described herein. For example, the network node 800 may include user interface equipment to allow input of information into the network node 800 and to allow output of information from the network node 800. This may allow a user to perform diagnostic, maintenance, repair, and other administrative functions for the network node 800. In some embodiments, the network node 800 may implement content creator module 555 and perform operations described herein.Virtualization Environment per Some Embodiments
[0136] Figure 9 is a block diagram illustrating a virtualization environment '900 in which functions implemented by some embodiments may be virtualized. In the present context, virtualizing means creating virtual versions of apparatuses or devices which may include virtualizing hardware platforms, storage devices and networking resources. As used herein, virtualization can be applied to any device described herein, or components thereof, and relates to an implementation in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 900 hosted by one or more of hardware nodes, such as a hardware computing device that operates as a network node, UE, core network node, or host. Further, in embodiments in which the virtual node does not require radio connectivity (e.g., a core network node or host), then the node may be entirely virtualized. In some embodiments, the virtualization environment 900 includes components defined by the O-RAN Alliance, such as an O-Cloud environment orchestrated by a Service Management and Orchestration Framework via an O-2 interface. Virtualization may facilitate distributed implementations of a network node, UE, core network node, or host.
[0137] Applications 902 (which may alternatively be called software instances, virtual appliances, network functions, virtual nodes, virtual network functions, etc.) are run in thevirtualization environment 950 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein. In some embodiments, Applications 902 may implement content creator module 555 and perform operations described herein.
[0138] Hardware 904 includes processing circuitry, memory that stores software and / or instructions executable by hardware processing circuitry, and / or other hardware devices as described herein, such as a network interface, input / output interface, and so forth. Software may be executed by the processing circuitry to instantiate one or more virtualization layers 906 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 908A and 908B (one or more of which may be generally referred to as VMs 908), and / or perform any of the functions, features and / or benefits described in relation with some embodiments described herein. The virtualization layer 906 may present a virtual operating platform that appears like networking hardware to the VMs 908.
[0139] The VMs 908 comprise virtual processing, virtual memory, virtual networking or interface and virtual storage, and may be run by a corresponding virtualization layer 906. Different embodiments of the instance of a virtual appliance 902 may be implemented on one or more of VMs 908, and the implementations may be made in different ways. Virtualization of the hardware is in some contexts referred to as network function virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which can be located in data centers, and customer premise equipment.
[0140] In the context of NFV, a VM 908 may be a software implementation of a physical machine that runs programs as if they were executing on a physical, non-virtualized machine. Each of the VMs 908, and that part of hardware 904 that executes that VM, be it hardware dedicated to that VM and / or hardware shared by that VM with others of the VMs, forms separate virtual network elements. Still in the context of NFV, a virtual network function is responsible for handling specific network functions that run in one or more VMs 908 on top of the hardware 904 and corresponds to the application 902.
[0141] Hardware 904 may be implemented in a standalone network node with generic or specific components. Hardware 904 may implement some functions via virtualization. Alternatively, hardware 904 may be part of a larger cluster of hardware (e.g., such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration 910, which, among others, oversees lifecycle management of applications 902. In some embodiments, hardware 904 is coupled to one or more radio units that each include one or more transmitters and one or more receivers that may be coupled to one or more antennas. Radio units may communicate directly with other hardware nodes via one or more appropriatenetwork interfaces and may be used in combination with the virtual components to provide a virtual node with radio capabilities, such as a radio access node or a base station. In some embodiments, some signaling can be provided with the use of a control system 912 which may alternatively be used for communication between hardware nodes and radio units.
[0142] Although the computing devices described herein (e.g., UEs, network nodes, hosts) may include the illustrated combination of hardware components, other embodiments may comprise computing devices with different combinations of components. It is to be understood that these computing devices may comprise any suitable combination of hardware and / or software needed to perform the tasks, features, functions and methods disclosed herein.Determining, calculating, obtaining or similar operations described herein may be performed by processing circuitry, which may process information by, for example, converting the obtained information into other information, comparing the obtained information or converted information to information stored in the network node, and / or performing one or more operations based on the obtained information or converted information, and as a result of said processing making a determination. Moreover, while components are depicted as single boxes located within a larger box, or nested within multiple boxes, in practice, computing devices may comprise multiple different physical components that make up a single illustrated component, and functionality may be partitioned between separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of the components may be partitioned between the processing circuitry and the communication interface. In another example, non-computationally intensive functions of any of such components may be implemented in software or firmware and computationally intensive functions may be implemented in hardware.
[0143] In certain embodiments, some or all of the functionality described herein may be provided by processing circuitry executing instructions stored on in memory, which in certain embodiments may be a computer program product in the form of a non-transitory computer- readable storage medium. In alternative embodiments, some or all of the functionalities may be provided by the processing circuitry without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hard-wired manner. In any of those particular embodiments, whether executing instructions stored on a non-transitory computer- readable storage medium or not, the processing circuitry can be configured to perform the described functionality. The benefits provided by such functionality are not limited to the processing circuitry alone or to other components of the computing device, but are enjoyed by the computing device as a whole, and / or by end users and a wireless network generally.Terms
[0144] References in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” and so forth, indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0145] The description and claims may use the terms “coupled” and “connected,” along with their derivatives. These terms are not intended as synonyms for each other. “Coupled” is used to indicate that two or more elements, which may or may not be in direct physical or electrical contact with each other, co-operate or interact with each other. “Connected” is used to indicate the establishment of wireless or wireline communication between two or more elements that are coupled with each other. A “set,” as used herein can refer to any whole number of items including one item.
[0146] An electronic device (such as the electronic device 502) stores and transmits (internally and / or with other electronic devices over a network) code (which is composed of software instructions and which is sometimes referred to as a computer program code or a computer program) and / or data using machine-readable media (also called computer-readable media), such as machine-readable storage media (e.g., magnetic disks, optical disks, solid state drives, read only memory (ROM), flash memory devices, phase change memory) and machine-readable transmission media (also called a carrier) (e.g., electrical, optical, radio, acoustical, or other form of propagated signals - such as carrier waves, infrared signals). Thus, an electronic device (e.g., a computer) includes hardware and software, such as a set of one or more processors (e.g., of which a processor is a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), other electronic circuitry, or a combination of one or more of the preceding) coupled to one or more machine-readable storage media to store code for execution on the set of processors and / or to store data. For instance, an electronic device may include non-volatile memory containing the code since the non-volatile memory can persist code / data even when the electronic device is turned off (when power is removed). When the electronic device is turned on, that part of the code that is to be executed by the processor(s) of the electronic device is typically copied from the slower non-volatile memory into volatile memory (e.g., dynamic random-access memory (DRAM), static random-access memory (SRAM)) of the electronic device. Typical electronic devices also include a set of one or more physical network interface(s)(NI(s)) to establish network connections (to transmit and / or receive code and / or data using propagating signals) with other electronic devices. For example, the set of physical NIs (or the set of physical NI(s) in combination with the set of processors executing code) may perform any formatting, coding, or translating to allow the electronic device to send and receive data whether over a wired and / or a wireless connection. In some embodiments, a physical NI may comprise radio circuitry capable of (1) receiving data from other electronic devices over a wireless connection and / or (2) sending data out to other devices through a wireless connection. This radio circuitry may include transmitter(s), receiver(s), and / or transceiver(s) suitable for radio frequency communication. The radio circuitry may convert digital data into a radio signal having the proper parameters (e.g., frequency, timing, channel, bandwidth, and so forth). The radio signal may then be transmitted through antennas to the appropriate recipient(s). In some embodiments, the set of physical NI(s) may comprise network interface controller(s) (NICs), also known as a network interface card, network adapter, or local area network (LAN) adapter. The NIC(s) may facilitate connecting the electronic device to other electronic devices allowing them to communicate with wire through plugging in a cable to a physical port connected to an NIC. One or more parts of an embodiment of the invention may be implemented using different combinations of software, firmware, and / or hardware.
[0147] The terms “module,” “logic,” and “unit” used in the present application, may refer to a circuit for performing the function specified. In some embodiments, the function specified may be performed by a circuit in combination with software such as by software executed by a general-purpose processor.
[0148] Any appropriate steps, methods, features, functions, or benefits disclosed herein may be performed through one or more functional units or modules of one or more virtual apparatuses. Each virtual apparatus may comprise a number of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessor or microcontrollers, as well as other digital hardware, which may include digital signal processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as read-only memory (ROM), random-access memory (RAM), cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory includes program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the respective functional unit to perform corresponding functions according to one or more embodiments of the present disclosure.
[0149] The term unit may have conventional meaning in the field of electronics, electrical devices, and / or electronic devices and may include, for example, electrical and / or electronic circuitry, devices, modules, processors, memories, logic solid state and / or discrete devices, computer programs or instructions for carrying out respective tasks, procedures, computations, outputs, and / or displaying functions, and so on, as such as those that are described herein.
Claims
CLAIMSWhat is claimed is:
1. A method for content creation in a media streaming application, comprising: generating semantic information based on a first set of media frames, the semantic information to identify one or more objects within the first set of media frames, wherein the first set of media frames are from a media stream of the media streaming application to be presented to a user; transmitting the semantic information generated based on the first set of media frames to a server through a wireless network; receiving updated semantic information from the server through the wireless network, wherein the updated semantic information is generated to anticipate movement of the user and is generated based on the semantic information; and generating a second set of media frames based on the updated semantic information using a generative artificial intelligence model, the second set of media frames to be included in the media stream.
2. The method of claim 1, wherein the semantic information generated based on the first set of media frames includes one or more of viewpoint information tracking viewpoint of the user and deformation information of the objects.
3. The method of claim 1 or 2, wherein generating the semantic information based on the first set of media frames is performed based on a machine learning model using a set of hyperparameters, and wherein the set of hyperparameters are selected based on characteristics of: the first set of media frames; and the user or user device.
4. The method of any of claims 1 to 3, wherein generating the semantic information based on the first set of media frames comprises creating a hierarchical structure of nested objects based on object identification.
5. The method of any of claims 1 to 4, wherein generating the semantic information based on the first set of media frames comprises identifying and grouping a set of objects of interest, wherein relative movement between objects groups or objects are tracked.
6. The method of any of claims 1 to 5, wherein the semantic information generated based on the first set of media frames includes a marker indicating the objects as identified.
7. The method of any of claims 1 to 6, wherein transmitting the semantic information comprises adding metadata to the semantic information to identify from where the media stream is transmitted.
8. The method of claim 7, wherein the metadata is used by the server to authenticate the semantic information.
9. The method of any of claims 1 to 8, wherein transmitting the semantic information comprises encrypting the semantic information to prevent eavesdropping during transmission.
10. The method of any of claims 1 to 9, wherein generating the second set of media frames based on the updated semantic information comprises performing error correction on an object identified based on the updated semantic information.1 1. The method of any of claims 1 to 10, wherein at least a subset of the first set of media frames is transmitted along with the semantic information to the server.
12. The method of any of claims 1 to 11, wherein the media streaming application is to be experienced by another user, and wherein the updated semantic information is generated further based on semantic information from media frames for the another user.
13. An electronic device (502), comprising: a processor (542) and non-transitory machine-readable storage medium (549) that provides instructions that, when executed by the processor (542), are capable of causing the processor to perform any of methods 1 to 12.
14. The electronic device (502), further comprising: a display subsystem for presenting media streams to a user; a camera subsystem for capturing an environment of the electronic device.
15. A non-transitory machine-readable storage medium (549) that provides instructions that, when executed by a processor (542), are capable of causing the processor to perform any of methods 1 to 12.
Citation Information
Patent Citations
Neural network processing for multi-object 3D modeling
US20190130639A1
Method for analysing media content
US20190251360A1
Content filtering in media playing devices
US20210329338A1
Method and electronic device for determining motion saliency and video playback style in video
US20220230331A1
Providing shared augmented reality environments within video calls
US20230300292A1