Network system employing distributed generative modeling with joint trained neural network communication paths
By implementing generative models on user devices and using jointly trained neural network paths, the efficiency problem of transmitting high-resolution content in generative modeling is solved, achieving efficient content transmission and generation under dynamic conditions and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-01
AI Technical Summary
In generative modeling, existing technologies struggle to efficiently transmit high-resolution content, especially when network channel bandwidth and latency are limited between the source and consumer, resulting in high resource consumption and an inability to adapt to the dynamic conditions of the UE.
A distributed generative modeling approach is adopted, which utilizes jointly trained neural network paths to implement generative models at the user device. Input prompts are efficiently transmitted between the network and the UE through encoder and decoder neural networks, and content is generated by combining local sensor data.
It enables efficient content transmission and generation under limited network resources, adapts to the dynamic conditions of the UE, and improves the quality and efficiency of user experience.
Smart Images

Figure CN121970068A_ABST
Abstract
Description
Background Technology
[0001] Generative modeling fosters rich user experiences by generating highly detailed, immersive content. Use cases for this type of generative modeling include augmented reality / virtual reality (AR / VR), holographic telepresence, natural language processing (NLP), and more. However, in many cases, the source of the generated content (e.g., an application server) and the consumer of the generated content (e.g., AR / VR goggles, personal computers, smartphones, or other user devices) are implemented separately and connected via one or more network channels. While bandwidth and latency in such network channels are constantly improving, the high data communication rates and low latency required for efficient transmission of content created by generative models between the source and consumer will continue to pose significant challenges to networking implementations. Attached Figure Description
[0002] This disclosure will be better understood by referring to the accompanying drawings, and many features and advantages of this disclosure will become apparent to those skilled in the art. The same reference numerals are used in different drawings to indicate similar or identical items.
[0003] Figure 1 This is a diagram illustrating an example wireless system employing a jointly trained neural network pathway with distributed generative modeling, according to some embodiments.
[0004] Figure 2 This illustrates some embodiments. Figure 1 A diagram showing an example configuration of user equipment in a wireless system.
[0005] Figure 3 This illustrates some embodiments. Figure 1 A diagram showing an example configuration of the network components of a wireless system.
[0006] Figure 4 This is a diagram illustrating a machine learning module, according to some embodiments, for use in a jointly trained neural network pathway with distributed generative modeling.
[0007] Figure 5 This illustrates a method for using according to some embodiments. Figure 1 A signal graph of example methods for remote generative cues and local generative modeling in wireless systems.
[0008] Figure 6 This illustrates a method for using according to some embodiments. Figure 1 Signal diagram of an example method for local sensor feedback in a wireless system.
[0009] Figure 7This is a flowchart illustrating an example method for jointly training a candidate set of neural network architecture configurations for distributed paths, according to some embodiments. Detailed Implementation
[0010] Typically, in networked systems utilizing generative modeling, a remote application server executes software implementing one or more generative models, such as generative adversarial networks (GANs) or generative pre-trained transformers (GPTs), to generate graphical, audio, and / or textual content based on a set of prompts or other inputs. The resulting generated content is transmitted to a user equipment (UE), which then processes the content to output it to a user via one or more user experience (UX) devices, such as AR goggles, flat panel displays, speakers, etc. In many cases, the generated content is represented by large amounts of data, such as high-resolution still images or high-resolution video, the transmission of which may exceed the communication capabilities of the network channel connecting the application server and the UE.
[0011] To deliver generated content to consumer devices more efficiently, in one implementation, the network utilizes a distributed generative modeling approach. This distributed generative modeling approach employs a jointly trained neural network path, where the generative model is implemented at the UE rather than entirely at a remote application server. This neural network path includes at least a jointly trained encoder neural network, a decoder neural network, and a generative model. The decoder neural network is implemented between the application server and the UE (e.g., at a base station and / or at core network functions used to wirelessly connect the UE to the remainder of the network) and operates to efficiently encode input prompts from the application server for efficient transmission to the UE. At the UE, the decoder neural network is used to efficiently decode the encoded prompts to recover a lossless or lossy representation of the original prompts. The decoded prompts are provided as input to the generative model at the UE for processing, in some implementations, along with one or more additional inputs such as local sensor data, to generate output content, which is then processed by one or more UX components associated with the UE. Therefore, instead of using a generative model at the application server to generate a large amount of data representing the generated content to be consumed by the UE and then transmitting that data from the application server to the UE via a potentially restricted network channel, the network can determine input prompts suitable for the generative model implemented by the UE to generate the desired output content, and then use a jointly trained encoder neural network at the network side and a decoder neural network at the UE side to efficiently transmit the prompts (which can typically be several orders of magnitude smaller than the generated content) to the UE for implementation at the local generative model.
[0012] Different UEs can have different capabilities for implementing generative models. Furthermore, the ability of a given UE to implement a specific generative model can vary over time due to changing circumstances for the UE (e.g., varying battery reserves or changes in current margins regarding thermal limits). Further, network channel capabilities can differ between UEs and can also vary over time for the same UE. Therefore, in some implementations, the system can access different combinations of neural network architectures with encoder neural networks, decoder neural networks, and generative models to reflect multiple candidate neural network paths for different implementation environments. Thus, the UE can report its current capabilities, its current context, and / or the current network conditions observed by the UE to the network. The network then uses this reported information to select a suitable neural network path from the candidate neural network paths for implementation in the path between the application server and the UE. The network can then instruct the base station or other network edge components to use the corresponding encoder neural network and instruct the UE to use the corresponding decoder neural network and generative model from the selected neural network path.
[0013] In some implementations, the generative model can be further trained to simultaneously receive local sensor data from the UE as input, to be processed by the local generative model along with prompts (or their decoded representations) supplied by the application server. To illustrate, in an implementation where the generative model operates to generate a video representation incorporating AR overlays of local objects, the UE can then provide the generative model with sensor data from location sensors relating to the position or motion of these local objects (e.g., video images of the local objects, geographic location information of the local objects, etc.) for use in generating the next frame sequence for AR overlay. The UE and network can also feed this local sensor data (and / or user input) back to the application server for processing via a feedback path in a neural network pathway, where the encoder neural network at the UE effectively encodes the local sensor data and / or user input data, and the output is transmitted to the base station or other edge components of the network. This edge component of the network implements a decoder neural network (which has been jointly trained with the UE's encoder neural network) to effectively decode the received input, thereby recovering the local sensor data or its representation.
[0014] For ease of illustration, the implementation of systems and techniques for network-based distributed generative modeling methods is described herein in an example scenario where the network is a cellular system, where the core network component or edge component is a base station (BS) or similar network edge component, and the UE is a cellular UE wirelessly connected to the BS or similar network edge component using cellular signaling. However, as also described below, these systems and techniques can also be employed in non-cellular networks, such as networks where the UE is a user device with more limited capabilities connected via wired or wireless connections (such as via Universal Serial Bus (USB), Bluetooth™ or other Wireless Personal Area Network (PAN) connections, Wi-Fi or other Wireless Local Area Network (WLAN) connections, etc.) to another device with more unrestricted capabilities. Therefore, unless otherwise stated, references to cellular networks or cellular-specific components equally apply to non-cellular networks or similar non-cellular-specific components.
[0015] Figure 1 An example wireless communication network 100 employing a distributed generative model scheme using jointly trained neural networks is illustrated according to some embodiments. In the depicted example, the wireless communication network 100 is a cellular network comprising a core network 102 coupled to one or more application servers 106 via one or more wide area networks (WANs) 104 or other packet data networks (PDNs) such as the Internet. The core network 102 includes multiple base stations or other network edge components—including the illustrated base station (BS) 108—to support wireless communication with one or more UEs, such as the illustrated UE 110, via radio frequency (RF) signaling using one or more applicable radio access technologies (RATs) specified by one or more communication protocols or standards. Thus, the BS 108 operates as a wireless interface between the UE 110 and various networks, as well as between the core network 102 and services provided by other networks such as packet-switched (PS) data services, circuit-switched (CS) services, etc. Conventionally and as used herein, signaling communication from BS 108 to UE 110 is referred to as “downlink” or “DL”, while signaling communication from UE 110 to BS 108 is referred to as “uplink” or “UL”.
[0016] BS 108 can employ any of a variety of RATs, such as a NodeB (or Base Transceiver Station (BTS)) operating as a Universal Mobile Telecommunications System (UMTS) RAT (also known as "3G"), an Enhanced NodeB (eNodeB) operating as a 3GPP Long Term Evolution (LTE) RAT, a 5G NodeB ("gNB") operating as a 3GPP 5th Generation (5G) New Radio (NR) RAT, and others. In this example cellular implementation, UE 110 can then implement any of a variety of electronic devices operable to communicate with BS 108 via a suitable RAT, including, for example, mobile cellular phones, cellular-enabled tablets or laptops, desktop computers, cellular-enabled video game systems, servers, cellular-enabled home appliances, cellular-enabled car communication systems, cellular-enabled smartwatches or other wearable devices, etc.
[0017] To achieve generative modeling functionality, conventional cellular networks or other conventional communication networks employ generative models entirely at the application server or other components of the network, or entirely at the UE device. In such approaches, input prompts are generated at or locally on the same component implementing the generative model ti. Where the generative model is implemented at the network component, considerable resources can be consumed in transmitting the resulting output to the UE for further processing. For example, a generative model operating to generate a series of video frames representing, for example, a video teleconference would require transmitting a considerable amount of video data over the network connection to the UE. Conversely, a conventional implementation of the generative model at the UE may result in the inefficient transmission of the same considerable amount of data, in the form of input prompts, between the network component and the UE. Furthermore, implementing the generative model at the UE in this manner results in a static implementation that may be unsuitable for dynamic conditions associated with the UE, such as changes in available resources at the UE, changes in network conditions at the UE, etc. (referred to herein as "changes in the UE's context").
[0018] Therefore, in the implementation, network 100 adopts a distributed generative modeling scheme based on neural networks, wherein UE 110 operates to provide a local generative model, while application servers or other network components operate to generate input prompts for use by the local generative model at the UE. This involves the joint operation of a set of jointly trained neural networks, including at least an encoder neural network at the network and a decoder neural network at the UE, to efficiently encode, transmit, and decode input prompts from network components for use by the local generative model at the UE. For illustration, as... Figure 1As depicted, UE 110 natively implements generative model 112, which has been trained and otherwise configured to receive input cue information representing one or more input prompts and generate output based on this input cue information. This output is then passed to one or more user experience (UX) components 114 of UE 110 (or associated with them) for processing to enhance the user experience. Generative model 112 may include any one or a combination of various generative models, including generative adversarial networks (GANs), variational autoencoders, Boltzmann machines, large language models (LLMs), diffusion probabilistic models, generative random networks (GSNs), variational autoencoders, diffusion networks, etc. One or more UX components 114 may include any of the various components that contribute to the user experience for UE 110, such as a display, speaker, haptic feedback generator, etc. For example, UE 110 may be a smartphone, and one or more UX components 114 may include the smartphone's display and speaker. As another example, UE 110 may be an XR headset, and one or more UX components 114 may include a near-eye AR display and / or an integrated speaker. One or more UX components 114 may be implemented at UE 110 or connected to UE 110 via a wired or wireless connection such as a USB connection, Bluetooth connection, or direct WiFi connection.
[0019] In some implementations, application server 106 (or other network component) generates input prompt information 116, which is provided to UE 110 via wireless transmission from BS 108 (or other edge network component) in some form. To facilitate the efficient transmission of this input prompt information 116, BS 108 and UE 110 jointly implement a jointly trained downlink path 118, wherein BS 108 (or other edge network component or core network component) implements a network encoder neural network 120, and UE 110 implements a UE decoder neural network 122. The network encoder neural network 120 and the UE decoder neural network 122 can be implemented as any of a variety of suitable neural networks, such as DNN or convolutional neural networks (CNN), and are jointly trained (with or without generative model 112) to facilitate the transmission of input prompt information 116 between BS 108 and UE 110. Therefore, the network encoder neural network 120 processes the input cue information 116 according to its trained configuration to generate an intermediate signal 124, which is essentially an encoded representation of the input cue information 116. The BS 108 wirelessly transmits the intermediate signal 124 to the UE 110 as part of cellular signaling between the BS 108 and the UE 110. At the UE 110, the UE decoder neural network 122 receives the intermediate signal 124 as input and processes it according to its trained configuration to generate a final signal 126, which is essentially a decoded or otherwise recovered representation of the input cue information 116. As described in more detail herein, in some cases, accurate and complete recovery of the original input cue information 116 is not necessary for satisfactory operation of the generative model 112, and therefore in such cases, the intermediate signal 124 can be a lossless representation (i.e., an inaccurate recovery of the input cue information 116). This lossless representation can be a conventional lossless process, such as a reduction in resolution, frame rate, etc. In other implementations, this lossless representation can be achieved, for example, through soft output or semantic representation. Since the final signal 126 represents a partial or complete version of the input cue information 116 (or even a soft output representation or semantic meaning representation), the final signal 126 represents input cue information with one or more input cues for input to the generative model 112. The generative model 112 then processes these one or more input cues according to its trained configuration to generate an output 128, which is passed to one or more UX components 114 for further processing to enhance the user experience at the UE 110.
[0020] To illustrate this with an example, in one implementation, the distributed generative modeling scheme supports a video conferencing use case where application server 106 generates input cues designed to reflect, for example, facial expressions, gesture changes, hand movements, etc., and generative model 112 at UE 110 uses these input cues, along with the current video of the corresponding participant (or avatar), to generate a sequence of video images of the corresponding participant (or avatar) performing these movements / gestures. As another example, application server 106 may include a game server supporting multiple UEs 110. Sensor information (including, for example, game controller input) from one UE 110 is transmitted to application server 106, which then generates input cues based on this sensor information to prompt generative model 112 at a second UE 110 to generate gameplay and a corresponding game video representing the state of the gameplay generated by this sensor information as game input.
[0021] In addition to employing a neural network-based downlink path 118, in some implementations, BS 108 and UE 110 also jointly employ a jointly trained uplink path 130 to facilitate efficient communication of information from UE 110 to one or more components of the network—such as application server 106. For uplink path 130, UE 110 implements a UE encoder neural network 132, and BS 108 implements a network decoder neural network 134. The UE encoder neural network 132 and network decoder neural network 134 can be implemented as any of a variety of suitable neural networks, such as DNNs or convolutional neural networks (CNNs), and are jointly trained to facilitate the transmission of uplink information between UE 110 and BS 108. In embodiments, this uplink information may include feedback information for use by application server 106 in controlling one or more subsequent iterations of input prompt information 116 or adapting it to the current context of UE 110. This current context may include the context of UE 110, such as UE 110's current attitude or location, current battery capacity, current computing power, current network bandwidth capacity, or current network latency observed by UE 110, one or more current parameters of one or more UX components 114 and / or one or more software applications utilizing one or more UX components 114, etc. Additionally or alternatively, this feedback information may include user input, such as user input provided via a game controller (in the case where a distributed generative method is used for a remote gaming application) or user input providing feedback for correcting or improving the generative process.
[0022] In at least one embodiment, this context of UE 110 is represented at least in part by sensor data 136 generated by sensor set 144 of UE 110. For example, sensor set 144 may include a network interface for sensing current network capability parameters, a battery sensor for measuring current battery capacity, a software-based or hardware-based sensor for measuring current computing capacity such as remaining available memory or current processor utilization, a screen activation sensor for determining the current state of the display, a thermal sensor for determining the current thermal state of one or more components of UE 110, attitude / position / or orientation sensors such as inertial management unit (IMU), gyroscope, global positioning system (GPS) sensor, global navigation satellite system (GNSS) sensor, and / or cellular / wireless triangulation system for determining the current attitude (position and / or orientation) of UE 110, and a thermal sensor for detecting thermal conditions of UE 110, etc.
[0023] Iterations of sensor data 136 collected from some or all of the sensors in sensor set 144 can occur on a periodic basis (e.g., according to a software-based timer), in response to a non-periodic trigger (e.g., in response to a UE capability query radio resource control (RRC) message from BS 108), or a combination thereof. The current iteration of sensor data 136 generated by sensor set 144 is provided as input to UE encoder neural network 132, which processes the current iteration of sensor data 136 according to its trained configuration to generate an intermediate signal 138, which is essentially an encoded representation of sensor data 136. UE 110 then transmits the intermediate signal 138 to BS 108. At BS108, the intermediate signal 138 is provided as input to the network decoder neural network 134, which processes the intermediate signal 138 according to its trained configuration to generate the final signal 140, depending on the implementation requirements and the jointly trained configuration of the neural networks 132 and 134 of the uplink path 130. The final signal is a lossy or lossless representation of the current iteration of the sensor data 136.
[0024] The final signal 140, representing the current iteration of sensor data 136, can then be passed to application server 106 or other components of the network for processing. In one implementation, application server 106 uses the current context of UE 110, represented by the recovered sensor data (in the final signal 140), to modify the next iteration of input prompt information 116 to be transmitted to UE 110 via downlink path 118. For illustration, a software application executing at UE 110 can be using generative model 112 to generate and render virtual reality (VR) content, and application server 106 can be operated to generate input prompts used by generative model 112 when generating this VR content. Sensor data 136 may include, for example, current attitude information of UE 110 determined by one or more attitude / position sensors of sensor set 144, and application server 106 may use this current attitude information to generate input cues to be included in the next iteration of input cue information 116. When processed by generative model 112, these input cues cause generative model 112 to generate and render a sequence of video frames that accurately reflect the attitude of UE 110 in the VR world from the angle corresponding to the current attitude of UE 110 in the real world.
[0025] In addition to or instead of providing sensor data, user input, and other feedback to application server 106 for use, for example, in determining input prompts, in some implementations, sensor data 142 from sensor set 144 is provided as input to generative model 112 for processing along with the current iteration of final signal 126 when generating output result 128. Therefore, sensor data 142 can be the same as sensor data 136, or a different set of sensor data from a different subset of sensors in sensor set 144 and / or sensor data captured at different times. For the purposes of the above example where application server 106 uses sensor data 136 to generate input prompts intended to guide generative model 112 in generating VR scene video based on the corresponding pose of the UE in the VR world, sensor data 142 may include updated (i.e., most recent) pose information of UE 110, which is processed by generative model 112 to further refine the rendered pose of the UE in the VR world. As another example, sensor data 142 may include current graphics processing unit (GPU) capabilities that can be used by generative model 112, for example, to scale the resolution of the rendered VR video image accordingly.
[0026] In some implementations, the jointly trained neural networks 120 and 122 of downlink path 118 and / or the jointly trained neural networks 132 and 134 of uplink path 130 are "fixed"; that is, their specific architectural configurations (e.g., weights, connections, layers, etc.) do not change based on the implementation or environment. However, in many cases, the potential current context of UE 110 can be wide-ranging. For example, one software application at UE 110 may have a set of operating parameters for generative model 112, while another software application at UE 110 may have a different set of operating parameters. Moreover, specific current context parameters such as current processing power and current network channel conditions can facilitate the use of specific architectural configurations for neural networks 120, 122, 132, and 134, while making other architectural configurations less feasible. Therefore, in other implementations, the selection of one or more neural networks among neural networks 120, 122, 132, and 134, and the architectural configuration for generative model 112, can be based on the current context of UE 110, the current context of BS 110 or other network components, the current context of application server 106, or a combination thereof. Thus, in some embodiments, to initiate this selection process, when a software application signals its intention to use generative model 112, UE 110 may transmit a request to BS 108 and provide a representation of its current context to BS 108 as, for example, a UE capability information RRC message or a UE assistance information RRC message. BS 108 (or other network components, such as servers in the core network) can then use the information from the request and the current context information of UE 110 to select a suitable set of jointly trained neural network architectures for downlink path 118 and / or uplink path 130 from a plurality of candidate sets of jointly trained neural networks, and instruct UE 110 to implement the corresponding neural architectures for neural networks 122, 132 and for generative model 112. Similarly, BS 108 implements the corresponding neural network architectures for neural networks 120 and 134. BS 108 can signal the neural network architecture to be used by UE 110 by transmitting data detailing the neural network architecture to be implemented (e.g., weights, nodes, layers, etc.) or by sending index values or other identifiers used by UE 110 to reference the selected neural network architecture from a store of candidate neural network architectures accessible to UE 110, as described in more detail below.
[0027] Figure 2 and Figure 3Example hardware configurations of UE 110 and BS 108 according to some embodiments are shown respectively. Note that the depicted hardware configurations represent the processing and communication components most directly related to the distributed generative model scheme described herein, and specific components commonly implemented in such electronic devices, such as displays, non-sensor peripherals, power supplies, etc., are omitted. Moreover, although references Figure 3 The hardware configuration of BS 108 is shown, but it should be understood that a similar hardware configuration can be used for another edge network component when it employs some or all of the functionalities of BS 108 that are part of the distributed generative modeling scheme described in this paper.
[0028] First refer to Figure 2 The hardware configuration of UE 110, in an implementation, includes one or more antenna arrays 202, each antenna array 202 having one or more antennas 203; and further includes an RF front-end 204, one or more processors 206, one or more non-transitory computer-readable media 208, and one or more UX components 114 and sensor sets 144 described above. The RF front-end 204 actually operates as a physical (PHY) transceiver interface to conduct and process signaling between one or more processors 206 and antenna arrays 202 to facilitate various types of wireless communication. Antenna 203 may include an array of multiple antennas configured to be similar or different from each other and may be tuned to one or more frequency bands associated with a corresponding RAT, such as a cellular RAT (e.g., the 3GPP 4G LTE RAT or the 3GPP 5G NR RAT), a WLAN RAT (e.g., an IEEE 802.11-based RAT), a WPAN RAT (e.g., a Bluetooth™ RAT), and the like.
[0029] One or more processors 206 may include, for example, one or more central processing units (CPUs), graphics processing units (GPUs), artificial intelligence (AI) accelerators, or other application-specific integrated circuits (ASICs). For illustration, processor 206 may include an application processor (AP) used by UE 110 to execute an operating system and various user-level software applications, and one or more processors utilized by a modem or baseband processor of RF front-end 204. Computer-readable medium 208 may include any of a variety of media used by electronic devices to store data and / or executable instructions, such as random access memory (RAM), read-only memory (ROM), cache, flash memory, solid-state drive (SSD), or other mass storage devices. For ease of illustration and brevity, given that system memory or other memory is frequently used to store data and instructions for execution by processor 206, computer-readable medium 208 is referred to herein as "memory 208," but it should be understood that unless otherwise stated, the reference to "memory 208" also applies to other types of storage media.
[0030] As described above, UE 110 may further include multiple sensors, referred to herein as sensor set 144, from which sensor data is obtained for transmission to application server 106 or for use by generative model 112, or both. Generally, the sensor captures of sensor set 144 represent sensor data reflecting the current context of UE 110, which may include a position / attitude context, resource availability context, resource usage context or resource capacity context, network channel state (e.g., bandwidth and / or latency), etc. Therefore, sensor set 144 may include, for example, GPS sensors, GNSS sensors, IMU sensors, gyroscopes, tilt sensors or other tiltmeters, ultra-wideband (UWB) based sensors. Sensor set 144 may also include imaging sensors, such as cameras for image capture by the user, cameras for face detection, cameras for stereo vision or visual odometry, light sensors for detecting objects approaching features of UE 110, etc. Sensor set 144 may further include user interface (UI) sensors (such as touchscreens, user-operable input / output devices (e.g., "buttons" or keyboards), or other touch / contact sensors), microphones or other voice sensors, thermal sensors (such as those for detecting proximity to a user), etc. Another example of sensors in sensor set 144 may include thermal sensors for determining the thermal state of one or more components of UE 110, a battery capacity sensor for determining the current battery capacity of UE 110, a network interface for determining the state of the network channel connecting UE 110 to BS 108, a memory monitor for determining current memory availability, a processor monitor for determining current processor capacity / utilization, etc.
[0031] UX component 114 includes components of UE 110 for providing aspects of user experience to users of UE 110. This may include hardware components such as a display, speaker, haptic feedback device, touch screen, etc.; and software components such as drivers for such hardware devices or software applications that utilize the output of generative model 112 to control the hardware UX components.
[0032] One or more memories 208 of UE 110 are used to store one or more executable software instruction sets and associated data that manipulate one or more processors 206 and other components of UE 110 to perform the various functions described herein and belonging to UE 110. The set of executable software instructions includes, for example, an operating system (OS) and various drivers (not shown), various software applications 210 (including at least one user application utilizing the output of generative model 112), implementations such as neural networks 122 and 132 (… Figure 1The UE neural network manager 212 for one or more neural networks of UE 110, and the generative model 112 implementing UE 110. Figure 1 It may be the same as or different from the UE neural network manager 212. The data stored in one or more memories 208 includes, for example, a set of one or more candidate neural network architecture configurations 216 for implementation at neural networks 122 and 132, and a set of one or more candidate generative model architecture configurations 218 for implementation at generative model 112.
[0033] One or more candidate neural network architecture configurations 216 include one or more data structures containing data and other information representing the corresponding architecture and / or parameter configuration of the corresponding neural network used by the UE neural network manager 212 to form the UE 110, such as in the UE decoder neural network 122 or the UE encoder neural network 132. The information included in the neural network architecture configuration 216 includes parameters specifying, for example, the following: fully connected layer neural network architecture, convolutional layer neural network architecture, recurrent neural network layers, the number of connected hidden neural network layers, input layer architecture, output layer architecture, the number of nodes utilized by the neural network, coefficients utilized by the neural network (e.g., weights and biases), kernel parameters, the number of filters utilized by the neural network, the stride / pooling configuration utilized by the neural network, the activation function of each neural network layer, the interconnections between neural network layers, neural network layers to be skipped, etc. Therefore, the neural network architecture configuration 216 includes any combination of neural network formation configuration elements that can be used to create a neural network formation configuration (e.g., a combination of one or more neural network formation configuration elements) that defines and / or forms a DNN. One or more candidate generative model architecture configurations 218 also include one or more data structures containing data and other information representing the corresponding architecture and / or parameter configurations used by the generative model manager 214 to form the corresponding generative model 112 of the UE 110.
[0034] Turning Figure 3An example hardware configuration of BS 108 is shown according to an embodiment. However, note that although the figures shown represent an implementation of BS 108 as a single network node (e.g., a 5G NR node B, a WiFi access point, a Bluetooth parent device, etc.), the functionality of BS 108 and therefore its hardware components can alternatively be distributed across multiple network nodes or devices, and can be distributed in a manner that performs the functions described herein. Like UE 110, BS 108 includes at least one array 302 of one or more antennas 303, an RF front end 304, one or more processors 306, and one or more non-transitory computer-readable storage media 308 (for brevity, computer-readable media 308 is referred to herein as "memory 308," similar to memory 208 in UE 110). These components operate in a manner similar to that described above with reference to the corresponding components of UE 110.
[0035] One or more memories 308 of BS 108 store one or more executable software instruction sets and associated data that manipulate one or more processors 206 and other components of BS 108 to perform the various functions described herein and belonging to BS 108. The set of executable software instructions includes, for example, an OS and various drivers (not shown), various software applications (not shown), BS manager 310, and BS neural network manager 312. BS manager 310 configures RF front-end 204 for communication with UE 110, and for communication via core network interface 314 with core network such as core network 102 and application server 106. BS neural network manager 312 implements one or more neural networks for BS 108, such as neural networks 120 and 134 for downlink path 118 and uplink path 130, respectively. Figure 1 ).
[0036] The data stored in one or more memories 308 of BS 108 includes one or more neural network architecture configurations 316, which represent one or more data structures containing data and other information representing the corresponding architecture and / or parameter configurations used by BS neural network manager 312 to form the corresponding neural network of BS 108. Similar to neural network architecture configuration 216 of UE 110, the information included in neural network architecture configuration 316 includes parameters specifying, for example, the following: fully connected layer neural network architecture, convolutional layer neural network architecture, recurrent neural network layers, number of connected hidden neural network layers, input layer architecture, output layer architecture, number of nodes utilized by the neural network, coefficients utilized by the neural network, kernel parameters, number of filters utilized by the neural network, stride / pooling configuration utilized by the neural network, activation function of each neural network layer, interconnections between neural network layers, neural network layers to be skipped, etc. Therefore, neural network architecture configuration 316 includes any combination of neural network formation configuration elements that can be used to create neural network formation configurations that define and / or form DNNs or other neural networks.
[0037] Figure 4 An example machine learning (ML) module 400 for implementing neural networks according to some embodiments is shown. As described herein, one or both of BS 108 and UE 110 implement one or more neural networks in downlink path 118 and uplink path 130, and UE 110 further implements generative model 112. Therefore, ML module 400 illustrates an example module for implementing one or more of these neural networks. For example, ML module 400 may be represented as being implemented as Figure 1 The ML module 400 can represent one of the following DNNs: 120, 122, 132, or 134. The ML module 400 can also represent one of the DNNs or other neural networks implemented as part of the generative model 112. For example, GANs utilize unstable DNN pairs: a generator and a discriminator. Thus, in an implementation where the generative model 112 is a GAN, one instance of the ML module 400 can be used to implement the generator DNN, and another instance of the ML module can be used to implement the discriminator DNN.
[0038] In the depicted example, the ML module 400 implements at least one deep neural network (DNN) 402, which has groups of connected nodes (e.g., neurons and / or perceptrons) organized into three or more layers. The nodes between layers can be configured in various ways, such as partially connected configurations, where a first subset of nodes in a first layer is connected to a second subset of nodes in a second layer; fully connected configurations, where each node in a first layer is connected to every node in a second layer, etc. Neurons process input data to produce continuous output values, such as any real number between 0 and 1. In some cases, the output value indicates how close the input data is to the desired category. Perceptrons perform linear classification on the input data, such as binary classification. Whether neurons or perceptrons, nodes can use various algorithms to generate output information based on adaptive learning. Using the DNN 402, the ML module 400 performs various different types of analysis, including unilinear regression, multiple linear regression, logistic regression, stepwise regression, binary classification, multi-class classification, multivariate adaptive regression splines, local estimation scatter smoothing, etc.
[0039] In some implementations, the ML module 400 performs adaptive learning based on supervised learning. In supervised learning, the ML module 400 receives various types of input data as training data. The ML module 400 processes the training data to learn how to map the inputs to the desired outputs. During the training process, the ML module 400 uses labeled or known data as input to the DNN 402. The DNN 402 uses nodes to analyze the inputs and generate corresponding outputs. The ML module 400 compares the corresponding outputs with real data and tunes the algorithm implemented by the nodes to improve the accuracy of the output data. Subsequently, the DNN 402 applies the tuned algorithm to unlabeled input data to generate corresponding output data. The ML module 400 uses one or both of statistical analysis and adaptive learning to map the inputs to the outputs. For example, the ML module 400 uses features learned from the training data to correlate unknown inputs with outputs that are statistically likely to be within a threshold range or value. This allows the ML module 400 to receive complex inputs and identify corresponding outputs. Some implementations train the ML module 400 with respect to the characteristics of communications transmitted over the wireless communication system (e.g., time / frequency interleaving, time / frequency deinterleaving, convolutional coding, convolutional decoding, power level, channel equalization, inter-symbol interference, quadrature amplitude modulation / demodulation, frequency division multiplexing / demultiplexing, and transmission channel characteristics). This allows the trained ML module 400 to receive signal samples as input (such as samples of downlink signals received at the UE) and recover information (such as binary data embedded in the downlink signals) from the downlink signals.
[0040] In the depicted example, DNN 402 includes an input layer 404, an output layer 406, and one or more hidden layers 408 located between the input layer 404 and the output layer 406. Each layer has an arbitrary number of nodes, wherein the number of nodes between layers can be the same or different. That is, the input layer 404 can have the same and / or different number of nodes as the output layer 406, the output layer 406 can have the same and / or different number of nodes as the one or more hidden layers 408, and so on.
[0041] Node 410 corresponds to one of the nodes included in the input layer 404, where these nodes perform separate, independent computations. As further described, a node receives input data and processes the input data using one or more algorithms to produce output data. Typically, the algorithms include weights and / or coefficients that change based on adaptive learning. Thus, the weights and / or coefficients reflect information learned by the neural network. In some cases, each node may determine whether to pass the processed input data to one or more subsequent nodes. For illustration, after processing the input data, node 410 may determine whether to pass the processed input data to one or both of nodes 412 and 414 in the hidden layer 408. Alternatively or additionally, node 410 passes the processed input data to nodes based on the layer connection architecture. This process may be repeated across multiple layers until DNN 402 generates output using nodes of the output layer 406 (e.g., node 416).
[0042] Neural networks can also employ a variety of architectures, which determine which nodes are connected, how data is advanced and / or preserved within the network, what weights and coefficients are used to process the input data, how the data is processed, and so on. These various factors collectively describe neural network architecture configurations, such as the neural network architecture configurations 216, 218, and 316 briefly described above. For illustration, recurrent neural networks (such as Long Short-Term Memory (LSTM) neural networks) form loops between node connections to preserve information from previous portions of the input data sequence. The recurrent neural network then uses the preserved information for subsequent portions of the input data sequence. As another example, feedforward neural networks pass information to forward connections without forming loops to preserve information. Although described in the context of node connections, it should be understood that neural network architecture configurations can include a variety of parameter configurations that influence how a DNN 402 or other neural networks process input data.
[0043] The neural network architecture configuration can be characterized by various architecture and / or parameter configurations. For illustration, consider the example of a DNN 402 implementing a convolutional neural network (CNN). Generally, a convolutional neural network corresponds to a type of DNN in which its layers use convolution operations to process data to filter the input data. Therefore, the CNN architecture configuration can be characterized by, for example, pooling parameters, kernel parameters, weights, and / or layer parameters.
[0044] Pooling parameters correspond to the parameters of a pooling layer within a specified convolutional neural network that reduces the dimensionality of the input data. For illustration, a pooling layer can combine the outputs of nodes in a first layer into the inputs of nodes in a second layer. Alternatively or additionally, pooling parameters specify how and where the neural network pools data in its data processing layers. For example, a pooling parameter indicating "max pooling" configures the neural network to pool by selecting the maximum value from the grouping of data generated by nodes in the first layer, and using that maximum value as the input to a single node in the second layer. A pooling parameter indicating "average pooling" configures the neural network to generate an average value from the grouping of data generated by nodes in the first layer, and using that average value as the input to a single node in the second layer.
[0045] Kernel parameters indicate the size of the filter (e.g., width and height) used to process the input data. Alternatively or additionally, kernel parameters specify the type of kernel method used to filter and process the input data. For example, support vector machines correspond to kernel methods that use regression analysis to identify and / or classify data. Other types of kernel methods include Gaussian processes, canonical correlation analysis, spectral clustering methods, etc. Therefore, kernel parameters can indicate the filter size and / or the type of kernel method to be applied in a neural network.
[0046] The weight parameters specify the weights and biases used by the algorithm within a node to classify the input data. In some implementations, the weights and biases are learned parameter configurations, such as those generated from the training data.
[0047] Layer parameters specify layer connections and / or layer types, such as fully connected layer types indicating which nodes in the first layer (e.g., output layer 406) are connected to each node in the second layer (e.g., hidden layer 408), partially connected layer types indicating which nodes in the first layer are disconnected from the second layer, activation layer types indicating which filters and / or layers are activated within the neural network, and so on. Alternatively or additionally, layer parameters specify the type of node layer, such as normalization layer type, convolutional layer type, pooling layer type, etc.
[0048] Although described in the context of pooling parameters, kernel parameters, weight parameters, and layer parameters, it should be understood that other parameter configurations can be used to form DNNs consistent with the guidelines presented herein. Therefore, neural network architecture configurations can include any suitable type of configuration parameters that can be applied to influence how a DNN processes input data to generate output data.
[0049] In some embodiments, the configuration of ML module 400 is based on the current context of UE 110, BS 108, or both, and the requirements or other parameters of one or more software applications 210 of the UE utilizing the output of generative model 112. For illustration, consider two jointly trained instances of ML module 400; one trained to efficiently compress or otherwise encode input prompts for wireless transmission, and the other trained to efficiently decompress or otherwise decode encoded input prompts from the first instance for input to generative model 112. The way software application 210 utilizes the output of generative model 112 generated from this input, and the quality of service (QoS) or quality of experience (QoE) requirements of software application 210 (such as minimum required bandwidth, maximum required latency, maximum required bit error rate), can inform the training of the two paired instances of ML module 400. For example, when the output of generative model 112 is audio output and software application 210 has relatively low QoS requirements, the paired ML module 400 can be trained to accept lower throughput, higher bit error rate, or lower audio resolution, thus benefiting reduced processing resource requirements at one or both of UE 110 or BS 108. Conversely, when the output of generative model 112 is high-resolution video and software application 210 has relatively high QoE requirements, the paired ML module 400 can be trained to higher throughput and accuracy standards at the cost of higher processor resource utilization.
[0050] Therefore, in some embodiments, the apparatus implementing ML module 400 generates and stores different neural network architecture configurations for different potential contexts such as different combinations of network conditions, available resource configurations, QoS / QoE requirements, and software application types. For this purpose, paired instances of ML module 400 can be trained for each combination of interest, and training can occur offline when no active communication exchanges occur, or online during active communication exchanges. The training of such paired instances of ML module 400 is described in detail below with reference to FIG8.
[0051] Figure 5A signal diagram 500 illustrating example operation of a network 100 for implementing a distributed generative modeling scheme according to some embodiments is shown. Operation of the distributed generative modeling scheme is typically initiated using a local application trigger 502 at UE 110, which may include, for example, a software application 210 using the output of generative model 112. Figure 2 The UE 110 initiates or otherwise initiates events such as startup, execution of specific subroutines, libraries, function calls, and application programming interfaces (APIs). In response, at action 504, the UE neural network manager 212 determines the current context of the UE 110, as it relates to the configuration of downlink path 118, uplink path 130, and / or generative model 112. This current context can be represented at least in part based on the current iteration of sensor data 136 obtained from sensor set 144, and can represent, for example, current network channel conditions such as available bandwidth or observed latency, signal-to-noise ratio (SNR) observed by the UE, signal-to-interference-plus-noise ratio (SINR), reference signal received power (RSRP), reference signal received quality (RSRQ), etc. Sensor data 136 can also further represent the current context of the UE 110's own operating state, such as available processing resources or processing resource utilization, available memory resources or memory resource utilization, and current battery capacity. Sensor data 136 may also include data representing the relationship between UE 110 and its physical environment, including attitude data, location data, radar data, lidar data, etc. Furthermore, the current context may include an identifier of the type of software application 210 that initiated the local application trigger 502 (e.g., video, audio, VR, XR, gaming, telepresence, etc.), and the software application 210's requirements regarding the output to be generated by the generative model 112, such as QoS requirements or QoE requirements.
[0052] Then, UE 110 wirelessly transmits a request 506 for instantiation of the generative model to BS 108, wherein this request 506 includes data representing or associated with the determined current context of UE 110. In one embodiment, one or both of the request 506 or the data representing the current context of UE 110 are transmitted to BS 108 as part of a UE Capability Information RRC message or a UE Assistance Information RRC message. For example, in one embodiment, UE 110 transmits request 506 to BS 108, and in response, BS 108 replies to UE 110 using a UE Capability Query RRC message. In response to the UE Capability Query RRC message, UE 110 determines data representing the current context and transmits this data as a UE Capability Information RRC message to BS 108. In other embodiments, request 506 and the data representing the current context of the UE are transmitted together as a UE Capability Information RRC message. At action 508, BS 108 can then determine its own current situation, such as network channel conditions observed by the BS, such as bandwidth, delay, bit error rate, SNR, SINR, RSRP, and / or RSRQ observed by the BS. The current situation of the BS may further include the current resource availability or utilization at BS 108.
[0053] See below for reference. Figure 7In a more detailed description, in the implementation, various combinations of parameter values for the parameters used in the UE and / or BS scenarios can be identified, and then, for each identified combination of parameter values, corresponding tuples for neural network architecture configurations for each of the generative model 112, the network encoder neural network 120, and the UE decoder neural network 122 are jointly trained to generate corresponding candidate neural network paths. For example, for a set of parameters for a UE scenario in which a specific type of generative model is specified (e.g., the type of output generated by the generative model), a QoE expectation for the output is specified, a specific set of network channel conditions is specified, and a specific set of UE resource availability is specified, corresponding candidate network encoder neural network architecture configurations, candidate UE decoder neural network architecture configurations, and candidate generative model neural network architecture configuration tuples can be jointly trained for downlink path 118 and generative model 112 under simulated conditions that satisfy these specified parameters, in order to generate a candidate neural network path for downlink path 118 and generative model 112. For different combinations of specified generative model types, different QoE expectations, different sets of specified network channel conditions, and / or different sets of specified UE resource availability, different candidate network encoder neural network architecture configurations, candidate UE decoder neural network architecture configurations, and candidate generative model neural network architecture configuration tuples can be jointly trained for this different set of specified parameters to generate another candidate neural network path, and so on. This same process can be applied to determine one or more jointly trained candidate neural network architecture configuration sets for uplink path 130. Therefore, as a result of different neural network training iterations for different combinations of context parameters (for UE 110, BS108, and / or the network channel connecting both), network 100 can generate a set of tuples of candidate jointly trained neural network architecture configurations for downlink path 118 and generative model 112 and / or uplink path 130.
[0054] The context parameters selected for joint training of the corresponding candidate set of neural network architectures can include any kind and combination of context parameters. For example, selected context parameters can include hardware capabilities indicating specific processing, memory, and / or storage capacity, such as maximum instruction throughput, maximum memory access speed, current instruction throughput, current memory access speed, current utilization of specific hardware resources, etc. Selected context parameters can also include one or more QoS or QoE requirements, such as QoS requirements related to jitter, throughput, latency, data loss, etc. Similarly, power and / or thermal parameters can be considered. This can include, for example, whether UE 110 is connected to a non-battery power source or a battery power source, and if the latter, the amount of remaining battery power. Another example could be parameters relating to the overall thermal state of UE 110 or one or more components of UE 110—such as current skin temperature and its relationship to a specified threshold, or the thermal state of one or more processing components of UE 110 given one or more corresponding thresholds. The selected context parameters may also include parameters related to the software application 210 that will consume content generated by the generative model 112 and / or the type of content to be generated, such as how the software application will consume the content (e.g., audio output, video output, still image output, text output, etc.), the priority of the software application 210, etc. Parameters related to the current conditions of the network channel connecting UE 110 and BS 108 may be selected, such as the observed bandwidth, latency, bit error rate, SNR, SINR, RSRP, and / or RSRQ mentioned above. Furthermore, parameters related to the relationship between UE 110 and the physical environment, such as location and / or attitude, the presence or absence of interfering objects, etc., can be used as training parameters.
[0055] Candidate neural network architecture configurations determined through such joint training can be used for implementation at BS 108 and UE 110 in any of a variety of ways. In some embodiments, network 100 may utilize a centralized repository of these candidate neural network architecture configurations, such as at a component of core network 102 or at application server 106, and BS 108 or UE 110 may access the identified neural network architecture configuration for implementation as a corresponding neural network by requesting the identified neural network architecture configuration using a corresponding index or other identifier, or by being supplied with the identified neural network architecture configuration as a result of the selection of the identified neural network architecture configuration based on analysis of one or both of the current context of UE 110 or BS 108. In other implementations, some or all of the trained neural network architecture configuration sets may be stored in a local cache at each of BS 108 and UE 110, and BS 108 and UE 110 may each access the corresponding neural network architecture configuration from the corresponding local cache for implementation based on an index or other identifier associated with the corresponding neural network architecture configuration.
[0056] In some embodiments, application server 106 operates to select a neural network architecture configuration to be implemented for downlink path 118, uplink path 130, and / or generative model 112. In such cases, at action 510, BS 108 forwards request 506 to application server 106 along with data representing the current context of UE 110 and / or BS 108. In other embodiments, BS 108 (or other network edge components) operates to select a neural network architecture configuration for downlink path 118, uplink path 130, and / or generative model 112. In either method, at action 512, BS 108 or application server 106 uses the current context of UE 110 and / or the current context of BS 108 to select a suitable jointly trained encoder / decoder neural network architecture configuration pair from a corresponding set of candidate jointly trained network architecture configurations for implementation, and to select a suitable neural network architecture configuration from a corresponding set of candidate neural network architecture configurations for implementation as generative model 112. BS 108 / Application Server 106 may use the current context of UE 110 and / or BS 108 as input, via a lookup table (LUT) or other similar selection structure using the same inputs to make this selection in an algorithmic manner, or the selection itself may utilize a trained neural network that takes the current context of UE 110 and / or BS 108 as input.
[0057] The selection of a specific candidate neural network path determines the neural network architecture configuration for the network encoder neural network 120, the UE decoder neural network 122, and the generative model 112 for downlink path 118, and also determines the neural network architecture configuration for the UE encoder neural network 132 and the network decoder neural network 134 if uplink path 130 is implemented. Therefore, at action 514, the component selecting the candidate neural network path sends an indication of the selected neural network architecture configuration to BS 108 (if BS 108 is not the component making the selection) and to UE 110, wherein BS 108 receives the indication of the neural network architecture configuration to be implemented for neural networks 120 and 134, and UE 110 receives the indication of the neural network architecture configuration to be implemented for neural networks 112, 122, and 132. As described above, these indications may include identifiers used to index into local or global stores for such neural network architecture configurations, the actual weights of the neural network itself, connections and other parameters, or combinations thereof. In response to receiving this instruction (or in response to selection from candidate neural network paths), at action 516, BS 108 implements the indicated neural network architecture configuration at network encoder neural network 120, and if uplink path 130 is in use, uses, for example, at network decoder neural network 134. Figure 4 The corresponding instance of the ML module 400 implements the indicated neural network architecture configuration. Similarly, at action 518, the UE 110 implements the indicated neural network architecture configuration at the UE decoder neural network 122 and the generative model 112, and if the uplink path 130 is in use, the corresponding instance of the ML module 400 implements the indicated neural network architecture configuration at the UE encoder neural network 132.
[0058] With generative model 112 and paths 118 and 130 configured and initialized, the distributed generative modeling scheme is ready to provide the generated content at UE 110. Therefore, at action 520, application server 106 generates an initial set of one or more input prompts based on various factors such as the application type of software application 210, the type of content to be generated, and content generation parameters, and provides this initial set of one or more input prompts as an initial instance of input prompt information 116 for transmission to BS 108. For example, if the distributed generative modeling scheme is implemented to generate AR graphical content for display at UE 110, one or more input prompts may include descriptors of the AR graphical content to be generated. At action 522, the network encoder neural network 120 of BS 108 receives the input prompt information 116 as input and processes this input according to its trained architecture configuration to generate an initial instance of intermediate signal 124, which is essentially an encoded representation of the input prompt information 116. At action 524, BS 108 transmits intermediate signal 124 for reception by UE 110.
[0059] At action 526, the UE decoder neural network 122 receives the intermediate signal 124 as input and processes this input according to its trained architecture configuration to generate a corresponding instance of the final signal 126. Simultaneously, at action 528, the sensor set 144 generates an instance of sensor data 142. At least in part due to the nature of the joint training of the encoder neural network 120, the decoder neural network 122, and the generative model 112, and due to the generative nature of the generative model 112, in some implementations, the final signal 126 does not need to be a lossless reconstruction of the input prompt information 116 or the intermediate signal 124. Therefore, when the UE decoder neural network 122 processes the intermediate signal 124 as input to generate the final signal 126, in some implementations, specific processes for confirming 100% recovery accuracy can be bypassed, such as by bypassing or ignoring the Cyclic Redundancy Check (CRC) process, the Hybrid Automatic Repeat Request (HARQ) process, the Radio Link Control (RLC) Acknowledgment (ACK) process, etc. Furthermore, as a supplement to or in addition to hard output, the UE decoder neural network 122 can be trained to provide the final signal 126 as a soft output, such as a specific probability distribution as a function of the softmax output. As another example, the UE decoder neural network 122 can be trained to provide the final signal 126 as a semantic representation or meaning of the intermediate signal 124. For example, in the case where the intermediate signal 124 is actually an encoded representation of speech groups, instead of being trained to recover the speech groups in their exact original form, the UE decoder neural network 122 can alternatively use previous speech groups, previous speech contexts or situations, and the encoded representation of the speech signal to generate the desired speech sequence, even with some errors.
[0060] At action 530, generative model 112 receives final signal 126 and sensor data 142 as inputs and processes these inputs according to its architectural configuration to generate an instance of output 128, which represents UX content generated based on one or more input prompts transmitted in some form via downlink path 118 and sensor data 142 from sensor set 144, which provides an indication of the current context of UE 110 (e.g., relative to the surrounding physical environment). This generated output 128 is then passed at action 532 to one or more UX components 114 for processing to enhance the user experience at UE 110, depending on the type of generated content, such as via display output, audio output, etc. In some embodiments, at least one of the one or more UX components 114 is wirelessly connected to UE 110, such as via Bluetooth™ or WiFi Direct™.
[0061] Referring again to action 528, as explained above, the generative model 112 of UE 110 can utilize sensor data generated by the local sensor set 144 of UE 110 when generating output content. Similarly, the application server 106 can utilize the same sensor data from the sensor set 144 or a different set of sensor data when generating the next instance of the input prompt information 116. Figure 6 Signal diagram 600 illustrates an example of this use of local sensor data as feedback for configuring subsequent instances of input prompt information 116, depending on the implementation. The following description assumes an implementation where the sensor data fed back to application server 106 is the same sensor data used by generative model 112 (i.e., sensor data 136 and sensor data 142 are identical), but this same approach can be employed using different sets of sensor data from the same or different sets of sensors, according to the guidelines provided below.
[0062] In action 528 ( Figure 5 In the case where the UE 110 obtains the current instance of sensor data 142 from sensor set 144, it can also feed this sensor data back to application server 106 via uplink path 130. Therefore, at action 602, the sensor data is input to UE encoder neural network 132, which processes the sensor data according to its trained architecture configuration to generate an instance of intermediate signal 138, which is essentially an encoded representation of the input sensor data. At action 604, UE 110 wirelessly transmits intermediate signal 138 to BS 108. At action 606, intermediate signal 138 is provided as input to network decoder neural network 134, which processes this input according to its trained architecture configuration to generate a corresponding instance of final signal 140, which is essentially a decoded representation of the effectively encoded representation of the sensor data obtained at UE 110. At action 608, the final signal 140 is transmitted from BS 108 to the application server, and at action 610, as part of the process of generating the next instance of input prompt information 116 representing one or more input prompts, the application server 106 processes the recovered sensor data represented by the final signal 140.
[0063] To illustrate, a distributed generative modeling scheme of network 100 can be employed to support video conferencing applications (an example of software application 210), and sensor data from sensor set 144 can include data from radar sensors, lidar sensors, or other depth sensors that identify the location and size of objects in the immediate vicinity of UE 110. Application server 106 can utilize this information to construct input cues designed to trigger the generation of AR coverage for these objects. As another example, sensor data can include attitude information from attitude sensors for UE 110, which is then used by application server 106 to construct input cues that trigger the generation of video images corresponding to the current perspective of UE 110. The generated input cues can then be used in… Figure 5 In another iteration of the actions 520 to 532 of the signal diagram 500, a distributed generative modeling scheme provided by the jointly trained downlink path 118 and generative model 112 is used to generate, transmit, and process corresponding instances of the input prompt information 116.
[0064] As explained above, in some embodiments, the distributed generative modeling scheme of network 100 utilizes one candidate neural network path selected from multiple candidate neural network paths for the jointly trained downlink path 118 and generative model 112, so that the distributed generative modeling scheme can efficiently distribute the generative modeling workload between application server 106 and UE 110. In this implementation, each candidate neural network path includes a corresponding neural network architecture configuration for each of the network encoder neural network 120, UE decoder neural network 122, and generative model 112, the combination of which has been jointly trained with respect to a corresponding set of context parameters representing specific context parameters for UE 110 and / or BS 108—such as power parameters, resource availability parameters, thermal parameters, software type parameters, QoS / QoE parameters, etc.
[0065] Figure 7An example method 700 for such joint training of candidate neural network architecture configurations for downlink path 118 and generative model 112, based on the implementation, is shown. Although described with reference to downlink path 118, the same or similar process can be used for the joint training of neural networks 132 and 134 for uplink path 130 using the guidelines provided herein. The iteration of method 700 begins at box 702, where the training system identifies the context parameters to be associated with the corresponding candidate neural network path to be jointly trained for this iteration. For illustration, for this iteration, the corresponding parameter set may include: a specific range of processing resource utilization at UE 110 in combination with a specific QoS / QoE specification, a specific range of remaining battery power for UE 110, a specific range of thermal conditions for UE 110, a range of network channel conditions observed by UE 110, a range of network channel conditions observed by BS 108, and a range of UEs currently served by BS 108. At box 704, the training system initializes training instances of each of the network encoder neural network 120, the UE decoder neural network 122, and the generative model 112 based on the context parameters identified at box 702.
[0066] At box 706, the training system obtains a batch of training datasets reflecting the corresponding set of context parameters, such as a batch of training datasets comprising multiple training data instances, each training data instance including known input cue information and corresponding known and appropriate generated content. At box 708, the training system processes each training dataset in the chain of encoder neural network 120, decoder neural network 122, and generative model 112. This may include: inputting known input cue information from the training dataset into a training instance of encoder neural network 120; transmitting the obtained output through a network channel with constraints representing the channel characteristic parameters identified at box 702 (or transmitting the obtained output simulated by a simulated version of such a network channel with similar constraints); and processing the received results at a training instance of decoder neural network 122 at an actual or simulated UE with constraints consistent with those of the UE context parameters identified at box 702 (e.g., the same available processing resources as specified, the same remaining battery level as specified, etc.). The output is then processed at a training instance of generative model 112 on an actual or simulated UE with the same identified UE context parameters, resulting in the generated output content.
[0067] After processing the entire training dataset in this manner, at box 710, the training system determines the performance of the candidate neural network path in processing the batch of training dataset. This may include, for example, calculating a joint loss function using performance metrics associated with one or more of the generated outputs for each training dataset (e.g., how closely the generated outputs reflect the expected outputs), performance metrics of the real or simulated operations of each neural network in the neural network (e.g., how closely the operations meet the specified QoS / QoE objectives, maintain an acceptable battery consumption rate, and utilize processing resources within acceptable limits), etc. At box 712, the training system determines whether this analyzed performance indicates that the candidate neural network path has been adequately trained. For example, whether one or more joint loss functions indicate errors within the corresponding acceptable range. If the training system determines that the candidate neural network path has been adequately trained, at box 714, the resulting trained architecture configuration for each neural network in the neural network path is extracted and stored for subsequent use as the neural network architecture configuration for the corresponding candidate neural network path. This may include, for each of the encoder neural network 120, decoder neural network 122, and generative model 112 being trained, extracting and storing parameters such as specifying the fully connected layer neural network architecture, convolutional layer neural network architecture, recurrent neural network layer, multi-connected hidden neural network layer, input layer architecture, output layer architecture, multiple nodes utilized by the neural network, coefficients (e.g., weights and biases) utilized by the neural network, kernel parameters, multiple filters utilized by the neural network, stride / pooling configuration utilized by the neural network, activation functions of each neural network layer, interconnections between neural network layers, neural network layers to be skipped, etc.
[0068] Otherwise, if the training system determines in box 712 that the candidate neural network path is not sufficiently trained, then in box 716, the training system updates the architectural configuration of the neural networks in the candidate neural network path being trained by updating the weights and other parameters of the training instances of neural networks 120, 122, and 112 based on the aforementioned loss function, for example, using gradient-based backpropagation techniques. After the update, method 700 returns to box 706, where another batch of training data is obtained, and this next batch of training data is used to perform another iteration of the training process.
[0069] In some embodiments, certain aspects of the techniques described above may be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and certain data that, when executed by the one or more processors, manipulate the one or more processors to perform one or more aspects of the techniques described above. The non-transitory computer-readable storage medium may include, for example, disk or optical disk storage devices, solid-state storage devices such as flash memory, caches, random access memory (RAM), or one or more other non-volatile memory devices. The executable instructions stored on the non-transitory computer-readable storage medium may be source code, assembly language code, object code, or other instruction formats interpreted or otherwise executed by one or more processors.
[0070] Computer-readable storage media can include any storage medium or combination of storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compressed optical discs (CDs), digital versatile optical discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, or magnetic hard disks), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS) based storage media. Computer-readable storage media can be embedded in a computing system (e.g., system RAM or ROM), fixedly attached to a computing system (e.g., magnetic hard disks), removably attached to a computing system (e.g., optical discs or flash memory based on Universal Serial Bus (USB)), or coupled to a computer system via a wired or wireless network (e.g., network accessible storage device (NAS)).
[0071] It should be noted that not all activities or elements described in the general description above are necessary; parts of a specific activity or apparatus may be unnecessary, and one or more additional activities may be performed, or additional elements may be included in addition to those described. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. Moreover, the concepts have been described with reference to specific embodiments. However, those skilled in the art will understand that various modifications and changes can be made without departing from the scope of this disclosure as set forth in the appended claims. Therefore, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are intended to be included within the scope of this disclosure.
[0072] The benefits, other advantages, and solutions to the problems have been described above with reference to specific embodiments. However, these benefits, advantages, solutions to the problems, and any features that may lead to or make more prominent any benefit, advantage, or solution should not be construed as key, essential, or essential features of any or all claims. Moreover, the specific embodiments disclosed above are merely illustrative, as the disclosed subject matter can be modified and practiced in different but equivalent ways that would be apparent to those skilled in the art benefiting from the teachings herein. No limitation is intended on the details of the constructions or designs shown herein other than those described in the appended claims. Therefore, it is apparent that the specific embodiments disclosed above can be altered or modified, and all such variations are considered to be within the scope of the disclosed subject matter. Therefore, the protection sought herein is precisely what is set forth in the appended claims.
Claims
1. A computer-implemented method in a user equipment, comprising: Receive the first signal indicating input prompt information at the network interface; The first signal is processed at a first neural network of the user equipment to generate a second signal; Output content is generated based on the processing of the second signal at the generative model at the user equipment. as well as The output content is passed to the user experience component of the user device.
2. The method as described in claim 1, wherein, The output content is generated by further processing sensor data from one or more sensors of the user device at the generative model.
3. The method as described in claim 2, wherein, The one or more sensors include: a position sensor, an attitude sensor, a camera, a microphone, a thermal sensor, a battery sensor, a network interface, a processing resource utilization sensor, a radar sensor, or a lidar sensor.
4. The method of claim 2 or 3, further comprising: The sensor data is processed at the second neural network to generate a third signal; as well as The third signal is transmitted via a network channel to a network component connected to the user equipment.
5. The method of claim 4, wherein, The network channel is a wireless network channel.
6. The method of claim 4, wherein, The second neural network and the third neural network of the network components are jointly trained.
7. The method as described in any one of the preceding claims, further comprising: In response to transmitting context information from the user equipment to the network component, the user equipment receives an instruction from the network component of the first neural network, the context information representing the current context of the user equipment; as well as In response to receiving the instruction, the first neural network is implemented at the user equipment.
8. The method of claim 7, wherein, Transmitting the context information includes transmitting the context information as at least one of a UE capability information radio resource control message or a UE assistance information radio resource control message.
9. The method of claim 7, wherein, The context information includes at least one of the following: The current capabilities of the user equipment; The application type to process the output content; Service quality parameters; Experience quality parameters; The current location of the user equipment; or Network conditions for the network channel between the network component and the user equipment.
10. The method according to any one of claims 7 to 9, wherein, The indication of the first neural network includes at least one of the following: The identifier of one of a plurality of candidate neural networks accessible to the user equipment; or This represents the neural network architecture configuration of the first neural network.
11. The method as claimed in any of the preceding claims, wherein, The first neural network and the second neural network at the network component are jointly trained.
12. The method of claim 11, wherein, The first neural network, the second neural network, and the generative model are jointly trained.
13. The method of any one of claims 11 or 12, wherein, The first neural network and the second neural network are deep neural networks (DNNs).
14. The method as described in any of the preceding claims, wherein, The generative model includes at least one of the following: Generative Adversarial Networks; Generative pre-trained transformers; Variational autoencoder; or Diffusion network.
15. The method as claimed in any of the preceding claims, wherein, The network interface is a wireless network interface.
16. The method as claimed in any of the preceding claims, wherein, The user experience component is wirelessly connected to the user device.
17. A user equipment, comprising: Radio frequency interface; At least one processor coupled to the radio frequency interface; as well as A non-transitory computer-readable medium storing a set of instructions configured to manipulate one or both of the at least one processor or the radio frequency interface to perform the method as described in any one of claims 1 to 17.
18. A computer-implemented method, comprising: Receive input prompts from the application server at the network component; The input cue information is processed at a first neural network of the network component to generate a first signal for use as one or more input cues for a generative model at a user device connected to the network component via a network channel; as well as The first signal is transmitted for reception by the user equipment.
19. The method of claim 18, further comprising: Receive a second signal from the user equipment, the second signal representing sensor data captured by one or more sensors of the user equipment; The second signal is processed at the second neural network of the network component to generate a first output; as well as The first output is provided to the application server.
20. The method of claim 19, wherein: The second neural network and the third neural network at the user equipment are jointly trained.
21. The method of claim 18, wherein, The first neural network and the second neural network of the user equipment are jointly trained.
22. The method of claim 21, further comprising: Receive context information representing the current context of the user equipment from the user equipment; as well as The second neural network is provided to the user equipment based on the context information.
23. The method of claim 22, wherein, The indication of the second neural network includes at least one of the following: The identifier of one of a plurality of candidate neural networks accessible to the user equipment; or This represents the neural network architecture configuration of the second neural network.
24. The method according to any one of claims 21 to 23, wherein, The first neural network and the second neural network include deep neural networks (DNNs).
25. The method according to any one of claims 21 to 24, wherein, The first neural network, the second neural network, and the generative model of the user device are jointly trained.
26. The method of any one of claims 18 to 25, wherein, Transmitting the first signal for reception by the user equipment includes wirelessly transmitting the first signal for reception by the user equipment.
27. A network component, comprising: Radio frequency interface; At least one processor coupled to the radio frequency interface; as well as A non-transitory computer-readable medium storing a set of instructions configured to manipulate one or both of the at least one processor or the radio frequency interface to perform the method as described in any one of claims 16 to 23.
28. A method in a cellular system, comprising: Based on one or both of the current context of the user equipment or the network conditions of the network components and the network channels of the user equipment, the network components are configured to implement a first neural network and the user equipment is configured to use a second neural network and a generative model. Data representing prompts is generated at the application server for use by generative models implemented at the user device; The data is processed at the first neural network of the network component to generate a first signal, and the first signal is transmitted to the user equipment; The first signal is processed at the second neural network of the user equipment to generate the second signal; The second signal is processed at the generative model of the user equipment to generate output content; as well as The output content is processed at the user device to enhance the user experience.
29. The method of claim 28, further comprising: In addition to processing the second signal to generate the output content, sensor data, which is captured by one or more sensors of the user equipment, is processed concurrently at the generative model.
30. The method of claim 29, further comprising: The sensor data is processed at the third neural network of the user equipment to generate a third signal, and the third signal is transmitted to the network component; The third signal is processed at the fourth neural network of the network component to generate the fourth signal; as well as The fourth signal is provided to the application server.
31. A system comprising: Network components and user equipment for performing the method as described in any one of claims 28 to 30.