Device and method for index coding and beamforming optimization between wireless apparatuses
A reinforcement learning model with hierarchical agents optimizes index coding and beamforming for wireless devices, addressing inefficiencies in current systems by minimizing transmission time and enhancing system performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
- Filing Date
- 2025-05-13
- Publication Date
- 2026-06-04
AI Technical Summary
Current wireless communication systems are inefficient in utilizing unique characteristics of wireless data traffic, such as predictable demand for popular content, and existing optimization techniques for index coding and beamforming are complex and suboptimal.
A computing device employing a reinforcement learning model with a hierarchical structure of agents to determine index coding and beamforming behaviors that minimize transmission time between wireless devices, using a first-level agent for discrete actions and a second-level agent for continuous actions.
Faster and more accurate optimization of index coding and beamforming is achieved, reducing transmission time and improving overall system efficiency.
Smart Images

Figure KR2025006423_04062026_PF_FP_ABST
Abstract
Description
Wireless Inter-device Index Encoding and Beamforming Optimization Device and Method
[0001] The present invention relates to an apparatus and method for index encoding and beamforming optimization between wireless devices.
[0002] Wireless data traffic has been increasing explosively over the past few years due to text, voice messaging, and video streaming. This upward trend is expected to intensify further due to services such as virtual reality, augmented reality, and holograms. While wireless data traffic possesses unique characteristics, such as preferences for popular content and predictable demand, current wireless communication systems have improved efficiency through additional wireless resources without reflecting these unique traits. However, methods for utilizing wireless resources have reached a certain limit, necessitating new technologies. For this reason, caching techniques have emerged in 5G wireless communication systems as a new technology, allowing network users to store a portion of their data in memory. Caching is known to significantly reduce overall network traffic. According to general caching methods, multicast data transmission is provided when user terminals request the same data, while unicast transmission is provided when users request different data.
[0003] Index coding techniques have been utilized to handle multicast data transmission for users requiring different data. Index coding is a method of transmitting data by XORing two different bit streams, referring to an encoding technique that satisfies all requests from multiple users with minimal transmission. The index coding problem is defined when a single central server possesses all requested file libraries; research is now being extended to the cross-device index coding problem, which involves finding index coding techniques for devices with limited file libraries. While existing studies have primarily focused on index coding in wired environments, extending this to wireless multi-antenna systems can yield greater gains in actual transmission time by utilizing spatial multiplexing benefits.
[0004] In consideration of this, the present invention proposes an index encoding technique between wireless devices with multiple antennas that minimizes transmission time based on reinforcement learning techniques to alleviate the high complexity of existing optimization techniques.
[0005] (Patent Document 1) Korean Published Patent No. 10-2024-0049112 (Title of Invention: Method and Apparatus for Transmitting Channel Status Information)
[0006] The present invention aims to solve the aforementioned problems and has the technical objective of providing an apparatus and method for performing index encoding and beamforming optimization between wireless devices using a reinforcement learning model.
[0007] However, the technical problem that this embodiment aims to solve is not limited to the technical problem described above, and other technical problems may exist.
[0008] As a technical means for solving the aforementioned technical problem, a computing device for performing index coding and beamforming optimization between wireless devices according to the first aspect of the present invention comprises a memory storing an optimization program and a processor executing said optimization program, wherein the optimization program inputs state information of wireless devices into a reinforcement learning model to determine an index coding behavior and a beamforming design behavior that minimizes the transmission time between each wireless device. And, the reinforcement learning model includes a first-level agent that determines whether to perform index coding for each wireless device and the type of index coding technique to be used for index coding as said index coding behavior, and a second-level agent that determines an optimal beamformer design plan for each wireless device as said beamforming design behavior by considering the index coding behavior determined by said first-level agent.
[0009] Additionally, a method for optimizing index coding and beamforming between wireless devices, performed by a computing device according to a second aspect of the present invention, comprises: receiving state information of each wireless device; and inputting the state information of each wireless device into a reinforcement learning model to determine an index coding behavior and a beamforming design behavior that minimizes the transmission time between each wireless device. The reinforcement learning model includes a first-level agent that determines whether to perform index coding of each wireless device and the type of index coding technique to be used for index coding as the index coding behavior, and a second-level agent that determines an optimal beamformer design for each wireless device as the beamforming design behavior by considering the index coding behavior determined by the first-level agent.
[0010] According to the aforementioned means for solving the problem of the present invention, it was confirmed that by applying reinforcement learning, faster and more accurate optimization actions are determined compared to conventional optimization techniques.
[0011] FIG. 1 illustrates a wireless communication system to which the present invention is applied.
[0012] FIG. 2 illustrates the detailed configuration of a computing device according to one embodiment of the present invention.
[0013] FIG. 3 illustrates the specific configuration of a reinforcement learning model according to one embodiment of the present invention.
[0014] FIG. 4 is a flowchart illustrating a wireless inter-device index encoding and beamforming optimization method of a computing device according to an embodiment of the present invention.
[0015] The present invention will be described in detail below with reference to the attached drawings. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification, and the technical concept disclosed in this specification is not limited by the attached drawings. In order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and the size, form, and shape of each component shown in the drawings may be varied in various ways. Identical or similar parts throughout the specification are denoted by identical or similar reference numerals.
[0016] Suffixes such as "module" and "part" for components used in the following description are assigned or used interchangeably solely for the sake of ease of drafting the specification, and do not inherently possess distinct meanings or roles. Furthermore, in describing the embodiments disclosed in this specification, detailed descriptions of related prior art have been omitted where it is determined that such detailed descriptions could obscure the essence of the embodiments disclosed in this specification.
[0017] Throughout the specification, when it is stated that a part is "connected (connected, contacted, or coupled)" to another part, this includes not only cases where they are "directly connected (connected, contacted, or coupled)," but also cases where they are "indirectly connected (connected, contacted, or coupled)" with other members interposed therebetween. Furthermore, when it is stated that a part "includes (provides, or provides)" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but rather allows for additional "included (provided, or provided)" of other components.
[0018] Terms indicating ordinal numbers, such as first, second, etc., used in this specification are used solely for the purpose of distinguishing one component from another and do not limit the order or relationship of the components. For example, the first component of the present invention may be named the second component, and similarly, the second component may be named the first component.
[0019] FIG. 1 illustrates a wireless communication system to which the present invention is applied, and FIG. 2 illustrates a detailed configuration of a computing device according to an embodiment of the present invention.
[0020] The wireless communication system (10) includes a plurality of user terminals (200 to 204) capable of wireless communication, and a computing device (100) that performs index encoding and beamforming optimization between each user terminal.
[0021] Each user terminal (200 to 204) is a wireless communication device that ensures portability and mobility, and may include all types of handheld-based wireless communication devices such as various smartphones, tablet PCs, and smartwatches. Additionally, each user terminal (200 to 204) is a device capable of communicating via a wireless data communication network, and is capable of communicating via 5G, 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, etc. In particular, each user terminal (200 to 204) supports at least one index coding technique and beamforming design.
[0022] A computing device (100) determines index coding behavior and beamforming design behavior that minimize transmission time between each wireless device using a reinforcement learning model in an environment that includes multiple user terminals (200 to 204). The computing device (100) may be configured in the form of a server and may operate in a cloud computing service model such as SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service).
[0023] The computing device (100) includes memory (120) and a processor (130), and may further include a communication module (110) and a database (140).
[0024] The computing device (100) is capable of wireless communication with each user terminal (200-204) and can be implemented as a computer or portable terminal capable of connecting to other computing devices through a network. A network refers to a connection structure capable of exchanging information between each node, such as terminals and devices, and includes a Local Area Network (LAN), a Wide Area Network (WAN), the World Wide Web (WWW), wired and wireless data communication networks, telephone networks, wired and wireless television communication networks, etc. Examples of wireless data communication networks include, but are not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, Bluetooth communication, infrared communication, ultrasonic communication, Visible Light Communication (VLC), and LiFi.
[0025] The communication module (110) may include a device comprising hardware and software necessary to transmit and receive signals, such as control signals or data signals, through a wired or wireless connection with another network device.
[0026] The memory (120) stores an optimization program, and the optimization program inputs state information of wireless devices into a reinforcement learning model to determine index coding behavior and beamforming design behavior that minimize the transmission time between each wireless device. At this time, the reinforcement learning model includes a first-level agent that determines whether to perform index coding for each wireless device and the type of index coding technique to be used for index coding as an index coding behavior, and a second-level agent that determines an optimal beamformer design plan for each wireless device as a beamforming design behavior by considering the index coding behavior determined by the first-level agent.
[0027] The term "memory" (120) should be interpreted as a collective term for a non-volatile storage device that retains stored information even when power is not supplied, and a volatile storage device that requires power to retain stored information. The memory (120) can perform the function of temporarily or permanently storing data processed by the processor (130). In addition to a volatile storage device that requires power to retain stored information, the memory (120) may include a magnetic storage medium or a flash storage medium, but the scope of the present invention is not limited thereto.
[0028] The processor (130) executes an optimization program stored in memory (120) and transmits the index encoding behavior and beamforming design behavior of each wireless device to each wireless device as a result of the execution.
[0029] In one example, the processor (130) may be implemented in the form of a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the invention is not limited thereto.
[0030] The database (140) manages status information of each wireless device included in the system. It also manages various programs and data required for the execution of the optimization program. In particular, it manages training data for training a reinforcement learning model, and information on the index encoding behavior and beamforming design behavior of each wireless device inferred by the reinforcement learning model.
[0031] The system (10) of the present invention can be described as follows. The system (10) includes a total of K user terminals (200 to 204) capable of wireless communication, and each user terminal has K-1 antennas. Additionally, each user terminal stores a portion of the entire file library in a cache, and a situation is assumed where a file that is not possessed by each user terminal is requested. Here, each file in the file library is stored in the cache of at least one user terminal. File transmission requested by each user terminal is performed through data transmission between user terminals, and to reduce transmission time, index encoding and multicast beamforming techniques are used for transmission. At this time, a half-duplex transmission method is assumed, which means a transmission method in which reception does not occur while transmitting, and transmission does not occur while receiving. In other words, each user terminal acts as a transmitting node and a receiving node, and two or more devices cannot act as transmitting nodes at the same time. Therefore, the signal (y) received by the k-th user terminal in the t-th communication round (the round in which the t-th user terminal transmits) can be expressed as follows.
[0032] [Mathematical Formula 1]
[0033]
[0034] Here, x represents the transmitted signal, H represents the channel gain, and z represents the noise. The set of encoded messages transmitted by the t-th user terminal , set of indices of encoded messages When that is the case, the transmitted signal (x) can be represented as follows.
[0035] [Mathematical Formula 2]
[0036]
[0037] v represents the transmit beamforming vector, and m represents the requested data. Each user terminal has a transmission power of P. Let be the index of the message containing the data requested by the k-th user terminal. The received signal of Equation 1 can be expressed by dividing it into a request signal and an interference signal as in Equation 3.
[0038] [Mathematical Formula 3]
[0039]
[0040] Therefore, the signal-to-interference-and-noise ratio (SINR) is expressed as follows.
[0041] [Mathematical Formula 4]
[0042]
[0043] For reference, in mathematical formula 4 is a set consisting of indices of encoded messages, a set of encoded messages ( It includes ).
[0044] If we define as the set of destination user terminals for messages sent by the t-th user terminal, the transmission time of the entire communication round can be expressed as follows.
[0045] [Mathematical Formula 5]
[0046]
[0047] Here, B represents the file size and W represents the bandwidth.
[0048] The present invention is an index encoding technique that minimizes the transmission time of the entire communication round defined in Equation 5. and Beamformer We intend to optimize . For reference, since the index encoding technique refers to how the encoded message is constructed, the previously defined set of encoded messages It is defined as a concept corresponding to.
[0049] Since the present invention assumes a half-duplex transmission situation, once the index coding technique is determined, it can be replaced with a multicast beamforming problem of finding the optimal beamformer for each transmission round. The optimization problem representing this is as follows.
[0050] [Mathematical Formula 6]
[0051]
[0052] In a situation where an index encoding technique is given first, beamformer design must proceed dependently thereafter; however, since index encoding and beamformers are actions in a discrete action space and a continuous action space, respectively, there is a disadvantage in that they are difficult to design using existing reinforcement learning algorithms. Therefore, in this invention, two different agents are designed to have a hierarchical structure within a single environment so that they are trained with the goal of convergence of the reinforcement learning algorithm.
[0053] FIG. 3 illustrates the specific configuration of a reinforcement learning model according to one embodiment of the present invention.
[0054] First, for each communication round, a Level 1 agent learned in a higher-level environment determines index encoding behavior, and a Level 2 agent learned in a lower-level environment determines beamformer design behavior.
[0055] The first-level agent uses a discrete action space to determine, as an index coding action, whether to perform index coding for each user terminal and, if so, which index coding technique to apply. The first-level agent receives an indicator as the state (St) input of the reinforcement learning environment, which includes the channel state information, transmission power, and information on user terminals whose transmission has not yet been completed for all user terminals. At this time, in addition to the environment state (St) information, environment goal (Gt) information may also be transmitted to the first-level agent. Based on this environment state information (St) and environment goal information (Gt), the first-level agent performs an index coding action (Action, Determines ).
[0056] The second-level agent receives not only the environment state information (St) and environment goal information (Gt) as input values from the first-level agent, but also the index encoding behavior determined by the first-level agent ( ) is received as input to the second-level agent, and based on this, the second-level agent performs beamformer design action (Action, Determines ). Once the behavior of the 2nd level agent is determined by the hierarchical structure of the 1st level agent and the 2nd level agent, the delay time is calculated based on the mathematical formula 5 described earlier using this.
[0057] [Mathematical Formula 5]
[0058]
[0059] Based on this, the first reward of the Level 1 Agent (Reward, ) and the 2nd reward of the 2nd level agent (Reward, ) is defined as a value obtained by multiplying the delay time by -1, and accordingly, a state-behavior to reduce the delay time is learned.
[0060] For every communication round, the first-level agent and the second-level agent hierarchically and cooperatively determine the actions of each round (index encoding and beamformer design), and when the Kth round ends, the episode ends.
[0061] Meanwhile, according to the embodiment, the first level agent may learn discrete behaviors by utilizing the conventionally known Dueling Deep Q-Learning Network algorithm, and the second level agent may learn continuous behaviors by utilizing the SAC (Soft actor-critic) algorithm, but it is also possible to utilize any reinforcement learning algorithm for discrete behaviors / continuous behaviors in the above hierarchical structure.
[0062] FIG. 4 is a flowchart illustrating a wireless inter-device index encoding and beamforming optimization method of a computing device according to an embodiment of the present invention.
[0063] First, a computing device (100) receives status information of each wireless device (S410).
[0064] At this time, the status information of each wireless device may include channel status information and information on transmission power.
[0065] Next, the state information of each wireless device is input into a reinforcement learning model to determine the index encoding behavior and beamforming design behavior that minimize the transmission time between each wireless device (S420).
[0066] At this time, the reinforcement learning model includes a first-level agent and a second-level agent with a hierarchical structure as previously discussed. The first-level agent determines whether to perform index coding for each wireless device and the type of index coding technique to use for index coding as an index coding behavior. Then, the second-level agent determines the optimal beamformer design for each wireless device as a beamforming design behavior, taking into account the index coding behavior determined by the first-level agent.
[0067] Meanwhile, although not illustrated in the drawing, the method may further include a step of transmitting the determined index encoding behavior and beamforming design behavior to each wireless device.
[0068] The wireless inter-device index encoding and beamforming optimization methods described above may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, as well as removable and inseparable media. Additionally, a computer-readable medium may include a computer storage medium. A computer storage medium includes both volatile and non-volatile, removable and inseparable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data.
[0069] A person skilled in the art to which the present invention pertains will understand that, based on the description above, other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims set forth below, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts should be interpreted as being included within the scope of the present invention.
[0070] [Explanation of the symbol]
[0071] 100: Computing device
[0072] 110: Communication module
[0073] 120: Memory
[0074] 130: Processor
[0075] 140: Database
[0076] 200: User terminal
Claims
1. A computing device that performs wireless inter-device index encoding and beamforming optimization, Memory where the optimization program is stored and A processor that executes the above optimization program, comprising The above optimization program inputs state information of wireless devices into a reinforcement learning model to determine index coding behavior and beamforming design behavior that minimize the transmission time between each wireless device, A computing device comprising: a first-level agent that determines, as an index encoding action, whether to perform index encoding of each wireless device and the type of index encoding technique to be used for index encoding; and a second-level agent that determines, as a beamforming design action, an optimal beamformer design for each wireless device by considering the index encoding action determined by the first-level agent.
2. In Paragraph 1, The first level agent determines the index encoding action based on channel state information, transmission power, and a first compensation of the wireless devices, and A computing device in which the second level agent determines the beamforming design behavior based on channel state information of the wireless devices, transmission power, index encoding behavior determined by the first level agent, and a second compensation.
3. In Paragraph 2, A computing device in which the first and second compensations are each defined as a value obtained by multiplying the delay time by -1.
4. In Paragraph 1, A computing device that transmits the above-determined index encoding behavior and beamforming design behavior to each wireless device.
5. A method for wireless inter-device index encoding and beamforming optimization performed by a computing device, A step of receiving status information of each wireless device; and The method includes the step of inputting state information of each of the above wireless devices into a reinforcement learning model to determine index encoding behavior and beamforming design behavior that minimize transmission time between each wireless device, A method for optimizing index encoding and beamforming between wireless devices, wherein the reinforcement learning model comprises a first-level agent that determines whether to perform index encoding of each wireless device and the type of index encoding technique to be used for index encoding as an index encoding action, and a second-level agent that determines an optimal beamformer design for each wireless device as a beamforming design action by considering the index encoding action determined by the first-level agent.
6. In Paragraph 5, The step of determining the above index encoding behavior and beamforming design behavior is The step of the first level agent determining the index encoding action based on channel state information, transmission power, and a first compensation of the wireless devices, and A method for index coding and beamforming optimization between wireless devices, wherein the second level agent determines the beamforming design behavior based on channel state information of the wireless devices, transmission power, index coding behavior determined by the first level agent, and a second compensation.
7. In Paragraph 6, A wireless inter-device index encoding and beamforming optimization method in which the first and second compensations are each defined as values obtained by multiplying the delay time by -1.
8. In Paragraph 5, A method for optimizing index encoding and beamforming between wireless devices, further comprising the step of transmitting the determined index encoding behavior and beamforming design behavior to each wireless device.
9. A non-transient computer-readable recording medium having a computer program recorded thereon for performing a wireless inter-device index encoding and beamforming optimization method according to any one of claims 5 to 8.