Communication management device and communication management method
The communication management device optimizes relay network routing using reinforcement and supervised learning to enhance communication efficiency and reduce latency by learning optimal paths for packet transmission.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERNET INITIATIVE JAPAN INC
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Conventional technologies do not consider the efficient routing of the relay network between the User Plane Function (UPF) of the core network and the cloud or edge location in Multi-access Edge Computing (MEC) systems, leading to suboptimal communication paths.
A communication management device and method that utilize reinforcement and supervised learning models to learn and optimize the communication route through a relay network, selecting optimal paths for packet transmission using a reward function and learning models to maximize efficiency.
Provides an optimal communication route for the relay network between the gateway of a mobile communication network and external network resources, enhancing communication efficiency and reducing latency.
Smart Images

Figure 2026068792000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a communication management device and a communication management method. [Background technology]
[0002] Conventionally, Multi-access Edge Computing (MEC), a method for providing cloud computing functions in a location physically close to the user terminal, has been known in mobile communication networks (see Patent Document 1).
[0003] For example, Patent Document 2 discloses a technology that, when a user terminal accesses a specific website, identifies the base station to which the user terminal connects and controls communication to communicate with the cloud at the cloud location that is physically closest to the identified base station.
[0004] However, the technologies described in Patent Documents 1 and 2 do not consider the efficient routing of the relay network between gateways such as the User Plane Function (UPF) of the core network and the cloud or edge location to which the user terminal is connected. For example, in the technology disclosed in Patent Document 2, in order for a user terminal to access a specific website, the path that is physically shorter between the base station where the user terminal is located, the UPF of the core network, and the cloud location where the specific website is hosted is applied to the user terminal. However, Patent Document 2 does not consider the efficient routing of the relay network, which consists of a group of relay nodes, that exists between the UPF and the cloud location. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2018-160813 [Patent Document 2] Patent No. 7481595 [Overview of the Initiative]
Problems to be Solved by the Invention
[0006] In the conventional technology, the optimal communication route of the relay network between the gateway of the mobile communication network and the external network resources has not been considered.
[0007] The present invention has been made to solve the above-described problems, and an object thereof is to provide an optimal communication route of a relay network between a gateway of a mobile communication network and external network resources.
Means for Solving the Problems
[0008] To solve the above-mentioned problems, the communication management device according to the present invention is a communication management device that manages a communication route from a user terminal to a specific external network resource via a relay network in which a plurality of relay nodes are connected, via a gateway provided in the core network of a mobile communication network, and includes an acquisition unit configured to acquire information on the gateway through which the user terminal is communicating, and information on the relay nodes through which packets from the user terminal pass at each time, as information on the node through which packets from the user terminal are currently passing, and calculates the selection of a route for relay nodes that packets from the user terminal should sequentially pass from the gateway through the relay network to reach the specific external network resource. The system comprises: a first learning unit configured to apply a reward function to the estimation results to update them so as to maximize the reward for packets from the user terminal to reach the specific external network resource, and to learn a first route selection policy for relay nodes that packets from the user terminal should sequentially traverse from the gateway, using a first reinforcement learning model; a second learning unit configured to learn the relationship between the node that packets from the user terminal are currently traversing and the first route selection policy for relay nodes that packets from the user terminal should sequentially traverse from the gateway, obtained through learning by the first learning unit, using a first supervised learning model; and a storage unit configured to store the learned first supervised learning model constructed by the second learning unit.
[0009] In the communication management apparatus according to the present invention, the acquisition unit acquires information on a gateway with which a management target user terminal is communicating and information on a relay node through which packets from the management target user terminal pass at each time as information on a node through which packets from the management target user terminal are currently passing. Further, the acquisition unit provides the learned first supervised learning model with the information on the node through which packets from the management target user terminal are currently passing as unknown input, performs calculation of the learned first supervised learning model, and is configured to output a first route selection policy for relay nodes through which packets from the management target user terminal should sequentially pass from the gateway. The communication management apparatus may further include a communication route management unit configured to notify the gateway and the plurality of relay nodes of route information determined based on the first route selection policy output by the calculation unit.
[0010] In the communication management apparatus according to the present invention, the acquisition unit acquires information on a relay node through which a packet addressed to the user terminal from the specific external network resource passes at each time as information on a node through which the packet addressed to the user terminal is currently passing. The first learning unit applies a reward function to an estimation result obtained by calculating a selection of a route for relay nodes through which a packet addressed to the user terminal should sequentially pass until the packet reaches the gateway via the relay network, and updates the reward so that the reward for the packet addressed to the user terminal reaching the gateway is maximized. The first learning unit learns a second route selection policy for relay nodes through which a packet addressed to the user terminal should sequentially pass from the specific external network resource using a second reinforcement learning model. The second learning unit learns the relationship between the node through which the packet addressed to the user terminal is currently passing and the second route selection policy for relay nodes through which the packet addressed to the user terminal should sequentially pass from the specific external network resource obtained by learning by the first learning unit using a second supervised learning model. The storage unit may store the learned second supervised learning model constructed by the second learning unit.
[0011] Furthermore, in the communication management device according to the present invention, the acquisition unit acquires information on the relay nodes that packets destined for the managed user terminal from the specific external network resource pass through at each time, as information on the node that packets destined for the managed user terminal are currently passing through; the calculation unit provides the information on the node that packets destined for the managed user terminal are currently passing through as unknown input to the trained second supervised learning model, performs calculations on the trained second supervised learning model to output the second route selection strategy for the relay nodes that packets destined for the managed user terminal should sequentially pass through from the specific external network resource; and the communication route management unit may notify the specific external network resource and the plurality of relay nodes of the route information determined based on the second route selection strategy output by the calculation unit.
[0012] Furthermore, in the communication management device according to the present invention, the user terminals are a plurality of user terminals to which common identification information is assigned, the gateways through which each user terminal communicates include different gateways from each other, the acquisition unit acquires information on the nodes through which packets from each user terminal are currently passing, the first learning unit learns the first route selection strategy for packets from each user terminal using the first reinforcement learning model, and the second learning unit may learn the relationship between the nodes through which packets from each user terminal are currently passing and the first route selection strategy for relay nodes through which packets from each user terminal should sequentially pass from the gateway through which each user terminal communicates, as learned by the first learning unit, using the first supervised learning model.
[0013] Furthermore, in the communication management device according to the present invention, the gateway is a user plane function, and further comprises a management information storage unit configured to store management information relating the common identification information, subscriber identification information of the plurality of user terminals, and identification information of a specific external network resource, and a setting unit configured to specify the management information based on a location registration request signal from each of the plurality of user terminals and to instruct the setting of the user plane function of each of the plurality of user terminals, wherein the management information storage unit further stores the identification information of each of the plurality of user terminals' user plane functions set according to the instructions of the setting unit, relating it to the management information, and the identification information of the user plane function may indicate information of the gateway with which each of the user terminals is communicating.
[0014] Furthermore, in the communication management device according to the present invention, the relay network has a hierarchical structure in which relay nodes with common network addresses are connected to each other, and each of the first route selection strategy and the second route selection strategy may include a probabilistic route selection strategy from any relay node in the relay node group having a first network address, which the packet will pass through at the first time step, to each relay node in the relay node group having a second network address, which the packet will pass through at the next second time step.
[0015] To solve the above-mentioned problems, the present invention provides a communication management method for managing a communication route from a user terminal to a specific external network resource via a relay network in which multiple relay nodes are connected, via a gateway provided in the core network of a mobile communication network, and comprising: an acquisition step of acquiring information on the gateway through which the user terminal is communicating, and information on the relay nodes through which packets from the user terminal pass at each time, as information on the node through which packets from the user terminal are currently passing; and a calculation of the selection of a route for relay nodes that packets from the user terminal should sequentially pass from the gateway through the relay network to reach the specific external network resource. The system includes: a first learning step in which a reward function is applied to the estimation results to update them so as to maximize the reward for packets from the user terminal to reach the specific external network resource, and a first route selection policy for relay nodes that packets from the user terminal should sequentially traverse from the gateway is learned using a first reinforcement learning model; a second learning step in which the relationship between the node that packets from the user terminal are currently traversing and the first route selection policy for relay nodes that packets from the user terminal should sequentially traverse from the gateway, obtained through learning in the first learning step, is learned using a first supervised learning model; and a storage step in which the learned first supervised learning model constructed in the second learning step is stored in a storage unit.
[0016] Furthermore, in the communication management method according to the present invention, the acquisition step may include: an acquisition step of information about the gateway through which the managed user terminal is communicating, and information about the relay nodes through which packets from the managed user terminal pass at each time, as information about the node through which packets from the managed user terminal are currently passing; an calculation step of providing the information about the node through which packets from the managed user terminal are currently passing as an unknown input to the trained first supervised learning model, performing calculations on the trained first supervised learning model to output the first route selection strategy for the relay nodes through which packets from the managed user terminal should sequentially pass from the gateway; and a communication route management step of notifying the gateway and the plurality of relay nodes of the route information determined based on the first route selection strategy output in the calculation step.
[0017] Furthermore, in the communication management method according to the present invention, the acquisition step acquires information on the relay nodes that packets destined for the user terminal from the specific external network resource pass through at each time, as information on the node that packets destined for the user terminal are currently passing through, and the first learning step applies a reward function to the estimation result of calculating the selection of the route that packets destined for the user terminal should sequentially pass through to reach the gateway via the relay network, and updates the information so as to maximize the reward for packets destined for the user terminal to reach the gateway, and packets destined for the user terminal The system learns a second route selection strategy for relay nodes that packets destined for the user terminal should sequentially traverse from the specific external network resource, using a second reinforcement learning model. The second learning step learns the relationship between the node that packets destined for the user terminal are currently traversing and the second route selection strategy for relay nodes that packets destined for the user terminal should sequentially traverse from the specific external network resource, obtained through learning in the first learning step, using a second supervised learning model. The storage step may store the learned second supervised learning model constructed in the second learning step in the storage unit.
[0018] Furthermore, in the communication management method according to the present invention, the acquisition step may acquire information on the relay nodes through which packets destined for the managed user terminal from the specific external network resource pass at each time, as information on the node through which the packets destined for the managed user terminal are currently passing; the calculation step may provide the information on the node through which the packets destined for the managed user terminal are currently passing as an unknown input to the trained second supervised learning model, perform calculations on the trained second supervised learning model to output the second route selection strategy for the relay nodes through which packets destined for the managed user terminal should sequentially pass from the specific external network resource; and the communication route management step may notify the specific external network resource and the plurality of relay nodes of the route information determined based on the second route selection strategy output in the calculation step.
[0019] Furthermore, in the communication management method according to the present invention, the user terminals are a plurality of user terminals to which common identification information is assigned, the gateways through which each user terminal communicates include different gateways from each other, the acquisition step acquires information on the nodes through which packets from each user terminal are currently passing, the first learning step learns the first routing strategy for packets from each user terminal using the first reinforcement learning model, and the second learning step may learn the relationship between the nodes through which packets from each user terminal are currently passing and the first routing strategy for relay nodes through which packets from each user terminal should sequentially pass from the gateway through which each user terminal communicates, as learned in the first learning step, using the first supervised learning model.
[0020] Furthermore, in the communication management method according to the present invention, the gateway is a user plane function, and the method further includes a management information storage step of storing management information in a management information storage unit that associates the common identification information, the subscriber identification information of the plurality of user terminals, and the identification information of a specific external network resource, and a setting step of specifying the management information based on a location registration request signal from each of the user terminals and instructing the setting of each of the plurality of user terminals' user plane functions, wherein the management information storage step further stores the identification information of each of the plurality of user terminals' user plane functions set in accordance with the instructions in the setting step, associating it with the management information, and the identification information of the user plane function may indicate information of the gateway with which each of the user terminals is communicating. [Effects of the Invention]
[0021] According to the present invention, the relationship between the node through which a packet from a user terminal is currently passing and a first route selection strategy for the relay nodes that the packet from the user terminal should sequentially pass through from the gateway is learned using a first supervised learning model. Therefore, it is possible to provide the optimal communication route for the relay network between the gateway of the mobile communication network and external network resources. [Brief explanation of the drawing]
[0022] [Figure 1] Figure 1 is a block diagram showing the configuration of a communication management system equipped with a communication management device according to an embodiment of the present invention. [Figure 2] Figure 2 is a block diagram showing the configuration of a communication management system equipped with a communication management device according to this embodiment. [Figure 3] Figure 3 is a diagram illustrating the structure of the first storage unit of the communication management device according to this embodiment. [Figure 4] Figure 4 is a diagram illustrating the structure of the first storage unit of the communication management device according to this embodiment. [Figure 5]Figure 5 is a diagram illustrating the learning process performed by the first learning unit of the communication management device according to this embodiment. [Figure 6] Figure 6 is a block diagram showing the configuration of the first learning unit included in the communication management device according to this embodiment. [Figure 7] Figure 7 is a diagram illustrating the learning process performed by the second learning unit of the communication management device according to this embodiment. [Figure 8] Figure 8 is a block diagram showing the hardware configuration of the communication management device according to this embodiment. [Figure 9] Figure 9 is a sequence diagram showing the operation of the communication management system according to this embodiment. [Figure 10] Figure 10 is a flowchart showing the operation of the communication management device according to this embodiment. [Figure 11] Figure 11 is a sequence diagram showing the operation of the communication management system according to this embodiment. [Figure 12] Figure 12 is a flowchart showing the first learning process of the communication management device according to this embodiment. [Figure 13] Figure 13 is a flowchart showing the first learning process of the communication management device according to this embodiment. [Figure 14] Figure 14 is a flowchart showing the operation of the communication management device according to this embodiment. [Figure 15] Figure 15 is a sequence diagram showing the operation of the communication management system according to this embodiment. [Modes for carrying out the invention]
[0023] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to Figures 1 to 15.
[0024] [Configuration of the communication management system] First, with reference to Figure 1, an overview of a communication management system comprising a communication management device 1 according to an embodiment of the present invention will be described.
[0025] The communication management system according to this embodiment comprises a communication management device 1, a user terminal 2, a base station 3, a core network 4, a relay network 5, and a cloud (external network resource) 6. The communication management system complies with the 5G mobile communication network and manages uplink and downlink communication routes from the user terminal 2, through the UPF 40 (gateway) provided by the core network 4 of the mobile communication network, and through the relay network 5, which is connected to multiple relay devices 50 (relay nodes), to a specific cloud 6.
[0026] The communication management device 1 is connected to the core network 4, the relay network 5, multiple relay devices 50, and the cloud 6 via the network NW.
[0027] User terminal 2 can be implemented as a mobile communication terminal such as a smartphone, a tablet computer, or a laptop computer. User terminal 2 is equipped with a SIM card, and the SIM contract profile stores the user's subscriber identification information, including identifier information such as the subscriber identification number (IMSI: International Mobile Subscriber Identity) assigned to the mobile phone line contract, the user's telephone number (MSISDN: Mobile Subscriber International Subscriber Directory Number), and the SIM card number (ICCID: Integrated Circuit Card Identifier). User terminal 2 is uniquely identified by its IMSI.
[0028] Each user terminal 2 is also assigned a terminal IP address that uniquely identifies the terminal. The IP address is assigned to the user terminal 2 via SMF41 after the session is established. In this embodiment, there are N user terminals 2 (where N is a positive integer of 2 or more), and each user terminal 2 is located in the communication area of a different base station 3. Each user terminal 2 accesses a specific cloud 6 via the relay network 5 from a UPF40 designated according to the base station 3 and cloud 6 in their respective area. Furthermore, the user terminal 2 receives data packets from the specific cloud 6 via the relay network 5 and UPF40. In this embodiment, multiple user terminals 2 are pre-grouped based on predetermined attributes and assigned a group ID (common identification information).
[0029] Base station 3 is a wireless base station compatible with the 5G system and relays communication between user terminals 2 located within the communication area and the core network 4. Base station 3 is connected to the core network 4 via a network such as a backhaul link. N base stations 3 (where N is an integer of 2 or more) are installed.
[0030] The core network 4 is connected to the communication management device 1 via a network NW such as a LAN or WAN. The core network 4 includes multiple UPFs (User Plane Functions) 40 within the U-plane. The core network 4 also includes nodes within the C-plane, namely an AMF (Access and Mobility Management Function) 41, an SMF (Session Management Function) 42, a PCF (Policy Control Function) 43, and a UDM (Unified Data Management) / UDR (Unified Data Repository) 44. Functional nodes within the U-plane and C-plane that are included in the core network 4 other than those mentioned above are not shown in the diagram.
[0031] The UPF40 is a user plane function that processes packets between the base station 3 and data networks such as the internet. The UPF40 functions as a gateway between the core network 4 and external data networks. Multiple UPF40s are provided in the core network 4. In this embodiment, each UPF40 transmits packets from the user terminal 2 to the cloud 6 via the relay network 5. Each UPF40 also forwards packets transmitted from the cloud 6 and transmitted via the relay network 5 to the user terminal 2. The user terminal 2 communicates with the designated UPF40 depending on the base station 3 it is located in and the cloud 6 it is accessing.
[0032] Each UPF40 has an IP address, which allows for unique identification of the UPF40. In this embodiment, as shown in Figure 2, multiple UPF40s located at the end of the core network 4, i.e., the U-plane of the mobile communication network, share a single common network address. Each UPF40 is connected to a relay device 50, which is a relay node constituting the relay network 5. Specifically, each UPF40 has a connection configuration that connects to each of the first-stage relay devices 50 of the relay network 5 in a so-called full-mesh configuration.
[0033] The AMF41 is an access and mobility management device that manages the registration and wireless connection of user terminals 2 that have moved to each communication area.
[0034] SMF42 is a session management function that establishes, modifies, and releases PDU (Packet Data Unit) sessions between user terminal 2 and data networks such as the Internet. Based on the PCC (Policy and Charging Control) policy from PCF43, SMF42 sets the appropriate communication path for data communication between user terminal 2 and UPF40.
[0035] PCF43 determines QoS and policies and provides them to SMF42. PCF43 applies PCC rules according to the 3GPP (registered trademark) specification and creates PCC policies for configuring the communication path of UPF40 that user terminal 2 communicates with, in response to instructions from communication management device 1.
[0036] The UDM / UDR44 manages subscriber profiles, performs authentication, and manages mobility. In this embodiment, the UDM / UDR44 stores the group ID assigned to each IMSI in the subscriber profile. In this embodiment, the UDM / UDR44 is shown as a single device in which the UDM and UDR are configured, but the UDM / UDR44 may also be a device in which the UDM and UDR are located separately.
[0037] The relay network 5 is a group of relay nodes connected to each other. The relay network 5 is located between UPF40 and the external cloud 6 and relays packets from user terminals 2 and cloud 6. Each relay node consists of a relay device 50, and each relay device 50 is uniquely identified by its IP address. Each relay device 50 is also connected to the communication management device 1 via the network NW. The relay device 50 is equipped with a CPU, memory, communication interface (WAN / LAN), routing table, etc., and performs receiving, routing, and forwarding processing of packets from user terminals 2 and cloud 6 based on route information from the communication management device 1.
[0038] As shown in Figure 2, the relay network 5 has a hierarchical structure in which relay devices 50 with common network addresses are connected in groups. For example, in the example in Figure 2, multiple UPFs 40 located at the end of the U-plane of the core network 4 have network (NW) address 1. The first-stage group of relay devices 50 connected to the UPFs 40 have NW address 2. Furthermore, the group of relay devices 50 with NW address 2 connects to the next second-stage group of relay devices 50, and finally connects to the m-th-stage group of relay devices 50 with NW address m. Furthermore, the group of relay devices 50 with NW address m connects to the cloud 6. Each relay device 50 for each network address is connected to each of the relay devices 50 for the next network address.
[0039] For example, a packet from user terminal 2 on the uplink passes through the base station 3 located in the area, starting at UPF 40 with NW address 1, and then through one of the n relay devices 50 with the next NW address 2. Subsequently, the packet passes through one of the n relay devices 50 with the next NW address 3. In this way, for each network address, one of the relay devices 50 with the next hop network address is selected as the relay device 50 to be traversed.
[0040] The communication management system manages the relay devices 50 for each network address to which packets from user terminal 2 are forwarded, from the UPF40 group at network address 1 to network addresses 2 through m in that order, until they reach cloud 6. Similarly, the communication management system manages the relay devices 50 for each network address to which packets from cloud 6 are forwarded, from the UPF40 group at network address m to network addresses m-1 through m-2 in that order, until they reach the UPF40 with which user terminal 2 is communicating.
[0041] Cloud 6 provides a predetermined website or web application. Cloud 6 has a cloud base, which is a geographical location or area where the physical equipment constituting the website is located. Cloud 6 can be an edge server such as a server, data center, or MEC server. In this embodiment, multiple user terminals 2 access a specific Cloud 6 from their respective base stations 3 and UPF 40. Packets from Cloud 6 are transmitted to each of the multiple user terminals 2, each connected to a different base station 3 and UPF 40.
[0042] The communication management system learns a route selection strategy using reinforcement learning, which indicates which relay device 50 at the next network address a packet should pass through, based on information about the relay device 50 and other nodes the packet is currently traversing. Furthermore, using the route selection strategies for each user terminal 2 obtained through reinforcement learning as training data, the system learns the relationship between the information of the node the packet is currently traversing and the route selection strategy for the next relay device 50 the packet should pass through, using a supervised learning model.
[0043] Furthermore, the communication management system, for both the uplink and downlink, uses the information of the node the packet is currently traversing as unknown input, performs calculations on a trained supervised learning model, and outputs a strategy for selecting the route to be taken by the relay devices 50 in sequence. The system then notifies the UPF 40, relay devices 50, or cloud 6, which are currently communicating, of the route information determined based on the outputted route selection strategy. The UPF 40, relay devices 50, and cloud 6 then forward the packet to the next relay device 50 based on the notified route information. As a result, the user terminal 2 can access cloud 6 via the optimal communication route and receive data from cloud 6.
[0044] [Functional blocks of the communication management device] As shown in Figure 1, the communication management device 1 comprises a first storage unit (management information storage unit) 10A, a setting unit 10B, an acquisition unit 11, a first learning unit 12, a second learning unit 13, a second storage unit (storage unit) 14, a third storage unit 15, a calculation unit 16, a determination unit 17, and a communication route management unit 18. The communication management device 1 manages the communication route of the relay network 5 in the uplink from UPF 40 to the cloud 6 for packets from the user terminal 2, and the communication route of the relay network 5 in the downlink from the cloud 6 to the UPF 40.
[0045] The first storage unit 10A stores management information that associates the group ID of the user terminal 2, the subscriber identification information of the user terminal 2, and the identification information of a specific cloud 6. Figures 3 and 4 show table 10a of the management information stored in the first storage unit 10A. Table 10a holds the values of "Group ID," "IMSI," "Cloud IP Address," and "UPF IP Address." The group ID is an ID that is assigned as common identification information to multiple user terminals 2 that have been grouped in advance. The group ID allows for the identification of the IMSI of each user terminal 2 belonging to the group. In addition, the destination cloud 6 is specified in advance for each group ID.
[0046] In Figure 3, Table 10a shows that the UPF40 communication path for user terminal 2's data communication has not yet been configured, so the value of "UPF IP address" is "null". On the other hand, in Figure 4, Table 10a shows that the UPF40 communication path for user terminal 2's data communication has already been configured by the configuration unit 10B, so the IP address of the UPF40 currently communicating is stored as the value of "UPF IP address".
[0047] The configuration unit 10B specifies management information based on the location registration request signal from each user terminal 2 and instructs the UPF 40 settings for each user terminal 2. Specifically, the configuration unit 10B instructs the user plane function settings for data communication of each user terminal 2 based on the group ID associated with the IMSI (subscriber identification information) included in the location registration request signal. In detail, the configuration unit 10B receives the location registration request signal from the user terminal 2 via the core network 4 and refers to the table 10a shown in Figure 3. The location registration request signal includes the IMSI of the user terminal 2. The location registration request signal is transmitted to the core network 4 via the base station 3 when the user terminal 2 moves within the communication area, as a periodic location update, and when the power is turned on, etc.
[0048] The configuration unit 10B requests the creation of a PCC policy from the PCF43 of the core network 4, specifying the group ID, IMSI, and the IP address of the communication management device 1. In response to the creation request, the PCF43, SMF42, and UDM / UDR44 of the core network 4 cooperate to set the appropriate communication path for UPF40 for the data communication of the user terminal 2. Once the UPF40 communication path is set for the data communication of the user terminal 2, as explained in Figure 4, the IP addresses of the UPF40 related to the communication path setting are stored for each IMSI in the table 10a of the first storage unit 10A.
[0049] The acquisition unit 11 acquires information about the UPF40 (multiple gateways) that the user terminal 2 is communicating with, and information about the relay devices 50 (relay nodes) through which packets from the user terminal 2 pass at each time, as information about the node through which packets from the user terminal 2 are currently passing in the uplink. The node through which a packet is currently passing is a node that has received the packet but has not yet forwarded it to the next destination node.
[0050] Specifically, the acquisition unit 11 obtains the IP address of the UPF 40 or relay device 50 that sent the query for the next destination node, as information about the node currently being traversed, from the query for the next destination node sent from the UPF 40 or relay device 50 that received the packet from the user terminal 2. The acquisition unit 11 can obtain information on which relay device 50 in which group of relay devices 50, among the multiple network addresses shown in Figure 2, is being traversed.
[0051] In managing downlink communication routes, the acquisition unit 11 acquires the IP addresses of the relay devices 50 that packets destined for user terminal 2 from cloud 6 pass through at each time point, as information about the node the packet destined for user terminal 2 is currently passing through. The acquisition unit 11 also acquires the IP address of the originating relay device 50 as information about the node the packet is currently passing through, based on the query for the next destination node sent from the relay device 50 that received the packet destined for user terminal 2. In the first route selection time step when the packet is forwarded from cloud 6 to the next-hop relay device 50, the acquisition unit 11 acquires the IP address of cloud 6.
[0052] The first learning unit 12 applies a reward function to the estimated result of calculating the selection of the route that packets from user terminal 2 should sequentially traverse through relay devices 50 from UPF 40 to cloud 6 via the relay network 5, and updates the result to maximize the reward for packets from user terminal 2 to reach cloud 6. The first learning unit learns a first route selection policy for relay nodes that packets from user terminal 2 should sequentially traverse from UPF 40 using a first reinforcement learning model.
[0053] The first learning unit 12 further applies a reward function to the estimated result of calculating the selection of the route that a packet destined for user terminal 2 should sequentially traverse through the relay network 5 to reach UPF 40, and updates it so as to maximize the reward for the packet destined for user terminal 2 to reach UPF 40. It then learns a second route selection strategy for the relay devices 50 that the packet should sequentially traverse from cloud 6 using a second reinforcement learning model. In the following, the first reinforcement learning model and first route selection strategy related to uplink communication routes, and the second reinforcement learning model and second route selection strategy related to downlinks, may be collectively referred to as the reinforcement learning model and route selection strategy, respectively.
[0054] In this embodiment, as a route selection strategy, action a is taken to route the packet from the UPF40, relay device 50, or cloud 6 that the packet immediately passed through to each of the n relay devices 50 that have the network address of the next hop. n The following configuration is adopted: Packets are forwarded from one of the relay devices 50 in a group of relay devices 50 having the same network address to one of the groups of relay devices 50 having different network addresses, which are connected downstream along the direction of the link.
[0055] Furthermore, the routing strategy includes a probabilistic routing strategy from the relay device 50 having the first network address, which the packet will traverse in the first time step, to each of the n relay devices 50 having the second network address, which the packet should traverse in the next second time step. Furthermore, it includes a probabilistic routing strategy from the selected relay device 50 having the second network address to each of the group of relay devices 50 having the third network address, which the packet should traverse in the next third time step.
[0056] The first learning unit 12 uses a neural network model including an input layer s, a hidden layer h, and an output layer q as shown in Figure 5, as the first and second reinforcement learning models. Furthermore, the neural network model includes a state s, which is the IP address of the node that the packet passes through at time t. tReceive all action value functions Q(s t , a1), Q(s t , a2), Q(s t , a3), ···, Q(s t , a n-1 ), Q(s t , a n ), and adopt Deep Q - Network (DQN), a neural network that outputs them.
[0057] More specifically, the first learning unit 12 gives the IP address of the node through which the packet is currently passing as the input of the neural network model and performs the operation of the neural network model. Then, the first learning unit 12 routes (selects a path) the packet to each of the n relay device 50 groups related to the network address of the next hop as the path selection of the relay device 50 to be passed through next by the packet from the currently passing node (UPF40, relay device 50, or cloud 6), and outputs the first estimated value Q1 of the action value function representing the expected value of the cumulative value of the future reward obtained when taking the action a n .
[0058] The reward is given by the reward function r = r(s, a, s'), where s is the state which is the IP address of the node (UPF40, relay device 50, or cloud 6) through which the packet is currently passing, a n is the action of routing the packet to a specific relay device 50 among the relay device 50 groups of the next hop, and s' is the next state which is the IP address of the next relay device 50 through which the packet passes.
[0059] In this embodiment, the reward function in the first reinforcement learning model includes as a variable the degree to which packets reach the cloud 6 from the UPF 40 with which the user terminal 2 is communicating. For example, if packets from user terminal 2 reach the cloud 6 via the shortest route through an action that routes them to a specific relay device 50, the reward, which is a scalar quantity, is set to a larger value. On the other hand, if packets from user terminal 2 move away from the physical location of the cloud 6, a negative reward value (for example, r = -1) can be assigned. Similarly, the reward function in the second reinforcement learning model includes as a variable the degree to which packets reach the UPF 40 with which the user terminal 2 is communicating.
[0060] Furthermore, the first learning unit 12 receives the IP address of the node the packet is currently passing through as input to the neural network model, performs calculations on the neural network model, and outputs a second estimate Q2 of the action-value function. The first learning unit 12 learns the weight parameters of the neural network model so that the first estimate Q1 becomes the target value calculated from the second estimate Q2.
[0061] If we denote the weight parameters of the neural network model as θ and the action-value function as Q(s,a;θ), the learning minimization loss function is given by the following equation (1). L(θ) = 1 / 2{r + γmax} a’ Q(s',a';θ)-Q(s,a;θ)} 2 ...(1)
[0062] In equation (1) above, r is the reward (immediate reward) and γ is the discount rate. Q(s,a;θ) corresponds to the first estimate Q1, and Q(s',a';θ) corresponds to the value of the action in state s', one step ahead, i.e., the second estimate Q2. The target value is r + γmax a’ It can be represented by Q(s',a';θ).
[0063] The first learning unit 12 can update the weight parameters of the neural network model by backpropagating the gradient of the loss function given by equation (1) above.
[0064] More specifically, the first learning unit 12 can employ a Fixed Target Q-Network using two neural networks, main QN121 and target QN123, as shown in Figure 6. Main QN121 selects the optimal action and updates the action-value function Q. Meanwhile, target QN123 estimates and evaluates the value of the action a' to be taken in the next state s' resulting from the action. Main QN121 and target QN123 have neural networks with the same layer structure, but the parameter of main QN121 is "θ" and the parameter of target QN123 is "θ". - It is given by ".
[0065] Main QN121 receives the IP address of the node the packet is currently passing through as state s from environment 120. Environment 120 is the mobile communication network system where user terminal 2 is located. Under this environment 120, by taking action a, which routes the packet to one of the relay devices 50 at each network address, the packet is forwarded to one of the relay devices 50 at the next network address, transitioning to the next state s' and simultaneously receiving a reward r from environment 120.
[0066] The first learning unit 12 inputs the state s related to the IP address of the node the packet is currently passing through to the main QN121 and calculates the action-value function Q(s,a;θ). The first learning unit 12 calculates the action a using, for example, the ε-greedy method, or the optimal action at the present moment using argmax. a We will find Q(s,a;θ). In environment 120, the argmax function is used to determine the optimal path selection at this time. a Perform Q(s,a;θ). Environment 120 performs the action of routing packets. aAs a result of Q(s,a;θ), the IP address of the routing destination relay device 50 is observed as the next state s', and the reward r is output. Experience data 124 stores the experience (s,a,r,s') output from environment 120.
[0067] The first learning unit 12 calculates the loss function L in the DQN loss calculation 122 and updates the weights of the main QN 121 using the gradient of the loss function L.
[0068] The first learning unit 12 periodically copies the weights of the main QN121 to the target QN123 and synchronizes them. The synchronization of the target QN123 is performed at a lower frequency than the update frequency of the weights of the main QN121. The first learning unit 12 extracts experience from the experience data 124, inputs the past state into the target QN123, and estimates the max value. a’ Q(s',a';θ - The first learning unit 12 outputs the estimated value max output by target QN123. a’ Q(s',a';θ - ) Target value r+γmax a’ Q(s',a';θ - Using this method, the weights of the main QN121 are trained using DQN loss calculation 122.
[0069] The first reinforcement learning model obtained by the first learning unit 12 through learning is a trained first reinforcement learning model for uplink, consisting of N trained first reinforcement learning models constructed for each of the N user terminals 2 that connect from multiple UPF 40s with different initial locations to a specific cloud 6 with the same final location. Similarly, a trained second reinforcement learning model for downlink is constructed for each of the N user terminals 2. These N trained second reinforcement learning models indicate the optimal route selection strategy for communication routes from the same initial location, cloud 6, to multiple UPF 40s with different final locations. These trained first and second reinforcement learning models are used as training data for supervised learning by the second learning unit 13.
[0070] The action-value function Q(s) is a path selection policy. t a1), Q(s t a2), Q(s t ,a3),...,Q(s t ,a n-1 ), Q(s t ,a n ) in action a1~a n This represents the value of the cumulative reward expected when routing is selected to each of the n relay devices 50 having the same network address. As shown in Figure 2, uplink packets are routed from UPF40 (IP address: IP1) at NW address 1 to one of the n relay devices 50 (IP addresses: IP21~IP2n) at NW address 2. For example, a1=0.1, a2=0.1, a3=0.1, ..., a5=0.6, ..., a n If the probability is 0.2, the IP26 relay device 50, which has the highest probability value, will be selected.
[0071] The second learning unit 13 learns the relationship between the node through which packets from user terminal 2 are currently traversing and the first route selection strategy for relay nodes that packets should sequentially traverse from UPF 40, obtained through learning by the first learning unit 12, using the first supervised learning model. In the case of downlink, the second learning unit 13 learns the relationship between the node through which packets destined for user terminal 2 are currently traversing and the second route selection strategy for relay devices 50 that packets destined for user terminal 2 should sequentially traverse from cloud 6, obtained through learning by the first learning unit 12, using the second supervised learning model. In the following, the first supervised learning model for uplink and the second supervised learning model for downlink may be collectively referred to as the supervised learning model.
[0072] Figure 7 shows the structure of a neural network model adopted as an example of the first supervised learning model and the second supervised learning model used by the second learning unit 13. The neural network model comprises an input layer x, a hidden layer h, and an output layer y. The second learning unit 13 provides the IP address of the node (UPF 40, relay device 50, or cloud 6) that the packet is currently passing through, and the IP address of the relay device 50 that the packet will pass through at each subsequent time t, to the input layer of the neural network model, applies an activation function to the weighted sum of the inputs, and passes the output determined by thresholding to the output layer. Each output node of the output layer outputs the model's predicted output corresponding to n action-value functions Q that route the packet to each of the n relay devices 50 that have a common network address.
[0073] The second learning unit 13 introduces the objective function E shown in equation (2) below, and learns the parameters of the neural network model so that the route selection policy for the relay devices 50 that the packets should sequentially pass through, which is a predicted value from the neural network model for the IP addresses of the nodes that the user terminal 2 packets are currently passing through, becomes the value of the optimized route selection policy for the relay devices 50 that the packets should sequentially pass through, which has been reinforced and learned by the first learning unit 12.
[0074]
number
[0075] In equation (2) above, y1, y2, ..., y n The predicted output values for each output node are shown. Also, Y1, Y2, ..., Y n is the correct label included in the training data, and here it is the n optimized action-value functions Q for the n relay devices 50 that the packet should pass through next, for the IP address of the node the packet is currently passing through, obtained by reinforcement learning by the first learning unit 12. Furthermore, as mentioned above, the first learning unit 12 performs reinforcement learning for each of the N user terminals 2.
[0076] In Uplink, N pre-trained first reinforcement learning models (first path selection policies) are constructed, ranging from N different initial locations UPF40 to a specific cloud 6. Each of these N first reinforcement learning models is further optimized for n action-value functions Q(s t a1), Q(s t a2), Q(s t ,a3),...,Q(s t ,a n-1 ), Q(s t ,a n ) is provided. In the downlink, N pre-trained second reinforcement learning models (second path selection policies) are constructed from Cloud 6, which corresponds to the same initial location, to UPF40, which corresponds to N different final locations. Similarly, n more action-value functions Q are provided for each of the N pre-trained second reinforcement learning models. The second learning unit 13 uses all of these action-value functions Q as training data for supervised learning in both the uplink and downlink.
[0077] In the example neural network model configuration shown in Figure 7, the training data is displayed to the right of each output node that outputs the model's predicted output value. Starting from the top output node, the training data Y1 to Y n n action-value functions Q are shown. The second learning unit 13 adjusts the weight parameters of the neural network related to the supervised learning model so that the objective function E is minimized, i.e., becomes 0. The second learning unit 13 can optimize the objective function E using methods such as backpropagation.
[0078] The second learning unit 13 can learn the first supervised learning model related to the uplink communication route using a single neural network model. Furthermore, for learning the second supervised learning model related to the downlink, the second learning unit 13 can learn using a single neural network model by providing the input node with the IP address of the final destination UPF40 along with the IP address of the node the packet is currently passing through.
[0079] The second memory unit 14 stores the trained first reinforcement learning model and the trained second reinforcement learning model constructed by reinforcement learning by the first learning unit 12. In this embodiment, the second memory unit 14 stores N trained first reinforcement learning models and N trained second reinforcement learning models.
[0080] The third memory unit 15 stores the trained first supervised learning model and the trained second supervised learning model constructed by the second learning unit 13 through supervised learning. In this embodiment, the third memory unit 15 stores one trained first supervised learning model and one trained second supervised learning model.
[0081] The calculation unit 16 provides information about the nodes that the packet from user terminal (managed user terminal) 2 is currently passing through as unknown input to a pre-trained first supervised learning model, performs calculations on the pre-trained first supervised learning model, and outputs a first uplink route selection strategy for the relay nodes that the packet from user terminal 2 should sequentially pass through starting from UPF 40. The managed user terminal 2 refers to the user terminal 2 that is the target of inference.
[0082] Furthermore, the arithmetic unit 16 provides information about the nodes that the packet destined for user terminal 2 is currently passing through as unknown input to a pre-trained second supervised learning model, performs calculations on the pre-trained second supervised learning model, and outputs a second downlink route selection strategy for the relay nodes that the packet destined for user terminal 2 should sequentially pass through from cloud 6.
[0083] The information about the node through which packets from the managed user terminal 2 are currently traversed, and the information about the node through which packets destined for the managed user terminal 2 are currently traversed, are the IP addresses of the UPF 40, relay device 50, or cloud 6 acquired by the acquisition unit 11.
[0084] The decision unit 17 determines the next relay device 50 that the packet should pass through, based on the first route selection strategy related to the uplink output by the calculation unit 16. More specifically, based on the optimal first route selection strategy output by the calculation unit 16, it determines the IP address of the UPF 40 or relay device 50 that the packet is currently communicating with, in state s t For each state s t The relay device 50 to be passed through sequentially is determined by selecting the relay device 50 related to action a, for which the value of the action-value function Q takes its maximum value.
[0085] The decision unit 17 further determines, based on the second route selection strategy related to the downlink, which relay device 50 the packet destined for the managed user terminal 2 should pass through next, from the node (cloud 6 or relay device 50) that the packet is currently passing through.
[0086] The communication route management unit 18 notifies the UPF 40 and the relay device 50 of the route information determined by the decision unit 17 based on the first uplink route selection strategy output by the calculation unit 16. The communication route management unit 18 also notifies the cloud 6 and the relay device 50 of the route information determined by the decision unit 17 based on the second downlink route selection strategy output by the calculation unit 16.
[0087] More specifically, the communication route management unit 18 notifies the UPF 40, relay device 50, or cloud 6, through the network NW, of route information indicating the next destination relay device 50. Upon receiving the notification, the UPF 40, relay device 50, or cloud 6 updates its routing table and forwards the received packet to the next-hop relay device 50.
[0088] [Hardware configuration of the communication management device] Next, an example of a hardware configuration for realizing the communication management device 1 having the functions described above will be explained using Figure 8.
[0089] As shown in Figure 8, the communication management device 1 can be implemented, for example, by a computer equipped with a processor 102, main memory 103, communication interface 104, auxiliary storage 105, and input / output I / O 106 connected via a bus 101, and a program to control these hardware resources. Furthermore, the communication management device 1 may include a display device 107 connected via the bus 101.
[0090] Processor 102 is implemented using CPUs, GPUs, FPGAs, ASICs, etc.
[0091] The main memory 103 contains pre-stored programs for the processor 102 to perform various controls and calculations. The processor 102 and the main memory 103 work together to realize the various functions of the communication management device 1, including the setting unit 10B, acquisition unit 11, first learning unit 12, second learning unit 13, calculation unit 16, determination unit 17, and communication route management unit 18 shown in Figure 1.
[0092] The communication interface 104 is an interface circuit for networking the communication management device 1 with various external electronic devices.
[0093] The auxiliary storage device 105 consists of a read / write storage medium and a drive device for reading and writing various information such as programs and data to the storage medium. The auxiliary storage device 105 can use semiconductor memory such as a hard disk or flash memory as the storage medium.
[0094] The auxiliary storage device 105 has a program storage area for storing the communication management program executed by the communication management device 1. It also has a program storage area for storing the reinforcement learning program executed by the communication management device 1. Furthermore, the auxiliary storage device 105 has an area for storing the supervised learning program. The auxiliary storage device 105 realizes the first storage unit 10A, the second storage unit 14, and the third storage unit 15 described in Figure 1. Furthermore, it may have, for example, a backup area for backing up the above-mentioned data and programs.
[0095] The I / O106 is an input / output device that accepts signals from external devices and outputs signals to external devices.
[0096] The display device 107 is composed of an organic EL display, a liquid crystal display, or the like. The display device 107 can display information related to the management of the communication route of the user terminal 2.
[0097] [Operation of the communication management system] Next, the operation of the communication management system equipped with the communication management device 1 having the above-described configuration will be explained with reference to the sequence shown in Figure 9.
[0098] Figure 9 shows a sequence of steps by the communication management system to configure the communication path for data communication between user terminal 2 and UPF40.
[0099] First, Cloud 6 transmits information to the communication management device 1 that associates the group ID, IMSI, and the IP address of Cloud 6 (step S100). Information regarding the grouping of user terminals 2 is stored in Cloud 6 in advance. Next, the first storage unit 10A of the communication management device 1 creates a table 10a (Figure 3) based on the received group ID, IMSI, and the IP address of Cloud 6 (step S101). Subsequently, the setting unit 10B of the communication management device 1 sends an instruction to the UDM / UDR44 to provision the group ID using the IMSI as the key (step S102).
[0100] Subsequently, when user terminal 2 starts data communication with cloud 6, it sends a location registration request signal to core network 4 (step S103). The location registration request signal includes the IMSI of user terminal 2. When UDM / UDR44 receives the location registration request signal, UDM / UDR44 associates a group ID with the IMSI included in the received location registration request signal from the information in table 10a (Figure 3) created in step S101, and forwards the location registration request signal to communication management device 1 (step S104).
[0101] Next, the setting unit 10B of the communication management device 1 requests the PCF43 to create a PCC policy by specifying the IMSI, group ID, and its own IP address included in the location registration request signal received in step S104 (step S105). Subsequently, the PCF42 creates a PCC policy based on the specified requirements (step S106). The PCF43 creates a PCC policy that includes information on which UPF40 the data communication from the user terminal 2 will pass through. The PCC policy specifies the optimal UPF40 as the communication path for data communication from the user terminal 2 to the cloud 6.
[0102] Next, PCF43 sends the created PCC policy to SMF42 (step S107). Then, SMF42 sends configuration information related to user plane functions, such as the setting of communication paths defined in the PCC policy, to UPF40, which is specified by the PCC policy (step S108). Next, UPF40 registers the received configuration information in memory (step S109).
[0103] Next, UPF40 sends an ACK to SMF42 to notify it that the communication path settings have been applied (step S110). Furthermore, SMF42 sends an ACK to communication management device 1 to notify it that the communication path settings have been applied (step S111). Subsequently, communication management device 1 sends an ACK to user terminal 2 (step S112). After that, a data communication path is established between user terminal 2 and UPF40 (step S113).
[0104] Next, the communication management device 1 performs communication management processing to manage the communication route of the relay network 5 in the uplink from the user terminal 2 to access the cloud 6 from UPF 40, and the communication route of the relay network 5 in the downlink from the user terminal 2 to receive data from the cloud 6 via UPF 40 (step S114). After that, based on the route information from the communication management device 1, data transfer processing is performed between UPF 40 and the cloud 6 using the optimal route (step S115).
[0105] [Operation of the communication management device] Next, with reference to Figures 10 and 11, the communication management processes performed by the communication management device 1 will be described in detail. Figure 10 is a sequence showing the learning process of the communication route from the uplink packet to the cloud 6. Figure 11 is a sequence showing the optimal route selection process for the uplink packet using the first supervised learning model that has been trained. Below, we will describe the case where each user terminal 2 accesses a specific cloud 6 from a different UPF 40.
[0106] First, in Figure 10, when the UPF 40, which user terminal 2 is currently communicating with, notifies the communication management device 1 of the group ID of user terminal 2, the IMSI, and the IP address of the UPF (self-device) 40, the following processing is executed in the communication management device 1.
[0107] First, the first storage unit 10A updates table 10a (Figure 3) by associating the IP address of UPF40 received from UPF40 with the IMSI in table 10a (Figure 4) (step S1). Next, the acquisition unit 11 acquires the IP address of the node the packet is currently passing through (step S2). The acquisition unit 11 acquires the IP address of UPF40 held in table 10a updated in step S1 as the transit node for the first time. Furthermore, for each subsequent time, the acquisition unit 11 acquires the IP address of the relay device 50 that received the packet and sent the query for the next transit node.
[0108] Next, the first learning unit 12 applies a reward function to the estimated result of calculating the route selection for the relay devices 50 that packets from user terminal 2 should sequentially traverse from UPF 40 to cloud 6 via the relay network 5. The unit updates the result to maximize the reward for packets from user terminal 2 to reach cloud 6, and learns a first route selection policy for the relay nodes that packets should sequentially traverse from UPF 40 using the first reinforcement learning model (first learning process) (step S3). Step S3 is performed for each of the grouped user terminals 2. Details of the first learning process will be described later.
[0109] Subsequently, the second memory unit 14 stores the first reinforcement learning model obtained in step S3 (step S4). Next, the second learning unit 13 learns the relationship between the node that the user terminal 2's packet is currently traversing and the first route selection strategy for the relay nodes that the packet should sequentially traverse from UPF 40, obtained through learning by the first learning unit 12, using a supervised learning model (second learning process) (step S5).
[0110] Specifically, the second learning unit 13 repeatedly adjusts and updates parameters such as weights and thresholds to minimize the error between the predicted output value of the first route selection policy for the relay devices 50 that packets should sequentially traverse, when the IP address of the node that a packet from the user terminal 2 is currently passing through, i.e., the current state, is given to the first supervised learning model as input, and the training data, thereby determining the values of these parameters.
[0111] In step S5, the second learning unit 13 can determine the parameters that minimize the objective function E using methods such as backpropagation. At the initial time, the packet passes through UPF40, so the IP address of UPF40 is input to the first supervised learning model. Furthermore, the IP addresses of the relay devices 50 that the packet passes through at each subsequent time point are input to the first supervised learning model, and supervised learning is performed.
[0112] Next, the third memory unit 15 stores the first supervised learning model that was built in step S5 (step S6). With these processes, the learning of the uplink communication route is completed.
[0113] Next, referring to the sequence in Figure 11, we will explain the process by which the communication management device 1 selects the optimal route for uplink packets using the first supervised learning model that has been trained.
[0114] First, user terminal 2 sends a packet to UPF40 (step S200). The packet from user terminal 2 includes the IMSI, the IP address of the source user terminal 2, and the IP address of the destination UPF40. Subsequently, UPF40, having received the packet, sends the group ID of user terminal 2, the IMSI, and the IP address of UPF40 to the communication management device 1 (step S201). Next, the first storage unit 10A of the communication management device 1 stores and updates the IP address of UPF40 received in step S201, associating it with the IMSI in table 10a (step S202).
[0115] Next, the acquisition unit 11 of the communication management device 1 acquires the IP address of the node the packet is currently passing through (step S203). The acquisition unit 11 acquires the IP address of UPF40, which is held in the table 10a updated in step S202, as the traversal node at the first time.
[0116] Subsequently, the arithmetic unit 16 provides the information of the nodes that the packet from user terminal 2 is currently passing through, obtained in step S203, as an unknown input to the trained first supervised learning model, performs calculations on the trained first supervised learning model, and outputs a first route selection strategy for the relay nodes that the packet from user terminal 2 should sequentially pass through from UPF 40 (step S204).
[0117] Next, the decision unit 17 determines the relay devices 50 that the packets should sequentially pass through by selecting the relay device 50 that takes the action a with the maximum value of the n action-value functions Q output in step S204 (step S205). Next, the communication route management unit 18 notifies the UPF 40 of the route information determined in step S205 (step S206). Specifically, the communication route management unit 18 sends the IP address (IP 21) of the next-hop relay device 50 as route information to the UPF 40.
[0118] Upon receiving notification of routing information, UPF40 forwards the packet to the designated destination relay device 50 (step S207). The packet to be forwarded includes the group ID of user terminal 2, IMSI, the IP address of the source user terminal 2, and the IP address (IP 21) of the destination relay device 50. Next, relay device 50 (IP 21), having received the packet from UPF40, sends the group ID and IMSI of user terminal 2 to communication management device 1 to inquire about the next transit node (step S208). After that, the process proceeds to step S203.
[0119] The acquisition unit 11 of the communication management device 1 acquires the IP address (IP 21) of the relay device 50 that sent the query including the group ID and IMSI in step S208 as the IP address of the node the packet is currently passing through (step S203). Furthermore, after the calculation processing in step S204 and the route determination processing in step S205, the communication route management unit 18 notifies the relay device 50 (IP 21) of the determined route information, which is the IP address (IP 31) of the next hop relay device 50 (step S209).
[0120] Subsequently, relay device 50 (IP 21) forwards the packet to the next relay device 50 (IP 31) (step S210). The packet contains the group ID of user terminal 2, IMSI, the IP address of the source user terminal 2, and the IP address (IP 31) of the destination relay device 50. In this way, packets from user terminal 2 are sequentially forwarded to the next-hop relay device 50 for each network address.
[0121] Subsequently, when the final relay device 50 (IP N1) receives the packet, it sends the group ID and IMSI of user terminal 2 to the communication management device 1 to inquire about the next traversal node (step S211). Then, the processes from steps S203 to S205 are executed, and the IP address of cloud 6 is notified to the relay device 50 (IP N1) as routing information (step S212). Then, the relay device 50 (IP N1) forwards the packet to cloud 6 (step S213). With these processes, the management of the uplink communication route is completed.
[0122] Next, the first learning process by the communication management device 1 (step S3 in Figure 10) will be explained using the flowcharts in Figures 12 and 13. First, as shown in Figure 12, the aforementioned step S2 (Figure 10) is executed. After that, the first learning unit 12 provides the IP address of the node that the packet from user terminal 2 is currently passing through, which is the current state obtained in step S2, as input to the neural network model, performs calculations on the neural network model, and outputs a first estimated value Q1 of the action value function, which represents the expected value of the cumulative value of future rewards obtained when taking each action of routing the packet from the node that the packet from user terminal 2 is currently passing through to each of the n relay devices 50 that the packet should next pass through (step S20).
[0123] Next, the acquisition unit 10 acquires the IP address of the relay device 50 that the packet passed through at the next time t as the next state s' (step S21). The IP address of the next relay device 50 that the packet passed through is based on the information acquired by the acquisition unit 11 at each time step. Furthermore, the first learning unit 12 provides the IP address of the next relay device 50 that the packet passed through, acquired in step S21, as input to the neural network model, performs calculations on the neural network model, and outputs the second estimated value Q2 of the action-value function (step S22).
[0124] Next, the first learning unit 12 calculates the target value from the second estimated value Q2 (step S23). Subsequently, the first learning unit 12 learns the weight parameters of the neural network model so that the first estimated value Q1 becomes the target value calculated from the second estimated value Q2 (step S24). Specifically, the first learning unit 12 updates the weight parameters of the neural network model to minimize the loss function in equation (1) above.
[0125] Subsequently, the process from step S3 to step S24 is repeated until a reinforcement learning model has been trained for all of the user terminals 2 (step S25: NO). After that, if training has been performed for all of the user terminals 2 (step S25: YES), the second memory unit 14 stores the trained first reinforcement learning model obtained in step S24 (step S6).
[0126] Next, referring to Figure 13, we will explain the first learning process performed by the first learning unit 12 when a Fixed Target Q-Network is adopted, which uses two neural networks: main QN121 and target QN123.
[0127] The processing in step S2 is the same as the processing in the first learning process described in Figure 10. Subsequently, the first learning unit 12 provides the main QN 121 with the IP addresses of the nodes that the packet from the user terminal 2 is currently passing through, which were obtained in step S2, as input, performs calculations on the neural network model, outputs the action-value function Q, and calculates the route selection a of the relay device 50 that the packet should next pass through (step S120).
[0128] Next, the first learning unit 12 returns the action to the environment 120 based on the route selection a obtained in step S120, and obtains the next state s', which is the IP address of the relay device 50 to which the packet was forwarded and the reward r (step S121).
[0129] The first learning unit 12 saves the experience (s, a, r, a') obtained in step S121 to the experience data 124 (step S122). Next, in the DQN loss calculation 122, the first learning unit 12 calculates the loss function L and updates the weights of the main QN 121 using the gradient of the loss function L (step S123). The first learning unit 12 repeats the process from step S120 to step S123 a set number of times.
[0130] Subsequently, the first learning unit 12 periodically copies the weights of the main QN121 to the target QN123 and synchronizes them (step S124). The synchronization of the target QN123 is performed at a lower frequency than the update frequency of the weights of the main QN121. Next, the first learning unit 12 extracts experience from the experience data 124, inputs the past state into the target QN123, and estimates the max a’ Q(s',a';θ - Output (step S126).
[0131] Next, the first learning unit 12 processes the estimated value max output by the target QN123. a’ Q(s',a';θ - ) Target value r+γmax a’ Q(s',a';θ -The first learning unit 12 calculates the target value (step S127). Next, the first learning unit 12 calculates the loss function L using the DQN loss calculation 122 with the target value calculated in step S127 (step S128). Next, the first learning unit 12 learns the weights of the main QN 121 to minimize the loss given by the loss function L (step S129). After that, the second storage unit 14 stores the learned first reinforcement learning models for all of the user terminals 2 (step S6).
[0132] Here, referring to the sequences in Figures 14 and 15, we will explain the management of downlink communication routes by the communication management device 1. Note that the following explanation will focus on processes that differ from the operation of the communication management device 1 in the uplink communication route management described in Figures 10 and 11.
[0133] Figure 14 is a flowchart showing the learning process for the downlink communication route. First, when Cloud 6 sends the group ID of user terminal 2, IMSI, and the IP address of Cloud 6 to the communication management device 1, the following processes are executed. First, the configuration unit 10B identifies the IP address of the UPF40 that user terminal 2 is communicating with based on the information received from Cloud 6 (step S30). Specifically, it identifies the IP address of the UPF40 by referring to table 10a (Figure 4).
[0134] Next, the acquisition unit 11 acquires the IP address of the node the packet is currently passing through (step S31). The acquisition unit 11 acquires the group ID of the user terminal 2 and the IP address of the cloud 6 that sent the IMSI as the IP address of the node at the initial time, and then acquires the IP address of the node at each subsequent time from the relay device 50 (step S31).
[0135] Next, the first learning unit 12 performs the first learning process (step S32). After that, the second storage unit 14 stores the learned second reinforcement learning model (step S33). Next, the second learning unit 13 performs the second learning process (step S34). Subsequently, the third storage unit 15 stores the learned second supervised learning model (step S35). The first learning process in step S32 can be performed as reinforcement learning using DQN and Fixed Target Q-Network, similar to the learning process for the uplink communication route described in Figures 12 and 13.
[0136] Here, Figure 15 shows a sequence illustrating the route selection process for downlink communication routes using a pre-trained second supervised learning model. As shown in Figure 15, Cloud 6 transmits the group ID, IMSI, and IP address of Cloud 6 of the managed user terminal 2 to the communication management device 1 (step S300). Subsequently, the setting unit 10B of the communication management device 1 identifies the IP address of the UPF 40 that the user terminal 2 is communicating with based on the information received in step S300 (step S301).
[0137] Next, the acquisition unit 11 acquires the IP address of the node through which the packet destined for the managed user terminal 2 is currently passing (step S302). Then, the calculation unit 16 performs calculations on the learned second supervised learning model (step S303). Next, the decision unit 17 determines the relay device 50 to which the packet will be forwarded based on the second route selection strategy obtained in step S303 (step S304). Subsequently, the communication route management unit 18 notifies the cloud 6, which is the current node of the packet, of the route information indicating the destination relay device 50 determined in step S304 (step S305). The route information includes the IP address (IP N1) of the next-hop relay device 50.
[0138] Cloud 6, having received the routing information, forwards the packet to the designated destination relay device 50 (step S306). The packet contains the user terminal 2's group ID, IMSI, the source cloud 6's IP address, and the destination relay device 50's IP address (IP N1). The relay device 50 (IP N1), having received the packet, then sends the user terminal 2's group ID and IMSI to the communication management device 1 to inquire about the next hop node (step S307). Subsequently, steps S302 to S304 are repeatedly executed. After that, the communication route management unit 18 of the communication management device 1 notifies the relay device 50 (IP N1) of the routing information, which is the IP address of the next hop relay device 50 (step S308). In this way, the packets are sequentially forwarded to the next hop relay device 50.
[0139] Subsequently, relay device 50 (IP 31) sends the packet to the next-hop relay device 50 (IP 21) according to the routing information (step S309). The packet includes the group ID of user terminal 2, IMSI, the IP address of the source cloud 6, and the IP address of the destination relay device 50 (IP 21). Upon receiving the packet, relay device 50 (IP 21) sends the group ID and IMSI of user terminal 2 to the communication management device 1 to inquire about the next traversal node (step S310). After that, the process moves from step S302 to step S304.
[0140] Next, the communication route management unit 18 notifies the relay device 50 (IP 21) of the IP address of the next hop UPF40, which is routing information (step S311). Then, the relay device 50 (IP 21) forwards the packet to UPF40 (step S310). After that, UPF40 sends the packet to user terminal 2 (step S313). Through the above process, the downlink packet reaches user terminal 2 from cloud 6.
[0141] As described above, the communication management device 1 according to this embodiment learns the optimal route selection strategy for the relay devices 50 that uplink and downlink packets should sequentially traverse through by reinforcement learning. Using the route selection strategy obtained through reinforcement learning as training data, the relationship between the IP address of the node that the packet from the user terminal 2 is currently traversing and the route selection strategy for the relay devices 50 that the packet should sequentially traverse is learned using a supervised learning model. Therefore, it is possible to provide the optimal communication route for the relay network between the gateway of the mobile communication network and a specific external network resource.
[0142] Furthermore, according to the communication management device 1 of this embodiment, when multiple user terminals 2, each assigned a group ID, transmit packets from different UPF 40s to a specific cloud 6, the optimal route selection strategy for the relay network 5 is learned through reinforcement learning. In addition, all route selection strategies from different UPF 40s to the specific cloud 6 are used as training data, and the relationship between the information of the node that the packets from the user terminals 2 are currently passing through and the route selection strategy is learned through supervised learning. As a result, multiple user terminals 2 can be grouped together, and the optimal communication route can be provided on a group basis.
[0143] In the embodiment described, the reinforcement learning model used by the first learning unit 12 is exemplified as a DQN related to a Fixed Target Q-Network composed of a multilayer neural network. However, other reinforcement learning models such as CNNs and multilayer perceptrons can be used. In addition to the DQN exemplified as a reinforcement learning model, Double DQN, Dueling DQN, Actor-Critic (AC) method, Soft Actor-Critic (SAC), Deep Deterministic Policy Gradient (DDPG), Q-learning, etc., can also be used.
[0144] Furthermore, in the embodiment described, the supervised learning model used by the second learning unit 13 was exemplified as a multilayer neural network. However, the supervised learning model can also be a multilayer perceptron, a decision tree-based model such as a random forest, or a support vector machine.
[0145] Although embodiments of the communication management device and communication management method of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications that a person skilled in the art can envision are possible within the scope of the invention described in the claims. [Explanation of Symbols]
[0146] 1...Communication management device, 10A...First memory unit, 10B...Setting unit, 11...Acquisition unit, 12...First learning unit, 13...Second learning unit, 14...Second memory unit, 15...Third memory unit, 16...Calculation unit, 17...Decision unit, 18...Communication route management unit, 2...User terminal, 101, 201...Bus, 102...Processor, 103...Main memory, 104...Communication interface, 105...Auxiliary memory, 106...Input / output I / O, 107...Display device, 120...Environment, 121...Main QN, 122...DQN loss calculation, 123...Target QN, 124...Experience data, NW...Network.
Claims
1. A communication management device that manages communication routes from a user terminal to a specific external network resource via a relay network consisting of multiple relay nodes connected to a gateway provided in the core network of a mobile communication network, An acquisition unit is configured to acquire information about the gateway through which the user terminal is communicating, and information about the relay nodes through which packets from the user terminal pass at each time, as information about the node through which packets from the user terminal are currently passing, among the multiple gateways provided by the core network. A first learning unit is configured to apply a reward function to an estimated result of calculating the route selection for relay nodes that a packet from the user terminal should sequentially traverse from the gateway through the relay network to reach the specific external network resource, and to update the result so as to maximize the reward for the packet from the user terminal to reach the specific external network resource, and to learn a first route selection strategy for relay nodes that a packet from the user terminal should sequentially traverse from the gateway using a first reinforcement learning model, A second learning unit is configured to learn, using a first supervised learning model, the relationship between the node through which packets from the user terminal are currently passing and the first route selection strategy for relay nodes that packets from the user terminal should sequentially pass from the gateway, which is obtained through learning by the first learning unit. A storage unit configured to store the first supervised learning model that has been trained by the second learning unit, and A communication management device equipped with the following features.
2. In the communication management device according to claim 1, The acquisition unit acquires information about the gateway through which the managed user terminal is communicating, and information about the relay nodes through which packets from the managed user terminal pass at each time, as information about the node through which packets from the managed user terminal are currently passing. Furthermore, the system includes a calculation unit configured to provide the information of the nodes that packets from the managed user terminal are currently passing through as unknown input to the trained first supervised learning model, perform calculations on the trained first supervised learning model, and output the first route selection strategy for the relay nodes that packets from the managed user terminal should sequentially pass through from the gateway, A communication route management unit configured to notify the gateway and the plurality of relay nodes of the route information determined based on the first route selection strategy output by the calculation unit, A communication management device equipped with the following features.
3. In the communication management device according to claim 2, The acquisition unit acquires information on the relay nodes through which packets destined for the user terminal from the specific external network resource pass at each time, as information on the node through which the packets destined for the user terminal are currently passing. The first learning unit applies a reward function to the estimated result of calculating the route selection for the relay nodes that a packet destined for the user terminal should sequentially traverse to reach the gateway via the relay network, and updates the result to maximize the reward for the packet destined for the user terminal to reach the gateway. The second learning unit learns a second route selection strategy for the relay nodes that a packet destined for the user terminal should sequentially traverse from the specific external network resource, using a second reinforcement learning model. The second learning unit learns, using a second supervised learning model, the relationship between the node through which the packet destined for the user terminal is currently traversing and the second route selection strategy obtained by the first learning unit regarding the relay nodes through which the packet destined for the user terminal should sequentially traverse from the specific external network resource. The memory unit stores the trained second supervised learning model constructed by the second learning unit. A communication management device characterized by the following features.
4. In the communication management device according to claim 3, The acquisition unit acquires information on the relay nodes through which packets destined for the managed user terminal from the specific external network resource pass at each time, as information on the node through which the packets destined for the managed user terminal are currently passing. The calculation unit provides the information of the nodes through which packets destined for the managed user terminal are currently passing as unknown input to the trained second supervised learning model, performs calculations on the trained second supervised learning model, and outputs the second route selection strategy for the relay nodes that packets destined for the managed user terminal should sequentially pass through from the specific external network resource. The communication route management unit notifies the specific external network resource and the multiple relay nodes of the route information determined based on the second route selection strategy output by the calculation unit. A communication management device characterized by the following features.
5. In the communication management device according to claim 1, The user terminals are a group of user terminals to which common identification information is assigned, and the gateways through which each user terminal communicates include different gateways from each other. The acquisition unit acquires information about the node through which packets from each user terminal are currently passing. The first learning unit learns the first route selection strategy for packets from each user terminal using the first reinforcement learning model, The second learning unit learns, using the first supervised learning model, the relationship between the node through which packets from each user terminal are currently traversing and the first route selection strategy for relay nodes through which packets from each user terminal should sequentially traverse from the gateway with which each user terminal is communicating, as learned by the first learning unit. A communication management device characterized by the following features.
6. In the communication management device according to claim 5, The aforementioned gateway is a user plane function, Furthermore, the system includes a management information storage unit configured to store management information that associates the common identification information, the subscriber identification information of the multiple user terminals, and the identification information of the specific external network resource. A setting unit configured to specify the management information based on the location registration request signals from each user terminal and to instruct the setting of the user plane functions of each user terminal. Equipped with, The management information storage unit further stores identification information of the user plane functions of each user terminal, which is set according to the instructions of the setting unit, in association with the management information. The identification information of the user plane function stored in the management information storage unit indicates the information of the gateway with which each user terminal is communicating. A communication management device characterized by the following features.
7. In the communication management device according to claim 3, The relay network has a hierarchical structure in which each group of relay nodes having a common network address is connected. Each of the first and second routing strategies includes a probabilistic routing strategy from any of the relay nodes in the group of relay nodes having a first network address, through which a packet will pass at the first time step, to each of the relay nodes in the group of relay nodes having a second network address, through which the packet will pass at the next second time step. A communication management device characterized by the following features.
8. A communication management method that manages a communication route from a user terminal to a specific external network resource via a relay network consisting of multiple relay nodes connected to a gateway provided in the core network of a mobile communication network, An acquisition step of acquiring information about the gateway through which the user terminal is communicating, and information about the relay nodes through which packets from the user terminal pass at each time, from among the multiple gateways provided in the core network, as information about the node through which packets from the user terminal are currently passing; A first learning step involves applying a reward function to an estimated result of calculating the route selection for relay nodes that a packet from the user terminal should sequentially traverse from the gateway through the relay network to reach the specific external network resource, updating the result so as to maximize the reward for the packet from the user terminal to reach the specific external network resource, and learning a first route selection strategy for relay nodes that a packet from the user terminal should sequentially traverse from the gateway using a first reinforcement learning model. A second learning step involves learning, using a first supervised learning model, the relationship between the node through which the packet from the user terminal is currently traversing and the first route selection strategy for the relay nodes through which the packet from the user terminal should sequentially traverse from the gateway, as obtained through learning in the first learning step. A memory step in which the trained first supervised learning model constructed in the second learning step is stored in the memory unit. A communication management method comprising the following features.
9. In the communication management method described in claim 8, The acquisition step acquires information about the gateway through which the managed user terminal is communicating, and information about the relay nodes through which packets from the managed user terminal pass at each time, as information about the node through which packets from the managed user terminal are currently passing. Furthermore, the calculation step involves providing the information of the nodes that packets from the managed user terminal are currently passing through as unknown input to the trained first supervised learning model, performing calculations on the trained first supervised learning model, and outputting the first route selection strategy for the relay nodes that packets from the managed user terminal should sequentially pass through from the gateway, A communication route management step which notifies the gateway and the plurality of relay nodes of the route information determined based on the first route selection strategy output in the calculation step, A communication management method comprising the following features.
10. In the communication management method described in claim 9, The acquisition step involves acquiring information about the relay nodes through which packets destined for the user terminal from the specific external network resource pass at each time, as information about the node through which the packets destined for the user terminal are currently passing. The first learning step involves applying a reward function to an estimated result of calculating the route selection for relay nodes that a packet destined for the user terminal should sequentially traverse to reach the gateway via the relay network, updating it so as to maximize the reward for the packet destined for the user terminal to reach the gateway, and learning a second route selection strategy for relay nodes that a packet destined for the user terminal should sequentially traverse from the specific external network resource using a second reinforcement learning model. The second learning step learns, using a second supervised learning model, the relationship between the node through which packets destined for the user terminal are currently traversed and the second route selection strategy obtained through learning in the first learning step, which concerns the relay nodes through which packets destined for the user terminal should sequentially traverse from the specific external network resource. The memory step involves storing the trained second supervised learning model constructed in the second learning step in the memory unit. A communication management method characterized by the following features.
11. In the communication management method described in claim 10, The acquisition step involves acquiring information on the relay nodes through which packets destined for the managed user terminal from the specific external network resource pass at each time, as information on the node through which the packets destined for the managed user terminal are currently passing. The calculation step involves providing the information of the nodes through which packets destined for the managed user terminal are currently traversing as unknown input to the trained second supervised learning model, performing calculations on the trained second supervised learning model, and outputting the second route selection strategy for the relay nodes that packets destined for the managed user terminal should sequentially traverse from the specific external network resource. The communication route management step notifies the specific external network resource and the multiple relay nodes of the route information determined based on the second route selection strategy output in the calculation step. A communication management method characterized by the following features.
12. In the communication management method described in claim 8, The user terminals are a group of user terminals to which common identification information is assigned, and the gateways through which each user terminal communicates include different gateways from each other. The acquisition step involves acquiring information about the node through which packets from each user terminal are currently passing. The first learning step involves learning the first routing strategy for packets from each user terminal using the first reinforcement learning model, The second learning step learns, using the first supervised learning model, the relationship between the node through which packets from each user terminal are currently traversing and the first routing strategy for relay nodes through which packets from each user terminal should sequentially traverse from the gateway with which each user terminal is communicating, as learned in the first learning step. A communication management method characterized by the following features.
13. In the communication management method described in claim 12, The aforementioned gateway is a user plane function, Furthermore, a management information storage step involves storing management information in a management information storage unit that associates the common identification information, the subscriber identification information of the multiple user terminals, and the identification information of the specific external network resource. A setting step which specifies the management information based on the location registration request signal from each user terminal and instructs the setting of the user plane function of each user terminal. Equipped with, The management information storage step further stores in the management information storage unit the identification information of the user plane functions of each user terminal, which was set according to the instructions in the setting step, in association with the management information. The identification information of the user plane function stored in the management information storage unit indicates the information of the gateway with which each user terminal is communicating. A communication management method characterized by the following features.
Citation Information
Patent Citations
Wireless communication system, wireless access network node, and communication method
JP2018160813A
Communication management device and communication management method
JP7481595B1