Client request control system and reinforced learning model

JP2025050396A5Active Publication Date: 2025-08-07KDDI CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023159165
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2025-08-07
Estimated Expiration
2043-09-22

AI Technical Summary

Technical Problem

The prior art is difficult to meet the request management of communication services in real-time and efficiency at the same time, especially how to dynamically decide whether to accept customer service usage requests under limited communication resources.

Method used

By creating pseudo-training data based on past information and building a network traffic prediction model, reinforcement learning technology is used to decide whether to accept customer communication network usage requests, and the training of reinforcement learning model is optimized through the reward value mechanism.

Benefits of technology

Real-time decision-making and resource management of communication network usage requests is realized, the efficiency and reliability of communication services are improved, and the reasonable allocation of network resources is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a client request control system capable of determining in real time that a client request to utilize a communication network is accepted.SOLUTION: A client request control system comprises: a client request management DB storing past network utilization information of a client as client request information; a pseudo request generation node which creates pseudo client request information on the basis of the network utilization information; a reinforced learning node which uses a reinforced learning model to determine whether or not a client request indicated in the pseudo client request information can be accepted and outputs information indicating a determination result; and a reward value determination node which predicts a traffic volume of every client using a traffic prediction model based on a traffic volume in the case where the client utilized a network in the past, and determines a reward value to be applied to the determination result acquired from the reinforced learning node based on a result of the prediction.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a customer request control system and a reinforcement learning model. [Background technology]

[0002] In recent years, communication carriers have been providing high-quality communication services to certain customers, which provide communication resources isolated from other customers, such as a virtual private network (VPN) or a dedicated line.

[0003] In 5G (5th generation mobile communication system) and the like, it will be possible to provide customers with mobile networks built with logically separated communication resources in the form of network slicing, which divides the physical network into multiple slices, and expectations for such services are increasing (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] 3GPP TS23.501 System architecture for the 5G System (5GS) [Non-Patent Document 2] 5G Network Slice Admission Control Using Optimization and Reinforcement Learning Summary of the Invention [Problem to be solved by the invention]

[0005] However, because communication resources such as routers, servers, and lines owned by communication carriers are limited, it is difficult for them to accept all customer requests at any time. Therefore, it is necessary for them to dynamically determine whether to provide communication services to customers when they receive such requests.

[0006] Until now, such decisions as to whether or not to provide the service have been made by the sales or technical staff involved in that work, but there is a demand for the introduction of a system that can make such judgments and decide whether or not to provide the service in real time.

[0007] In recent years, attempts have been made to apply AI (Artificial Intelligence) technology to control communication service usage requests from customers, and in particular, methods using reinforcement learning have been developed (Non-Patent Document 2).

[0008] Thus, while it is becoming clear that it is theoretically possible to use reinforcement learning to control customer requests for communication services, the specific configuration for introducing reinforcement learning into an operational system has not yet been clarified.

[0009] In particular, in reinforcement learning, determining the training scenario and the reward for the agent's decisions are important elements, and it is essential to introduce it into an actual operational system.

[0010] In addition, the training scenarios (simulated data sets) need to be created based on information from past customer requests. Furthermore, when determining reward values ​​in reinforcement learning, it is necessary to predict customer network usage and confirm that it does not exceed the actual network bandwidth.

[0011] The present invention has been made in consideration of the above circumstances, and aims to provide a customer request control system that creates pseudo training data from past information and builds a predictive model of network traffic volume, thereby making it possible to determine in real time whether a customer request to use a communication network will be accepted. [Means for solving the problem]

[0012] (1) In order to achieve the above object, the present invention provides the following means. That is, the customer request control system of the present invention is a customer request control system that determines whether a customer request to use a communication network is acceptable and performs reinforcement learning based on the determination result, and includes a customer request management DB that stores past network usage information of customers as customer request information, a pseudo request generation node that creates pseudo customer request information based on the network usage information, a reinforcement learning training node that uses a reinforcement learning model to determine whether a customer request shown in the pseudo customer request information is acceptable and outputs information indicating the determination result, and a reward value determination node that predicts the traffic volume of each customer using a traffic prediction model based on the traffic volume when the customer used the network in the past and determines a reward value to be assigned to the determination result obtained from the reinforcement learning training node based on the prediction result, and the reinforcement learning training node performs reinforcement learning training for the reinforcement learning model based on the reward value determined by the reward value determination node.

[0013] (2) Furthermore, the customer request control system of the present invention is characterized in that it further includes a customer request judgment node that uses AI (Artificial Intelligence) to judge whether a customer request to use a communication network is acceptable or not based on the trained reinforcement learning model.

[0014] (3) Furthermore, in the customer request control system of the present invention, the pseudo request generation node is characterized in that it calculates statistical information regarding network usage from the network usage information, and calculates the pseudo customer request information using the calculated statistical information.

[0015] (4) Furthermore, the customer request control system of the present invention further includes a traffic management DB that stores the traffic volume when each customer used the network in the past, and the reward value determination node trains a traffic prediction model based on the stored traffic volume and determines the reward value using the trained prediction model.

[0016] (5) Furthermore, the reinforcement learning model of the present invention is characterized in that it is trained by a customer request control system described in any one of (1) to (4) and is used to determine, by AI (Artificial Intelligence), whether a customer request to use a communication network is acceptable or not. Effect of the Invention

[0017] According to the present invention, it is possible to provide a customer request control system that makes it possible to determine in real time whether a customer request to use a communication network is accepted. [Brief description of the drawings]

[0018] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a customer request control system. [Diagram 2] FIG. 2 is a diagram showing an example of a configuration of customer request information. [Diagram 3] FIG. 11 is a diagram showing an example of customer request information to which a determination result is added. [Figure 4] 1 is a graph showing data on the amount of traffic used by a customer in the past at a certain point in time. [Diagram 5] FIG. 1 is a sequence diagram showing a procedure for implementing model training in reinforcement learning. [Figure 6] FIG. 13 is a diagram showing a cumulative distribution function (statistical information) indicating the distribution of network utilization requests of customers. [Figure 7] FIG. 1 is a diagram showing the regional distribution (statistical information) of sending nodes and receiving nodes used by customers on the network. [Figure 8] FIG. 11 is a flow chart showing the steps of a method for determining a reward value. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0019] The inventors focused on the fact that it is not possible to efficiently allocate customer network usage in real time, and discovered that by creating pseudo training data from past information and building a predictive model for network traffic volume, it is possible to determine in real time whether a customer request to use a communication network is accepted, thereby arriving at the present invention.

[0020] That is, the present invention is a customer request control system that judges whether a customer request to use a communication network is acceptable or not and performs reinforcement learning based on the judgment result, and includes a customer request management DB that stores past network usage information of customers as customer request information, a pseudo request generation node that creates pseudo customer request information based on the network usage information, a reinforcement learning training node that judges whether the customer request indicated in the pseudo customer request information is acceptable or not using a reinforcement learning model and outputs information indicating the judgment result, and a reward value judgment node that predicts the traffic volume of each customer using a traffic prediction model based on the traffic volume when the customer used the network in the past and determines a reward value to be assigned to the judgment result obtained from the reinforcement learning training node based on the prediction result, and is characterized in that the reinforcement learning training node trains the reinforcement learning model in reinforcement learning based on the reward value determined by the reward value judgment node.

[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In order to facilitate understanding of the description, the same reference numerals are used to refer to the same components in the drawings of the respective embodiments, and duplicated descriptions will be omitted.

[0022] [1] Overview of customer request control system 1 is a diagram showing a schematic configuration of a customer request control system 100. The customer request control system 100 includes at least a network 1, a customer request transmission node 3, a customer request determination node 5, a customer request management DB 7, a pseudo request generation node 9, a reinforcement learning training node 11, a traffic management DB 13, and a reward value determination node 15 as components.

[0023] Network 1 is a network that is built and operated by a telecommunications carrier for communication services, and is the network that is the target of usage requests from customers. Examples of the network type include, but are not limited to, a mobile network consisting of base stations such as 5G and mobile core equipment, and an IP network consisting of routers such as FTTH (Fiber To The Home).

[0024] The customer request transmission node 3 is a node that transmits information regarding a network usage request from a customer (hereinafter, also referred to as request information) to the customer request determination node 5. In addition, the customer request transmission node 3 receives a determination result on whether or not the network usage request from the customer can be accepted from the customer request determination node 5, and transmits the received determination result to the customer.

[0025] The customer request determination node 5 is a node that determines whether or not to accept the request information received from the customer request transmission node 3. In this embodiment, the customer request determination node 5 automatically determines whether or not to accept the request information using a learning model trained by a reinforcement learning agent, and transmits the determination result to the customer request transmission node 3.

[0026] The customer request management DB7 is a database for storing usage information (also referred to as past network usage information) transmitted from customers up to now as customer request information. Fig. 2 is a diagram showing an example of the configuration of customer request information. As shown in Fig. 2, the customer request information includes information such as the reception time when the request information was received from the customer, the cancellation time when the request was canceled, the start time of network usage, the end time of network usage, the bandwidth used, the transmitting node, and the receiving node.

[0027] The pseudo requirement generation node 9 is a node that creates a pseudo customer requirement information list (also simply called pseudo customer requirement information) required for training of reinforcement learning. It is a node that creates a statistically close pseudo customer requirement information list based on the customer requirement information stored in the customer requirement management DB 7. The pseudo customer requirement information has, for example, the same configuration as that shown in FIG. 2.

[0028] The reinforcement learning training node 11 is a node for training reinforcement learning for determining whether a customer request is acceptable or not. The node 11 determines whether a customer request is acceptable or not for the pseudo customer request information list received from the pseudo request generation node 9 by reinforcement learning, and transmits the determination result to the reward value determination node 15. FIG. 3 is a diagram showing an example of customer request information to which a determination result is assigned. As shown in FIG. 3, the pseudo customer request information received from the pseudo request generation node 9 is assigned a response "Accept the request" or "Reject the request" indicating the determination result. The reinforcement learning training node 11 further trains reinforcement learning from the reward value determined by the reward value determination node 15. The reward value determination node 15 will be described in detail later.

[0029] The traffic management DB 13 is a database for storing data on the amount of traffic of each customer when the customer has used the network up to now (in the past). Fig. 4 is a graph showing, as an example, data on the amount of traffic used by a customer in the past at a certain point in time.

[0030] Reward value determination node 15 is a node that determines a reward value for the determination result transmitted from reinforcement learning training node 11 by the reinforcement learning agent. When determining the reward value, it is necessary to determine whether the network resources are exceeded. For this reason, the traffic usage volume of each customer is predicted using a traffic prediction model learned from the actual traffic volume in the past obtained from traffic management DB 13, and based on the prediction result, it is determined whether the available resources of the network are exceeded, and the reward value is determined.

[0031] In this way, by determining a reward value using pseudo training data (pseudo customer request information) for reinforcement learning created from network usage request information from past customers, and a resource prediction model created from the amount of network resources used by past customers, it is possible to appropriately determine whether or not the network can be used, and to determine in real time whether a customer request to use the communication network is accepted.

[0032] [2. Reinforcement learning model training procedure] Next, a procedure for carrying out model training of reinforcement learning in the customer request control system according to this embodiment will be described. Fig. 5 is a sequence diagram showing a procedure for carrying out model training of reinforcement learning.

[0033] First, the pseudo requirement generation node requests the customer requirement management DB to transmit past customer requirement information (FIG. 2) stored in the database in order to create a pseudo customer requirement information list required for training reinforcement learning (step S1). If past customer requirement information has already been received, there is no need to obtain it again, so step S2 can be omitted and the process can proceed to step S3.

[0034] Next, the customer request management DB acquires a customer request information list from within the DB based on a transmission request for customer request information from the pseudo request generation node, and transmits the list to the pseudo request generation node (step S2).

[0035] Next, the pseudo request generation node calculates various statistical information using the customer request information list obtained from the customer request management DB. The statistical information includes, for example, the cumulative distribution function (CDF) of the distribution of the time from the time when the customer requests network use (Fig. 2: Reception time (t_Request)) to the time when use starts (Fig. 2: Use start time (t_Start)) as shown in Fig. 6, and the regional distribution of sending nodes and receiving nodes as shown in Fig. 7.

[0036] Then, a pseudo customer request information list to be used for training of reinforcement learning, which is in accordance with the actual statistical information, is created using statistical information calculated from the list of customer request information received from the customer (step S3). In this way, by creating a pseudo customer request information list using statistical information calculated from the list of customer request information actually received from the customer, reinforcement learning can be trained with a customer request information list close to the actual one, and the accuracy of the resulting training model can be improved.

[0037] Next, the pseudo requirement generating node transmits the pseudo customer requirement information list created in step S3 to the reinforcement learning training node (step S4).

[0038] Next, the reinforcement learning training node judges whether each request in the pseudo customer request information list received from the pseudo request generation node is to be accepted (Accept) or rejected (Reject), and assigns the judgment result (step S5). The judgment is assumed to be performed using a model of reinforcement learning or deep reinforcement learning, but is not limited to this. FIG. 3 is a diagram showing an example of the customer request information list to which the judgment result has been assigned.

[0039] Next, the reinforcement learning training node transmits the result of the customer request information list determined in step S5 to the reward value determination node (step S6).

[0040] The reward value determination node determines a reward value for each point in time, which is the usage time (from the usage start time to the usage end time) of each customer request information, for the customer request information list to which the results of the reinforcement learning judgments received from the reinforcement learning training node have been added. In order to determine the reward value, it is first necessary to predict the traffic volume of each customer. In order to build the prediction model, the reward value determination node requests the traffic management DB to transmit past traffic data (step S7). Here, if the acquisition of past traffic data and the training of the AI ​​model for traffic prediction have recently been completed, the list and the AI ​​model may be used, in which case steps S8 and S9 may be omitted and the process may proceed to step S10.

[0041] Next, the traffic management DB acquires the traffic data requested by the reward value determination node from the database, and transmits the acquired traffic data to the reward value determination node (step S8). Figure 4 is a diagram showing the configuration of traffic data.

[0042] Next, in step S8, the reward value determination node trains an AI model that predicts the traffic volume that each customer will actually use, using the traffic data acquired from the traffic management DB (step S9). It is assumed that an existing AI model such as LSTM (Long Short-Term Memory) is used to predict the traffic, but is not limited to this.

[0043] Next, the reward value judgment node judges whether the request of each customer can be satisfied by using the AI ​​model of traffic prediction obtained in step S9 in order to determine the reward value for the judgment at each time point for the customer request information list to which the judgment result of acceptance or rejection by reinforcement learning, received from the reinforcement learning training node (step S10). In particular, in a network, network resources (e.g., band width) are limited, and if the total value of the predicted traffic of all the accepted customer request information in a certain time period is equal to or less than the upper limit of the network resource, the reward value is set to a positive value, and conversely, if the upper limit of the network resource is exceeded by a certain customer request information, a refund is given to each customer as a penalty, so the reward value is set to a negative value.

[0044] (How to determine reward value) Here, a method for determining the reward value will be specifically described. Fig. 8 is a flow diagram showing the procedure of the method for determining the reward value.

[0045] First, from the customer request information list acquired from the reinforcement learning training node, one piece of customer request information with the oldest reception time is acquired from among the customer request information for which a reward value has not yet been determined (step S10-1).

[0046] Next, the acceptability (action) of the customer request given to each piece of customer request information in step S5 is confirmed (step S10-2). Specifically, first, the action ("Accept the request" or "Reject the request") of the acquired customer request information is confirmed.

[0047] If the action is "Accept", "1" is input to the variable (action) (step S10-3). On the other hand, if the action is "Reject", "-1" is input to the variable (action) (step S10-4). As will be described in detail later, the value input to the variable (action) becomes a coefficient used when determining the reward value.

[0048] Next, the traffic volume used by this customer is predicted using the traffic prediction model trained in step S9 (step S10-5).

[0049] Next, the predicted traffic data is added to a traffic prediction graph (step S10-6). The traffic prediction graph is a traffic graph that accumulates predicted traffic volumes for customer request information for which the action is "Accept" in the customer request information list that has been processed up to now.

[0050] The total sum of the traffic volume predicted by the current customer request information is compared with the current network source volume (step S10-7). In step S10-7, if the total sum of the traffic volume predicted by the current customer request information and the current network source volume is equal to or less than the upper limit of the network resource volume at all points in the traffic prediction graph, the reward value base (reward_base) is set to a positive value (100 in this embodiment) (step S10-8).

[0051] On the other hand, in step S10-7, if there is one or more points in the traffic prediction graph where the total of the traffic volume predicted by the current customer request and the current network source volume exceeds the upper limit of the network resource volume at all times, the base of the reward value is set to a negative value (-100 in this embodiment) (step S10-9). In this embodiment, as an example of the value to be set for the reward value base (reward_base), a positive value is set to "100" and a negative value is set to "-100", but this is not limited thereto. In addition, in this embodiment, as an example, a static value is used as the base of the reward value, but it is also possible to implement, for example, the service usage fee actually obtained from the customer as the reward value, or the reward value according to the amount of refund to the customer when the network resource volume is exceeded, and in that case, there is an advantage that the model is more likely to accept customer requests that are more likely to improve profits.

[0052] Finally, the reward value (reward_base) is multiplied by the variable (action) to determine the reward value (reward) (step S10-10). For example, if the action determined as "Accept" is within the network resources, the reward value will be a positive value, which indicates that the service fee that the telecommunications carrier can obtain will increase. If the action determined as "Reject" exceeds the amount of network resources, the reward value will also be a positive value by multiplying a negative value by a negative value, which indicates that the telecommunications carrier can determine not to perform a refund process for customers present during that time period due to the resource excess, and the refund paid by the telecommunications carrier will be reduced.

[0053] In this manner, the information to which the reward value determined in step S10 (S10-1 to S10-10) is assigned is transmitted to the reinforcement learning training node (step S11).

[0054] The reinforcement learning training node trains the model based on the reward value received from the reward value determination node (step S12). As in step S5, the model is trained using a learning method such as reinforcement learning or deep reinforcement learning, but is not limited thereto.

[0055] The reinforcement learning training node repeatedly executes steps S1 to S12 until a predetermined prediction accuracy is satisfied or a predetermined number of training runs is completed.

[0056] Finally, the reinforcement learning training node transmits the trained model to the customer request judgment node (step S13). The customer request judgment node uses this model to judge the customer request actually received.

[0057] As described above, according to the above embodiment, by determining a reward value using pseudo training data (pseudo customer request information) for reinforcement learning created from network usage request information from past customers and a resource prediction model created from the amount of network resources used by past customers, it is possible to appropriately determine whether or not the network can be used, and to determine in real time whether or not a customer request to use a communication network is accepted. [Explanation of symbols]

[0058] 100 Customer Demand Control System 1 Network 3 Customer request sending node 5 Customer requirement decision node 7 Customer requirement management DB 9. Pseudo Requirement Generation Node 11 Reinforcement learning training node 13 Traffic Management DB 15 Reward value judgment node

Claims

1. A customer request control system that determines whether a customer request to use a communication network is acceptable or not, and performs reinforcement learning based on the determination result, a customer request management DB that stores past network usage information of customers as customer request information; a pseudo-request generation node that generates pseudo customer request information based on the network usage information; a reinforcement learning training node that uses a reinforcement learning model to determine whether the customer request indicated in the pseudo customer request information is acceptable or not, and outputs information indicating the determination result; a reward value determination node that predicts the traffic volume of each customer using a traffic prediction model based on the traffic volume when the customer used the network in the past, and determines a reward value to be assigned to the determination result obtained from the reinforcement learning training node based on the result of the prediction, A customer request control system, characterized in that the reinforcement learning training node performs reinforcement learning training for the reinforcement learning model based on the reward value determined by the reward value determination node.

2. The customer request control system according to claim 1, further comprising a customer request determination node that uses AI (Artificial Intelligence) to determine whether a customer request to use a communication network is acceptable based on the trained reinforcement learning model.

3. The customer request control system according to claim 2, characterized in that the pseudo-request generation node calculates statistical information regarding network usage from the network usage information, and calculates the pseudo-customer request information using the calculated statistical information.

4. A traffic management DB is further provided to store the traffic volume when each customer has used the network in the past, 4. The customer request control system according to claim 3, wherein the reward value determination node trains a traffic prediction model based on the stored traffic volume and determines the reward value using the trained prediction model.

5. A reinforcement learning model trained by the customer request control system according to any one of claims 1 to 4 and used for determining by AI (Artificial Intelligence) whether a customer request to use a communication network is acceptable.