Information processing device, information processing method, and information processing program

The information processing device optimizes resource options and prices using reinforcement learning to address the challenges of finite resources, enhancing corporate profits by ensuring desirable offerings and preventing shortages.

JP7772212B2Active Publication Date: 2025-11-18NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024527910
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-11-18
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

Existing technologies fail to optimize resource options and prices for markets dealing with finite resources, such as taxi platforms and cloud computing, due to the inability to manage resource availability and customer demand effectively, leading to reduced corporate profits.

Method used

An information processing device that formulates and solves an optimization problem to determine optimal resource options and prices using reinforcement learning, incorporating both discrete and continuous variables through an improved Wolpertinger Architecture.

Benefits of technology

Optimizes resource options and prices, ensuring desirable offerings to customers while preventing resource shortages, thereby enhancing corporate profits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007772212000041
    Figure 0007772212000041
  • Figure 0007772212000042
    Figure 0007772212000042
  • Figure 0007772212000043
    Figure 0007772212000043
Patent Text Reader

Abstract

An information processing device being provided with: a setting unit that defines a problem by formulating optimization of options for a resource and the prices of the options to be presented to each user in a market dealing with a plurality of types of finite resources; and an action determination unit that determines the options for the resource and the prices of the options to be presented to the user by solving the problem.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] Markets for multiple types of finite resources include various services such as taxi platforms that use taxis as a resource and cloud computing that uses CPUs as a resource. In such markets, market providers must present (a) options and (b) prices for resources to customers who appear in online services. For example, in a taxi platform, when a customer appears and requests transportation from their current location to their destination through a ride-hailing service, the provider must present the customer with (a) options for taxis in the area and (b) the prices of each option. In this case, the company's profits fluctuate depending on the options and prices in (a) and (b).

[0003] First, regarding (a), if a company continues to offer the same options to multiple customers, it will run out of resources, and the number of options it can offer to customers will decrease. Conversely, if it tries to offer customers only options that are convenient from a resource perspective, the number of customers presented with undesirable options will increase. These problems can result in the company being unable to provide appropriate service to customers in terms of resources, or in customers no longer using services, reducing the company's profits. Next, regarding (b), if a resource that is in high demand is priced too low, customers will choose only that resource, running out of resources, and the number of options it can offer to customers will decrease. Conversely, if a particular resource is priced too high, it will end up with a surplus of that resource. This will also reduce the company's profits.

[0004] There is a technology that maximizes corporate profits by optimizing the product lineup and prices of multiple products in the market (see, for example, Non-Patent Document 1). However, although the technology disclosed in Non-Patent Document 1 can be applied to businesses such as retail, where the quantity of products can be controlled by the amount of production, it does not take into consideration the finiteness of resources and therefore cannot be used for services where the number of resources cannot be changed in the short term (for example, a taxi platform that allocates a finite number of taxis to customers, or cloud computing that rents out a finite number of servers to customers).

[0005] There is also a technology that optimizes the price of a finite resource so as to maximize a company's profits while preventing demand from exceeding the amount of the resource at a certain time (see, for example, Non-Patent Document 2). However, while the technology disclosed in Non-Patent Document 2 can be applied to markets for electricity and gas where resources do not have different characteristics, it cannot be applied to markets where multiple products exist and where resources have different characteristics for each customer (for example, close or far distance). [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Ali Aouad, Vivek Farias, and Retsef Levi. Assortment optimization under consider-then-choose choice models. Management Science, Vol.67, No.6, pp.3368-3386, 2021. [Non-patent document 2] Christian Borgs, Ozan Candogan, Jennifer Chayes, Ilan Lobel, and Hamid Nazerzadeh. Optimal multiperiod pricing with service guarantees. Management Science, Vol.60, No.7, pp. 1792-1811, 2014. Summary of the Invention [Problem to be solved by the invention]

[0007] This invention has been made in light of the above circumstances, and in one aspect, aims to provide a technology that enables optimization of resource options and prices offered to each customer in a market that handles multiple types of finite resources. [Means for solving the problem]

[0008] In order to solve the above problem, one aspect of the information processing device of the present invention comprises a setting unit that defines a problem by formulating the optimization of resource options and the prices of the options to be presented to each user in a market that handles multiple types of finite resources, and an action determination unit that determines the resource options and the prices of the options to be presented to each user by solving the problem. [Effects of the Invention]

[0009] According to one aspect of the present invention, it is possible to optimize the resource options and prices offered to each customer. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a server according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an outline of the information processing executed by the server according to the embodiment. [Figure 3] FIG. 3 is a flowchart showing the processing procedure and processing contents of information processing executed by the server according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the process content of information processing executed by the server according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [Embodiment] (Configuration example) FIG. 1 is a block diagram showing an example of the configuration of a server 1 according to the embodiment. The server 1 is an electronic device that collects and processes data, and includes a computer.

[0012] The server 1 is an electronic device that includes a processor 11, a main memory 12, an auxiliary storage device 13, and a communication interface 14. The components that make up the server 1 are connected to each other so that signals can be input and output. In FIG. 1, the interface is written as "I / F."

[0013] The processor 11 corresponds to the central part of the server 1. The processor 11 is an element that constitutes the computer of the server 1. For example, the processor 11 is a CPU (Central Processing Unit), but is not limited to this. The processor 11 may be composed of various circuits. The processor 11 loads a program that is pre-stored in the main memory 12 or the auxiliary storage device 13 into the main memory 12. The program is a program that causes the processor 11 of the server 1 to realize or execute each unit described below. The processor 11 performs various operations by executing the program loaded into the main memory 12.

[0014] The main memory 12 corresponds to the main storage portion of the server 1. The main memory 12 is an element that constitutes the computer of the server 1. The main memory 12 includes a non-volatile memory area and a volatile memory area. The main memory 12 stores an operating system or programs in the non-volatile memory area. The main memory 12 uses the volatile memory area as a work area where data is rewritten by the processor 11 as appropriate. For example, the main memory 12 includes a ROM (Read Only Memory) as a non-volatile memory area. For example, the main memory 12 includes a RAM (Random Access Memory) as a volatile memory area. The main memory 12 stores programs.

[0015] The auxiliary storage device 13 corresponds to the auxiliary storage portion of the server 1. The auxiliary storage device 13 is an element that constitutes the computer of the server 1. The auxiliary storage device 13 is an EEPROM (registered trademark) (Electric Erasable Programmable Read-Only Memory), an HDD (Hard Disc Drive), an SSD (Solid State Drive), or the like. The auxiliary storage device 13 stores the above-mentioned programs, data used by the processor 11 when performing various processes, and data generated by the processes in the processor 11. The auxiliary storage device 13 stores the above-mentioned programs.

[0016] The communication interface 14 includes various interfaces that connect the server 1 to other electronic devices via a network in accordance with a predetermined communication protocol so that they can communicate with each other.

[0017] The hardware configuration of the server 1 is not limited to the above configuration, and the server 1 allows the omission or modification of the above components and the addition of new components as appropriate.

[0018] Each unit implemented in the above-mentioned processor 11 will now be described. The processor 11 realizes a setting unit 100, an input unit 110, a continuous value determination unit 111, an action determination unit 112, and an output unit 113. Each unit realized in the processor 11 can also be referred to as each function. Each unit realized in the processor 11 can also be referred to as being realized in a control unit including the processor 11 and the main memory 12.

[0019] The setting unit 100 defines a problem by formulating resource options and option price optimization to be presented to each user in a market dealing with multiple types of finite resources. Resources include products or services distributed in the market. For example, resources are taxis in a taxi market that provides a service of dispatching taxis to customers. A customer may be interpreted as a user or a person. In this case, resource options include, for example, taxis located in different areas. Resource options are, for example, taxis in area 1, taxis in area 2, taxis in area 3, etc. The option price is, for example, the price of the resource option. The option price is, for example, the minimum taxi fare. Optimizing resource options and option prices includes, for example, providing resource options and option prices that maximize the reward of the resource provider. The problem defined by the setting unit 100 is, for example, maximizing the reward of the resource provider. The resource provider is, for example, a company. The resource provider is, for example, a taxi company. The problem involves maximizing the total reward by repeating multiple times the following processes: observing a user who emerges from a set of multiple users based on a probability distribution; presenting the emerged user with multiple options contained in multiple resources and the prices of the options; obtaining a reward from the resource provider when one option is selected from the multiple options based on the probability distribution; reducing the remaining amount of resource for the selected option by 1; and changing the remaining amount of resource for each option based on a value that follows the probability distribution.

[0020] The input unit 110 inputs a state based on a vector representing the users who have appeared, a vector representing the remaining amount of each resource, and a vector representing the current number of iterations. The number of iterations is the number of times the process by the setting unit 100 is repeated. The current number of iterations is the number of times the process by the setting unit 100 has been repeated up to the present time.

[0021] The continuous value determination unit 111 determines continuous values ​​of options and prices from the state using a mapping.

[0022] The action determination unit 112 determines resource options and their prices to be presented to each user by solving the problem defined by the setting unit 100. The action determination unit 112 determines resource options and their prices to be presented to each user through reinforcement learning for the problem. The action determination unit 112 determines one option as one action based on the continuous value determined by the continuous value determination unit 111. The action is, for example, an optimal option included in multiple resource options. The optimal option is, for example, an option that maximizes reward among multiple resource options. The action indicates, for example, each combination of options and their prices to be presented to each user. For continuous values, the action determination unit 112 determines one action by using mapping from a set obtained by extracting a predetermined number of neighbors in the discrete part of the action space. The action space indicates the entire set of possible actions. The action space is a set of combinations of vectors composed of discrete variables representing which options to present as options and vectors composed of continuous variables representing the price of each option.

[0023] The output unit 113 outputs the action determined by the action determining unit 112. In the following description, "output" may be read as "transmit".

[0024] (Example of server information processing) FIG. 2 is a diagram schematically showing the details of information processing executed by the server 1 according to the embodiment.

[0025] Figure 2 shows the sequence of events that occur after a customer appears in the target market. For each resource i=1,2,...,m, a residual quantity ri is defined. In (i), a customer v from a set V of customers appears under an unknown probability distribution D V In (ii), for the customers who appear, the choice set K ⊆ L and the price vector

number

number

number

[0026] This series of steps will be explained using the example of the taxi platform market. Assume V:={customer departing from area 1, customer departing from area 2, customer departing from area 3}. (i) represents a situation in which a customer departing from a certain area appears. Define L={taxi in area 1, taxi in area 2, taxi in area 3}. In (ii), processor 11 determines which area's taxis the taxi service provider will present as options to the customer and the fare for each option. Here, the upper and lower fare limits are specified by u and l, respectively. (iii) represents a situation in which the customer chooses either one of the options or none. Based on the taxi selected by the customer, processor 11 obtains the taxi service provider's remuneration as (fare) + (negative profit, such as gasoline, resulting from dispatching the taxi). (iv) represents the increase or decrease in the number of taxis other than those allocated to customers. The increase or decrease in the number of taxis other than those allocated to customers includes the arrival and departure of drivers. Consider maximizing the following company profits when steps (i)-(iv) above are repeated n times.

number

[0027] where β is a parameter that determines how much future value is discounted, and R(t) is the amount of reward obtained in the t-th iteration. At each t-th iteration, the processor 11 selects an appropriate choice set K⊆L and a price vector

number

[0028] As a method for efficiently solving the above quantified problem, we will explain the processing procedure when applying a solution method based on reinforcement learning.

[0029] In this example, the resource type is m, and the possible values ​​of the price vector are

number

number

number

number

number

number

number

number

number

number

[0030] For example, the Bellman equation

number

number

number

number

[0031] The problem with applying the Bellman equation is that the number of actions is a combination of continuous and discrete values, and the number of possible combinations of discrete values ​​is enormous. To address this problem, it is possible to improve and apply reinforcement learning with the structure of the Wolpertinger Architecture. The Wolpertinger Architecture is a framework for applying reinforcement learning to problems with a large discrete action space.

[0032] We will explain the processing procedure when using a method that improves the Wolpertinger Architecture so that it can also handle continuous values.

[0033] First, an action (continuous value) is calculated from the state s using a (learned) mapping.

number

number

[0034] Next, the action

number

number

number

number

number

number

[0035] (Server operation example) The processing procedure performed by the server 1 will now be described. In the following description, which focuses on the server 1, the server 1 may be read as the processor 11.

[0036] The processing procedures described below are merely examples, and each process may be modified as much as possible. Furthermore, steps may be omitted, replaced, or added as appropriate depending on the embodiment.

[0037] FIG. 3 is a flowchart showing the processing procedure and processing contents of information processing executed by the server 1 according to the embodiment.

[0038] In the following example, the processor 11 determines an action at each iteration by learned reinforcement learning. Reinforcement learning can be realized, for example, by a method in which the known Wolpertinger Architecture, which is one of the frameworks, is improved as described above. The input unit 110 receives possible values ​​X of the price and a state s t As a result, the appeared v and the remaining amount r for each resource i (i=1, 2,..., m), and input the number of current iterations t (step S1).

[0039] The market state at (ii) in Figure 2 for each iteration t is

number

number

number

number

number

number

number

number

number

number

number

[0040] FIG. 4 is a diagram schematically illustrating an example of the processing content of information processing executed by the server 1 according to the embodiment.

[0041] Figure 4 shows the reinforcement learning process in the taxi market example. First, let T=1 and K=3. The continuous value determination unit 111 determines the continuous values ​​of the prices and options according to the options.

number

[0042] For example, when the continuous value determination unit 111 is given a state s that represents the number of users and the relative positions of each user and each taxi, it determines the price for taxi 1 to be "20 dollars" and the continuous value of the options to be "0.5", the price for taxi 2 to be "10 dollars" and the continuous value of the options to be "0.7", and the price for taxi 3 to be "15 dollars" and the continuous value of the options to be "0.4".

[0043] Next, the action decision unit 112 inputs (1,1,0), (0,1,0), and (1,1,1), which are neighbors of the continuous value (0.5,0.7,0.4), which corresponds to the discrete part in the original action space, into the DNN (Deep Neural Network). In this case, the feature quantities are the state s and the price vector x. The action determination unit 112 selects (1,1,0) as the optimal action. This indicates that Taxi 1 and Taxi 2 are presented as options (the corresponding element is "1"), and Taxi 3 is not presented as an option (the corresponding element is "0").

[0044] The output unit 113 outputs the action (1, 2). At this time, processor 11 calculates the reward for taking the action (1, 2, 20 dollars, 10 dollars) in state s.

number

[0045] Thereafter, processor 11 performs learning by determining options through feedback. Processor 11 also performs learning by feedback on the price vector and the continuous values ​​of the options.

[0046] (effect) As described above in detail, this embodiment makes it possible to optimize the resource options and prices offered to each customer in a market that handles multiple types of finite resources. This embodiment makes it possible to present desirable options to each customer and to reduce the likelihood of resources running out, thereby increasing corporate profits.

[0047] Although the present embodiment has been described using an example assuming the provision of resources and prices in a taxi market, the present embodiment is not limited thereto and can also be applied to various services that provide resources and prices to customers.

[0048] The information processing device may be realized by one device as explained in the above example, or may be realized by multiple devices with distributed functions.

[0049] The program may be transferred in a state where it is stored in an electronic device, or in a state where it is not stored in an electronic device. In the latter case, the program may be transferred via a network, or in a state where it is recorded on a recording medium. The recording medium is a non-transitory tangible medium. The recording medium is a computer-readable medium. The form of the recording medium is not important as long as it is a medium that can store the program and is computer-readable, such as a CD-ROM or a memory card.

[0050] Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.

[0051] In short, this invention is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0052] 1...Server 11...Processor 12...Main memory 13...Auxiliary storage device 14...Communication interface 100...Settings section 110...input section 111...Continuous value determination unit 112...Action decision section 113...Output section

Claims

1. a problem setting unit that defines a problem by formulating resource options and the optimization of the prices of the options to be presented to each user in a market dealing with multiple types of finite resources; an action decision unit that solves the problem to decide resource options and prices of the options to be presented to each user; an output unit that outputs a combination of resource options and prices of each option to be presented to each user, determined by the action determination unit; An information processing device comprising:

2. The information processing apparatus according to claim 1 , wherein the problem is to maximize a reward for a resource provider.

3. 2. The information processing device of claim 1, wherein the problem involves maximizing the sum of the rewards by repeating multiple processes of observing a user that emerges from a set of multiple users based on a probability distribution, presenting the emerged user with multiple options contained in multiple resources and the prices of the multiple options, earning a reward from the resource provider when one option is selected from the multiple options based on the probability distribution, reducing the remaining amount of resource for the selected option by 1, and changing the remaining amount of resource for each option based on a value that follows the probability distribution.

4. The information processing apparatus according to claim 1 , wherein the action determination unit determines resource options and prices of the options to be presented to each user through reinforcement learning for the problem.

5. an input unit for inputting a state based on a vector representing the appeared users, a vector representing the remaining amount of each resource, and a vector representing the current number of iterations; The system further includes a continuous value determination unit that determines continuous values ​​of options and prices from the state using a mapping, The information processing device according to claim 3 , wherein the action determination unit determines the one option as one action based on the continuous value.

6. the action determination unit determines the one action by using a mapping from a set of a predetermined number of neighbors in a discrete part of an action space for the continuous value; The information processing device according to claim 5 .

7. An information processing method executed by an information processing device, A problem is defined by formulating the optimization of resource options and the prices of options presented to each user in a market dealing with multiple types of finite resources; determining resource options and option prices to present to each user by solving the problem; outputting the determined combination of resource options and prices to be presented to each user; An information processing method comprising:

8. On the computer, A problem is defined by formulating the optimization of resource options and the prices of options presented to each user in a market dealing with multiple types of finite resources; determining resource options and option prices to present to each user by solving the problem; outputting the determined combination of resource options and prices to be presented to each user; An information processing program for executing the above.

Citation Information

Patent Citations

  • Support system for car dispatching planning

    JP2001344317A

  • Transaction support system for order receiving and order placing

    JP2002007764A

  • Shared vehicle management server, and shared vehicle management program

    JP2019159685A

  • Pick-up / drop-off position determination method, pick-up / drop-off position determination device, and pick-up / drop-off position determination system

    WO2019220205A1