Learning method and program for value calculation model, and selection probability estimation method

By correlating attribute values with selection probabilities using a neural network, the learning method addresses the challenge of calculating option values, enabling accurate estimation of selection probabilities and optimizing transportation fares.

JP7799186B2Active Publication Date: 2026-01-15FUJITSU LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022086979
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2026-01-15
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Existing methods struggle to calculate the value of options using attribute values measured on different scales, as only selection probabilities are observed, making it difficult to train a value calculation model effectively.

Method used

A learning method that adjusts a value calculation model by correlating attribute values with selection probabilities, using a neural network to learn the relationship between values and probabilities, allowing the model to accurately estimate selection probabilities.

Benefits of technology

Enables the training of a value calculation model that accurately calculates the value of means used in actions, facilitating precise estimation of selection probabilities and optimizing transportation fares to manage congestion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799186000003
    Figure 0007799186000003
  • Figure 0007799186000004
    Figure 0007799186000004
  • Figure 0007799186000005
    Figure 0007799186000005
Patent Text Reader

Abstract

To learn a value calculation model by using selection probability being a measurement value even though learning data does not includes a value.SOLUTION: An information processing device acquires selection probability (an observation value) of each means (a vehicle, an electric train, and a bus) available in moving between O and D, and data on attribute values of each means when the selection probability is acquired. Further, the information processing device calculates a relation (a ratio P1 / P2) between selection probabilities of each means in each combination of two means extractable from a plurality of means. Then, the information processing device adjusts (learns) a value calculation model such that a relation (a ratio V1 / V2) between values V1, V2 to be calculated when attribute values of each means are inputted to the value calculation mode, and the relation (the ratio P1 / P2) between the selection probabilities come close to each other.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning method and program for a value calculation model, and a selection probability estimation method. [Background technology]

[0002] It is desirable to control people's movements in order to reduce CO2 emissions and alleviate traffic congestion. For example, if there are multiple means of transportation for a pair (OD) of origin (O) and destination (D), changing the fare for each means of transportation will change the means of transportation people choose. Therefore, by setting the fare for each means of transportation appropriately, it is possible to appropriately control people's movements.

[0003] In the past, when predicting the proportion of choices that would be made for options with attribute values ​​measured on different scales, such as cost or time, the attribute values ​​were input into a predetermined formula (for example, a linear formula) to obtain a value (value) that could be expressed on a single scale for each option.Then, the relative relationship between the values ​​of each option thus obtained was used to predict the degree to which each option would be selected (selection probability). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-114988 Summary of the Invention [Problem to be solved by the invention]

[0005] In order to train a value calculation model that calculates value from attribute values ​​using machine learning or the like, data that associates the attribute values ​​of multiple options with the values ​​of each option is required as training data. However, the only observed value that can be obtained for each option is the rate at which each option is selected (selection probability). Even if the selection probability of each option is obtained, the value of each option cannot be calculated from the selection probability, so selection probability alone is insufficient as training data.

[0006] In one aspect, the present invention aims to provide a method and a program for learning a value calculation model that can learn a value calculation model that calculates the value of means used when a person takes action. Another object of the present invention is to provide a selection probability estimation method that can accurately estimate the selection probability of multiple means. [Means for solving the problem]

[0007] In one aspect, the learning method for a value calculation model is a learning method for a value calculation model that calculates the value of a means used by a person when performing an action from the attribute values ​​of the means, and is a learning method for a value calculation model in which a computer executes the following processes: acquires input data that corresponds a selection probability indicating the proportion of times each means is selected from a plurality of means, with the attribute values ​​of the plurality of means when the selection probability is obtained; acquires from the input data, for each combination of two means that can be extracted from the plurality of means, the relationship between the selection probabilities of each of the two means included in each combination; and adjusts the value calculation model so that the relationship between the values ​​calculated when each of the attribute values ​​of the two means included in each combination approaches the relationship between the selection probabilities corresponding to each of the combinations. [Effects of the Invention]

[0008] It is possible to train a value calculation model that calculates the value of the means used when a person acts, and to accurately estimate the probability of selecting multiple means. [Brief explanation of the drawings]

[0009] [Figure 1] 1(a) to 1(c) are diagrams for explaining an outline of the processing executed by an information processing apparatus according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing apparatus according to an embodiment. [Figure 3] FIG. 3 is a functional block diagram of the information processing device of FIG. 2. [Figure 4] FIG. 4(a) is a diagram showing an example of movement data, and FIG. 4(b) is a diagram showing an example of selection probability data. [Figure 5] FIG. 10 is a diagram illustrating an example of attribute value data. [Figure 6] FIG. 10 is a diagram illustrating an example of learning data. [Figure 7] FIG. 10 is a diagram showing the input and output of a value calculation model. [Figure 8] FIG. 8(a) is a diagram showing an example of attribute value data (target OD), and FIG. 8(b) is a diagram showing an example of target selection probability data. [Figure 9] FIG. 10 is a diagram for explaining an overview of learning by a model learning unit. [Figure 10] FIG. 2 is a diagram illustrating an overview of a learning device of a model learning unit. [Figure 11] 10 is a flowchart illustrating an example of a learning process of a value calculation model. [Figure 12] 12 is a flowchart showing detailed processing of step S16 in FIG. 11. [Figure 13] 12 is a flowchart showing detailed processing of step S20 in FIG. 11. [Figure 14] 10 is a flowchart illustrating an example of a charge amount determination process. [Figure 15] FIG. 10 is a diagram for explaining an overview of learning by a model learning unit according to a modified example. [Figure 16] FIG. 10 is a diagram illustrating an overview of a learning device of a model learning unit according to a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0010] An embodiment will be described in detail below with reference to FIGS.

[0011] 1(a) to 1(c) are diagrams for explaining an outline of the processing executed by the information processing device 10 of this embodiment. For example, as shown in FIG. 1(a), there is a pair (OD) of a departure point (O) and a destination (D), and there are means (car), means (train), and means (bus) as means of transportation between the departure point and the destination. Also, as shown in FIG. 1(b), cost and time are respectively set as attribute values ​​for each means. In this case, it is assumed that 50% of people traveling between OD select the means (car), 30% select the means (train), and 20% select the means (bus). In the example of FIG. 1(b), many people select the means (car), which results in road congestion.

[0012] The information processing device 10 of this embodiment is a device that determines and outputs an appropriate toll (charge amount) when a user wants to set road pricing (tolls) to alleviate road congestion. For example, when a user inputs that they want to make the selection probabilities of methods 1 to 3 the same (33%) as shown in FIG. 1(c), the information processing device 10 calculates and outputs the cost (charge amount) required to travel the road that will make the selection probabilities the same.

[0013] FIG. 2 shows the hardware configuration of the information processing device 10. As shown in FIG. 2, the information processing device 10 includes a central processing unit (CPU) 90, a read-only memory (ROM) 92, a random access memory (RAM) 94, storage (a solid-state drive (SSD) or a hard disk drive (HDD)) 96, a network interface 97, a display unit 93, an input unit 95, and a portable storage medium drive 99. These components of the information processing device 10 are connected to a bus (data transmission path) 98. In the information processing device 10, the CPU 90 executes a program (including a learning program for a value calculation model) stored in the ROM 92 or the storage 96, or a program read by the portable storage medium drive 99 from the portable storage medium 91, thereby realizing the functions of the components shown in FIG. 3. Note that FIG. 3 also shows various storage units stored in the storage 96, etc. The functions of the units in FIG. 3 may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0014] 3, when the CPU 90 executes a program, the information processing device 10 functions as a movement data acquisition unit 20, a selection probability calculation unit 22, an attribute value acquisition unit 24, a learning data generation unit 26, a model learning unit 28, a target selection probability acquisition unit 30, an optimal billing amount calculation unit 32, and an output unit 34. Each unit will be described in detail below.

[0015] The travel data acquisition unit 20 acquires travel data as shown in FIG. 4(a). Here, the travel data in FIG. 4(a) is, as an example, a record of which means (car, train, bus) a person who traveled three types of ODs used (selected) to travel. Note that in FIG. 4(a), the "selected means" is recorded in association with the "personal ID," but the "personal ID" does not have to be recorded. In other words, the form of the travel data does not matter as long as it can indicate the number of times each means was selected.

[0016] The selection probability calculation unit 22 calculates the proportion of people who selected each means of transportation (car, train, bus) for each OD from the travel data in FIG. 4(a), and generates selection probability data as shown in FIG. 4(b). The selection probability data in FIG. 4(b) shows that, for example, of people who traveled on OD1, 50% selected car, 33% selected train, and 17% selected bus. The selection probability calculation unit 22 stores the generated selection probability data (FIG. 4(b)) in the selection probability recording unit 40.

[0017] The attribute value acquisition unit 24 acquires attribute values ​​(cost and time in this embodiment) of each means of transportation in each OD. Here, cost refers to the fare when using a train or bus, or the road toll when using a car. Time refers to the time required to travel between ODs. The attribute value acquisition unit 24 acquires, for example, attribute value data input by the user as shown in FIG. 5 and stores it in the attribute value recording unit 42. In the case of the attribute value data in FIG. 5, for example, the attribute values ​​of a car in OD1 are cost = 100 yen and time = 10 minutes, the attribute values ​​of a train are cost = 200 yen and time = 6 minutes, and the attribute values ​​of a bus are cost = 500 yen and time = 3 minutes. Note that the selection probabilities of the selection probability data in FIG. 4(b) and the attribute values ​​of the attribute value data in FIG. 5 exist for each means of transportation in each OD, and therefore can be said to have a 1:1 correspondence. In other words, the selection probability data and the attribute value data can be said to be input data that associates a selection probability with the attribute value of each means of transportation when the selection probability is obtained.

[0018] The learning data generation unit 26 generates learning data using the selection probability data (FIG. 4(b)) stored in the selection probability recording unit 40 and the attribute value data (FIG. 5) stored in the attribute value recording unit 42. The learning data is data as shown in FIG. 6. For OD1, the learning data generation unit 26 generates learning data corresponding to each of the two-mode combinations (car / bus, car / train, bus / train) (learning data IDs = 001 to 003). Similarly, for OD2 and OD3, the learning data generation unit 26 generates learning data corresponding to each of the two-mode combinations (car / bus, car / train, bus / train) (learning data IDs = 004 to 006, 007 to 009). In each learning data, the attribute values ​​of the two modes included in the combination are associated with the ratio of the selection probabilities of the two modes. For example, for a combination of OD1's means (car) and means (bus) (learning data ID=001), the ratio of the car selection probability (50%) to the bus selection probability (17%) is 50 / 17. The learning data generation unit 26 stores the generated learning data (FIG. 6) in the learning data recording unit 44. Note that FIG. 6 also includes information on "data source (notes)," but this information is for reference only and does not need to be included in the actual learning data.

[0019] The model learning unit 28 executes a learning process for the value calculation model using the learning data stored in the learning data recording unit 44. FIG. 7 is a diagram showing the input to the value calculation model and the output of the value calculation model. As shown in FIG. 7, the value calculation model is a model that can calculate and output a value (value) expressed in a single scale by inputting attribute values ​​of different scales, such as cost and time. In this embodiment, the value calculation model is a model that uses a neural network called an MLP (Multi-Layer Perceptron). As the MLP, a three-layer perceptron with two input layer nodes, one output layer node, and six hidden layer nodes can be used. The two input layer nodes correspond to the attribute values ​​(cost and time) of the means, respectively, and the one output layer node corresponds to the value of the means. Details of the model learning unit 28 will be described later. The model learning unit 28 records the parameters of the value calculation model obtained by the learning process in the model parameter recording unit 46.

[0020] The target selection probability acquisition unit 30 acquires the target value of the selection probability of each means in a certain OD (target OD) input by the user. The target value data input by the user is target selection probability data as shown in Figure 8(b).

[0021] The optimal charge amount calculation unit 32 acquires attribute value data (see FIG. 8(a)) of each means of travel in the target OD from the attribute value recording unit 42, and calculates what the cost (toll) of the means (car) should be so that the selection probability of each means becomes the target value (FIG. 8(b)). For example, in FIG. 7, suppose that the attribute values ​​of each means (car, train, bus) are input into the value calculation model and the output value values ​​are V1=25, V2=15, and V3=10. In this case, the optimal charge amount calculation unit 32 calculates the selection probability P1 of the means (car) using the relative evaluation formula shown in the following formula (1). P1=V1 / (V1+V2+V3) …(1)

[0022] In the example of FIG. 7, P1 is calculated as 25 / (25+15+10)=0.5=50%. Similarly, the selection probabilities P2 and P3 of the means (train) and the means (bus) are calculated as P2=30% and P3=20%. The optimal billing amount calculation unit 32 calculates the cost (optimal billing amount) of the means (car) such that the values ​​of P1, P2, and P3 match the target values. The optimal billing amount calculation unit 32 notifies the output unit 34 of the calculated optimal billing amount.

[0023] The output unit 34 outputs the optimum billing amount notified by the optimum billing amount calculation unit 32 onto the display unit 93 .

[0024] (Overview of learning in model learning unit 28) Here, an overview of the learning performed by model learning unit 28 will be described.

[0025] As shown in FIG. 7, the value calculation model used in this embodiment is a model that inputs the attribute values ​​of each means and outputs the value of each means. Therefore, in order to train the value calculation model, data on combinations of attribute values ​​and values ​​is required as training data. However, in this embodiment, it is not possible to obtain value values ​​as observed values; only the rate at which each means is actually selected (selection probability) can be obtained as observed values. The value of the selection probability and the value value do not necessarily match (see FIG. 7), and the calculation to obtain the selection probability from the value (relative evaluation in FIG. 7) is an irreversible operation, so it is not possible to obtain the value from the selection probability. Therefore, if the attribute values ​​and selection probability are simply used as training data, it is not possible to machine-learn the parameters of the value calculation model.

[0026] As a result of intensive research, the inventors have noticed that the relationship (ratio) between values ​​can be determined from the selection probability, which is an observed value. For example, as shown in FIG. 9, the ratio of the selection probability of the means (car) to the selection probability of the means (bus) is 50 / 20 (times), but the ratio of the value of the means (car) to the value of the means (bus) is also 50 / 20 (times). Similarly, the ratio of the selection probability of the means (train) to the selection probability of the means (bus) is 30 / 20, and the ratio of the value of the means (train) to the value of the means (bus) is also 30 / 20. Based on the above findings, the inventors decided to machine-learn a value calculation model so that the relationship (ratio) between values ​​output from the value calculation model would approach the relationship (ratio) of the selection probabilities.

[0027] FIG. 10 shows an overview of the learning device of the model learning unit 28. As shown in FIG. 10, the model learning unit 28 first inputs the attribute values ​​of the two means included in the learning data (FIG. 6) into the value calculation model. Then, the model learning unit 28 calculates the relationship (ratio V1 / V2) between the values ​​V1 and V2 output from the value calculation model. The model learning unit 28 also calculates the relationship (ratio P1 / P2) between the selection probabilities P1 and P2 as observed values. Then, the model learning unit 28 calculates the difference (residual (V1 / V2)-(P1 / P2)) between the relationship between the values ​​(ratio V1 / V2) and the relationship between the selection probabilities (ratio P1 / P2). The model learning unit 28 calculates the residual using all the learning data, and updates the parameters of the value calculation model so that the sum of all the residuals is equal to or less than a threshold. In this way, the model learning unit 28 can learn the value calculation model.

[0028] (Regarding the processing of the information processing device 10) Next, a detailed description will be given of the processing of the information processing device 10. The information processing device 10 executes a "learning preparation and learning process" shown in Fig. 11 (and Figs. 12 and 13) and a "charge amount determination process" using a value calculation model shown in Fig. 14.

[0029] (About learning preparation and learning processing) Fig. 11 is a flowchart showing the learning preparation and learning process of the value calculation model. The process of Fig. 11 is executed, for example, at predetermined time intervals or whenever a predetermined amount of movement data, which will be described later, is accumulated.

[0030] 11 starts, first, in step S10, the movement data acquisition unit 20 reads movement data for multiple ODs (see FIG. 4(a)), and the attribute value acquisition unit 24 reads attribute value data of means (see FIG. 5). The movement data acquisition unit 20 passes the read movement data to the selection probability calculation unit 22. In addition, the attribute value acquisition unit 24 stores the read attribute value data in the attribute value recording unit 42.

[0031] Next, in step S12, the selection probability calculation unit 22 calculates the selection probability of each means in each OD by referring to the travel data (FIG. 4(a)). The selection probability calculation unit 22 stores the calculated selection probability of each means in each OD in the selection probability recording unit 40 as selection probability data (FIG. 4(b)).

[0032] Next, in step S14, the learning data generation unit 26 selects one unselected OD. If there are three ODs (OD1 to OD3) as shown in Figures 4(a) to 5, the learning data generation unit 26 selects one of them (for example, OD1).

[0033] Next, in step S16, the learning data generating unit 26 executes a process of generating learning data. In step S16, a process is executed in accordance with the flowchart of FIG.

[0034] (Learning data generation process (S16)) 12, first, in step S30, the learning data generation unit 26 selects one unselected combination from the combinations of two modes of transportation. For example, the learning data generation unit 26 selects the combination of mode (car) and mode (bus).

[0035] Next, in step S32, the learning data generation unit 26 acquires the attribute values ​​of each of the two modes of transportation. The learning data generation unit 26 references the attribute value recording unit 42 and acquires, for example, the attribute values ​​(cost, time) of the mode of transportation (car) and the mode of transportation (bus) of OD1 from the attribute value data in FIG.

[0036] Next, in step S34, the learning data generation unit 26 calculates the ratio of the selection probabilities of the two means of travel. The learning data generation unit 26 refers to the selection probability recording unit 40, obtains the selection probabilities (50%, 17%) of the means of travel OD1 (car) and the means of travel (bus) from the selection probability data in FIG. 4(b), and calculates the ratio (50 / 17).

[0037] Next, in step S36, the learning data generation unit 26 records the ratio between the acquired attribute value and the calculated selection probability as learning data. In the above example, the learning data generation unit 26 records the data of learning data ID="001" in FIG. 6 in the learning data recording unit 44.

[0038] Next, in step S38, the learning data generation unit 26 determines whether all combinations of means have been selected. If the determination in step S38 is negative, the process returns to step S30 and repeats the processes from step S30 onward. On the other hand, if the determination in step S38 is positive, the process proceeds to step S18 in FIG. 11.

[0039] 11, the learning data generation unit 26 determines whether all ODs have been selected. If the determination in step S18 is negative, the process returns to step S14, and the processes in steps S14 and S16 are repeatedly executed. On the other hand, if the determination in step S18 is positive, the process proceeds to step S20. Note that, at the stage of proceeding to step S20, all of the learning data in FIG. 6 has been prepared.

[0040] When the process proceeds to step S20, the model learning unit 28 executes a learning process for the value calculation model. In step S20, the process is executed in accordance with the flowchart of FIG.

[0041] (Learning process of value calculation model (step S20)) When the processing of FIG. 13 starts, first, in step S40, the model learning unit 28 sets the value calculation model to MLP and initializes the parameters.

[0042] Next, in step S42, the model learning unit 28 selects one piece of unselected learning data. For example, the model learning unit 28 selects the first piece of learning data (learning data ID=001) in FIG.

[0043] Next, in step S44, the model learning unit 28 inputs the attribute values ​​of the selected learning data into the value calculation model to calculate the value of each means (V1, V2 in FIG. 10).

[0044] Next, in step S46, the model learning unit 28 calculates the ratio (V1 / V2) of the value of each means.

[0045] Next, in step S48, the model learning unit 28 calculates the difference between the ratio of the values ​​of each means and the ratio of the selection probability of the selected learning data ((V1 / V2)-(P1 / P2)), and records it as a residual.

[0046] Next, in step S50, model learning unit 28 determines whether all learning data have been selected. If the determination in step S50 is negative, the process returns to step S42, and steps S42 to S50 are repeatedly executed until residuals for all learning data have been calculated. On the other hand, if the determination in step S50 is positive, model learning unit 28 proceeds to step S52.

[0047] In step S52, the model learning unit 28 determines whether the sum of the residuals calculated in step S48 is within a threshold value. If the determination in step S52 is negative, the value calculation model needs to be adjusted, and the process proceeds to step S54.

[0048] When the process proceeds to step S54, the model learning unit 28 updates the parameters of the value calculation model. Furthermore, the model learning unit 28 unselects all learning data and deletes all recorded residuals. Thereafter, the model learning unit 28 repeatedly executes the processes of steps S42 to S54 using the updated value calculation model. Then, when the total residuals falls within the threshold, the determination in step S52 becomes positive, and the model learning unit 28 proceeds to step S56.

[0049] When the process proceeds to step S56, the model learning unit 28 records the parameters of the value calculation model in the model parameter recording unit 46. This ends the process of Fig. 13, and the entire process of Fig. 11 also ends.

[0050] (Regarding the billing amount determination process) Next, the charge amount determination process will be described with reference to the flowchart in FIG. 14. This charge amount determination process is a process for determining a road toll using the value calculation model learned by the learning process in FIG. 11. For example, it is assumed that the user has selected "OD1" as the OD to be considered (target OD). It is also assumed that the user has input the target selection probability data as shown in FIG. 8(b) as information on the target selection probability. In this case, the optimal charge amount calculation unit 32 calculates the cost of the means "car" so that the selection probability of each means of OD1 matches the selection probability in FIG. 8(b), and outputs the calculated cost as the optimal charge amount.

[0051] 14 starts, first, in step S70, the optimal billing amount calculation unit 32 reads the attribute value data of each means of the OD (e.g., OD1) under consideration and the target selection probability data (target selection probability data). Note that the optimal billing amount calculation unit 32 reads the target selection probability data (FIG. 8(b)) entered by the user via the target selection probability acquisition unit 30.

[0052] Next, in step S72, the optimum billing amount calculation unit 32 selects one of the unselected means. For example, the optimum billing amount calculation unit 32 selects the means (car) from the means (car), means (train), and means (bus).

[0053] Next, in step S74, the optimum billing amount calculation unit 32 calculates and records the value of the selected means using the value calculation model that has been trained through the processes of FIGS.

[0054] Next, in step S76, the optimum billing amount calculation unit 32 determines whether all means have been selected. If the determination in step S76 is negative, the process returns to step S72, and steps S72 to S76 are repeated until the values ​​of all means have been calculated. If the determination in step S76 is positive, the optimum billing amount calculation unit 32 proceeds to step S78.

[0055] In step S78, the optimal billing amount calculation unit 32 calculates (estimates) the selection probability of each means from the calculated value of each means. Specifically, the optimal billing amount calculation unit 32 calculates (estimates) the selection probability of each means using the above formula (1).

[0056] Next, in step S80, the optimal billing amount calculation unit 32 determines whether the calculated selection probability matches the target selection probability. The optimal billing amount calculation unit 32 may consider the calculated selection probability and the target selection probability to match when the difference between them falls within a predetermined range. If the determination in step S80 is negative, the process proceeds to step S82, where the optimal billing amount calculation unit 32 updates the cost of the means (vehicle), and the process returns to step S72. Thereafter, the optimal billing amount calculation unit 32 repeats the processes from step S72 onwards until the determination in step S80 is positive. If the determination in step S80 is positive, the process proceeds to step S84.

[0057] In step S84, the output unit 34 outputs the cost of the means (car) when the determination in step S80 is affirmative as the optimal charge amount. By checking the output optimal charge amount, the user can confirm what the appropriate car toll should be set to in order to make the selection probability of each means match the target selection probability.

[0058] As described above in detail, the information processing device 10 of this embodiment acquires the selection probability of each means available when traveling between ODs and the data of the attribute values ​​of each means when the selection probability is obtained (FIGS. 4(b) and 5). Furthermore, the information processing device 10 calculates the relationship (ratio) of the selection probabilities of each means for each combination of two means that can be extracted from a plurality of means. Then, the information processing device 10 adjusts (learns) the value calculation model so that the relationship (ratio) of each value calculated when the attribute values ​​of each means are input into the value calculation model approaches the relationship (ratio) of the selection probabilities. As a result, in this embodiment, even if the learning data does not include value values, the value calculation model can be learned from the ratio of selection probabilities, which is an observed value. Furthermore, by machine learning the value calculation model, a value calculation model that can accurately calculate the value of a means can be obtained.

[0059] In addition, the value calculation model used in this embodiment is a neural network (MLP, etc.) whose input is the attribute value of each means and whose output is the value of each means. This allows the user to automatically learn the value calculation model without having to assume a linear formula or the like in advance as the value calculation model.

[0060] Furthermore, in this embodiment, the model learning unit 28 calculates the difference (residual) between the relationship (ratio) of the value of each means obtained from each learning data (learning data ID=001 to 009) and the relationship (ratio) of the selection probability. Then, the model learning unit 28 adjusts the parameters of the value calculation model so that the sum of each difference is equal to or less than a threshold value (S42 to S54 in FIG. 13). This makes it possible to obtain a value calculation model that can calculate the value of each means with high accuracy.

[0061] In this embodiment, the optimal billing amount calculation unit 32 calculates the value of each means by inputting the attribute values ​​of each means into the value calculation model learned by the processes of Figures 11 to 13 (S74 in Figure 14).The optimal billing amount calculation unit 32 then calculates the selection probability of each means based on the calculated value of each means (S78).This makes it possible to accurately calculate the selection probability of each means.

[0062] In this embodiment, the optimal charge calculation unit 32 adjusts at least a part of the attribute values ​​of each means so that the estimated selection probability of each means approaches the target selection probability (S82). As a result, for example, by adjusting the cost of the means (car) so that the selection probability of the means (car) becomes smaller, it becomes possible to determine the optimal toll (road pricing) to alleviate road congestion.

[0063] In the above embodiment, the optimal charge amount calculation unit 32 optimizes road tolls, but the present invention is not limited to this. Train and bus fares (costs) may be adjusted so that the selection probability of each means approaches a target selection probability. The means of transportation may include other means of transportation (motorcycles, ships, airplanes, etc.) in addition to cars, trains, and buses, or instead of at least one of cars, trains, and buses.

[0064] (Variation) In the above embodiment, the case where the relative evaluation in Fig. 7 is performed based on the above formula (1) has been described, but this is not limiting, and for example, a logit model that is frequently used in behavioral selection models can also be used for the relative evaluation, as shown in Fig. 15. When a logit model is used for the relative evaluation, the selection probability Pi of each means can be calculated from the following formula (2).

[0065]

number

[0066] In this case, the relationship between P1 and P2 can be expressed as in the following equation (3).

[0067]

number

[0068] From the above formula (3), we can see that when using a logit model for relative evaluation, the difference in value can be used as the value relationship. We can also see that the value calculation model should be trained so that the difference in value approaches the relationship between selection probabilities (the difference between the natural logarithms of the selection probabilities).

[0069] FIG. 16 shows an overview of the learning device of the model learning unit 28 according to this modification. As shown in FIG. 16, in this modification, as in the above embodiment, the model learning unit 28 inputs the attribute values ​​of two means into the value calculation model during learning. Then, the model learning unit 28 calculates the relationship between the values ​​(difference (V1-V2)) from the values ​​V1 and V2 output from the value calculation model. The model learning unit 28 also calculates the relationship (lnP1-lnP2) between the selection probabilities P1 and P2 as observed values. The model learning unit 28 then calculates the difference ((lnP1-lnP2)-(V1-V2)) between the relationship between the values ​​(V1-V2) and the relationship between the selection probabilities (lnP1-lnP2). The model learning unit 28 calculates the difference (residual) using all learning data and updates the parameters of the value calculation model so that the sum of the differences (residuals) is equal to or less than a threshold. In this way, the value calculation model can be trained in this modification as well.

[0070] In this modified example, as shown in the above formula (3), the difference between V1 and V2 is used, so calculations can be performed even if the values ​​V1 and V2 are 0. In this way, since the machine learning loss function does not have a singular point, calculations in machine learning can be stabilized.

[0071] In the above embodiment, the value calculation model is a neural network model such as MLP, but the present invention is not limited to this. A linear expression such as V=w1×cost+w2×time (w1 and w2 are weighting coefficients, and V is value) can also be used as the value calculation model.

[0072] In the above embodiment, transportation means (car, train, bus) have been described as examples of means used by people when performing actions, but the present invention is not limited to this. There are various means used by people when performing actions, and for example, online shopping and brick-and-mortar stores used when shopping also fall under the category of means used by people when performing actions. In other words, in situations where a person must select from multiple means when performing some action, and the value of each means and the selection probability of each means are to be calculated, the above embodiment can be modified as appropriate and used. In the above embodiment, the attribute values ​​are described as cost and time, but the attribute values ​​may be something other than cost and time.

[0073] In the above embodiment, the case where the information processing device 10 used by the user has the functions shown in Fig. 3 has been described, but the present invention is not limited to this. For example, a server device connected to the information processing device 10 used by the user via a network or the like may have the functions shown in Fig. 3.

[0074] The above processing functions can be realized by a computer. In this case, a program is provided that describes the processing contents of the functions that the processing device should have. By executing the program on a computer, the above processing functions are realized on the computer. The program that describes the processing contents can be recorded on a computer-readable recording medium (excluding carrier waves).

[0075] When distributing a program, it is sold in the form of a portable recording medium, such as a DVD (Digital Versatile Disc) or a CD-ROM (Compact Disc Read Only Memory), on which the program is recorded. Alternatively, the program can be stored in a storage device of a server computer and transferred from the server computer to other computers via a network.

[0076] A computer that executes a program stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer in its own storage device. The computer then reads the program from its own storage device and executes processing in accordance with the program. Note that the computer can also read the program directly from a portable recording medium and execute processing in accordance with that program. The computer can also execute processing in accordance with the program received each time a program is transferred from the server computer.

[0077] The above-described embodiment is a preferred example of the present invention, but the present invention is not limited to this and can be modified in various ways without departing from the spirit of the present invention.

[0078] In addition, the following supplementary notes are disclosed regarding the above-described embodiment and modified examples. (Supplementary Note 1) A learning method for a value calculation model that calculates the value of a means used by a person when taking an action from an attribute value of the means, comprising: acquiring input data that associates a selection probability indicating the rate at which each means is selected from among a plurality of means with an attribute value of the plurality of means when the selection probability is obtained; For each combination of two means that can be extracted from the plurality of means, the relationship between the selection probabilities of the two means included in each combination is obtained from the input data, and the value calculation model is adjusted so that the relationship between the values ​​calculated when the attribute values ​​of the two means included in each combination are input into the value calculation model is brought closer to the relationship between the selection probabilities corresponding to each combination. A method for learning a value calculation model, characterized in that processing is executed by a computer. (Supplementary Note 2) The learning method for a value calculation model according to Supplementary Note 1, wherein the value calculation model is a neural network that receives the attribute value as an input and outputs the value. (Appendix 3) The method for learning a value calculation model described in Appendix 1 or 2, characterized in that the adjustment process calculates, for all combinations of two means that can be extracted from the plurality of means, the difference between the relationship between each of the values ​​of the two means included in the combination and the relationship between the selection probabilities corresponding to the combination, and adjusts the value calculation model so that the sum of the differences for all the combinations is smaller than a predetermined value. (Appendix 4) A learning method for a value calculation model described in any one of Appendices 1 to 3, characterized in that the relationship between each of the values ​​is the ratio of one value to the other value, and the relationship between the selection probabilities is the ratio of one selection probability to the other selection probability. (Appendix 5) A learning method for a value calculation model described in any one of Appendices 1 to 3, characterized in that the relationship between each of the values ​​is the difference between one value and the other value, and the relationship between the selection probabilities is the difference between the natural logarithm of one selection probability and the natural logarithm of the other selection probability. (Supplementary Note 6) The value calculation model is trained using the learning method for the value calculation model according to any one of Supplementary Notes 1 to 5, Calculating the value of each of the multiple means by inputting attribute values ​​of the multiple means into the value calculation model; estimating the selection probability of each means based on the calculated values ​​of the plurality of means; A selection probability estimation method characterized in that the processing is executed by a computer. (Appendix 7) The selection probability estimation method described in Appendix 6, characterized in that the computer executes a process of adjusting at least some of the attribute values ​​of the multiple means so that the estimated selection probability of each means approaches a target selection probability. (Appendix 8) A learning program for a value calculation model that calculates the value of a means used when a person acts from an attribute value of the means, On the computer, acquiring input data that associates a selection probability indicating the rate at which each means is selected from among a plurality of means with an attribute value of the plurality of means when the selection probability is obtained; For each combination of two means that can be extracted from the plurality of means, the relationship between the selection probabilities of the two means included in each combination is obtained from the input data, and the value calculation model is adjusted so that the relationship between the values ​​calculated when the attribute values ​​of the two means included in each combination are input into the value calculation model is brought closer to the relationship between the selection probabilities corresponding to each combination. A learning program for a value calculation model, characterized by executing processing. (Supplementary Note 9) The learning program for a value calculation model according to Supplementary Note 8, wherein the value calculation model is a neural network that receives the attribute value as an input and outputs the value. (Appendix 10) A learning program for a value calculation model described in Appendix 8 or 9, characterized in that the adjustment process calculates, for all combinations of two means that can be extracted from the plurality of means, the difference between the relationship between each of the values ​​of the two means included in the combination and the relationship between the selection probabilities corresponding to the combination, and adjusts the value calculation model so that the sum of the differences for all the combinations is smaller than a predetermined value. (Appendix 11) A learning program for a value calculation model described in any one of Appendices 8 to 10, characterized in that the relationship between each of the values ​​is the ratio of one value to the other value, and the relationship between the selection probabilities is the ratio of one selection probability to the other selection probability. (Appendix 12) A learning program for a value calculation model described in any one of Appendices 8 to 10, characterized in that the relationship between each of the values ​​is the difference between one value and the other value, and the relationship between the selection probabilities is the difference between the natural logarithm of one selection probability and the natural logarithm of the other selection probability. [Explanation of symbols]

[0079] 10. Information processing equipment 20. Mobile data acquisition unit 22 Selection probability calculation unit 24 Attribute value acquisition section 26 Learning data generation unit 28 Model Learning Department 30 Target selection probability acquisition unit 32 Optimal billing amount calculation section 34 Output section 40 Selection probability recording section 42 Attribute Value Recording Section 44 Learning data recording unit 46 Model parameter recording section

Claims

1. A learning method for a value calculation model that calculates the value of a means used by a person when taking action from an attribute value of the means, comprising: acquiring input data that associates a selection probability indicating the rate at which each means is selected from among a plurality of means with an attribute value of the plurality of means when the selection probability is obtained; For each combination of two means that can be extracted from the plurality of means, the relationship between the selection probabilities of the two means included in each combination is obtained from the input data, and the value calculation model is adjusted so that the relationship between the values ​​calculated when the attribute values ​​of the two means included in each combination are input into the value calculation model approaches the relationship between the selection probabilities corresponding to each combination. A method for learning a value calculation model, characterized in that processing is executed by a computer.

2. 2. The learning method for a value calculation model according to claim 1, wherein the value calculation model is a neural network that receives the attribute value as an input and outputs the value.

3. The method for learning a value calculation model described in claim 1, characterized in that the adjustment process calculates the difference between the relationship between the values ​​of the two means included in all combinations of two means that can be extracted from the multiple means and the relationship between the selection probabilities corresponding to the combinations, and adjusts the value calculation model so that the sum of the differences for all combinations is smaller than a predetermined value.

4. 2. A learning method for a value calculation model as described in claim 1, characterized in that the relationship between each of the values ​​is the ratio of one value to the other value, and the relationship between the selection probabilities is the ratio of one selection probability to the other selection probability.

5. The learning method for a value calculation model described in claim 1, characterized in that the relationship between each of the values ​​is the difference between one value and the other value, and the relationship between the selection probabilities is the difference between the natural logarithm of one selection probability and the natural logarithm of the other selection probability.

6. The value calculation model is trained using the value calculation model training method according to any one of claims 1 to 5, Calculating the value of each of the multiple means by inputting attribute values ​​of the multiple means into the value calculation model; estimating the selection probability of each means based on the calculated values ​​of the plurality of means; A selection probability estimation method characterized in that the processing is executed by a computer.

7. The method for estimating a selection probability according to claim 6, characterized in that the computer executes a process of adjusting at least some of the attribute values ​​of the plurality of means so that the estimated selection probability of each means approaches a target selection probability.

8. A learning program for a value calculation model that calculates the value of a means used by a person when taking action from an attribute value of the means, On the computer, acquiring input data that associates a selection probability indicating the rate at which each means is selected from among a plurality of means with an attribute value of the plurality of means when the selection probability is obtained; For each combination of two means that can be extracted from the plurality of means, the relationship between the selection probabilities of the two means included in each combination is obtained from the input data, and the value calculation model is adjusted so that the relationship between the values ​​calculated when the attribute values ​​of the two means included in each combination are input into the value calculation model approaches the relationship between the selection probabilities corresponding to each combination. A learning program for a value calculation model, characterized by executing processing.

Citation Information

Patent Citations

  • Processing device, processing method, and program

    JP2015114988A