An induced path recommendation method based on the just noticeable difference of passenger preference ranking

By constructing passenger portraits and using JND dictionary sequence preference model and DDPG reinforcement learning algorithm, the problem of lack of personalized path recommendations in rail transit is solved, the precise induction of passenger travel paths is achieved, and the accuracy of recommendations and personalized services are improved.

CN115391641BActive Publication Date: 2025-07-08BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210898410.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-07-08
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

The existing rail transit passenger flow induction path recommendations lack personalization and fail to effectively consider passengers' individual preferences, resulting in poor recommendation results.

Method used

By constructing a passenger portrait label system, using spectral clustering method to extract passenger travel preferences, combining the JND dictionary sequence preference model and DDPG reinforcement learning algorithm, the path selection model is optimized, and personalized path recommendations are provided.

Benefits of technology

It improves the accuracy and operation management level of rail transit passenger information services, and realizes accurate recommendations of passenger personalized induction paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391641B_ABST
    Figure CN115391641B_ABST
Patent Text Reader

Abstract

The present invention provides an induced path recommendation method based on the Just Noticeable Difference (JND) passenger preference ranking. The method includes: obtaining passenger travel information, analyzing the passenger travel process, and constructing a passenger portrait label system including direct attributes and indirect attributes; refining the passenger travel information from two dimensions of the origin-destination (OD) level and the time period level based on the passenger portrait label system, and using the spectral clustering method to extract passenger travel preferences; combining the preference ranking of passenger travel, constructing a lexicographic preference path selection model based on JND, obtaining a path that meets the passenger preferences, and recommending it to the passenger. The method of the present invention proposes a multi-dimensional passenger preference recognition method based on AFC data and passenger portraits, solves the matching problem between passengers and paths, and optimizes the parameters of the passenger path selection model by combining Deep Deterministic Policy Gradient (DDPG), improves the accuracy of personalized passenger guidance, and provides a decision-making reference for accurate passenger flow guidance in rail transit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of urban rail transit passenger flow management, and particularly to an induced path recommendation method based on the least noticeable difference (JND) of passenger preference ranking. Background Art

[0002] With the increase in people's travel activities and the rapid rise in traffic demand, rail transit has developed rapidly, the road network scale has been increasing day by day, and passenger flow congestion has been increasing. To alleviate passenger flow congestion, in addition to passenger flow control means, passenger flow guidance has become a new way and means. However, the existing research on the induced path recommendation in the field of rail transit has poor effect, lacking personalized and accurate guidance for passengers. In the aspect of passenger travel path generation, most only consider the group characteristics of passengers, adopt the same set of model parameters for all passengers, and do not consider the individual preferences of passengers.

[0003] Therefore, it is crucial to generate a personalized induced information release strategy by recommending an induced path based on passenger preference analysis. Summary of the Invention

[0004] The embodiments of the present invention provide an induced path recommendation method based on the least noticeable difference (JND) of passenger preference ranking, which not only proposes a personalized induced path recommendation method based on the least noticeable difference (JND) of passenger preference analysis, but also provides theoretical and technical references for improving the level of rail transit passenger information service and operation management.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] An induced path recommendation method based on the least noticeable difference (JND) of passenger preference ranking, comprising:

[0007] Obtain passenger travel information, analyze the passenger travel process, and construct a passenger portrait label system including direct attributes and indirect attributes;

[0008] Refine the passenger travel information from two dimensions of the origin-destination (OD) level and the time period level based on the passenger portrait label system, and adopt the spectral clustering method to extract the passenger travel preferences;

[0009] Combine the preference ranking of passenger travel, construct a lexicographic preference path selection model based on JND, obtain the path that meets the passenger preferences, and recommend it to the passengers.

[0010] Preferably, the obtaining passenger travel information, analyzing the passenger travel process, and constructing a passenger portrait label system including direct attributes and indirect attributes comprises:

[0011] Define the road network using interval attributes and path attributes. The interval attributes include interval running time and interval congestion situation, and the path attributes include: path travel time, path waiting time, number of path transfers, and path congestion level.

[0012] Set the mathematical model formula for the six-tuple defining the passenger travel process as:

[0013] x m =(id, t in , t out , st o , st d , r, tp)

[0014] In the formula, m — a certain passenger, id — passenger ID, used to identify the passenger; t in — entry time; t out — exit time; st o — starting station; st d — terminal station; r — the travel path of this passenger is the r-th feasible path for this OD; tp — travel period;

[0015] The passenger portrait label system includes: direct attributes and indirect attributes.

[0016] The direct attributes mainly include card type, number of trips, travel OD distribution, travel time distribution, and travel path distribution. The indirect attributes include average number of transfers, average travel time, average congestion level of travel paths, average waiting time, and label preference distribution.

[0017] Preferably, refine the passenger travel information from two dimensions of the OD level and the time period level based on the passenger portrait label system, and use the spectral clustering method to identify and extract passenger travel preferences, including:

[0018] Divide the travel record set X of all passengers according to OD, and screen out the passenger travel data X o,d with a certain OD as the travel OD, to achieve the division of the spatial dimension. The calculation formula of X o,d is as follows:

[0019] X o,d ={x m |st o =o, st d =d}

[0020] Divide X o,d according to the travel period to obtain the subset X o,d,τ of period τ, to achieve the division of the time dimension. The calculation formula of X o,d,τ is as follows:

[0021] X o,d,τ ={x m |st o =o, st d =d, tp = τ}

[0022] Divide X o,d,τ by passengers to obtain subsets of different passengers at different OD and different time periods The calculation formula is as follows:

[0023]

[0024]

[0025] Filter out passengers with more than 3 trips, calculate the indirect attributes of each passenger at different OD and different time periods, and obtain the travel individual characteristic attributes of passengers at the time period τ of (o, d) Form a set U of travel individual characteristic attributes of passengers at different OD and different time periods o,d,τ , and perform spectral clustering on U o,d,τ to determine the optimal clustering cluster C according to the silhouette coefficient and CH score. The calculation formula is as follows:

[0026] C = {C1, C2,..., C k}

[0027] According to the clustering center c of each class of passengers k , determine the passenger preferences of each class. The formula is as follows:

[0028]

[0029] In the formula, Q k is the passenger preference reflected by the clustering center c k . The passenger preference Q of passenger c c is the passenger preference presented by the clustering center of the class to which it belongs

[0030] Preferably, the preference degree ranking of passenger travel is combined to construct a lexicographic preference path selection model based on JND, obtain paths that meet passenger preferences, and recommend them to passengers, including:

[0031] Set the JND threshold for passengers' perception of different attributes of paths. When the travel time difference between multiple paths is less than the JND threshold, it is considered that there is no difference in travel time among these multiple paths; when the travel time difference between multiple paths is greater than the JND threshold, it is considered that there are differences in travel time among these multiple paths;

[0032] Let be the passenger preference attribute q iThe perceivable stimulus quantity change ratio, that is, if the attribute q of two paths i The difference ratio is less than Then the difference in this attribute between these two paths is not perceived, and then these two paths are considered to have no difference in this attribute. The calculation formula is as follows:

[0033]

[0034] In the formula, Is the optimal value of attribute q i , In percentage form. The above formula means that the perceivable change ratio of the passenger for attribute q i Is Path And path The better value of this attribute in is When the passenger is comparing these two paths, if the difference between the two paths is within , the passenger will think that the performance of these two paths on attribute q i Is the same, and either path can be selected;

[0035] Model assumptions of the lexicographic preference model based on JND:

[0036] It is assumed that the lexicographic preference set Q of passenger c for a certain OD pair (o, d) is known c , and the path attribute set of OD at time ω

[0037] Assume that the passenger's perception of attribute changes follows Weber's law, there is a JND threshold, and assume that the set of perceivable change ratios of the passenger for attribute changes is β, and the calculation formula is as follows:

[0038]

[0039] In the formula ——The perceivable change ratio of the preference attribute q i ;

[0040] ——The perceivable change ratio of the travel time attribute;

[0041] ——The perceivable change ratio of the transfer times attribute;

[0042] ——The perceivable change ratio of the waiting time attribute;

[0043] ——The perceivable change ratio of the path congestion degree attribute.

[0044] Randomly continuousize the β set, and use the DDPG algorithm to automatically explore the parameters;

[0045] Obtain the passenger preference degree ranking according to the passenger travel preferences, and construct a lexicographic preference path selection model based on JND according to the passenger preference degree ranking. The calculation process of the lexicographic preference model based on JND is as follows:

[0046] Step1: Input the ordered preference set of passengers, the perceived change rate, and the path set matrix of OD. Initialize the recommended path set and the temporary path set. Compare the attribute values in the feasible path set in sequence according to the preference order, that is, traverse the path set matrix of OD from top to bottom in row order;

[0047] Step2: Sort them from small to large, and calculate the upper limit of the perceptible difference of passengers If the temporary path set is not empty, then the paths in it that are greater than the upper limit are taken out and stored in the recommended path set from large to small in sequence; otherwise, jump to step 4;

[0048] Step3: After traversing the path set matrix of OD, if is empty, then directly jump to step 4; otherwise, take out the paths in and store them in from large to small in sequence;

[0049] Step 4: is an ordered path set. The later the sorting, the more it meets the passenger preferences. Reverse the order of to obtain an ordered path set that meets the passenger preferences The in it is the matching path l best , and is used as the path that meets the passenger preferences and is recommended to the passenger.

[0050] Preferably, the method further includes: considering the sensitivity differences of different passengers to path attributes, and optimizing the lexicographic preference path selection model based on JND by combining the DDPG reinforcement learning algorithm, including:

[0051] Step 1: Randomly initialize the main network parameters θ μ and θ Q and the target network parameters θ μ′ and θ Q′ , where θ μ′ = θ μ , θ Q′ = θQ ; Then initialize the sample storage buffer R and set a preset number of iterations;

[0052] Step 2: Initialize a random noise ε t and obtain the current state s t ;

[0053] Step 3: According to the output of the actor network and the noise ε t select the action a t , and the calculation formula is as follows; The environment executes the action a t , and obtains the reward r t and the new state s t+1 , and the actor network will (s t , a t , r t , s t+1 ) be stored in the sample storage buffer R as a set of data, serving as the dataset for training the network;

[0054] a t = μ(s t |θ μ ) + ε t

[0055] Step 4: Randomly sample N groups of (s t , a t , r t , s t+1 ) data from R as the training data for the actor main network and the critic main network, calculate the gradient of the critic main network, and update the main network. The update formula is as follows:

[0056]

[0057] In the formula, L is the loss function of the Critic network, which is the mean square error between the predicted Q value and the target Q value. Calculate the gradient of the actor main network and update the main network. The update formula is as follows:

[0058]

[0059] In the formula ——The gradient of J β (μ), that is, when s follows the ρ β distribution, the expected value of;

[0060] Step 5: Update the target network parameters. The calculation formula is as follows:

[0061] θ Q′ = τθ Q+(1 - τ)θ Q′

[0062] θ μ′ = τθ μ +(1 - τ)θ μ′

[0063] where τ is the update coefficient, and the value in the present invention is 0.01. After reaching the iteration times, the iteration ends; otherwise, it jumps to Step 2.

[0064] As can be seen from the technical solutions provided by the embodiments of the present invention above, the embodiments of the present invention disclose an induced path recommendation method and system based on JND (Just Noticeable Difference) passenger preference analysis. The method proposes an OD-based passenger preference recognition method based on AFC (Automatic Fare Collection) data and passenger portraits, and refines passenger travel information from two dimensions: the OD level and the time period level; considering the passenger preference degree ranking, constructs a lexicographic preference path selection model based on JND to solve the matching problem between passengers and paths, and combines the DDPG reinforcement learning method to optimize the parameters of the passenger path selection model, improve the accuracy of personalized induction of passengers, and provide a decision-making reference for accurate passenger flow induction in rail transit.

[0065] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0067] Figure 1 It is a flowchart of an induced path recommendation method based on the lexicographic preference of passengers based on the just noticeable difference provided by the embodiments of the present invention.

[0068] Figure 2 It is a flowchart of passenger preference recognition provided by the embodiments of the present invention.

[0069] Figure 3 It is a framework diagram of a lexicographic preference model based on JND provided by the embodiments of the present invention.

[0070] Figure 4 It is a schematic diagram of the DDPG reinforcement learning algorithm provided by the embodiments of the present invention. Detailed Embodiments

[0071] The following describes in detail the embodiments of the present invention. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0072] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of any of the one or more associated listed items.

[0073] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted with an idealized or overly formal meaning unless defined as here.

[0074] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments as examples in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.

[0075] Embodiment 1

[0076] The processing flow of an induced path recommendation method based on the just noticeable difference passenger preference ranking provided by the embodiment of the present invention is as Figure 1 shown, including the following processing steps:

[0077] Step S10: Obtain passenger travel data, analyze the passenger travel process, and construct a passenger portrait label system based on the user portrait idea.

[0078] The data used in the present invention includes Guangzhou subway network topology data, Guangzhou subway AFC data, historical full load rate data, station platform waiting time data, train timetable data, and path set data.

[0079] The road network is defined by using interval attributes and path attributes. The interval attributes include the interval running time and the interval congestion situation, and the path attributes include: the path travel time, the path waiting time, the number of path transfers, and the path congestion degree.

[0080] The mathematical model formula of the six-tuple defining the passenger travel process is as follows:

[0081] x m =(id,t in ,t out ,st o ,st d ,r,tp)

[0082] In the formula, m represents a certain passenger, id represents the passenger ID, which is used to identify the passenger; t in —— the boarding time; t out —— the alighting time; st o —— the starting station; st d —— the terminal station; r represents that the travel path of this passenger is the r-th feasible path of this OD; tp represents the travel period.

[0083] The passenger portrait label system includes: direct attributes and indirect attributes.

[0084] The direct attributes mainly include the card type, the number of trips, the travel OD (Origin to Destination) distribution, the travel time distribution, and the travel path distribution.

[0085] (1) Travel time distribution

[0086] The travel time distribution refers to the frequency statistics of passengers traveling in each period. The travel time distribution can reflect the travel rules of passengers to a certain extent. Commuting passengers mostly travel during the morning and evening rush hours, so their travel frequencies during the morning and evening rush hours are relatively high. The travel frequency f of passenger c in period τ t c,τ is calculated as follows. The travel frequencies f in each period t c,τ constitute the travel time distribution set F t c .

[0087]

[0088]

[0089] In the formula, |A| represents the total number of elements in set A, and the same applies hereinafter;

[0090] —— The set of travel records of passenger c during time period τ;

[0091] X c —— The set of all travel records of passenger c.

[0092] (2) Travel OD Distribution

[0093] The travel OD distribution refers to the frequency statistics of passengers at each travel OD, which can reflect the characteristics of passengers' work and residence locations. The higher the travel frequency of a certain OD, the closer the passenger's workplace and residence are to the OD. The travel OD distribution can be further divided into the origin station distribution and the destination station distribution. The higher the travel frequency with a certain station as O or D, the closer the passenger's residence or workplace is to that station. The calculation formula for the travel OD distribution of passenger c is as follows:

[0094]

[0095]

[0096] In the formula —— The travel frequency of passenger c with (o, d) as the travel OD;

[0097] —— The set of travel records of passenger c with (o, d) as the travel OD.

[0098] The travel frequencies of each OD constitute the travel OD distribution set which only contains the top three ODs sorted in descending order of travel frequency.

[0099]

[0100]

[0101] In the formula —— The travel frequency of passenger c with o as the origin station;

[0102] —— The set of travel records of passenger c with o as the origin station.

[0103] The travel frequencies of each origin station constitute the origin station distribution set which only contains the top three origin stations sorted in descending order of travel frequency.

[0104]

[0105]

[0106] In the formula —— The travel frequency of passenger c with d as the terminal

[0107] —— The set of travel records of passenger c with d as the terminal

[0108] The travel frequencies of each terminal constitute the terminal distribution set It only includes the top three terminals sorted in descending order of travel frequency.

[0109] (3) Travel path distribution

[0110] The travel path distribution refers to the frequency statistics of passengers on each travel path, which can reflect the travel patterns of passengers. The calculation formula is as follows:

[0111]

[0112]

[0113] where f r c,o,d,k —— The travel frequency of passenger c with the k-th path from the starting station o to the terminal d as the travel path;

[0114] —— The set of travel records of passenger c with the k-th path from the starting station o to the terminal d as the travel path.

[0115] The travel frequencies of each travel path constitute the travel path distribution set It only includes the top three paths sorted in descending order of travel frequency.

[0116] The indirect attributes include the average number of transfers, average travel time, average congestion degree of travel paths, average waiting time, and label preference distribution.

[0117] (1) Average number of transfers

[0118] The average number of transfers measures the number of transfers of passengers during travel. The calculation formula is as follows:

[0119]

[0120] where —— The average number of transfers of passenger c;

[0121] —— The number of transfers of the k-th path from the starting station o to the terminal d.

[0122] (2) Average travel time

[0123] The average travel time represents the average travel duration of passengers, and the calculation formula is as follows:

[0124]

[0125]

[0126] In the formula —— The average travel time of passenger c

[0127] —— The travel time of the k-th path from the origin station o to the destination station d, which is the difference between the departure time and the arrival time.

[0128] (3) Average crowding degree of travel path

[0129] The average crowding degree of the travel path can measure the crowding degree of each passenger's travel and is represented by the proportion of crowded intervals in the path. Usually, the full-load rate of the interval is used to represent the crowding degree of the interval. According to the value of the full-load rate of the interval, the crowding is divided into four categories: uncrowded (below 80%), slightly crowded (80% - 100%), moderately crowded (100% - 120%), and severely crowded (above 120%). Considering the different impacts of different crowding degrees, the present invention proposes to use the weighted proportion of crowded intervals as a measure of the crowding degree of the travel path, and the calculation formula is as follows:

[0130]

[0131]

[0132]

[0133] In the formula —— The weighted proportion of crowded intervals of the k-th path from the origin station o to the destination station d at time ω;

[0134] t ij —— The running time of each interval of the path;

[0135] λ —— The weight corresponding to the interval crowding degree;

[0136] —— The full-load rate of the interval at time ω;

[0137] —— The average crowding degree of passenger c's travel path.

[0138] (4) Average waiting time

[0139] The average waiting time measures the waiting duration of passengers at the starting station and transfer stations for each travel, and the calculation formula is as follows:

[0140]

[0141] where —— the average waiting time of passenger c;

[0142] —— the platform waiting time of the k-th path from the origin station o to the destination station d at time ω.

[0143] The described label preference distribution reflects the factors considered by passengers when choosing a path, including path travel time, number of path transfers, path waiting time, and path congestion level; the value of the label is 0 or 1, where 0 indicates that the path does not have the label feature and 1 indicates that the path has the label feature, depicting the attribute features of the travel path. The calculation formula is as follows:

[0144]

[0145]

[0146]

[0147]

[0148]

[0149] where —— the path attribute of path i at t in time;

[0150] —— the path attributes of all feasible paths composed of in at t time;

[0151] —— the travel time of path i, which is the sum of the interval running time and the path waiting time;

[0152] —— the waiting time of path i at t in time, which consists of the platform waiting time at the starting station and the platform waiting time at the transfer station and can be obtained from the station platform waiting time table;

[0153] —— the number of transfers of path i, which can be obtained from the path set table;

[0154] —— the congestion level of path i at t in time;

[0155] —— the set of attribute labels of path k at t in time;

[0156] —— The "least travel time" label value of path k at time t in ;

[0157] —— The "least waiting time" label value of path k at time t in ;

[0158] —— The "least transfers" label value of path k at time t in ;

[0159] —— The "least crowded" label value of path k at time t in ;

[0160] Calculate the preference degree of passengers for each label to obtain their label preference degree set. The preference degree index is the Target Group Index (TGI), which can reflect the preference of passengers for a certain feature. The larger the TGI value, the stronger the preference of passengers for a certain feature; TGI greater than 100 indicates that passengers have a strong preference for a certain feature, equal to 100 indicates that passengers' preference for a certain feature is at the average level, and less than 100 indicates that passengers have a weak preference for a certain feature. The calculation formula of TGI is as follows:

[0161]

[0162] Step S20: Propose a passenger preference recognition method based on the spectral clustering algorithm to refine passengers' travel information from two dimensions: the OD level and the time period level.

[0163] The described OD-based passenger preference recognition method refines passengers' travel information from two dimensions: the OD level and the time period level, and uses spectral clustering to cluster passengers to identify and extract passengers' travel preferences. The method flow chart is as Figure 2 shown.

[0164] The calculation steps of the OD-based passenger preference recognition method are as follows:

[0165] Step 1: Divide the travel record set X of all passengers according to OD, and screen out the passenger travel data X o,d with a certain OD as the travel OD to achieve the division of the spatial dimension. X o,d The calculation formula is as follows:

[0166] X o,d ={x m |st o =o, st d =d}

[0167] Step 2: Divide X o,d by travel time periods to obtain a subset X o,d,τ of period τ, achieving division in the time dimension. X o,d,τ The calculation formula is as follows:

[0168] X o,d,τ = {x m | st o = o, st d = d, tp = τ}

[0169] Step 3: Divide X o,d,τ by passengers to obtain subsets of different passengers at different OD and different time periods The calculation formula is as follows. Select passengers with more than 3 trips, calculate the indirect attributes of each passenger at different OD and different time periods, and obtain the travel individual characteristic attributes of passengers at time period τ of (o, d) Form a set U of travel individual characteristic attributes of passengers at different OD and different time periods o,d,τ .

[0170]

[0171]

[0172] Step 4: Perform spectral clustering on U o,d,τ and determine the optimal clustering clusters C according to the silhouette coefficient and CH score. The calculation formula is as follows:

[0173] C = {C1, C2,..., C k}

[0174] Step 5: Determine the passenger preferences of each category according to the clustering center c k of each category of passengers. The formula is as follows:

[0175]

[0176] In the formula, Q k is the passenger preference reflected by the clustering center c k . The passenger preference Q c of passenger c is the passenger preference presented by the clustering center of the category to which it belongs

[0177] Step S30: Considering the sorting of passenger preference degrees, construct a lexicographic preference path selection model based on JND to solve the matching problem between passengers and paths

[0178] When the JND-based lexicographic preference model compares paths, it compares them in turn according to the preference degrees of passengers for different attributes. First, it compares the attribute with the highest preference degree among the paths, and believes that the optimal path of this attribute better meets the needs of passengers; if the attribute with the highest preference degree cannot distinguish different paths, then it compares the second-highest attribute, and so on. Through this sequential comparison, the path that better meets the passengers' preferences can be recommended to passengers. The lexicographic preference model does not need to assign weights to attributes, but takes meeting the most important attribute as the main matching criterion.

[0179] The present invention believes that there is a JND threshold for passengers' perception of different attributes of paths. Suppose the travel time of path 1 is 20 minutes, the travel time of path 2 is 24 minutes, the travel time of path 3 is 40 minutes, and the JND threshold is 5 minutes. When passengers make comparisons, since the difference in travel time between path 2 and path 1 is less than the JND threshold, it is considered that there is no difference in travel time between path 1 and path 2. However, the difference in travel time between path 3 and path 1 is greater than the JND threshold, so it is considered that there is a difference in travel time between path 1 and path 3, and based on the travel time judgment, path 1 is better than path 3.

[0180] Let be the perceivable change ratio of the stimulus amount of the passenger preference attribute q i That is, if the difference ratio of the attribute q i of two paths is less than then the difference in this attribute between these two paths is not perceived, and then these two paths are considered to have no difference in this attribute. The calculation formula is as follows:

[0181]

[0182] In the formula, is the optimal value of the attribute q i , is in percentage form. The above formula means that the perceivable change ratio of the passenger for the attribute q i is The better value of this attribute in path and path is When passengers compare these two paths, if the difference between the two paths is within the range of , passengers will think that the performances of these two paths on the attribute q i are the same, and either path can be chosen.

[0183] Model assumptions of the JND-based lexicographic preference model:

[0184] Suppose the known lexicographic preference set of passenger c for a certain OD pair (o, d) is Qc The path attribute set of OD at time ω

[0185] Assume that passengers' perception of attribute changes follows Weber's law and there is a JND threshold. Assume that the set of perceivable change ratios of attribute changes for passengers is β, and the calculation formula is as follows:

[0186]

[0187] In the formula —— The perceivable change ratio of the preference attribute q i ;

[0188] —— The perceivable change ratio of the travel time attribute;

[0189] —— The perceivable change ratio of the transfer times attribute;

[0190] —— The perceivable change ratio of the waiting time attribute;

[0191] —— The perceivable change ratio of the path congestion degree attribute.

[0192] Randomly continuousize the β set and use the DDPG algorithm to automatically explore the parameters. The model solving objective is to determine the optimal path l that satisfies passengers' preferences best . The model uses as the judgment condition to compare and sort the path set L of OD o,d , and obtain the ordered path set that satisfies passengers' preferences The path ranked higher in satisfies passengers' preferences more. Therefore the path ranked first in best is l

[0193] The lexicographic preference model framework based on JND is as Figure 3 shown, and the model calculation process is as follows:

[0194] Step1: Input the ordered preference set of passengers, the perception change rate, and the path set matrix of OD, and initialize the recommended path set and the temporary path set. Compare the attribute values in the feasible path set in order of preference, that is, traverse the matrix row by row from top to bottom.

[0195] Step2: Sort from small to large and calculate the upper limit of the perceivable difference for passengers If the temporary path set is not empty, then the paths greater than the upper limit are taken out and stored in the recommended path set in descending order

[0196] ; otherwise, jump to step 4.

[0197] Step3: After traversing the matrix, if is empty, directly jump to step 4; otherwise, take out the paths in and store them in in descending order;

[0198] Step 4: is an ordered path set. The later the sorting, the more it meets the passenger preferences. Reverse to get an ordered path set that meets the passenger preferences in is the matching path l best , and is output as the model calculation result and recommended to the passenger.

[0199] Step S40: Considering the sensitivity differences of different passengers to path attributes, the parameters of the passenger path selection model are optimized by combining the DDPG reinforcement learning algorithm to improve the accuracy of passenger personalized guidance.

[0200] Basic elements of the DDPG reinforcement learning algorithm:

[0201] (1) State

[0202] Since the effect of parameter optimization is directly reflected in the change of the accuracy of the path recommended to the passenger, in the present invention, the state of the model is the accuracy rate δ of the recommendation, and the calculation formula is as follows. The state space is 1.

[0203]

[0204] In the formula, N——total number of trips of the passenger;

[0205] N1——number of trips in which the path recommended to the passenger is the same as the actual travel path of the passenger.

[0206] (2) Action

[0207] The action in the present invention is the set of perceived change ratios β of the lexicographic preference model based on JND, and the calculation formula is as follows. Since there are four parameters in β, the action space is 4.

[0208]

[0209] (3) Reward function

[0210] The calculation formula of the reward function is as follows. After taking a certain action, if the accuracy rate increases, the reward function is positive, and the agent will receive positive feedback and learn in this direction; if the accuracy rate decreases, the reward function is negative, and the agent will receive negative feedback and avoid learning in this direction.

[0211] r = δ′ - δ

[0212] In the formula, δ —— the accuracy rate before taking a certain action;

[0213] δ′ —— the accuracy rate after taking a certain action.

[0214] The present invention intends to randomize and serialize the β set, and use the DDPG algorithm to learn and calibrate the parameters. The core of the DDPG algorithm lies in using a deep neural network to simulate a function and training it using deep learning methods. The principle of the DDPG algorithm is as Figure 4 shown, and it is mainly divided into two modules, the Actor module and the Critic module.

[0215] The Actor module is responsible for updating the policy function and selecting actions, and can also be called the policy network. Its goal is to determine an optimal policy to maximize the cumulative reward. The parameters in the Actor network are defined as θ μ , and the calculation formula of the objective function is as follows:

[0216]

[0217] In the formula, Q μ (s,a) —— the expected value of the cumulative reward;

[0218] ρ β —— the distribution function of state s;

[0219] a t —— the action of the agent at time t;

[0220] s t —— the environmental state at time t;

[0221] r(s t ,a t ) —— the reward value obtained after state s t executes a t ;

[0222] γ —— the decay coefficient of the reward value of the next state, γ ∈ [0,1].

[0223] The Critic module is responsible for evaluating the current policy and outputting the Q function, that is, Q μ (s,μ(s)) in the Actor network. The parameters of the Critic network are defined as θQ , whose goal is to minimize the loss function of the network. The calculation formula of the objective function is as follows:

[0224] y t = r t + γQ′(s t+1 | μ′)| θ μ′ )

[0225]

[0226] In the formula, y i —— target Q value;

[0227] L —— loss function of the Critic network, which is the mean square error of the predicted Q value and the target Q value.

[0228] Combined with the goals and update processes of the Actor network and the Critic network, the goal of the DDPG algorithm is to maximize J β (μ) and minimize L. To improve the stability and efficiency of the algorithm, DDPG creates two neural networks for the Actor network and the Critic network respectively, which are responsible for training calculation and parameter update.

[0229] The schematic diagram of the DDPG reinforcement learning algorithm is as Figure 4 shown, and its algorithm process is as follows:

[0230] Step 1: Randomly initialize the parameters θ μ and θ Q of the main network, as well as the parameters θ μ′ and θ Q′ of the target network, where θ μ′ = θ μ , θ Q′ = θ Q ; Then initialize the sample storage buffer R and give a preset number of iterations.

[0231] Step 2: Initialize a random noise ε t and obtain the current state s t ;

[0232] Step 3: Select an action a t according to the output of the actor network and the noise ε t , and the calculation formula is as follows; The environment executes the action a t , obtains the reward r t and the new state s t+1 , and the actor network will (s t , a t , r t , s t+1) is stored as a set of data in the sample storage buffer R as a data set for training the network.

[0233] a t =μ(s t |θ μ )+ε t

[0234] Step 4: Randomly sample N groups (s t , a t , r t ,s t+1 ) data as the training data of the actor main network and the critic main network, calculate the gradient of the critic main network, and update the main network. The update formula is as follows:

[0235]

[0236] Where L is the loss function of the Critic network, which is the mean square error between the predicted Q value and the target Q value. Calculate the gradient of the actor main network and update the main network. The update formula is as follows:

[0237]

[0238] In the formula The gradient of s according to ρ β When distributed, expected value.

[0239] Step 5: Update the target network parameters. The calculation formula is as follows:

[0240] θ Q′ =τθ Q +(1-τ)θ Q′

[0241] θ μ′ =τθ μ +(1-τ)θ μ′

[0242] Where τ is the update coefficient, which is 0.01 in the present invention. When the number of iterations is reached, the iteration ends, otherwise it jumps to Step 2.

[0243] Embodiment 2

[0244] Taking the passenger with the card number 1759753155543242 as an example, the JND-based lexicographic preference model is used to perform individual-based passenger induced route recommendation instance analysis.

[0245] According to the travel information, using the above-mentioned passenger portrait establishment method, calculate the indicators of direct attributes and indirect attributes respectively, and obtain the passenger portrait as shown in Table 1 below:

[0246]

[0247] Table 1

[0248] Select the parameters of the lexicographic preference model based on JND, analyze the passengers' perception of travel time changes from three aspects: travel time, transfer times, and waiting time, and calculate the attributes of each path on this basis. For this passenger, during the morning rush hour, the travel preference ranking for this OD is travel time > waiting time. Therefore, when comparing, first compare the travel time, and then compare the waiting time. The calculation results of the lexicographic model based on JND are shown in Table 2 below. Analyzing the perceivable differences of the passenger for the three types of path attributes, it can be seen that Path 1 better meets his travel preferences.

[0249]

[0250] Table 2

[0251] Calculate all travel records using the same method and calculate the accuracy rate. The calculation results are shown in Table 3 below. The results show that among 30 travel records, 8 paths recommended are inconsistent with the passengers' actual travel, and the accuracy rate is 73%. The lexicographic preference model based on JND takes into account the preference degrees of passengers for different influencing factors during the path selection process, and is more in line with the decision-making process of passengers when choosing paths.

[0252]

[0253] Table 3

[0254] Based on the above example, briefly describe the calculation process of parameter optimization for the DDPG reinforcement learning algorithm. First, select each parameter required for the DDPG reinforcement learning algorithm, and the calculated optimization results are shown in Table 4 below. The results show that after the DDPG parameter optimization, among 30 travel records, 2 paths recommended are inconsistent with the passengers' actual travel, and the algorithm accuracy rate is 93.33%. For the current example, the accuracy rate of the model with optimized parameters has increased by 20% compared to the model without optimized parameters.

[0255]

[0256] Table 4

[0257] In summary, in an embodiment of the present invention, an induced path recommendation method and system based on the least perceptible difference passenger preference ranking verifies through an example that the induced path recommendation method and system based on the least perceptible difference passenger preference ranking proposed by the present invention can truly and effectively provide theoretical and technical references for improving the rail transit passenger information service and operation management level, and also provide certain help for the sound operation of urban rail transit.

[0258] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0259] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0260] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0261] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement without creative labor.

[0262] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An induced path recommendation method based on the just noticeable difference of passenger preference ranking, characterized in that, Including: Obtain passenger travel information, analyze the passenger travel process, and construct a passenger portrait label system including direct attributes and indirect attributes; Refine passenger travel information from two dimensions of the origin-destination (OD) level and the time period level based on the passenger portrait label system, and use the spectral clustering method to extract passenger travel preferences; Combine the preference ranking of passenger travel, construct a lexicographic preference path selection model based on the just noticeable difference (JND), obtain a path that meets the passenger's preferences, and recommend it to the passenger; Set the mathematical model formula of the six-tuple defining the passenger travel process as: x m =(id,t in ,t out ,st o ,st d ,r,tp) Where m represents a certain passenger, id represents the passenger ID for identifying the passenger; t in —— the entry time; t out —— the exit time; st o —— the starting station; st d —— the terminal station; r represents the r-th feasible path of the passenger's travel route from the starting point to the ending point; tp - Travel time period; The passenger portrait label system includes: direct attributes and indirect attributes; The refining of passenger travel information from two dimensions of the origin-destination (OD) level and the time period level based on the passenger portrait label system, and using the spectral clustering method to extract passenger travel preferences includes: The travel record set X of all passengers is divided according to their origin-destination OD, and the travel data X of passengers with a certain OD as the travel OD is screened out, o,d , realizing the division in the spatial dimension, X o,d The calculation formula is as follows: X o,d = {x m | st o = o, st d = d} Divide X o,d into subsets according to travel time periods, obtaining subsets X o,d,τ of period τ, achieving the division in the time dimension. X o,d,τ The calculation formula is as follows: X o,d,τ = {x m | st o = o, st d = d, tp = τ} Divide X o,d,τ According to passengers, subsets of different passengers at different OD and different time periods are obtained The calculation formula is as follows: Filter out passengers with more than 3 trips, calculate the indirect attributes of each passenger under different OD and different time periods, and obtain the travel individual characteristic attributes of the passenger at time period τ of (o, d). Form the set U of travel individual characteristic attributes of passengers for different OD and different time periods o,d,τ , for U o,d,τ Perform spectral clustering, and determine the optimal clustering cluster C according to the silhouette coefficient and CH score. The calculation formula is as follows: C = {C1, C2,..., C k} Based on the cluster center c of each type of passenger k , determine the passenger preferences of each type, and the formula is as follows: where Q k is the passenger preference reflected by the clustering center c k The passenger preference Q of passenger c c is the passenger preference presented by the clustering center of the category to which it belongs.

2. The method according to claim 1, characterized in that, The obtaining of passenger travel information, analyzing the passenger travel process, and constructing a passenger portrait label system including direct attributes and indirect attributes includes: Define the road network using interval attributes and path attributes. The interval attributes include interval running time and interval congestion situation, and the path attributes include: path travel time, path waiting time, number of path transfers, and path congestion level; The direct attributes include card type, number of trips, OD distribution of trips, time distribution of trips, and path distribution of trips. The indirect attributes include average number of transfers, average travel time, average path congestion level, average waiting time, and label preference degree distribution.

3. The method according to claim 1, wherein The combining of the preference ranking of passenger travel, constructing a lexicographic preference path selection model based on the JND, obtaining a path that meets the passenger's preferences, and recommending it to the passenger includes: Set the JND threshold for the passenger's perception of different attributes of the path. When the difference in travel time between multiple paths is less than the JND threshold, it is considered that there is no difference in travel time among these multiple paths; when the difference in travel time between multiple paths is greater than the JND threshold, it is considered that there are differences in travel time among these multiple paths; Let be the perceivable stimulus quantity change ratio of passenger preference attribute q i That is, if the difference ratio of attribute q i of two paths is less than then the difference in this attribute between these two paths is not perceived, and these two paths are considered to have no difference in this attribute. The calculation formula is as follows: Wherein, is the optimal value of attribute q i , in percentage form, the perceivable change ratio of the passengers for attribute q is i . The better values of this attribute in path and path are . When passengers are comparing these two paths, if the difference between the two paths is within the range of , passengers will consider that the performances of these two paths in attribute q are the same, and either path can be chosen; i ​ Model assumptions of the lexicographic preference path selection model based on JND: Let the known lexicographic preference set of passenger c for a certain OD pair (o, d) be Q c , and the path attribute set of OD at time ω Assume that the passenger's perception of attribute changes follows Weber's law and there is a JND threshold. Assume that the set of perceivable change ratios of the passenger's perception of attribute changes is β, and the calculation formula is as follows: where —— the perceivable change ratio of the preference attribute q i ; ——Perceivable change ratio of travel time attribute; ——Perceivable change ratio of the transfer times attribute; —— Perceivable change ratio of waiting time attribute; ——Perceivable change ratio of path congestion degree attribute; Perform random continuousization on the β set and use the Deep Deterministic Policy Gradient (DDPG) algorithm to automatically explore the parameters; Obtain the passenger preference degree ranking according to the passenger travel preferences, and construct a lexicographic preference path selection model based on the JND according to the passenger preference degree ranking. The calculation process of the lexicographic preference path selection model based on the JND is as follows: Step1: Input the ordered preference set of the passenger, the perception change rate, and the path set matrix of the OD. Initialize the recommended path set and the temporary path set, and compare the attribute values in the feasible path set in sequence according to the preference order, that is, traverse the path set matrix of the OD row by row from top to bottom; Step 2: Take Sort them in ascending order and calculate the upper limit of the perceptible difference for passengers If the temporary path set is not empty, then Take out the paths in it that are greater than the upper limit and store them in the recommended path set in descending order; otherwise, jump to step 4; Step 3: After traversing the path set matrix of OD, if is empty, directly jump to step 4; otherwise, take out the paths in descending order and store them in in turn; Step 4: is an ordered path set. The later the sorting, the more it meets the passenger preferences. Perform a reverse reordering to obtain an ordered path set that meets the passenger preferences in which is the matching path l best , and serve as the path that meets the passenger preferences and is recommended to the passenger.

4. The method according to claim 1, wherein The method further includes: considering the sensitivity differences of different passengers to path attributes, and optimizing the lexicographic preference path selection model based on the JND by combining the DDPG reinforcement learning algorithm, including: A: Randomly initialize the main network parameters θ μ and θ Q as well as the target network parameters θ μ′ and θ Q′ , where θ μ′ = θ μ , θ Q′ = θ Q ; then initialize the sample storage buffer R and given a preset number of iterations; B: Initialize a random noise ε t and obtain the current state s t ; C: Based on the output of the actor network and the noise ε t Select action a t , the calculation formula is as follows; the environment performs action a t , get reward r t and the new state s t+1 , the actor network will (s t ,a t ,r t ,s t+1 ) is stored as a set of data in the sample storage buffer area R as a data set for training the network; a t = μ(s t |θ μ ) + ε t D: Randomly sample N groups of (s t , a t , r t , s t+1 ) data from R as the training data for the actor main network and the critic main network, calculate the gradients of the critic main network, and update the main network; E: Update the target network parameters, and the calculation formula is as follows: θ Q′ = τθ Q + (1 - τ)θ Q′ θ μ′ = τθ μ +(1 - τ)θ μ′ In the formula, τ is the update coefficient, with a value of 0.

01. After reaching the number of iterations, the iteration ends; otherwise, jump to B.

Citation Information

Patent Citations

  • Door-to-door travel path scheme personalized recommendation method based on heuristic search

    CN106203725A

  • Method and system for precisely inducing passenger flow in multiple scenes of urban rail transit

    CN110428117A