Multi-agent cooperative reasoning and restoring method and system for inter-city highway-railway multi-type travel route

A multi-agent system decouples path search and mode recognition tasks using reinforcement learning to accurately reconstruct intercity travel paths from sparse data, enhancing travel pattern analysis.

CN120317469AActive Publication Date: 2025-07-15SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510481743.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-15
Estimated Expiration
2045-04-17

Smart Images

  • Figure CN120317469A_ABST
    Figure CN120317469A_ABST
Patent Text Reader

Abstract

The invention discloses an inter-city highway-railway multi-mode travel path multi-agent cooperative reasoning and restoring method and system. Based on inter-city travel sparse trajectory observation data of travelers, reasoning and restoring of real travel paths of the travelers are achieved. The method comprises the following steps: acquiring and preprocessing original track observation data and traffic network data; rasterizing and storing the road network data and the track observation data; designing and training a path search agent, and realizing reasoning of a reasonable restoration path sequence; designing and training a mode identification agent to realize reasoning of a reasonable travel mode sequence; and the two trained intelligent agents perform cooperative reasoning to realize the restoration of the real path of the inter-city highway-railway multi-type travel. According to the method, a complex inter-city highway-railway multi-type travel path restoration problem is decoupled into two sub-problems of path search and mode identification, and two intelligent agents are constructed to respectively solve the two sub-problems, so that reasoning restoration of an inter-city travel real path based on a sparse observation trajectory is realized; and a new thought and a new method are provided for the field of path restoration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and intelligent transportation, and specifically to a multi-agent collaborative reasoning and restoration method and system for intercity public-rail multi-modal travel paths. Background Technique

[0002] With the rapid development of intelligent transportation systems and the construction of China's modern transportation network, the comprehensive transportation network has become increasingly perfect. The railway network, highway network, national and provincial trunk line network, and low-grade roads together constitute a multi-level intercity travel traffic network structure. The intercity travel paths of residents are gradually increasing, and travel modes are becoming more diverse. In this context, in-depth understanding of residents' intercity travel patterns is of great significance for improving the efficiency of transportation resource allocation. However, the existing observed data of residents' intercity travel trajectories often have problems such as sparse spatio-temporal sampling granularity and limited spatial accuracy, which cannot support the accurate identification of residents' travel modes and travel paths, bringing great challenges to the in-depth analysis of residents' intercity travel patterns.

[0003] Up to now, many scholars have carried out research on path inference and restoration, aiming to recover the real travel paths of residents from sparse and biased trajectory observation data. Traditional path restoration methods based on hidden Markov models rely on the Markov assumption and model the path restoration process as a double stochastic process for path recovery. However, its computational complexity increases exponentially with the length of the travel path. Facing the characteristics of current residents' intercity travel, such as long trips, wide spatial scope involved, and diverse feasible paths, its path restoration ability is significantly limited. Supervised learning methods are another mainstream path restoration method. This type of method decouples multi-dimensional features from labeled data and directly learns the high-dimensional mapping relationship between observed trajectories and real paths to achieve path restoration. However, the learning effect of its model depends too much on the quality and quantity of labeled data, and the model generalization and migration ability are also affected by the distribution of labeled data. Semi-supervised learning and unsupervised learning path restoration methods have lower requirements for labeled data, but relevant research is still less at present and lacks corresponding theoretical support.

[0004] Considering the limitations of existing path restoration methods, the present invention aims to explore a new path restoration method that can adapt to the characteristics of current residents' intercity travel, and on the basis of having a high path restoration accuracy, minimize the cost as much as possible, so as to provide efficient technical support for the intelligent transportation system.

[0005] Differences compared with the prior art are as follows:

[0006] Technical comparison with patent CN118966349A "A multi-agent future behavior topology inference method, device, equipment, medium and product for autonomous driving"

[0007] Patent CN118966349A proposes a method, device, equipment, medium and product for inferring the future behavior topology of autonomous driving multi - agents, belonging to the field of autonomous driving technology; this patent proposes a method and system for collaborative inference and restoration of multi - agent inter - city public - rail multi - modal travel paths, belonging to the fields of artificial intelligence and intelligent transportation. There are essential differences in the technical fields involved.

[0008] Patent CN118966349A uses a collaborative learning model to iteratively complete the prediction of the future behavior of autonomous driving multi - agents based on a preset fitness function; this patent constructs a multi - agent model and completes the inference and restoration of real travel paths through feedback learning based on the output results of the agents. There are essential differences in the methods adopted.

[0009] Patent CN118966349A constructs a method, device, equipment, medium and product for inferring the future behavior topology of autonomous driving multi - agents, focusing on solving the uncertainty and heterogeneous interaction problems in multi - agent driving scenarios; this patent proposes a method and system for collaborative inference and restoration of multi - agent inter - city public - rail multi - modal travel paths, mainly solving the problem of path restoration under sparse observation data of inter - city travel trajectories. There are essential differences in the purposes and functions.

[0010] Technical comparison with patent CN102904343A "State monitoring system and method based on distributed multi - agents"

[0011] Patent CN102904343A proposes a state monitoring system and method based on distributed multi - agents, belonging to the field of substation state detection technology; this patent proposes a method and system for collaborative inference and restoration of multi - agent inter - city public - rail multi - modal travel paths, belonging to the fields of artificial intelligence and intelligent transportation. There are essential differences in the technical fields involved.

[0012] Patent CN102904343A uses multiple homogeneous state - detection agents to achieve collaborative detection of substation states based on a multi - agent network; this patent constructs two types of agents, namely path search and mode identification, and completes the inference and restoration of real travel paths through feedback learning based on each other's output results. There are essential differences in the methods adopted.

[0013] Patent CN102904343A proposes a state monitoring system and method based on distributed multi - agents, mainly used to solve the defects of isolated state monitoring and facilitate the organic integration of online monitoring of the entire power grid; this patent proposes a method and system for collaborative inference and restoration of multi - agent inter - city public - rail multi - modal travel paths, mainly solving the problem of path restoration under sparse observation data of inter - city travel trajectories. There are essential differences in the purposes and functions. Summary of the Invention

[0014] In the face of the characteristics of long travel distances, wide spatial scope involved, and diverse feasible paths in the current intercity travel of residents, the present invention proposes a multi-agent collaborative reasoning and restoration method and system for intercity public rail multi-modal travel paths, which jointly realizes the reasoning and restoration of the real paths of travelers' intercity travel through the learning and training of multiple agents, reduces the dependence on high-quality labeled samples while maintaining a high path restoration accuracy rate.

[0015] To achieve the above object, the technical solution adopted by the present invention is:

[0016] A multi-agent collaborative reasoning and restoration method for intercity public rail multi-modal travel paths, comprising the following steps:

[0017] Step S10, obtaining and preprocessing the original trajectory observation data and traffic network data;

[0018] Step S20, rasterizing and storing the road network data and trajectory observation data;

[0019] Step S30, designing and training a path search agent to infer a reasonable restored path sequence given a trajectory observation sequence and the road network corresponding to a specific transportation mode;

[0020] Step S40, designing and training a mode identification agent to infer a reasonable travel mode sequence given a trajectory observation sequence and the road matching degree feedback by the path search agent;

[0021] Step S50, the path search agent and the mode identification agent are trained asynchronously and reason collaboratively, and jointly realize the reasoning and restoration of the real path corresponding to the sparse trajectory observation sequence after the reinforcement learning training is completed.

[0022] As a further improvement of the present invention, the step S10 specifically includes:

[0023] Step S11, obtaining and preprocessing the trajectory observation data;

[0024] Obtain the trajectory observation data containing the intercity travel information of travelers through a specific software or mobile phone operator. The obtained data should at least include timestamp information and longitude and latitude information. Then, preprocess the original trajectory observation data of the residents' intercity travel, including data filling and duplicate removal operations;

[0025] Step S12, obtaining and preprocessing the road network data;

[0026] Obtain the traffic network data of the research area. This data should at least include the location information of each road. Based on the type of the road, divide the overall road network data into multiple subsets, which respectively correspond to different transportation modes and contain the complete road network data of the corresponding transportation modes. Then, screen out the road data subset involved in intercity travel for use.

[0027] As a further improvement of the present invention, the step S20 specifically includes:

[0028] Step S21, rasterizing the study area;

[0029] The study area is divided into grids of appropriate size using ArcGIS software, and the grids are arranged in rows and columns in the form of (locx i ,locy i ) in the form of numbering, where locx i is the column number of the grid, locy i The row number of the grid;

[0030] Step S22, rasterizing the road network data and storing it;

[0031] The road networks of various modes of transportation are mapped to numbered grids through ArcGIS software. If there is a road corresponding to a certain mode of transportation in the grid, the corresponding value is 1, otherwise it is 0. After all are completed, the traffic network information of each grid itself is a string of numbers with a length of the number of types of transportation and a value of 1 or 0. In order to reduce storage space, the 1 / 0 string is converted into a decimal number. At this point, the traffic network information of each grid itself is completely stored. When the traffic network information of each grid is complete, the traffic network information of the surrounding eight grids is retrieved and stored, that is, each grid finally stores the traffic network information of 9 grids including itself. The order of the grids is: upper left, directly above, upper right, directly left, itself, directly right, lower left, directly below, and lower right.

[0032] Step S23, rasterizing and storing the trajectory observation data;

[0033] The location information of the current trajectory observation data is still expressed as longitude and latitude. All location information in the trajectory observation data is mapped to the grid through ArcGIS software, and converted from longitude and latitude storage to storage in the rows and columns of the grid. The converted trajectory observation data includes the ID, timestamp, column, and row corresponding to each trajectory observation point.

[0034] As a further improvement of the present invention, the step S30 specifically includes:

[0035] Step S31, designing a path search agent;

[0036] At the tth time step in the path restoration process, the path reasoning agent performs a Markov decision process according to the strategy, starting from the current state s t Take action a t Transfer to the next state and obtain the reward r at this momentT (s t , a t ), where the state s obtained by the agent t includes the road network information corresponding to the given travel mode, the location information of the agent at this moment, and the location information of the target trajectory point; the action a taken by the agent t is one of the action spaces, and the action space includes moving to the upper left, directly above, upper right, directly left, directly right, lower left, directly below, and lower right; the reward r obtained by the agent T (s t , a t ) is the weighted sum of the following road reward and the approaching end point reward;

[0037] Step S32, training the path search agent;

[0038] The path search agent is trained using the policy-based SAC algorithm. During the training process, the policy is modeled as π by a neural network θ , where θ is the set of parameters of the neural network. Then, the policy parameters θ are learned by the gradient descent method to obtain the optimal policy. The optimization objective of θ is as follows:

[0039]

[0040] Since the single policy network parameters may not be able to adapt to the state differences between different trajectory points, target-oriented reinforcement learning is used to enhance the agent's generalization ability for diverse state spaces. The position information of the agent itself and the target trajectory point in the agent's state is mapped to the relative position information starting from the origin of coordinates, so as to transform the path reasoning problem between different trajectory points into a series of similar tasks starting from the same trajectory point and reaching different target points;

[0041] During the training process, an experience replay pool is also used to improve the training efficiency of the agent: in each round of sampling phase, the state, action, reward, and state of the next time step of the path reasoning agent at the current time step will be stored in the experience replay pool in the form of a tuple <s, a, r, s'>; in the gradient update phase, several data will be randomly sampled from the experience replay pool for gradient update of the policy network, which greatly improves the data utilization efficiency.

[0042] As a further improvement of the present invention, the step S40 specifically includes:

[0043] Step S41, designing a mode identification agent;

[0044] Similar to the path search agent, at the t-th time step in the mode identification process, the mode identification agent performs a Markov decision process according to the policy, starting from the current state s t taking the action at Transfer to the next state and obtain the reward r at this moment M (s t ,a t ). Among them, the state s obtained by the agent t contains the local dynamic and static features of the observation sequence at this moment, the road network information of the agent's location, and the road network information of the grid around its location; the action a taken by the agent t is the identified transportation mode; the reward r obtained by the agent M (s t ,a t ) is set to the product of the road matching degree feedback by the path search agent and the reward obtained according to the motion feature identification method;

[0045] Step S42, train the mode identification agent;

[0046] Different from the policy-based reinforcement learning process of the path inference agent, the mode identification agent is trained using the value-based DQN algorithm. By learning the state-action value function Q(s,a), the optimal selection of traffic mode labels for actions in different states is achieved. In each learning and training process, the update rule of the agent's state-action value function Q(s,a) is:

[0047] Q(s t ,a t )←Q(s t ,a t )+α[r M +γmax a’ Q(s t+1 ,a’)-Q(s t ,a t )]

[0048] To effectively find the optimal Q function, a neural network is used to model it; to analyze the differences in taking different actions in the same state, the state-action value function Q is expressed as the sum of the state value function V and the advantage function A, enabling the agent to better handle action selection in similar states. The following formula is the expression of the state-action value function Q, where V(s t ) is the expected return that the agent can obtain in the state s t , and A(s t ,a t ) is the advantage function of the agent taking different actions a t in the state s t , and ω is the parameter set of the neural network;

[0049] Q ω (s t ,a t )=V(st ) + A(s t , a t )

[0050] To obtain the deep connection between trajectory points and adjacent trajectory points, the state-action value neural network Q of the mode identification agent ω (s t , a t ) replaces the fully connected layer with a long short-term memory network layer. In addition, a target network is introduced during the training process Use a dual neural network for training, and train the network Q ω (s, a) is updated at each step during training, while the target network uses historical parameters and synchronizes with the training network every C steps to avoid the impact of overestimated network valuation on the mode identification performance.

[0051] The multi-agent collaborative inference and restoration system for intercity rail and road multi-modal travel paths of the present invention includes an original data acquisition and preprocessing module, a data grid and storage module, a path search agent design and training module, a mode identification agent design and training module, and a result output module, where:

[0052] The original data acquisition and preprocessing module is used to acquire and preprocess intercity travel trajectory observation data and road network data, so as to provide a data source for subsequent rasterization processing;

[0053] The data grid and storage module is used to rasterize and store the road network data corresponding to different transportation modes and the observation trajectory data containing the true path, which is convenient for the next step of path search and restoration;

[0054] The path search agent design and training module is used to infer a reasonable restoration path sequence based on a given trajectory observation sequence and the road network corresponding to a specific transportation mode, so as to calculate the road matching degree as gain information to support the training of the mode identification agent;

[0055] The mode identification agent design and training module is used to infer a reasonable travel mode sequence based on a given trajectory observation sequence and the road matching degree feedback by the path search agent, so as to provide the correct road network information for the path search agent and support it to infer and restore the true path corresponding to the observation trajectory;

[0056] The result output module is used to jointly complete the inference restoration and result output of its true path by the trained path search agent and mode identification agent after a given trajectory observation sequence.

[0057] Beneficial effects:

[0058] The multi-agent collaborative reasoning and restoration method and system for intercity public railway multi-modal travel paths proposed by the present invention decouple the complex problem of restoring intercity public railway multi-modal travel paths into two sub-problems: path search and mode identification, and solve the sub-problems by constructing two reinforcement learning agents respectively, realizing the efficient reasoning and restoration of the real travel paths of travelers during intercity travel while reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flowchart of the present invention;

[0060] Figure 2 is an architecture diagram of the acquisition and preprocessing of the original data of the present invention;

[0061] Figure 3 is an architecture diagram of the rasterization and storage of the data of the present invention;

[0062] Figure 4 is an architecture diagram of the collaborative training of the path search agent and the mode identification agent of the present invention;

[0063] Figure 5 is the grid division and road network display of the research area in the embodiment of the present invention;

[0064] Figure 6 is the display of the rasterization processing result of the trajectory observation data in the embodiment of the present invention;

[0065] Figure 7 is the display of the training process of the path search agent in the embodiment of the present invention;

[0066] Figure 8 is the display of the training process of the mode identification agent in the embodiment of the present invention;

[0067] Figure 9 is the display of the path restoration result output in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0069] The applicable area of the present invention is wide. When the administrative division information and road network information of the research area are complete, the real travel paths can be inferred and restored based on the sparse trajectory observation data of the intercity travel of the residents in the area. In the following embodiments, Jiangsu Province is used as the research area to realize the inference and restoration of the real travel paths of the residents in the area for intercity travel based on high-speed railways, ordinary railways, highways, and national and provincial roads.

[0070] Embodiment

[0071] As Figure 1As shown in the figure, the multi-agent collaborative reasoning and restoration method and system for the intercity rail-road multi-modal travel path in this embodiment include the following steps:

[0072] Step S10, obtain and preprocess the original trajectory observation data and traffic network data;

[0073] Step S20, perform rasterization processing and storage on the road network data and trajectory observation data;

[0074] Step S30, design and train a path search agent to infer a reasonable restored path sequence after a given trajectory observation sequence and the road network corresponding to a specific traffic mode;

[0075] Step S40, design and train a mode identification agent to infer a reasonable travel mode sequence after a given trajectory observation sequence and the road matching degree feedback by the path search agent;

[0076] Step S50, the path search agent and the mode identification agent are trained asynchronously and collaborate in reasoning. After the reinforcement learning training is completed, they jointly realize the inference and restoration of the real path corresponding to the sparse trajectory observation sequence.

[0077] As Figure 2 shown, the specific steps of step S10 include:

[0078] Step S11, obtain and preprocess the trajectory observation data;

[0079] Obtain the trajectory observation data - mobile phone signaling data containing the intercity travel information of travelers in Jiangsu Province through the mobile phone operator "China Unicom", which includes the time stamp reflecting the time information and the longitude and latitude reflecting the location information. However, due to the large spatio-temporal sampling interval of the mobile phone signaling data, or the signal interference and other abnormal conditions of the data collection device, the mobile phone signaling data often shows sparsity, and there are problems such as data missing and a large amount of duplicate data. Therefore, preprocess the obtained mobile phone signaling data, including data filling, duplicate removal, etc., to improve the usability and consistency of the data;

[0080] Step S12, obtain and preprocess the road network data;

[0081] The transportation network data of Jiangsu Province is obtained from the OpenStreetMap website or other open source resource libraries. The data should at least include the location information of each road. However, the original transportation network data is a collection of all road data in the study area, which is large in volume and contains redundant data. Therefore, based on the type of road, the overall road network data is divided into multiple subsets, corresponding to different modes of transportation, including complete road network data of the corresponding modes of transportation, and road data subsets such as high-speed railways, ordinary railways, highways, national and provincial roads involving intercity travel are selected for use.

[0082] like Figure 3 As shown, the step S20 specifically includes:

[0083] Step S21, rasterizing the study area;

[0084] The Jiangsu province was divided into grids using ArcGIS software. The overall grid size was 529 rows × 564 columns. The horizontal and vertical lengths of the unit grid were 1000 meters. The grids were divided into rows and columns in the form of (locx i ,locy i ) format, where locx i is the column number of the grid, locy i is the row number of the grid. The grid division and road network display of Jiangsu Province can be found in Figure 5 ;

[0085] Step S22, rasterizing the road network data and storing it;

[0086] The road networks of various modes of transportation are mapped to numbered grids through ArcGIS software. If the grid has a road corresponding to a certain mode of transportation, the corresponding value is 1, otherwise it is 0. After all is completed, the transportation network information of each grid is a string of numbers with a length of the number of transportation modes and a value of 1 or 0. In order to reduce storage space, the 1 / 0 digital string is converted into a decimal number. At this point, the transportation network information of each grid is completely stored. When each grid has complete transportation network information corresponding to itself, the transportation network information of the eight surrounding grids is retrieved and stored, that is, each grid finally stores the transportation network information of 9 grids including itself, and the order of the grids is: upper left, directly above, upper right, directly left, itself, directly right, lower left, directly below, and lower right;

[0087] Step S23, rasterizing and storing the trajectory observation data;

[0088] The position information of the current trajectory observation data still appears as longitude and latitude. To facilitate path search and restoration, all position information in the trajectory observation data is mapped to a grid using ArcGIS software, converting the storage from longitude and latitude to the row and column of the grid to which it belongs. After conversion, the trajectory observation data includes the ID, timestamp, column, and row corresponding to each trajectory observation point. The formats of the trajectory observation data before and after rasterization are shown in Figure 6 。

[0089] As Figure 4 shown, step S30 specifically includes:

[0090] Step S31, design a path search agent;

[0091] At the t-th time step during path restoration, the path inference agent performs a Markov decision process according to the policy, starting from the current state s t taking action a t to transfer to the next state and obtaining the reward r T (s t , a t ) at this moment. Among them, the state s t obtained by the agent includes the road network information corresponding to the given travel mode, the agent's current position information, and the position information of the target trajectory point; the action a t taken by the agent is one of the action spaces, and the action space includes moving to the upper left, directly above, upper right, directly left, directly right, lower left, directly below, and lower right; the reward r T (s t , a t ) is the weighted sum of the following-road reward and the approaching-endpoint reward;

[0092] Step S32, train the path search agent;

[0093] The path search agent is trained using the policy-based SAC algorithm. During training, the policy is modeled as π θ by a neural network, where θ is the set of parameters of the neural network. Then, the policy parameters θ are learned through gradient descent to obtain the optimal policy. The optimization objective of θ is as follows:

[0094]

[0095] Since the single policy network parameters may not be able to adapt to the state differences between different trajectory points, target-oriented reinforcement learning is used to enhance the generalization ability of the agent to diverse state spaces, and the position information of the agent itself and the target trajectory point in the agent state is mapped into relative position information starting from the origin of coordinates. Thus, the path inference problem between different trajectory points is transformed into a series of similar tasks starting from the same trajectory point and reaching different target points;

[0096] During the training process, an experience replay pool is also used to improve the training efficiency of the agent: in each round of the sampling stage, the state, action, reward, and state of the next time step of the path inference agent at the current time step will be stored in the experience replay pool in the form of a tuple <s,a,r,s’>; in the gradient update stage, several data will be randomly sampled from the experience replay pool for gradient update of the policy network, which greatly improves the data utilization efficiency. The training process display of the path search agent is shown in Figure 7 .

[0097] Such as Figure 4 shown, the specific steps of step S40 include:

[0098] Step S41, design a mode identification agent;

[0099] Similar to the path search agent, at the t-th time step in the mode identification process, the mode identification agent performs a Markov decision process according to the policy, and takes an action a t from the current state s t to transfer to the next state and obtain the reward r M (s t ,a t ). Among them, the state s t obtained by the agent includes the local dynamic and static features of the observation sequence at this moment (including speed, acceleration, steering angle, etc., calculated based on the time and space information of the trajectory observation sequence), the road network information of the agent's location, and the road network information of the grid around its location; the action a t taken by the agent is the identified traffic mode; the reward r M (s t ,a t ) is set to the product of the road matching degree feedback by the path search agent and the reward obtained by identifying the mode according to the motion characteristics;

[0100] Step S42, train the mode identification agent;

[0101] Different from the policy-based reinforcement learning process of the path inference agent, the mode identification agent is trained using the value-based DQN algorithm. By learning the state-action value function Q(s,a), the optimal selection of traffic mode label actions in different states is achieved. In each learning and training process, the update rule of the agent's state-action value function Q(s,a) is as follows:

[0102] Q(s t ,a t )←Q(s t ,a t )+α[r M +γmax a’ Q(s t+1 ,a’)-Q(s t ,a t )]

[0103] To effectively find the optimal Q function, a neural network is used to model it; to analyze the differences in taking different actions in the same state, the state-action value function Q is expressed as the sum of the state value function V and the advantage function A, enabling the agent to better handle action selection in similar states. The following formula is the expression of the state-action value function Q, where V(s t ) is the expected return that the agent can obtain in state s t , and A(s t ,a t ) is the advantage function of the agent taking different actions a t in state s t , and ω is the parameter set of the neural network;

[0104] Q ω (s t ,a t )=V(s t )+A(s t ,a t )

[0105] To obtain the deep connection between the trajectory points and the neighboring trajectory points, the state-action value neural network Q ω (s t ,a t ) of the mode identification agent uses a long short-term memory network layer to replace the fully connected layer. In addition, a target network is introduced during the training process and double neural networks are used for training. The training network Q ω (s,a) is updated at each step during training, while the target network uses historical parameters and is synchronized with the training network every C steps, thus avoiding the impact of overestimated network valuation on the mode identification performance. The training process demonstration of the mode identification agent is shown in Figure 8 .

[0106] In this embodiment, the sparse trajectory observation data of intercity trips of residents within Jiangsu Province are used to restore the real paths. For partial result display, see Figure 9 .

[0107] The above are only the preferred embodiments of the present invention, and are not any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.

Claims

1. Intercity rail and road multi-modal travel path multi-agent collaborative reasoning and restoration method, characterized in that: It includes the following steps: Step S10, obtain and preprocess the original trajectory observation data and traffic network data; Step S20, rasterize and store the road network data and trajectory observation data; Step S30, design and train a path search agent to infer a reasonable restored path sequence given a trajectory observation sequence and the road network corresponding to a specific traffic mode; Step S40, design and train a mode identification agent to infer a reasonable travel mode sequence given a trajectory observation sequence and the road matching degree feedback by the path search agent; Step S50, the path search agent and the mode identification agent are trained asynchronously and cooperate in inference. After the reinforcement learning training is completed, they jointly realize the inference and restoration of the true path corresponding to the sparse trajectory observation sequence.

2. The multi-agent collaborative reasoning and restoration method for the intercity rail and road multi-modal travel path according to claim 1, characterized in that: The specific steps of step S10 include: Step S11, obtain and preprocess the trajectory observation data; Obtain the trajectory observation data containing the intercity travel information of travelers through a specific software or mobile phone operator. The obtained data should at least include timestamp information and longitude and latitude information. Then, preprocess the original trajectory observation data of residents' intercity travel, including data filling and duplicate removal operations; Step S12, obtain and preprocess the road network data; Obtain the traffic network data of the research area. This data should at least include the location information of each road. Based on the type of the road, divide the overall road network data into multiple subsets, corresponding to different traffic modes respectively, including the complete road network data of the corresponding traffic mode, and select the road data subset involving intercity travel for use.

3. The multi-agent collaborative reasoning and restoration method for the intercity rail and road multi-modal travel path according to claim 1, wherein: The specific steps of step S20 include: Step S21, rasterize the research area; Use the ArcGIS software to select an appropriate size to divide the study area into grids, and number the grids in the form of the row and column positions they belong to, such as (locx i , locy i ), where locx i is the column number where the grid is located, and locy i is the row number where the grid is located; Step S22, rasterize and store the road network data; Use the ArcGIS software to map the road networks of various traffic modes into numbered grids respectively. If there is a road corresponding to a certain traffic mode in the grid, the corresponding value is 1, otherwise it is 0. After all are completed, the traffic network information of each grid itself is a string of numbers with a length equal to the number of traffic mode types, and the values are 1 or 0. To reduce the storage space, convert this 1 / 0 digital string into a decimal number. At this point, the traffic network information of each grid itself is completely stored. When the traffic network information of each grid corresponding to itself is complete, retrieve and store the traffic network information of the surrounding eight grids, that is, what each grid finally stores is the traffic network information of 9 grids including itself. The grid order is: upper left, directly above, upper right, directly left, itself, directly right, lower left, directly below, lower right; Step S23, rasterize and store the trajectory observation data; The position information of the current trajectory observation data still shows longitude and latitude. Use the ArcGIS software to map all the position information in the trajectory observation data into the grid, and convert the storage from longitude and latitude to the row and column of the grid to which it belongs. After conversion, the trajectory observation data includes the ID, timestamp, column, and row corresponding to each trajectory observation point.

4. The multi-agent collaborative reasoning and restoration method for the intercity rail and road multi-modal travel path according to claim 1, wherein: The specific steps of step S30 include: Step S31, design a path search agent; At the t-th time step during path restoration, the path inference agent performs a Markov decision process according to the policy, starting from the current state s t and taking an action a t to transfer to the next state and obtain the reward r at this moment T (s t , a t ). Among them, the state s obtained by the agent t includes the road network information corresponding to the given travel mode, the agent's current location information, and the location information of the target trajectory point; the action a taken by the agent t is one of the action spaces, and the action space includes moving to the upper left, directly above, upper right, directly left, directly right, lower left, directly below, and lower right; the reward r obtained by the agent T (s t , a t ) is the weighted sum of the following road reward and the approaching end reward; Step S32, train the path search agent; The path search agent is trained using the policy-based SAC algorithm. During training, the policy is modeled as π by a neural network θ , where θ is the set of parameters of the neural network. Then, the policy parameters θ are learned by the gradient descent method to obtain the optimal policy. The optimization objective of θ is as follows: Since the single policy network parameters may not be able to adapt to the state differences between different trajectory points, target-oriented reinforcement learning is used to enhance the generalization ability of the agent to diverse state spaces. The position information of the agent itself and the target trajectory point in the agent state is mapped into relative position information starting from the origin of coordinates, thus transforming the path inference problem between different trajectory points into a series of similar tasks starting from the same trajectory point and reaching different target points; During the training process, an experience replay pool is also used to improve the training efficiency of the agent: in each round of sampling phase, the state, action, reward, and state of the next time step of the path inference agent at the current time step will be stored in the experience replay pool in the form of a tuple <s,a,r,s’>; in the gradient update phase, several data will be randomly sampled from the experience replay pool for gradient update of the policy network, which greatly improves the data utilization efficiency.

5. The method for collaborative reasoning and restoration of the multi-agent of the intercity railway and highway multi-modal travel path according to claim 1, wherein: The specific steps of step S40 include: Step S41, design a mode identification agent; Similar to the path search agent, at the t-th time step during the mode identification process, the mode identification agent performs a Markov decision process according to the policy, starting from the current state s t to take an action a t and transfer to the next state, and obtain the reward r at this moment M (s t , a t ). Among them, the state s obtained by the agent t includes the local dynamic and static features of the observation sequence at this moment, the road network information of the location where the agent is located, and the road network information of the grid around its location; the action a taken by the agent t is the identified traffic mode; the reward r obtained by the agent M (s t , a t ) is set to the product of the road matching degree feedback by the path search agent and the reward obtained by identifying the mode according to the motion characteristics; Step S42, train the mode identification agent; Different from the policy-based reinforcement learning process of the path inference agent, the mode identification agent is trained using the value-based DQN algorithm. By learning the state-action value function Q(s,a), the optimal selection of traffic mode label actions under different states is realized. In each learning and training process, the update rule of the state-action value function Q(s,a) of the agent is: Q(s t ,a t ) ← Q(s t ,a t ) + α[r M + γ max a’ Q(s t+1 ,a’) - Q(s t ,a t )] To effectively find the optimal Q-function, a neural network is used to model it; to analyze the differences in taking different actions in the same state, the state-action value function Q is expressed as the sum of the state value function V and the advantage function A, enabling the agent to better handle action selection in similar states. The following is the expression of the state-action value function Q, where V(s t ) is the expected return that the agent can obtain in state s t , A(s t , a t ) is the advantage function of the agent taking different actions a t in state s t , and ω is the parameter set of the neural network; Q ω (s t ,a t ) = V(s t ) + A(s t ,a t ) To obtain the deep connection between the trajectory points and the neighboring trajectory points, the state-action value neural network Q of the manner recognition agent ω (s t ,a t ) replaces the fully connected layer with a long short-term memory network layer. In addition, a target network Q ω -(s,a) is introduced during the training process, and a dual neural network is used for training. The training network Q ω (s,a) is updated at each step during the training, while the target network Q ω -(s,a) uses historical parameters and is synchronized with the training network every C steps to avoid the impact of overestimated network valuation on the manner recognition performance.

6. The multi-agent collaborative reasoning and restoration system for intercity railway and highway multi-modal travel paths according to any one of claims 1-5, characterized in that: It includes an original data acquisition and preprocessing module, a data grid and storage module, a path search agent design and training module, a mode identification agent design and training module, and a result output module, where: The original data acquisition and preprocessing module is used to acquire and preprocess the intercity travel trajectory observation data and road network data, so as to provide a data source for subsequent rasterization processing; The data grid and storage module is used to rasterize and store the road network data corresponding to different traffic modes and the observation trajectory data containing the true path, which is convenient for the next step of path search restoration; The path search agent design and training module is used to infer a reasonable restored path sequence based on the given trajectory observation sequence and the road network corresponding to a specific traffic mode, so as to calculate the road matching degree as the gain information to support the training of the mode identification agent; The mode identification agent design and training module is used to infer a reasonable travel mode sequence based on the given trajectory observation sequence and the road matching degree feedback by the path search agent, so as to provide the correct road network information for the path search agent and support it to infer and restore the true path corresponding to the observation trajectory; The result output module is used to complete the inference restoration and result output of its true path by the trained path search agent and mode identification agent after a given trajectory observation sequence.

Citation Information

Patent Citations

  • Unmanned ship multipath optimization method for large-scale dynamic search and rescue tasks

    CN119472269A

  • Training policy neural networks using path consistency learning

    US20190332922A1