Travel chain identification method and device, electronic equipment and storage medium

By constructing a base station relationship diagram and evaluating the base station relationship using reinforcement learning model, the problem of low accuracy in travel chain identification caused by base station drift and ping-pong handover is solved, and more accurate travel chain identification is achieved.

CN120541454APending Publication Date: 2025-08-26CHINA MOBILE GROUP SHANDONG +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510546362.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When identifying user travel chains, the prior art faces the problems of base station drift and ping-pong switching, resulting in low recognition accuracy.

Method used

By obtaining the mobile signaling data and base station data of the target user, a base station relationship diagram is constructed, and the training reinforcement learning model is used to evaluate the continuity, smoothness, backtrackability and connectivity of the base station relationship based on the reward function, the real base station relationship is selected, the target movement trajectory is output, and the travel chain is identified.

Benefits of technology

It improves the accuracy of travel chain identification, can filter out problems such as base station drift or ping-pong switching, ensures the authenticity and continuity of the movement trajectory, and avoids the limitations of fixed parameters and data quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541454A_ABST
    Figure CN120541454A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a trip chain identification method and device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and can improve the accuracy of identifying a trip chain of a user. Comprising the following steps: acquiring target user data in a plurality of base stations related to a moving track of a target user, and determining a relation graph of the base stations according to the target user data; the relation graph is input into a trained reinforcement learning model, an action strategy for judging whether the base station relation corresponds to the position movement of the target user or not is executed through the reinforcement learning model, an execution result is obtained, and the reinforcement learning model is obtained based on reward function training; outputting a target moving track of the target user through a reinforcement learning model according to the execution result; and when the user staying time of a track point in the target moving track exceeds a preset dynamic threshold value, determining the track point as a staying point of the target user, and identifying a trip chain of the user according to the staying point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device and storage medium for identifying a travel chain. Background Art

[0002] In the field of intelligent transportation systems and urban planning, the identification of users' travel chains is very critical. The identification of travel chains mainly involves obtaining the user's movement trajectory from the user signaling data and extracting the complete travel path from the movement trajectory.

[0003] When the base station corresponding to the user signaling data has problems with base station drift and ping-pong switching, the relevant technology usually uses the distance speed threshold method to identify the travel chain, that is, the relative relationship between the sampling interval distance and speed of several trajectory points is continuously sampled to determine whether there is a problem with the base station. However, when faced with a changing and complex base station environment, the base station with problems cannot be accurately identified according to fixed thresholds or rules, resulting in low accuracy in identifying the user's travel chain. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a method, device, electronic device, and storage medium for identifying a travel chain, so as to solve the problem of low accuracy in identifying a user's travel chain.

[0005] To solve the above technical problems, the embodiments of the present application are implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for identifying a travel chain, comprising: obtaining target user data from multiple base stations related to a target user's movement trajectory, determining a relationship graph corresponding to the base stations based on the target user data, wherein the relationship graph is used to represent the adjacency relationship between each of the base stations, wherein the target user data includes: mobile phone signaling data of the target user and base station data of the corresponding base stations; inputting the relationship graph into a trained reinforcement learning model, and executing an action strategy through the reinforcement learning model to determine whether the base station relationship corresponds to the target user's location movement, thereby obtaining an execution result, wherein the base station relationship represents the connection relationship between the nodes of the target user in the relationship graph, and the reinforcement learning model is trained based on a reward function, wherein the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: a change in continuity, a change in smoothness, a change in retrospectiveness, and a change in connectivity of the movement trajectory after executing the action strategy; outputting a target movement trajectory of the target user through the reinforcement learning model based on the execution result; when the user stay time at a trajectory point in the target movement trajectory exceeds a preset dynamic threshold, determining the trajectory point as a stay point of the target user, and identifying the travel chain of the target user based on the stay point.

[0007] In the second aspect, an embodiment of the present application provides a travel chain identification device, including: an acquisition module, used to acquire target user data from multiple base stations related to the target user's movement trajectory, and determine the relationship graph corresponding to the base station based on the target user data, and the relationship graph is used to characterize the adjacency relationship between each of the base stations, and the target user data includes: the target user's mobile phone signaling data and the corresponding base station data of the base station; a reward module, used to input the relationship graph into a trained reinforcement learning model, and execute the action strategy of judging whether the base station relationship corresponds to the position movement of the target user through the reinforcement learning model to obtain an execution result, wherein the base station relationship characterizes the target user in The connection relationship of the nodes in the relationship graph, the reinforcement learning model is obtained by training based on the reward function, and the reward function includes: an evaluation index and the weight of the evaluation index; the evaluation index includes: the continuity change, smoothness change, retrospective change and connectivity change of the movement trajectory after executing the action strategy; an output module is used to output the target movement trajectory of the target user through the reinforcement learning model according to the execution result; a determination module is used to determine the trajectory point as the target user's stay point when the user's stay time at the trajectory point in the target movement trajectory exceeds a preset dynamic threshold, and identify the target user's travel chain based on the stay point.

[0008] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory electrically connected to the processor, wherein the memory stores a computer program, and the processor is configured to call and execute the computer program from the memory to implement the above-mentioned travel chain identification method.

[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program can be executed by a processor to implement the above-mentioned method for identifying a travel chain.

[0010] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the above-mentioned travel chain identification method.

[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for identifying a travel chain.

[0012] The technical solution of the embodiment of the present application is adopted to obtain target user data from multiple base stations related to the target user's movement trajectory. Based on the target user data, a relationship graph corresponding to the base station is determined. The relationship graph is used to characterize the adjacency relationship between each base station. The target user data includes: the target user's mobile phone signaling data and the base station data of the corresponding base station; the relationship graph is input into the trained reinforcement learning model, and the reinforcement learning model is used to execute an action strategy for determining whether the base station relationship corresponds to the target user's position movement to obtain an execution result, wherein the base station relationship represents the connection relationship between the target user's nodes in the relationship graph. The reinforcement learning model is trained based on a reward function, and the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: the continuity change, smoothness change, retrospective change and connectivity change of the movement trajectory after executing the action strategy; based on the execution result, the target movement trajectory of the target user is output through the reinforcement learning model; when the user's stay time at a trajectory point in the target movement trajectory exceeds a preset dynamic threshold, the trajectory point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

[0013] A base station relationship graph is determined based on the target user data collected by the base station. This graph is then input into a trained reinforcement learning model. Based on the graph, an action strategy is executed to determine whether the base station relationship corresponds to the target user's location movement, yielding an execution result. Because the reinforcement learning model is trained based on a reward function, after executing the action strategy, the target user's movement trajectory exhibits no issues with continuity, smoothness, retrospectiveness, or connectivity. Furthermore, due to the high authenticity of the selected execution results, problematic base stations that cause changes in the movement trajectory can be screened out. Based on the execution results, the actual base stations in the relationship graph are determined, and the target movement trajectory of the target user is output using the reinforcement learning model. When the user's dwell time at a point in the target movement trajectory exceeds a preset dynamic threshold, the point is identified as the target user's dwell point, and the target user's travel chain is identified based on the dwell point. When identifying the target user's movement trajectory, the trained reinforcement learning model focuses on each step of the target user's base station selection and evaluates the action strategy executed at each step using the evaluation metric in the reward function. This allows the user to identify base stations that the user did not actually pass through when there are issues such as drift or ping-pong switching, resulting in accurate execution results. Compared with fixed pattern or clustering methods, it can focus on the actions performed at each step and is not restricted by factors such as fixed parameters and data size, making the final target movement trajectory more accurate and solving the problem of low accuracy in identifying the user's travel chain. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate one or more embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 is a schematic flow chart of a method for identifying a travel chain according to an embodiment of the present application;

[0016] Figure 2 is a schematic diagram of a process of screening sample base stations in a sample movement trajectory of a sample user according to an embodiment of the present application;

[0017] Figure 3 is a schematic diagram of the interaction relationship between an agent and an environment according to an embodiment of the present application;

[0018] Figure 4 is a schematic flow chart of a method for identifying a travel chain according to another embodiment of the present application;

[0019] Figure 5 is a schematic block diagram of a travel chain identification device according to an embodiment of the present application;

[0020] Figure 6 This is a schematic diagram of the hardware structure of a travel chain identification device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] Embodiments of the present application provide a method, device, electronic device, and storage medium for identifying a travel chain, to solve the problem of low accuracy in identifying a user's travel chain.

[0022] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0023] The travel chain identification method provided in the embodiments of the present application can be executed by an electronic device or by software installed in the electronic device. Specifically, the electronic device can be a terminal device or a server device. The terminal device can include a smartphone, a laptop computer, a smart wearable device, an in-vehicle terminal, etc. The server device can include an independent physical server, a server cluster consisting of multiple servers, or a cloud server capable of cloud computing.

[0024] The following describes in detail a method for identifying a travel chain provided by an embodiment of the present application through specific embodiments and application scenarios in conjunction with the accompanying drawings.

[0025] Figure 1 A schematic flow chart of a method for identifying a travel chain provided by an embodiment of the present invention is shown. The method includes the following steps:

[0026] S102, obtaining target user data from multiple base stations related to the target user's movement trajectory, and determining a relationship graph corresponding to the base stations based on the target user data. The relationship graph is used to represent the adjacency relationship between the base stations. The target user data includes: mobile phone signaling data of the target user and base station data of the corresponding base stations.

[0027] The target user is the user who needs to determine whether each trajectory point passed in the movement trajectory is real.

[0028] The target user's movement trajectory includes the path generated during the target user's movement. During the target user's movement, the target user connects to multiple base stations through mobile phone signaling data. Each base station generates data when the target user is connected. The data generated by multiple base stations is determined as the target user data, and the data can be collected through the data collection module.

[0029] Specifically, the target user data collected from the base station includes: the target user's mobile phone signaling data, base station data for multiple base stations corresponding to the mobile phone signaling data, geographic grid data, and data on the correspondence between base stations and geographic grids. For example, the target user's mobile phone signaling data includes: user ID, base station code, time of entry and exit from the base station, etc.; base station data includes: base station code, base station longitude, and base station latitude, etc.; geographic grid data includes: grid code, grid type, grid size, grid center longitude, and grid center latitude, etc.; and data on the correspondence between base stations and geographic grids includes: grid code and base station code. A geographic grid is a data format that divides space into a regular grid, with each grid cell being called a unit, and assigning corresponding attribute values ​​to each unit to represent entities. For example, a geographic grid can be used to represent the geographic locations of multiple base stations related to a mobile trajectory in equal proportions.

[0030] Based on the acquired target user data, the base station code of the base station to which the target user is connected in the target user's mobile phone signaling data can be determined, which corresponds to the base station code in the base station data. Since the base station code in the base station data can correspond to the base station code in the geographic grid data, the grid code can be determined, and the geographical location of the base station can be mapped on the geographic grid. Then, the base stations passed by the target user during movement can be mapped in the geographic grid, and the target user's movement trajectory can also be displayed based on the geographic grid.

[0031] Furthermore, using an intelligent filtering module, the target user's timestamps for connecting and leaving each base station are determined based on the time they enter and leave the base station in the mobile phone signaling data. The base stations are then connected based on these timestamps to determine the adjacency relationships between the base stations. Based on these adjacency relationships, a relationship graph is generated to represent the relationships between the base stations. The nodes in the relationship graph represent base stations, and the lines between the nodes represent the distances between the base stations.

[0032] S104: input the relationship graph into the trained reinforcement learning model, and execute the action strategy of determining whether the base station relationship corresponds to the position movement of the target user through the reinforcement learning model to obtain an execution result.

[0033] Among them, the base station relationship represents the connection relationship of the target user's nodes in the relationship graph. The reinforcement learning model is trained based on the reward function. The reward function includes: evaluation indicators and the weights of evaluation indicators; the evaluation indicators include: the continuity change, smoothness change, retrospective change and connectivity change of the movement trajectory after the action is executed.

[0034] The trained reinforcement learning model is based on a common reinforcement learning method. This involves using reward and loss functions to ensure that the trained reinforcement learning model correctly identifies and executes the action strategy for determining whether base station relationships correspond to the target user's location and movement. Inputting the relationship graph into the trained reinforcement learning model allows it to determine whether the connected nodes in the graph are real. Specifically, it can determine whether the movement trajectory in the graph corresponds to the correct base station, thereby filtering out erroneous drift data.

[0035] Among them, after inputting the relationship graph into the trained reinforcement learning model, the reinforcement learning model can obtain the first base station corresponding to the first node in the relationship graph, and use the first base station as the current base station, that is, the position of the target user in the relationship graph as the current base station, and the next position to which the target user will move as the next base station. The base station relationship can represent the connection relationship between the current base station and the next base station, and represents the connection relationship between the current node where the target user is located and the next node in the relationship graph. It should be understood that the relationship in which multiple base stations related to the target user's movement trajectory are connected in sequence is a base station relationship, and represents the connection relationship between multiple nodes related to the target user's movement trajectory in the relationship graph.

[0036] Because the reinforcement learning model is trained based on a reward function, the reward function includes evaluation metrics and their weights. These metrics include the continuity, smoothness, retracing, and connectivity of the trajectory after executing the action strategy. This allows the reinforcement learning model to execute the action strategy that determines whether the base station relationship corresponds to the target user's location, resulting in a more accurate execution of the strategy and a lesser impact on the overall trajectory.

[0037] The execution result is used to determine whether the base station relationship corresponding to the node in the relationship graph corresponds to the location movement of the target user, that is, to determine whether the node in the relationship graph is the base station that the target user actually passed by.

[0038] As an example, the relationship graph is input into the trained reinforcement learning model, and the base station where the target user is located in the relationship graph at the current moment is used as the current base station. The reinforcement learning model is used to execute an action strategy to determine whether the base station relationship of the current base station corresponds to the position movement of the target user, and an execution result is obtained. The execution result includes: the base station relationship between the current base station and the next base station corresponds to the position movement of the target user, or the base station relationship between the current base station and the next base station does not correspond to the position movement of the target user.

[0039] S106: Output the target movement trajectory of the target user through the reinforcement learning model according to the execution result.

[0040] Based on the trained reinforcement learning model in S104, an execution result is obtained. Multiple execution results corresponding to all nodes in the relationship graph are obtained based on the relationship graph. All execution results in the relationship graph are integrated, and the reinforcement learning model is used to filter the target user's actual trajectory for each step, determining the base stations the target user actually passed during movement. The target user's target movement trajectory is output according to the order in the relationship graph. The target movement trajectory refers to the target user's final actual movement path.

[0041] S108 , when the user's stay time at a track point in the target movement track exceeds a preset dynamic threshold, the track point is determined as a stay point of the target user, and a travel chain of the target user is identified based on the stay point.

[0042] The trajectory point is a specific location point of the target user in the target movement trajectory, for example, a base station that the target user passes by in the target movement trajectory.

[0043] The preset dynamic threshold includes a preset and adjustable dwell time of the target user at a trajectory point. For example, the target movement trajectories corresponding to multiple users including the target user are obtained, and the multiple users are segmented according to the number of trajectory points in the target movement trajectories corresponding to the multiple users. The dwell time of the user at the trajectory point is adjusted according to the number of trajectory points in each segment, and the adjusted dwell time of the target user at the trajectory point is used as the preset dynamic threshold.

[0044] When the target user's dwell time at a point in the target trajectory exceeds a preset dynamic threshold, the point is considered a dwell point for the target user. Two consecutive dwell points in the target trajectory constitute a trip chain. If there are multiple dwell points in the target trajectory, the multiple dwell points are connected in chronological order to obtain the target user's complete trip chain.

[0045] The technical solution of the embodiment of the present application is adopted to obtain target user data from multiple base stations related to the target user's movement trajectory. Based on the target user data, a relationship graph corresponding to the base station is determined. The relationship graph is used to characterize the adjacency relationship between each base station. The target user data includes: the target user's mobile phone signaling data and the base station data of the corresponding base station; the relationship graph is input into the trained reinforcement learning model, and the reinforcement learning model is used to execute an action strategy for determining whether the base station relationship corresponds to the target user's position movement to obtain an execution result, wherein the base station relationship represents the connection relationship between the target user's nodes in the relationship graph. The reinforcement learning model is trained based on a reward function, and the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: the continuity change, smoothness change, retrospective change and connectivity change of the movement trajectory after executing the action strategy; based on the execution result, the target movement trajectory of the target user is output through the reinforcement learning model; when the user's stay time at a trajectory point in the target movement trajectory exceeds a preset dynamic threshold, the trajectory point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

[0046] A base station relationship graph is determined based on the target user data collected by the base station. This graph is then input into a trained reinforcement learning model. Based on the graph, an action strategy is executed to determine whether the base station relationship corresponds to the target user's location movement, yielding an execution result. Because the reinforcement learning model is trained based on a reward function, after executing the action strategy, the target user's movement trajectory exhibits no issues with continuity, smoothness, retrospectiveness, or connectivity. Furthermore, due to the high authenticity of the selected execution results, problematic base stations that cause changes in the movement trajectory can be screened out. Based on the execution results, the actual base stations in the relationship graph are determined, and the target movement trajectory of the target user is output using the reinforcement learning model. When the user's dwell time at a point in the target movement trajectory exceeds a preset dynamic threshold, the point is identified as the target user's dwell point, and the target user's travel chain is identified based on the dwell point. When identifying the target user's movement trajectory, the trained reinforcement learning model focuses on each step of the target user's base station selection and evaluates the action strategy executed at each step using the evaluation metric in the reward function. This allows the user to identify base stations that the user did not actually pass through when there are issues such as drift or ping-pong switching, resulting in accurate execution results. Compared with fixed pattern or clustering methods, it can focus on the actions performed at each step and is not restricted by factors such as fixed parameters and data size, making the final target movement trajectory more accurate and solving the problem of low accuracy in identifying the user's travel chain.

[0047] In one embodiment, based on the target user data, a relationship graph corresponding to a base station is determined (i.e., S102), and the following steps A1-A2 may be performed:

[0048] Step A1: Determine the base station location of the base station corresponding to the mobile phone signaling data according to the mobile phone signaling data and the base station data of the corresponding base station in the target user data. The base station data includes the base station code and the base station location.

[0049] According to the target user data acquired from the base station in S102, the mobile phone signaling data in the target user data includes: a user identifier, a base station code, a time of entry into the base station, and a time of departure from the base station. Based on the base station code in the mobile phone signaling data, the base station corresponding to the mobile phone signaling data, i.e., the base station corresponding to the target user, can be determined. The unique target user is determined based on the user identifier in the mobile phone signaling data, and the target user's stay time at the base station is determined based on the time of entry into the base station and the time of departure from the base station. Based on the base station data of multiple base stations corresponding to the mobile phone signaling data, the base station code, base station longitude, and base station latitude of each base station are determined. The base station code can correspond to the grid code in the mobile phone signaling data and the corresponding geographic grid data. Based on the base station longitude and base station latitude, the base station location is determined. The target user data also includes: geographic grid data and data on the correspondence between the base station and the geographic grid. The correspondence between the base station and the geographic grid is used to determine the grid position of the base station in the geographic grid data based on the base station location.

[0050] Step A2: Connect the base stations in chronological order according to the base station locations to obtain a relationship diagram with adjacency relationships between the base stations; wherein the relationship diagram includes node features and edge features, the node features represent the features of the base station locations, and the edge features represent the distances between the base stations.

[0051] Based on the mobile phone signaling data and base station data in step A1, the same base station code is determined, and the user signaling data and the base station are associated. Specifically, the mobile phone signaling data records the base station code of the base station to which the target user is connected when the mobile phone signaling occurs, and the base station data contains the base station code and base station location of each base station. By matching the base station code, the location of the base station to which the target user is connected at a specific point in time can be determined, thereby obtaining the target user's movement trajectory based on the base station, that is, the sequence and base station location of the base stations to which the target user is connected during the movement. For example, based on the mobile phone signaling data, the timestamp information of the time when the target user enters the base station and the time when the base station leaves the base station when the mobile phone signaling occurs, as well as the base station location of the corresponding base station, is determined.

[0052] Based on the time sequence in the timestamp information of the mobile phone signaling data, the time sequence in which the target user enters each base station in turn is determined. According to the base station position of each base station passed by the target user, the base stations are connected and processed according to the time sequence of the mobile phone signaling. The switching relationship between the target user and different base stations can be identified, and a relationship graph with an adjacency relationship can be constructed. The switching relationship includes the process of the target user selecting the next base station from the current base station.

[0053] The relationship graph includes node features and edge features. Nodes represent base stations, while edges can also represent handoff relationships between base stations, forming a relationship graph between base stations based on mobile phone signaling data. Node features in the relationship graph are determined by mobile phone signaling data, such as base station coding features, base station entry time features, and base station dwell time features. The base station entry time is determined by the earliest data point in the mobile phone signaling data that connected the base station, and the dwell time represents the total length of time the target user spent at the base station. Edge features are determined by base station data, such as the coding features of the edge connecting two base stations and the straight-line distance features between two nodes.

[0054] In this embodiment, user data from the base station is collected to determine the target user's mobile phone signaling data and base station data. Based on the timestamp information in the target user's mobile phone signaling data and the base station location recorded in the base station data, the base station location corresponding to the mobile phone signaling data is determined. The base stations corresponding to the mobile phone signaling data are then connected based on the time information, resulting in a relationship graph containing adjacency relationships of node features and edge features. Based on this relationship graph, a preliminary and rough determination of the target user's movement trajectory is made.

[0055] In one embodiment, the relationship graph is input into a trained reinforcement learning model, and the reinforcement learning model is used to execute an action strategy for determining whether the base station relationship corresponds to the target user's location movement. After obtaining an execution result (i.e., S104), the following steps B1-B5 may be executed:

[0056] Step B1: Input the relationship graph into the trained reinforcement learning model, determine the position of the target user in the relationship graph based on the adjacency relationship between the base stations in the relationship graph, and determine the position of the target user in the relationship graph as the current base station.

[0057] The reinforcement learning model is used to verify whether the switching of base stations in the relationship graph corresponds to the actual movement trajectory of the target user.

[0058] The relationship graph is input into the trained reinforcement learning model, and the reinforcement learning model verifies whether the connection relationship between base stations is correct in turn. The current base station refers to the position of the target user verified by the reinforcement learning model in the relationship graph. At this position, an action strategy is executed to determine whether the base station relationship corresponds to the position movement of the target user. For example, when verifying that the target user is located at the current base station, it is determined whether the position movement of the target user corresponding to the base station relationship is real. The current base station can be the initial base station in the relationship graph, or it can be the base station after multiple actions are subsequently executed.

[0059] Step B2, based on the current base station, determine the neighborhood information of the current base station; the neighborhood information includes: in the relationship diagram, the characteristics of the node corresponding to the current base station, the characteristics of the first-order nodes directly connected to the node corresponding to the current base station, and the characteristics of the second-order neighboring nodes connected to the first-order nodes.

[0060] Based on the adjacency relationships of each base station in the relationship graph, the neighborhood information of the current base station is determined. This neighborhood information can also be represented as the current state of the reinforcement learning environment. The neighborhood information includes the node features of the current base station, the node features corresponding to the base stations directly connected to the current base station as the features of the first-order nodes, and the node features of the base stations connected to the first-order nodes as the features of the second-order neighboring nodes. This helps the reinforcement learning model better capture the local structure and potential relationships of the environment.

[0061] Step B3: Use a graph autoencoder to convert the neighborhood graph corresponding to the neighborhood information of the current base station into a feature vector.

[0062] Based on the neighborhood information in step B2, the neighborhood graph corresponding to the neighborhood information is converted into a feature vector representation through a graph autoencoder. The conversion process can be expressed as Formula 1. The graph autoencoder is a deep neural architecture used to map nodes in the neighborhood information to a latent feature space and decode the information of the neighborhood graph from the latent representation.

[0063] f VGAE (G) = GCN(X, A) Formula 1

[0064] Among them, f VGAE (G) represents the result of converting the neighborhood graph into a feature vector using a graph autoencoder. G represents the neighborhood graph, and Graph Convolutional Networks (GCN) are used. X and A represent the node feature matrix and adjacency matrix of the neighborhood graph, respectively. GCN is a general deep learning method commonly used for graph data processing. It can effectively learn the structural relationships between nodes in the relationship graph by aggregating the neighborhood information of the current node, namely the current base station. Based on GCN, the features of each node in the neighborhood graph are integrated with the semantic information of its neighborhood, thus more comprehensively representing the local and global structural relationships in the relationship graph.

[0065] Step B4, based on the feature vector corresponding to the neighborhood information of the current base station, the reinforcement learning model is used to calculate the value of the action strategy for determining whether the base station relationship of the feature vector corresponds to the location movement of the target user. The value includes: a numerical value calculated based on the state value function and the advantage function.

[0066] Reinforcement learning primarily involves an agent and an environment. Learning occurs through interaction between the agent and the environment, allowing the agent to make optimal decisions within the environment. Therefore, a trained reinforcement learning model can leverage the agent's perception of the environment to execute corresponding action strategies. The environment can be the neighborhood information corresponding to a feature vector. Based on this neighborhood information, the agent can calculate the value of executing an action strategy to determine whether the base station relationship corresponds to the target user's location movement. For example, when a feature vector is input into the agent in the reinforcement learning model, the target user is at the base station location of the current base station. Based on the feature vector, the neighborhood information of the current base station where the target user is located is determined. The agent then calculates the value of executing an action strategy to determine whether the base station relationship corresponds to the target user's location movement at that location, and selects the base station with the highest value in the neighborhood information as the next base station selected by the target user.

[0067] Based on the reinforcement learning method, the intelligent agent calculates the value of executing each action strategy in turn according to the neighborhood information of each current base station. Specifically, there are many reinforcement learning methods, among which the improved reinforcement learning algorithm (Dueling DeepQ-Network, Dueling DQN) is a general reinforcement learning method. It calculates the value of the action strategy of judging whether the base station relationship of the neighborhood information corresponds to the position movement of the target user in the current state, and combines the greedy strategy to select the optimal execution action. It should be understood that the current state refers to the neighborhood information corresponding to the base station position of the current base station where the target user is located. The optimal execution action refers to the intelligent agent in the current state, according to the calculated execution action strategy value, determining the node with the highest value in the neighborhood information, and connecting it as the next base station connected to the current base station. The value formula for calculating the executed action strategy is shown in Formula 2:

[0068]

[0069] Here, V(s) represents the state value function, which estimates the value of the current state s itself; A(s,a) represents the advantage function, which measures the relative importance of action a relative to other actions in the current state s; and a′ represents the set of all possible actions. The calculated value is defined as the Q-value, which decomposes the Q-value function into a state value function and an advantage function.

[0070] DQN's reinforcement learning approach allows it to focus on distinguishing the relative contribution of each selected action when evaluating the value of an action strategy, rather than relying solely on the value of the overall trajectory. This reduces the estimation bias caused by the execution of a specific action.

[0071] Step B5, determining an execution result based on the value, the execution result including: the base station relationship of the current base station corresponds to the position movement of the target user, or the base station relationship of the current base station does not correspond to the position movement of the target user.

[0072] Based on the value of the agent's actions calculated in step B4 and the greedy strategy, the agent, when making decisions, maximizes its benefits and achieves a global optimal solution through local optimization. For example, based on the current state, the agent calculates the value of the action strategy to determine whether the base station relationship corresponds to the target user's location and selects the action with the highest value for execution. Specifically, when switching from the current base station to the next base station, if the base station determined by the agent's value calculation differs from the next base station in the target user's trajectory, this indicates that the next base station in the target user's trajectory is not the base station in the target user's actual trajectory. The target user should then select the base station with the higher value calculated by the agent based on neighborhood information, thereby filtering out problematic base stations in the target user's trajectory.

[0073] Therefore, the execution results include: the base station relationship of the current base station corresponds to the target user's location movement, or the base station relationship of the current base station does not correspond to the target user's location movement. In other words, if the value calculated by the current base station is high, it means that in the executed action strategy, the base station relationship of the current base station corresponds to the target user's location movement, and the next base station connected by the current base station is a base station in the target user's actual trajectory. Alternatively, if the value calculated by the current base station is low, it means that in the executed action strategy, the base station relationship of the current base station does not correspond to the target user's location movement, and the next base station connected by the current base station is not a base station in the target user's actual trajectory.

[0074] In this embodiment, the relationship graph is input into a trained reinforcement learning model. The reinforcement learning model determines the neighborhood information of the current base station based on the neighborhood relationships between base stations in the relationship graph. The reinforcement learning model then calculates and executes the action strategy to determine whether the base station relationships of each feature vector correspond to the target user's location movement. Based on this value, it is determined whether the next base station to be connected is the base station that the target user actually passed through. This can determine the authenticity of each base station passed by the target user in the relationship graph corresponding to the movement trajectory, thereby improving the authenticity of the entire target user's movement trajectory.

[0075] In one embodiment, the method further includes training a reinforcement learning model, which may specifically include performing the following steps C1-C5:

[0076] Step C1: Obtain a sample relationship graph of sample base stations corresponding to multiple sample users, as well as sample user movement trajectories of the sample users. Calculate the sample feature vector corresponding to the sample neighborhood information of the sample nodes in the sample relationship graph through a graph data processing method. The sample neighborhood information is used to characterize the adjacency relationship between the sample nodes.

[0077] The sample user is a user used to verify the authenticity of the sample movement trajectory corresponding to the sample user. The sample user data of each sample user on the relevant base station side is obtained through the data acquisition module. The sample user data includes: the sample mobile phone signaling data of the sample user and the corresponding sample base station data.

[0078] The sample user movement trajectory includes: the actual movement trajectory of the sample user. The graph data processing method includes: the deep learning method of GCN. The sample node includes: the node corresponding to the sample base station in the sample relationship graph.

[0079] The sample neighborhood information includes: in the sample relationship graph, the characteristics of the sample node corresponding to the current sample base station, the characteristics of the sample first-order nodes directly connected to the sample node corresponding to the current sample base station, and the characteristics of the sample second-order neighboring nodes connected to the sample first-order node.

[0080] According to the graph data processing method, the sample neighborhood information of each sample node in the sample relationship graph is converted into a sample feature vector of the adjacency relationship between the sample nodes.

[0081] In step C2, the sample feature vector corresponding to the sample node in the sample relationship graph is input into the reinforcement learning model to be trained, and the sample value of the sample action strategy for determining whether the sample base station relationship corresponds to the position movement of the sample user is calculated and executed by the intelligent agent of the reinforcement learning model to be trained, wherein the sample base station relationship represents the connection relationship between the node sample points of the sample user in the sample relationship graph.

[0082] According to step C1, the sample feature vector corresponding to the current sample node is determined, wherein the sample domain information corresponding to the sample feature vector can be used as the environment corresponding to when the intelligent agent in the reinforcement learning model to be trained performs the sample action.

[0083] Reinforcement learning models include: models based on DQN reinforcement learning method and reward function setting.

[0084] The sample feature vector is input into the reinforcement learning model to be trained. The intelligent agent in the reinforcement learning model to be trained executes the sample action strategy of judging whether the sample base station relationship corresponds to the position movement of the sample user based on the sample feature vector and the sample node in the current sample relationship graph, and calculates the sample value of executing the action strategy. The calculation formula of the sample value is shown in Formula 2. One sample node in the sample relationship graph corresponds to one sample base station.

[0085] The agent in the reinforcement learning model to be trained calculates the sample values ​​of all sample nodes in the sample relationship graph in sequence according to the order of the obtained sample feature vectors.

[0086] Step C3: determining a sample execution result based on the sample value. The sample execution result includes: the sample base station relationship of the sample node corresponds to the position movement of the sample user, or the base station relationship of the sample node does not correspond to the position movement of the target user.

[0087] Based on the sample value in step C2, the agent determines the sample value of all sample nodes in the sample neighborhood information for the sample current node. Each sample node in the sample neighborhood information is used as a verification node for the agent to execute the sample action strategy for determining whether the sample base station relationship corresponds to the location movement of the sample user. The sample value of the agent's execution of the sample action strategy for the sample current node is calculated. Using a greedy strategy, the verification node with the highest sample value is selected as the target sample node. Based on the target sample node, the agent selects the target sample node as the next sample node connected to the sample current node. The target sample node is the sample next base station of the sample current base station corresponding to the sample current node, corresponding to the sample base station that the sample user actually passed through.

[0088] The sample execution result includes whether the sample node's base station relationship corresponds to the sample user's location movement, or whether the sample node's base station relationship does not correspond to the target user's location movement; and also includes the agent's selection of the node to be verified. Specifically, if the base station relationship of the sample's current base station corresponds to the sample user's location movement, the node to be verified is the target sample node. Otherwise, if the base station relationship of the sample's current base station does not correspond to the sample user's location movement, the node to be verified is not the target sample node.

[0089] In step C4, the sample action strategy is evaluated by the sample reward function set in the agent interaction environment to obtain a sample evaluation result. Based on the sample evaluation result, the sample target movement trajectory of the sample user is output through the reinforcement learning model to be trained.

[0090] The interactive environment includes sample neighborhood information corresponding to multiple sample nodes and a sample reward function. The sample reward function includes sample evaluation metrics and their weights. The sample evaluation metrics include the continuity, smoothness, retroactivity, and connectivity of the sample movement trajectory corresponding to the sample relationship graph after executing the sample action strategy. The sample evaluation metrics and their weights can vary depending on the scenario.

[0091] The sample action strategy is evaluated to obtain a sample evaluation result, that is, the sample action strategy executed by the agent for the sample node in the sample neighborhood information is evaluated. When the sample execution result of the agent is that the sample base station relationship corresponding to the sample node corresponds to the position movement of the sample user, the reward function is used to strengthen the reward of the sample action strategy executed by the agent; when the sample execution result of the agent is that the base station relationship corresponding to the sample node does not correspond to the position movement of the target user, the reward function is used to negatively reward the sample action strategy executed by the agent.

[0092] The agent executes a sample action strategy to determine whether the sample base station relationship corresponds to the location movement of the sample user. The sample execution results obtained are used to evaluate the sample reward function to enhance the accuracy of the agent's selection, and the sample execution results are fed back to the interactive environment for further training of the reinforcement learning model to be trained.

[0093] Through the interaction between the sample execution results and the sample reward function, the sample base stations that the sample user actually passed through in the correct sample movement trajectory are screened out. The reinforcement learning model to be trained generates and outputs the sample target movement trajectory of the sample user according to the time sequence in the sample relationship graph.

[0094] In step C5, a loss function is calculated according to the sample reward functions corresponding to the multiple sample users. Based on the loss function, the movement trajectory of the sample users, and the movement trajectory of the sample targets, the parameters of the reinforcement learning model to be trained are adjusted. The reinforcement learning model to be trained is trained to obtain a trained reinforcement learning model.

[0095] The calculation formula of the loss function L(θ) is shown in Formula 3 below:

[0096] L(θ)=E (s,a,r,s′)~D [(yQ(s,a)) 2 ] Formula 3

[0097] Where s represents the current state of the sample node, a represents the agent’s execution of the sample action strategy based on the sample node, r represents the sample reward function calculated before and after the sample action strategy is executed, s′ represents the state of the agent after executing the sample action strategy based on the sample node, and D represents the experience pool, which is all the sample nodes in multiple sample relationship graphs and contains a sample set of interactions between the agent and the environment. (s,a,r,s′)~D represents the set of s, a, r, and s′ of all sample nodes in the entire sample relationship graph. Q is the sample value calculated by the agent based on the sample action strategy executed. y represents the target Q value, and the specific calculation formula is shown in Formula 4:

[0098] y=r+γmax a′ Q(s,a′;θ') Formula 4

[0099] Among them, γ is represented as a discount factor, which is used to control the influence of future rewards on the current decision, that is, the currently executed action; θ and θ' represent the parameters of the current Q network and the target Q network, respectively.

[0100] The sample reward function is calculated based on the sample relationship graph of multiple sample users, and the loss function is calculated based on the sample reward function. By adjusting the formula of the sample reward function, it is applicable to determining the sample target movement trajectory of the sample user in all scenarios. The loss function is used to adjust the loss of the reinforcement learning model to be trained, so that the overall loss is minimized. This can improve the authenticity of the trained reinforcement learning model in judging the movement trajectory of the sample target and adapt it to a wide range of scenarios. The similarity between the movement trajectory of the sample user and the movement trajectory of the sample target is achieved as expected, such as when the two are consistent.

[0101] It should be noted that the reward function generated by the sample target movement trajectory finally generated by each sample user can be calculated to comprehensively evaluate the sample evaluation indicators of the sample target movement trajectory. If a positive reward is generated, the sample action strategy adopted by the agent is strengthened; if a negative reward is generated, the sample action strategy adopted by the agent is adjusted.

[0102] As an example, Figure 2 Figure 2 illustrates the process of screening sample base stations in sample mobility trajectories. Sample mobile phone signaling data and sample base station data are correlated to construct a sample relationship graph, generating initial signaling data. This initial signaling data is then fed into the reinforcement learning model to be trained. The agent in the model then determines the rationality of each signaling data point. Specifically, the state of the sample node in the sample relationship graph, which includes all features of the current node's 2 neighbors (i.e., the sample neighborhood information corresponding to the sample node), is used to determine whether the executed sample action strategy is correct. The sample value of the executed sample action strategy is calculated to determine whether the sample action strategy is approved or disapproved. A sample reward function is then used to reward the sample execution result after executing the sample action strategy. After the nth evaluation of the sample action strategy corresponding to the signaling in the initial signaling data, the actual sample target mobility trajectory of the sample user after multiple simulations is obtained. The sample reward function also includes evaluating changes in the sample mobility trajectory before and after executing the sample action strategy, such as continuity, smoothness, retroactivity, and connectivity. Among them, the interaction between the environment and the agent in the reinforcement learning model, such as Figure 3The figure shows a schematic diagram of the interaction between the agent and the environment. The agent collects sample neighborhood information of the sample nodes corresponding to each sample base station in the environment, and determines the sample execution result of the executed sample action strategy by calculating the sample value of the executed sample action strategy. The sample execution result includes: agreeing to execute the sample action strategy and opposing the execution of the sample action strategy, and calculating the corresponding sample reward function, and feeding back the reward of the calculated sample reward function to the environment.

[0103] It should be noted that after completing the detection of all user data for a sample user and generating a sample target movement trajectory for that sample user, an overall sample reward is generated. This sample reward can be placed in the experience pool for subsequent optimization training of the reinforcement learning model to be trained, or for adjusting the parameters of the trained reinforcement learning model. For example, after completing the monitoring of all signaling changes for a sample user, the sample target movement trajectory for that sample user is generated, marking the end of a round of "game." At this time, a sample reward is also generated. The calculation of this sample reward requires a comprehensive evaluation of the continuity, smoothness, retrospectiveness, and connectivity of the sample target movement trajectory to comprehensively measure the effectiveness of the sample action strategy executed by the agent. (s, a, r, s′) is stored in the experience pool, but since s′ does not exist at this time, the calculation formula for the target Q value is simplified to y = r.

[0104] In this embodiment, the sample relationship graphs corresponding to the multiple sample users obtained are input into the reinforcement learning model to be trained. By calculating the sample value of executing the sample action strategy, the sample reward function and the loss function of executing the sample action strategy, the sample action strategy executed by the intelligent agent in the reinforcement learning model to be trained is trained. By adjusting the parameters in the reinforcement learning model to be trained, such as the sample reward function, etc., the sample target movement trajectory and sample movement trajectory output by the trained reinforcement learning model meet the expectations.

[0105] In one embodiment, the sample action strategy is evaluated by using the sample reward function set in the agent interaction environment to obtain a sample evaluation result (i.e., step C4). Specifically, the following steps D1-D6 can be performed:

[0106] Step D1: In the sample relationship diagram, based on the set of average distances of the edges connecting the sample base stations before the sample action strategy is executed, and the set of average distances of the edges connecting the sample base stations after the sample action strategy is executed, determine the continuity change of the sample movement trajectory of the sample user after the sample action strategy is executed.

[0107] After the reinforcement learning model executes the sample action strategy, the sample action strategy is evaluated using the sample reward function to further determine the authenticity of the sample action strategy execution. It should be understood that executing the sample action strategy is the intelligent agent in the reinforcement learning model. After executing the sample action strategy, the adjacency relationship between the sample base stations in the sample relationship graph may change. In other words, after the intelligent agent in the reinforcement learning model executes the sample action strategy, the corresponding environment will change, generating a new immediate reward. The immediate reward includes the new sample reward function. Therefore, the sample reward function can be used to determine the overall change in the sample user's movement trajectory after executing the sample action strategy.

[0108] In the sample relationship graph, we first calculate the average distances of the edges connecting each sample base station before the reinforcement learning model executes the sample action strategy, and then calculate the average distances of the edges connecting each sample base station after the sample action strategy is executed. The continuity of the sample movement trajectory is calculated based on the average distance of the edges, as shown in Formula 5:

[0109]

[0110] Among them, F continuity Indicates the continuous change of the sample movement trajectory, E G It represents the set of edges connected to each sample base station before the sample action strategy is executed, E G′ After the sample action strategy is executed, the set of edges connected to each sample base station is d(v i ,v j ) represents the Euclidean distance of the edges of each sample base station.

[0111] Step D2: Determine the smoothness of the sample movement trajectory of the sample user after executing the sample action strategy based on the turning angles of two adjacent sides connected by the sample base station.

[0112] In the sample relationship graph, the angles of the two adjacent edges connecting each sample base station before the reinforcement learning model executes the sample action strategy are first calculated, and then the angles of the two adjacent edges connecting each sample base station after the sample action strategy is executed are calculated. Based on the angles of the two adjacent edges connecting the sample base stations, the continuity change of the sample user's movement trajectory after the sample action strategy is executed is determined, as shown in Formula 6

[0113]

[0114] Among them, Fsmoothess represents the change in the smoothness of the sample movement trajectory after the sample action strategy is executed, which is measured by the average turning angle of the two edges, e i Represents two edges in the set of edges of each sample base station, Indicates the turning angle between two adjacent edges connected by the sample base station.

[0115] Step D3: determining the retrospective changes of the sample movement trajectory of the sample user after executing the sample action strategy based on the number of visits to the sample base station before executing the sample action strategy and the number of visits to the sample base station after executing the sample action strategy.

[0116] In the sample relationship diagram, the number of visits to each sample base station before the reinforcement learning model executes the sample action strategy is first calculated, and then the number of visits to each sample base station after the sample action strategy is executed is calculated. Based on the number of visits to the sample base station, the retrospective changes in the sample movement trajectory after the sample action strategy is executed are determined, as shown in Formula 7:

[0117]

[0118] Frevisit represents the retrospective change of the sample movement trajectory after the sample action strategy is executed, which is measured by the number of visits to repeated nodes. V G Represents the set of nodes before the sample action strategy is executed, V G′ represents the set of sample nodes after the sample action strategy is executed, and count(v) represents the number of visits to the sample node v. The current sample node in the sample relationship graph corresponds to the current sample base station. By counting the number of visits to each sample base station, it is determined whether there are repeated visits to the sample base station.

[0119] Step D4: determining the connectivity change of the sample movement trajectory of the sample user after executing the sample action strategy based on the connectivity subgraph of the base station before executing the sample action strategy and the connectivity subgraph of the sample base station after executing the sample action strategy.

[0120] In the sample relationship graph, the number of connected subgraphs between two sample nodes in the sample relationship graph corresponding to each sample base station before the reinforcement learning model executes the sample action strategy is first calculated. Then, the number of connected subgraphs between two sample nodes in the sample relationship graph corresponding to each sample base station after the sample action strategy is executed is calculated. Based on the number of connected subgraphs of the sample base stations, the connectivity change of the sample movement trajectory after the sample action strategy is executed is determined. Connectivity refers to the existence of a continuous path between two sample nodes in the sample relationship graph, as shown in Formula 8:

[0121] F connectivity =|C G |-|C G′ |Formula 8

[0122] Among them, Fconnectivity represents the connectivity change of the sample movement trajectory after the sample action strategy is executed, |C G | represents the number of connected subgraphs before the sample action strategy is executed, |C G′| represents the number of connected subgraphs after the sample action strategy is executed.

[0123] In step D5, the continuity change, smoothness change, retrospective change, and connectivity change are determined as sample evaluation indicators of the sample reward function, and the weights of the sample evaluation indicators are adjusted based on the changes in the sample evaluation indicators.

[0124] The sample reward function's sample evaluation metrics are determined based on multiple indicators, including continuity, smoothness, retrospective, and connectivity. These metrics assess the impact of the agent's execution of a sample action strategy on the user's sample movement trajectory. The sample evaluation metrics are weighted to reflect the impact of each metric.

[0125] In step D6, a sample reward function is determined based on the sample evaluation index and the weight of the sample evaluation index. Based on the sample reward function, the sample action strategy is evaluated to obtain a sample evaluation result. The sample evaluation result is used to train the sample execution result of the intelligent agent.

[0126] According to the continuity changes, smoothness changes, retrospective changes, and connectivity changes of the sample movement trajectory in the sample evaluation indicators of the sample reward function, the changes in the sample movement trajectory caused by the intelligent agent in the reinforcement learning model executing the sample action strategy of selecting the next base station of the sample corresponding to the current base station of the sample are judged. At the same time, the sample evaluation indicator is given a weight for the changes in each sample evaluation indicator, and the executed sample action strategy is comprehensively evaluated to obtain the sample evaluation result.

[0127] The calculation formula of the reward function is shown in Formula 9:

[0128]

[0129] Among them, r is the value calculated by the sample reward function, w1, w2, w3 and w4 represent the weights corresponding to the sample evaluation indicators, a1 indicates that the agent agrees to execute the sample action strategy, and the calculated sample reward function is 0; a2 indicates that the agent disagrees to execute the sample action strategy and calculates the value of the sample reward function.

[0130] For the weight of the sample evaluation index in the reward function, determine the impact of the data in each sample evaluation index on the reward function. When calculating the weight of the sample evaluation index, the calculation of the weight of the sample evaluation index is shown in Formulas 10 and 11:

[0131] g i =Max(X i )-Min(X i ) Formula 10

[0132]

[0133] Among them, the weight of the sample evaluation index is adjusted based on the variation range of each index in the sample evaluation index, g i Indicates the range of change of a certain indicator i, Max(X i ) and Min(X i ) represent a vector X i The maximum and minimum values ​​of w i is the weight of the sample evaluation index corresponding to the indicator in the sample evaluation index, such as the weights w1, w2, etc. corresponding to w1, w2, etc. in formula 9, g j It represents the set variation range of all sampled indicators j, and n is the number of sample evaluation indicators corresponding to the N sample users sampled in total.

[0134] It represents the initial weight of each indicator in the calculated sample evaluation index. The specific calculation process is shown in Formula 12:

[0135]

[0136] Among them, Formula 12 is a relationship diagram of randomly selecting N sample users from the data source in advance, selecting N sample users to execute the sample action strategy of the sample next base station corresponding to the sample current base station according to the random action strategy, and calculating each indicator in the sample evaluation index of the sample reward function corresponding to each sample action strategy. The initial weight of the i-th indicator in the sample evaluation index is expressed as For example, the initial weights corresponding to w1, w2, etc. in formula 9 etc. Mean(·) represents the average value of a vector, such as Mean(X i ) represents X i The average value of all the values ​​calculated by the i-th indicator is recorded as vector X i .X j It indicates that all the sampled indicators are all the calculated values ​​of j, and n is the number of sample evaluation indicators corresponding to the N sample users sampled in total.

[0137] It should be understood that the initial weights and the weights of the calculated sample evaluation indicators are calculated and determined during the training process of the reinforcement learning model. Finally, the sample reward function generated by combining various indicators can be directly applied when using the sample reward function to evaluate the sample action strategy executed by the agent, thereby obtaining the sample evaluation result, and the weights and other parameters of the sample reward function can be adjusted at any time according to the sample execution results of the agent.

[0138] The sample evaluation results are used to determine the authenticity of the agent's execution of the sample action strategy of selecting the sample next base station corresponding to the sample current base station.

[0139] In this embodiment, the reward function is used to judge the impact of the sample action strategy executed by the reinforcement learning model on the sample movement trajectory, and the sample evaluation result is determined. If the sample movement trajectory can be made more continuous and smooth, or backtracking and disconnection can be reduced, positive reward feedback is given to the sample action strategy executed by the agent, indicating that the authenticity of the sample action strategy is relatively strong; if there are trajectory breaks, sharp turns, path redundancy or multiple disconnected subgraphs, the agent is punished by negative rewards, indicating that the authenticity of the sample action strategy is relatively weak, that is, the next base station of the selected sample may not be real. By adjusting the parameters of the reinforcement learning model to be trained, the reinforcement learning model to be trained is trained so that the agent of the reinforcement learning model executes the correct sample action strategy to obtain a true sample movement trajectory.

[0140] In one embodiment, when the user's stay time at a track point in the target movement trajectory exceeds a preset dynamic threshold, the track point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point (i.e., S108). The following steps E1-E4 may be executed:

[0141] Step E1: Obtain the number of trajectory points in the target movement trajectory corresponding to multiple users, and obtain the initial threshold of the user stay time of multiple users at the trajectory point. According to the number of trajectory points of multiple users, the multiple users are segmented and sorted to obtain the segmented sorting results.

[0142] After obtaining the target user's target movement trajectory, traditional stop point recognition algorithms have limitations when dealing with frequent travel scenarios. For example, drivers or riders often have disorganized movements, often aiming to deliver customers or goods to a specific location and then quickly depart. Due to the short duration of a trip, stop point recognition methods using fixed thresholds struggle to accurately identify the user's origin-destination travel chain. Therefore, an adaptive threshold algorithm can be employed to dynamically adjust the threshold based on the target user's trajectory characteristics, enabling more precise identification of travel chains in complex scenarios.

[0143] Specifically, the initial threshold is the time a user initially spends at a trajectory point, as set by the human operator. When a user's stay time exceeds the initial threshold, the point is considered a stop. By obtaining the actual target movement trajectories of a large number of users, the number of trajectory points in each user's target movement trajectory is determined. Users are then segmented and sorted by the number of trajectory points, resulting in a segmented sorting result. This segmented sorting result includes grouping users with similar numbers of trajectory points into the same segment.

[0144] As an example, calculate the number of track points of the target movement trajectories of multiple users and sort the users by the number of track points. Divide the users into five segments based on the number of track points, and record them in ascending order as {seg1, seg2, seg3, seg4, seg5}. The multiple users may or may not include the target user.

[0145] Step E2: dynamically adjust the initial threshold of the target user according to the segment sorting result to obtain a preset dynamic threshold.

[0146] Based on the segment sorting result obtained in step E1, the initial threshold in the segment corresponding to the target user is adjusted according to the segment sorting to obtain a preset dynamic threshold corresponding to the number of trajectory points in the target movement trajectory of the target user.

[0147] Following the example in step E1, in the adaptive stay point identification algorithm, the initial threshold is denoted as: T base , according to the segmented sorting results of the number of trajectory points, the initial threshold corresponding to each segment of users is dynamically adjusted. The number of trajectory points is seg1, and the corresponding threshold can be T base ×100%; the number of trajectory points is seg2, and the corresponding preset dynamic threshold can be T base ×80%; the number of trajectory points is seg3, and the corresponding preset dynamic threshold can be T base ×60%; the number of trajectory points is seg4, and the corresponding preset dynamic threshold can be T base ×40%; the number of trajectory points is seg5, and the corresponding preset dynamic threshold can be T base ×20%.

[0148] The greater the number of trajectory points, the longer the user's target trajectory and the longer their itinerary chain, indicating that the user may have stopped at more locations, and the corresponding smaller the preset dynamic threshold value. Based on the number of trajectory points corresponding to the target user's target trajectory, the target user is assigned to the segmented sorting results, and the initial threshold for the target user is dynamically adjusted to obtain the target user's preset dynamic threshold value.

[0149] Step E3: When the user's stay time at a track point in the target movement track exceeds a preset dynamic threshold, the track point is determined as the target user's stay point.

[0150] Step E4: Determine two consecutive stopover points as the target user's travel chain.

[0151] Through the preset dynamic threshold, abnormal travel chains with abnormal conditions such as short travel time and short travel distance in the target movement trajectory of the target user are filtered out to ensure the accuracy and reliability of the identified travel chain of the target user.

[0152] It should be noted that based on the target user's movement trajectory, a first trajectory point corresponding to the starting stop point and a second trajectory point corresponding to the ending stop point are determined. The distance and time interval between the first and second trajectory points are calculated to determine the driving speeds of the starting and ending points. Based on the driving speeds, the transportation mode of each pair of starting and ending points is recorded according to the speed ranges of different travel modes. The starting and ending points are two consecutive stop points.

[0153] In this embodiment, the target user's stay point is determined based on the number of trajectory points in the target user's target movement trajectory by dynamically adjusting the initial threshold, and two consecutive stay points are determined as the target user's travel chain. In addition, the initial threshold is dynamically adjusted by the adaptive threshold to reduce the initial threshold of the target user with more trajectory points, so as to identify more detailed travel chains, which can meet a variety of travel scenarios.

[0154] Figure 4 is a schematic flow chart of a method for identifying a travel chain according to another embodiment of the present application, such as Figure 4 As shown, the method includes the following steps:

[0155] S401: Acquire multiple base stations related to the movement trajectory of a target user, and acquire target user data in each base station.

[0156] S402: Determine the base station location of the base station corresponding to the mobile phone signaling data according to the mobile phone signaling data in the target user data and the base station data of the corresponding base station.

[0157] S403 , connecting the base stations in chronological order according to the mobile phone signaling data and the corresponding base station locations, and obtaining a relationship diagram with adjacency relationships between the base stations.

[0158] S404: Input the relationship graph into the trained reinforcement learning model, and determine the neighborhood information of the current base station in the relationship graph based on the adjacency relationship between the base stations in the relationship graph.

[0159] The neighborhood information includes: in the relationship graph, the characteristics of the node corresponding to the current base station, the characteristics of the first-order nodes directly connected to the node corresponding to the current base station, and the characteristics of the second-order neighboring nodes connected to the first-order nodes.

[0160] S405 , using a graph autoencoder to convert a neighborhood graph corresponding to the neighborhood information of the current base station into a feature vector, and input the feature vector into the agent in the reinforcement learning model.

[0161] S406, the agent calculates the value of executing the action strategy for determining whether the base station relationship corresponds to the location movement of the target user based on the neighborhood information of the current base station, the state value function and the advantage function.

[0162] S407 , determining an execution result based on the value, where the execution result includes: the base station relationship of the current base station corresponds to the position movement of the target user, or the base station relationship of the current base station does not correspond to the position movement of the target user.

[0163] S408: Determine the target base station according to the execution result, and output the target movement trajectory of the target user through the reinforcement learning model based on the target base station.

[0164] S409: When the user's stay time at a track point in the target movement track exceeds a preset dynamic threshold, the track point is determined as a stay point of the target user, and the target user's travel chain is identified based on the stay point.

[0165] Continuous stop points are determined as a travel chain. In the target movement trajectory, the target user's travel chain can have multiple segments.

[0166] The specific process from S401 to S409 has been described in detail in the above embodiment and will not be repeated here.

[0167] The technical solution of the embodiment of the present application is adopted to obtain target user data from multiple base stations related to the target user's mobile trajectory. Based on the target user data, a relationship graph corresponding to the base station is determined. The relationship graph is used to characterize the adjacency relationship between each base station. The target user data includes: the target user's mobile phone signaling data and the base station data of the corresponding base station; the relationship graph is input into the trained reinforcement learning model, and the reinforcement learning model is used to execute an action strategy for determining whether the base station relationship corresponds to the target user's position movement to obtain an execution result, wherein the base station relationship represents the connection relationship between the target user's nodes in the relationship graph. The reinforcement learning model is trained based on a reward function, and the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: the continuity change, smoothness change, retroactive change and connectivity change of the mobile trajectory after executing the action strategy; based on the execution result, the target mobile trajectory of the target user is output through the reinforcement learning model; when the user's stay time at a trajectory point in the target mobile trajectory exceeds a preset dynamic threshold, the trajectory point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

[0168] A base station relationship graph is determined based on the target user data collected by the base station. This graph is then input into a trained reinforcement learning model. Based on the graph, an action strategy is executed to determine whether the base station relationship corresponds to the target user's location movement, yielding an execution result. Because the reinforcement learning model is trained based on a reward function, after executing the action strategy, the target user's movement trajectory exhibits no issues with continuity, smoothness, retrospectiveness, or connectivity. Furthermore, due to the high authenticity of the selected execution results, problematic base stations that cause changes in the movement trajectory can be screened out. Based on the execution results, the actual base stations in the relationship graph are determined, and the target movement trajectory of the target user is output using the reinforcement learning model. When the user's dwell time at a point in the target movement trajectory exceeds a preset dynamic threshold, the point is identified as the target user's dwell point, and the target user's travel chain is identified based on the dwell point. When identifying the target user's movement trajectory, the trained reinforcement learning model focuses on each step of the target user's base station selection and evaluates the action strategy executed at each step using the evaluation metric in the reward function. This allows the user to identify base stations that the user did not actually pass through when there are issues such as drift or ping-pong switching, resulting in accurate execution results. Compared with fixed pattern or clustering methods, it can focus on the actions performed at each step and is not restricted by factors such as fixed parameters and data size, making the final target movement trajectory more accurate and solving the problem of low accuracy in identifying the user's travel chain.

[0169] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.

[0170] The above is a method for identifying a travel chain provided in an embodiment of the present application. Based on the same idea, an embodiment of the present application also provides a device for identifying a travel chain.

[0171] Figure 5 FIG. 1 is a schematic diagram of a structure of a travel chain identification device according to an embodiment of the present invention. Figure 5 As shown, the identification device of the travel chain includes: an acquisition module 51, a reward module 52, an output module 53, and a determination module 54:

[0172] An acquisition module 51 is configured to acquire target user data from multiple base stations related to the target user's movement trajectory, and determine a relationship graph corresponding to the base stations based on the target user data. The relationship graph is used to represent the adjacency relationship between the base stations. The target user data includes: mobile phone signaling data of the target user and base station data of the corresponding base stations;

[0173] A reward module 52 is configured to input the relationship graph into a trained reinforcement learning model, execute an action strategy for determining whether the base station relationship corresponds to the target user's location movement through the reinforcement learning model, and obtain an execution result, wherein the base station relationship represents the connection relationship between the target user's nodes in the relationship graph, and the reinforcement learning model is trained based on a reward function, wherein the reward function includes: an evaluation index and a weight of the evaluation index; and the evaluation index includes: a change in continuity, a change in smoothness, a change in retrospectiveness, and a change in connectivity of the movement trajectory after executing the action strategy;

[0174] An output module 53 is used to output the target movement trajectory of the target user through the reinforcement learning model according to the execution result;

[0175] The determination module 54 is configured to determine the trajectory point as a target user's stay point when the user's stay time at the trajectory point in the target movement trajectory exceeds a preset dynamic threshold, and identify the target user's travel chain based on the stay point.

[0176] In one embodiment, the acquisition module 51 is specifically used to: determine the base station position of the base station corresponding to the mobile phone signaling data based on the mobile phone signaling data in the target user data and the base station data of the corresponding base station, the base station data including the base station code and the base station position of the base station; according to the base station position, the base stations are connected in chronological order to obtain a relationship graph with adjacency relationships between the base stations; wherein the relationship graph includes node features and edge features, the node features represent the features of the base station position of the base station, and the edge features represent the distance between the base stations.

[0177] In one embodiment, the reward module 52 is specifically used to: input the relationship graph into the trained reinforcement learning model, determine the position of the target user in the relationship graph based on the adjacency relationship between the base stations in the relationship graph, and determine the position of the target user in the relationship graph as the current base station; determine the neighborhood information of the current base station based on the current base station; the neighborhood information includes: in the relationship graph, the node features corresponding to the current base station, the features of the first-order nodes directly connected to the node corresponding to the current base station, and the features of the second-order neighboring nodes connected to the first-order nodes; use the graph autoencoder to convert the neighborhood graph corresponding to the neighborhood information of the current base station into a feature vector; based on the feature vector corresponding to the neighborhood information of the current base station, calculate the value of the action strategy of executing the judgment whether the base station relationship of the feature vector corresponds to the position movement of the target user through the reinforcement learning model, and the value includes: the numerical value calculated based on the state value function and the advantage function; determine the execution result based on the value, and the execution result includes: the base station relationship of the current base station corresponds to the position movement of the target user, or the base station relationship of the current base station does not correspond to the position movement of the target user.

[0178] In one embodiment, the apparatus further comprises a training module, the training module comprising:

[0179] an acquisition unit, configured to acquire a sample relationship graph of sample base stations corresponding to a plurality of sample users, and sample user movement trajectories of the sample users, and calculate a sample feature vector corresponding to sample neighborhood information of the sample nodes in the sample relationship graph by using a graph data processing method, wherein the sample neighborhood information is used to characterize the adjacency relationship between the sample nodes;

[0180] a computing unit, configured to input sample feature vectors corresponding to sample nodes in the sample relationship graph into the reinforcement learning model to be trained, and to compute, through the agent of the reinforcement learning model to be trained, a sample value of a sample action strategy for determining whether a sample base station relationship corresponds to a position movement of a sample user, wherein the sample base station relationship represents a connection relationship between nodes of the sample user in the sample relationship graph;

[0181] An execution unit, configured to determine a sample execution result according to the sample value, the sample execution result including: a sample base station relationship of the sample node corresponds to a position movement of the sample user, or a base station relationship of the sample node does not correspond to a position movement of the target user;

[0182] An evaluation unit is used to evaluate the execution of a sample action strategy using a sample reward function set in the agent's interaction environment, obtain a sample evaluation result, and output a sample target movement trajectory of the sample user through the reinforcement learning model to be trained based on the sample evaluation result;

[0183] The loss unit is used to calculate the loss function according to the sample reward functions corresponding to multiple sample users, adjust the parameters of the reinforcement learning model to be trained based on the loss function, the movement trajectory of the sample users and the movement trajectory of the sample targets, train the reinforcement learning model to be trained, and obtain the trained reinforcement learning model.

[0184] In one embodiment, the evaluation unit is further used to determine, in the sample relationship graph, the continuity change of the sample movement trajectory of the sample user after the sample action strategy is executed based on the set of average distances of the edges connected to each sample base station before the sample action strategy is executed, and the set of average distances of the edges connected to each sample base station after the sample action strategy is executed; determine the smoothness change of the sample movement trajectory of the sample user after the sample action strategy is executed based on the turning angles of two adjacent edges connected to the sample base station; determine the backtracking of the sample movement trajectory of the sample user after the sample action strategy is executed based on the number of visits to the sample base station before the sample action strategy is executed, and the number of visits to the sample base station after the sample action strategy is executed. The method comprises the following steps: determining the connectivity change of the sample movement trajectory of the sample user after executing the sample action strategy based on the connectivity subgraph of the base station before executing the sample action strategy and the connectivity subgraph of the sample base station after executing the sample action strategy; determining the continuity change, smoothness change, retroactive change and connectivity change as the sample evaluation indicators of the sample reward function, and adjusting the weight of the sample evaluation indicator based on the change of the sample evaluation indicator; determining the sample reward function based on the sample evaluation indicator and the weight of the sample evaluation indicator, and evaluating the sample action strategy based on the sample reward function to obtain the sample evaluation result, which is used to train the sample execution result of the intelligent agent.

[0185] In one embodiment, the determination module 54 is used to: obtain the number of trajectory points in the target movement trajectory corresponding to multiple users, and obtain the initial threshold of the user residence time of multiple users at the trajectory point; perform segmented sorting processing on the multiple users according to the number of trajectory points of the multiple users to obtain segmented sorting results; dynamically adjust the initial threshold of the target user according to the segmented sorting results to obtain a preset dynamic threshold; when the user residence time of a trajectory point in the target movement trajectory exceeds the preset dynamic threshold, determine the trajectory point as the target user's residence point; and determine two consecutive residence points as the target user's travel chain.

[0186] The technical solution of the embodiment of the present application is adopted to obtain target user data from multiple base stations related to the target user's mobile trajectory. Based on the target user data, a relationship graph corresponding to the base station is determined. The relationship graph is used to characterize the adjacency relationship between each base station. The target user data includes: the target user's mobile phone signaling data and the base station data of the corresponding base station; the relationship graph is input into the trained reinforcement learning model, and the reinforcement learning model is used to execute an action strategy for determining whether the base station relationship corresponds to the target user's position movement to obtain an execution result, wherein the base station relationship represents the connection relationship between the target user's nodes in the relationship graph. The reinforcement learning model is trained based on a reward function, and the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: the continuity change, smoothness change, retroactive change and connectivity change of the mobile trajectory after executing the action strategy; based on the execution result, the target mobile trajectory of the target user is output through the reinforcement learning model; when the user's stay time at a trajectory point in the target mobile trajectory exceeds a preset dynamic threshold, the trajectory point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

[0187] A base station relationship graph is determined based on the target user data collected by the base station. This graph is then input into a trained reinforcement learning model. Based on the graph, an action strategy is executed to determine whether the base station relationship corresponds to the target user's location movement, yielding an execution result. Because the reinforcement learning model is trained based on a reward function, after executing the action strategy, the target user's movement trajectory exhibits no issues with continuity, smoothness, retrospectiveness, or connectivity. Furthermore, due to the high authenticity of the selected execution results, problematic base stations that cause changes in the movement trajectory can be screened out. Based on the execution results, the actual base stations in the relationship graph are determined, and the target movement trajectory of the target user is output using the reinforcement learning model. When the user's dwell time at a point in the target movement trajectory exceeds a preset dynamic threshold, the point is identified as the target user's dwell point, and the target user's travel chain is identified based on the dwell point. When identifying the target user's movement trajectory, the trained reinforcement learning model focuses on each step of the target user's base station selection and evaluates the action strategy executed at each step using the evaluation metric in the reward function. This allows the user to identify base stations that the user did not actually pass through when there are issues such as drift or ping-pong switching, resulting in accurate execution results. Compared with fixed pattern or clustering methods, it can focus on the actions performed at each step and is not restricted by factors such as fixed parameters and data size, making the final target movement trajectory more accurate and solving the problem of low accuracy in identifying the user's travel chain.

[0188] Those skilled in the art should understand that Figure 5The travel chain identification device in can be used to implement the travel chain identification method described above, and the detailed description should be similar to the description of the method part above. To avoid tediousness, it will not be repeated here.

[0189] Based on the same technical concept, an embodiment of the present application further provides an electronic device for executing the above-mentioned travel chain identification method. Figure 6 The following is a schematic diagram of the structure of an electronic device for implementing various embodiments of the present application. Electronic devices may vary significantly due to different configurations or performances, and may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call a computer program stored in the memory 630 and executable on the processor 610 to perform the following steps:

[0190] Obtain target user data from multiple base stations related to the target user's movement trajectory. Based on the target user data, determine a relationship graph corresponding to the base stations. The relationship graph is used to represent the adjacency relationship between the base stations. The target user data includes: the target user's mobile phone signaling data and the base station data of the corresponding base stations;

[0191] The relationship graph is input into the trained reinforcement learning model. The reinforcement learning model is used to execute an action strategy to determine whether the base station relationship corresponds to the target user's position movement, and an execution result is obtained. The current base station represents the position of the target user in the relationship graph. The reinforcement learning model is trained based on a reward function. The reward function includes: an evaluation index and its weight; the evaluation index includes: the continuity change, smoothness change, retroactivity change, and connectivity change of the movement trajectory after the action is executed;

[0192] Based on the execution results, the target user's target movement trajectory is output through the reinforcement learning model;

[0193] When the user's stay time at a track point in the target movement trajectory exceeds a preset dynamic threshold, the track point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

[0194] The technical solution of the embodiment of the present application is adopted to obtain target user data from multiple base stations related to the target user's mobile trajectory. Based on the target user data, a relationship graph corresponding to the base station is determined. The relationship graph is used to characterize the adjacency relationship between each base station. The target user data includes: the target user's mobile phone signaling data and the base station data of the corresponding base station; the relationship graph is input into the trained reinforcement learning model, and the reinforcement learning model is used to execute an action strategy for determining whether the base station relationship corresponds to the target user's position movement to obtain an execution result, wherein the base station relationship represents the connection relationship between the target user's nodes in the relationship graph. The reinforcement learning model is trained based on a reward function, and the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: the continuity change, smoothness change, retroactive change and connectivity change of the mobile trajectory after executing the action strategy; based on the execution result, the target mobile trajectory of the target user is output through the reinforcement learning model; when the user's stay time at a trajectory point in the target mobile trajectory exceeds a preset dynamic threshold, the trajectory point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

[0195] A base station relationship graph is determined based on the target user data collected by the base station. This graph is then input into a trained reinforcement learning model. Based on the graph, an action strategy is executed to determine whether the base station relationship corresponds to the target user's location movement, yielding an execution result. Because the reinforcement learning model is trained based on a reward function, after executing the action strategy, the target user's movement trajectory exhibits no issues with continuity, smoothness, retrospectiveness, or connectivity. Furthermore, due to the high authenticity of the selected execution results, problematic base stations that cause changes in the movement trajectory can be screened out. Based on the execution results, the actual base stations in the relationship graph are determined, and the target movement trajectory of the target user is output using the reinforcement learning model. When the user's dwell time at a point in the target movement trajectory exceeds a preset dynamic threshold, the point is identified as the target user's dwell point, and the target user's travel chain is identified based on the dwell point. When identifying the target user's movement trajectory, the trained reinforcement learning model focuses on each step of the target user's base station selection and evaluates the action strategy executed at each step using the evaluation metric in the reward function. This allows the user to identify base stations that the user did not actually pass through when there are issues such as drift or ping-pong switching, resulting in accurate execution results. Compared with fixed pattern or clustering methods, it can focus on the actions performed at each step and is not restricted by factors such as fixed parameters and data size, making the final target movement trajectory more accurate and solving the problem of low accuracy in identifying the user's travel chain.

[0196] The specific execution steps can refer to the various steps of the embodiment of the travel chain identification method described above, and can achieve the same technical effect. To avoid repetition, they will not be described here.

[0197] It should be noted that the electronic devices in the embodiments of the present application include: servers, terminals, or other devices other than terminals.

[0198] The above electronic device structure does not constitute a limitation of the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, the input unit may include a graphics processing unit (GPU) and a microphone, and the display unit may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. to configure the display panel. The user input unit includes a touch panel and at least one of other input devices. The touch panel is also called a touch screen. Other input devices may include but are not limited to a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0199] The memory can be used to store software programs and various data. The memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory may include a volatile memory or a non-volatile memory, or the memory may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct rambus random access memory (DRRAM).

[0200] The processor may include one or more processing units; optionally, the processor may integrate an application processor and a modem processor, wherein the application processor primarily handles operations related to the operating system, user interface, and application programs, and the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into the processor.

[0201] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned travel chain identification method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0202] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0203] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned travel chain identification method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0204] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0205] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the processor is used to run the program or instructions to implement the various processes of the above-mentioned product recommendation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0206] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0207] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0208] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for identifying a travel chain, characterized in that: The method comprises: Obtain target user data from multiple base stations related to the target user's movement trajectory, and determine a relationship graph corresponding to the base stations based on the target user data, wherein the relationship graph is used to represent the adjacency relationship between the base stations, the target user data including: mobile phone signaling data of the target user and base station data of the corresponding base stations; Inputting the relationship graph into a trained reinforcement learning model, executing an action strategy for determining whether the base station relationship corresponds to the location movement of the target user through the reinforcement learning model, and obtaining an execution result, wherein the base station relationship represents the connection relationship between the nodes of the target user in the relationship graph, and the reinforcement learning model is trained based on a reward function, the reward function including: an evaluation index and a weight of the evaluation index; the evaluation index including: a change in continuity, a change in smoothness, a change in retrospectivity, and a change in connectivity of the movement trajectory after executing the action strategy; Outputting a target movement trajectory of the target user through the reinforcement learning model according to the execution result; When the user's stay time at a track point in the target movement track exceeds a preset dynamic threshold, the track point is determined as the target user's stay point, and the target user's travel chain is identified based on the stay point.

2. The method according to claim 1, characterized in that The determining, according to the target user data, a relationship graph corresponding to the base station includes: Determining the base station location of the base station corresponding to the mobile phone signaling data according to the mobile phone signaling data in the target user data and the base station data of the corresponding base station, wherein the base station data includes a code of the base station and the base station location of the base station; According to the base station positions, the base stations are connected in chronological order to obtain the relationship graph having the adjacency relationship between the base stations; wherein the relationship graph includes node features and edge features, the node features represent the features of the base station positions of the base stations, and the edge features represent the distances between the base stations.

3. The method according to claim 1, characterized in that Inputting the relationship graph into a trained reinforcement learning model, executing an action strategy for determining whether the base station relationship corresponds to the position movement of the target user through the reinforcement learning model, and obtaining an execution result, includes: Inputting the relationship graph into the trained reinforcement learning model, determining the position of the target user in the relationship graph according to the adjacency relationship between the base stations in the relationship graph, and determining the position of the target user in the relationship graph as the current base station; Determining, based on the current base station, neighborhood information of the current base station; the neighborhood information including: characteristics of a node corresponding to the current base station, characteristics of a first-order node directly connected to the node corresponding to the current base station, and characteristics of a second-order neighboring node connected to the first-order node in the relationship graph; Using a graph autoencoder, converting a neighborhood graph corresponding to the neighborhood information of the current base station into a feature vector; calculating, based on the feature vector corresponding to the neighborhood information of the current base station, a value of the action strategy for determining whether the base station relationship of the feature vector corresponds to the location movement of the target user through the reinforcement learning model, the value comprising: a numerical value calculated based on a state value function and an advantage function; The execution result is determined according to the value, and the execution result includes: the base station relationship of the current base station corresponds to the position movement of the target user, or the base station relationship of the current base station does not correspond to the position movement of the target user.

4. The method according to claim 1, wherein The training of the reinforcement learning model includes: Obtain a sample relationship graph of sample base stations corresponding to multiple sample users, as well as sample user movement trajectories of the sample users, and calculate sample feature vectors corresponding to sample neighborhood information of sample nodes in the sample relationship graph using a graph data processing method, where the sample neighborhood information is used to characterize the adjacency relationship between the sample nodes; Inputting the sample feature vector corresponding to the sample node in the sample relationship graph into the reinforcement learning model to be trained, and calculating, by the agent of the reinforcement learning model to be trained, the sample value of executing the sample action strategy for determining whether the sample base station relationship corresponds to the position movement of the sample user, wherein the sample base station relationship represents the connection relationship of the sample user to the sample node in the sample relationship graph; Determining a sample execution result according to the sample value, the sample execution result including: whether the sample base station relationship of the sample node corresponds to the position movement of the sample user, or whether the base station relationship of the sample node does not correspond to the position movement of the target user; Evaluate the execution of the sample action strategy by using a sample reward function set in the agent interaction environment to obtain a sample evaluation result, and output a sample target movement trajectory of the sample user through the reinforcement learning model to be trained based on the sample evaluation result; According to the sample reward functions corresponding to the multiple sample users, a loss function is calculated; based on the loss function, the movement trajectory of the sample users and the movement trajectory of the sample targets, the parameters of the reinforcement learning model to be trained are adjusted; the reinforcement learning model to be trained is trained to obtain the trained reinforcement learning model.

5. The method according to claim 4, characterized in that The sample action strategy is evaluated by using a sample reward function set in the agent interaction environment to obtain a sample evaluation result, including: In the sample relationship graph, determining the continuity change of the sample movement trajectory of the sample user after executing the sample action strategy based on a set of average distances of edges connected to the sample base stations before executing the sample action strategy and a set of average distances of edges connected to the sample base stations after executing the sample action strategy; determining, according to the turning angles of two adjacent sides connected by the sample base stations, a smoothness change of the sample movement trajectory of the sample user after executing the sample action strategy; Determining the retrospective change of the sample movement trajectory of the sample user after executing the sample action strategy based on the number of visits to the sample base station before executing the sample action strategy and the number of visits to the sample base station after executing the sample action strategy; Determining, based on the connectivity subgraph of the base station before executing the sample action strategy and the connectivity subgraph of the sample base station after executing the sample action strategy, a connectivity change of the sample movement trajectory of the sample user after executing the sample action strategy; Determining the continuity change, the smoothness change, the retroactive change, and the connectivity change as sample evaluation indicators of the sample reward function, and adjusting the weight of the sample evaluation indicator based on the change of the sample evaluation indicator; The sample reward function is determined according to the sample evaluation index and the weight of the sample evaluation index. Based on the sample reward function, the sample action is evaluated to obtain the sample evaluation result. The sample evaluation result is used to train the sample execution result of the agent.

6. The method according to claim 1, characterized in that When the user's stay time at a track point in the target movement track exceeds a preset dynamic threshold, the track point is determined as a stay point of the target user, and the travel chain of the target user is identified according to the stay point, including: Obtaining the number of trajectory points in the target movement trajectory corresponding to multiple users, and obtaining initial thresholds of the user stay times of the multiple users at the trajectory points, and performing segmented sorting processing on the multiple users according to the number of the trajectory points of the multiple users to obtain segmented sorting results; Dynamically adjusting the initial threshold of the target user according to the segment sorting result to obtain the preset dynamic threshold; When the user's stay time at a track point in the target movement track exceeds the preset dynamic threshold, the track point is determined as the target user's stay point; Two consecutive stay points are determined as the travel chain of the target user.

7. A device for identifying a travel chain, characterized in that: include: an acquisition module, configured to acquire target user data from multiple base stations related to a target user's movement trajectory, and determine a relationship graph corresponding to the base stations based on the target user data, wherein the relationship graph is used to represent the adjacency relationship between the base stations, the target user data including: mobile phone signaling data of the target user and base station data of the corresponding base stations; A reward module is configured to input the relationship graph into a trained reinforcement learning model, execute an action strategy for determining whether the base station relationship corresponds to the location movement of the target user through the reinforcement learning model, and obtain an execution result, wherein the base station relationship represents the connection relationship between the nodes of the target user in the relationship graph, and the reinforcement learning model is trained based on a reward function, wherein the reward function includes: an evaluation index and a weight of the evaluation index; the evaluation index includes: a change in continuity, a change in smoothness, a change in retrospectiveness, and a change in connectivity of the movement trajectory after executing the action strategy; an output module, configured to output a target movement trajectory of the target user through the reinforcement learning model according to the execution result; The determination module is configured to determine the trajectory point as the target user's stay point when the user's stay time at the trajectory point in the target movement trajectory exceeds a preset dynamic threshold, and identify the target user's travel chain based on the stay point.

8. An electronic device, characterized in that: The invention comprises a processor and a memory electrically connected to the processor, wherein the memory stores a computer program, and the processor is used to call and execute the computer program from the memory to implement a travel chain identification method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The storage medium is used to store a computer program, and the computer program can be executed by a processor to implement the method for identifying a travel chain according to any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for identifying a travel chain according to any one of claims 1 to 6.