User identification method, device, computer equipment, and storage medium

By determining base station clusters on high-speed rail lines and screening user sets, target high-speed rail users can be identified, solving the problem of poor data accuracy in high-speed rail simulation road tests and improving the accuracy of high-speed rail simulation road test data.

CN116033468BActive Publication Date: 2025-09-23GUANGDONG HAIGE ICREATE TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211595280.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-09-23
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

High-speed rail simulation road tests cannot accurately identify high-speed mobile users, resulting in poor data accuracy.

Method used

By identifying multiple base station clusters on the target high-speed rail line, obtaining base station cluster access data, screening out a set of users that meet preset conditions, identifying them as target high-speed rail users, and using their communication data to conduct simulated road tests.

Benefits of technology

The accuracy of high-speed rail simulation road test data is improved, ensuring that only communication data of high-speed rail users is used for simulation road tests, eliminating interference from other types of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116033468B_ABST
    Figure CN116033468B_ABST
Patent Text Reader

Abstract

The present application relates to a user identification method, apparatus, computer equipment, storage medium and computer program product. The method includes: determining multiple base station clusters on a target high-speed rail line; the target area range is an area range with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius; obtaining base station cluster access data of each base station cluster, and determining a user set to be filtered based on the base station cluster access data; the terminal where the user in the user set to be filtered is located has visited at least two base station clusters; filtering the user set to be filtered to obtain a filtered user set; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the filtered user set is located meets a preset first filtering condition; identifying the user in the filtered user set as a target high-speed rail user associated with the target high-speed rail line; the communication data corresponding to the terminal where the target high-speed rail user is located is used to simulate a road test on the target high-speed rail line. The use of this method can improve the accuracy of high-speed rail simulation road test data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of mobile communication technology, and in particular to a user identification method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0002] With the development of science and technology, transportation is also developing rapidly. High-speed rail has become a bridge and link for economic exchanges between cities, and has also become one of the main means of transportation for users.

[0003] However, due to the high speeds of high-speed rail and the significant Doppler effect, the quality of mobile communication networks on high-speed rail is a serious issue. Collecting line information through drive testing is an essential step in optimizing high-speed rail communication networks. To improve drive testing efficiency, simulated drive testing is often used instead of on-site testing.

[0004] However, prior art uses cell-level performance metrics during high-speed rail simulation tests. These metrics include both high-speed mobile users on the train and local low-speed mobile and stationary users near the line, the majority of which are local low-speed mobile and stationary users. The simulated test data required for high-speed rail simulation tests is precisely the communication data of high-speed mobile users on the train. This approach cannot accurately identify high-speed rail users, resulting in poor accuracy in the simulated test data.

[0005] Therefore, traditional technologies have the problem of poor accuracy of high-speed rail simulation road test data. Summary of the Invention

[0006] Based on this, it is necessary to provide a user identification method, device, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of high-speed rail user identification in response to the above technical problems.

[0007] In a first aspect, the present application provides a user identification method. The method comprises:

[0008] Determine multiple base station clusters on a target high-speed rail line; the base station cluster is a collection of base station cells within a target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius;

[0009] Acquiring base station cluster access data of each base station cluster, and determining a user set to be screened based on the base station cluster access data; wherein a terminal of a user in the user set to be screened has accessed at least two of the base station clusters;

[0010] Filtering the user set to be filtered to obtain a filtered user set; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the filtered user set is located meets a preset first filtering condition;

[0011] Identify users in the filtered user set as target high-speed rail users associated with the target high-speed rail line; and use communication data corresponding to the terminal where the target high-speed rail user is located to perform simulated road testing on the target high-speed rail line.

[0012] In one embodiment, the base station cluster access data includes a time point at which a terminal of a user in the to-be-screened user set accesses the base station cluster; and screening the to-be-screened user set to obtain a screened user set includes:

[0013] Determine the actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the access time point of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located;

[0014] Determine a predicted occupancy time of a base station cluster corresponding to a terminal where a user in the set of users to be screened is located; the predicted occupancy time of the base station cluster matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line;

[0015] The users in the user set to be screened are screened according to the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located and the corresponding predicted occupancy time of the base station cluster to obtain the screened user set.

[0016] In one embodiment, the actual base station cluster occupancy time includes the actual base station cluster occupancy period; the predicted base station cluster occupancy time includes the predicted base station cluster occupancy period; the filtering of users in the user set to be filtered based on the actual base station cluster occupancy time corresponding to the terminal where the user is located and the corresponding predicted base station cluster occupancy time to obtain the filtered user set includes:

[0017] Determine the total actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the actual occupancy time period of the base station cluster;

[0018] Determine, based on the actual occupied period of the base station cluster and the corresponding predicted occupied period of the base station cluster, the overlapping occupied period of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located;

[0019] Determine the total overlapping occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the overlapping occupancy time period of the base station cluster;

[0020] From the set of users to be screened, users whose terminal's corresponding target duration ratio is greater than a preset ratio threshold are screened out to obtain the screened user set; the target duration ratio is the ratio between the total overlapping duration of the base station cluster occupancy and the total actual occupancy duration of the corresponding base station cluster.

[0021] In one embodiment, the method further comprises:

[0022] Determining a user set to be clustered from the set of users to be screened; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the set of users to be clustered is not satisfied with the preset first screening condition;

[0023] Performing clustering processing on the users in the to-be-clustered user set to obtain a clustered user set; base station cluster access characteristics corresponding to terminals where users in the clustered user set are located meet a preset second screening condition;

[0024] Identify the users in the clustered user set as the target high-speed rail users.

[0025] In one embodiment, the base station cluster access data includes a time point at which a terminal of a user in the set of users to be screened accesses a base station cluster; and clustering the users in the set of users to be clustered to obtain a clustered user set includes:

[0026] Obtaining equivalent time intervals corresponding to the multiple base station clusters; the equivalent time intervals are obtained by dividing the total occupied time period of the target base station cluster corresponding to the target high-speed rail line into equal values; the total occupied time period of the target base station cluster is determined based on the earliest access base station cluster time point and the latest access base station cluster time point among the access base station cluster time points;

[0027] Clustering the users in the user set to be clustered according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the users in the user set to be clustered are located, to obtain multiple user clusters;

[0028] Taking a user cluster with a user number greater than a preset number threshold among the multiple user clusters as a target user cluster;

[0029] Add users in the target user cluster to the clustered user set.

[0030] In one embodiment, clustering the users in the user set to be clustered according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the user in the user set to be clustered is located to obtain multiple user clusters includes:

[0031] Determining access feature similarities between terminals where users in the set of users to be clustered are located, based on equivalent time intervals where time points at which terminals where users in the set of users to be clustered access the base station cluster belong;

[0032] In the set of users to be clustered, users whose terminals have access feature similarities greater than a preset similarity threshold are added to the same user cluster.

[0033] In a second aspect, the present application further provides a user identification device. The device comprises:

[0034] A base station cluster determination module is configured to determine multiple base station clusters on a target high-speed rail line; the base station cluster is a collection of base station cells within a target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius;

[0035] A user determination module is configured to obtain base station cluster access data of each base station cluster and determine a user set to be screened based on the base station cluster access data; wherein a terminal of a user in the user set to be screened has accessed at least two base station clusters;

[0036] A screening module is configured to screen the user set to be screened to obtain a screened user set; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the screened user set is located satisfies a preset first screening condition;

[0037] An identification module is used to identify users in the filtered user set as target high-speed rail users associated with the target high-speed rail line; and communication data corresponding to the terminal where the target high-speed rail user is located is used to perform simulated road testing on the target high-speed rail line.

[0038] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0039] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0040] In a fifth aspect, the present application further provides a computer program product, which includes a computer program that implements the steps of the above method when executed by a processor.

[0041] The above-mentioned user identification method, device, computer equipment, storage medium and computer program product, by determining multiple base station clusters on the target high-speed rail line; the base station cluster is a collection of base station cells within the target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius; the signal coverage of the base station cells of the target high-speed rail line can be effectively planned; by obtaining the base station cluster access data of each base station cluster, and determining the user set to be screened based on the base station cluster access data; the terminal where the user in the user set to be screened is located has visited at least two base station clusters; since the terminal where the high-speed rail user is located has usually visited at least two base station clusters on the target high-speed rail line, and even if the high-speed rail user has only visited one base station cluster, the communication data of the high-speed rail user cannot provide a reference for the high-speed rail simulation road test data, therefore, through the base station cluster access data corresponding to the base station cluster on the target high-speed rail line, the user whose terminal has visited at least two base station clusters is determined, and among the users whose terminals have visited the base station cluster on the target high-speed rail line, the corresponding communication data that cannot be used for the high-speed rail simulation road test can be accurately excluded.

[0042] By filtering the set of users to be filtered, a filtered user set is obtained; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the filtered user set is located meets the preset first filtering condition; the user in the filtered user set is identified as a target high-speed rail user associated with the target high-speed rail line; the communication data corresponding to the terminal where the target high-speed rail user is located is used to simulate a road test on the target high-speed rail line. In this way, since the actual occupancy time of the base station cluster corresponding to the base station cluster on the target high-speed rail line for the terminal where the high-speed rail user is located meets certain conditions, the users whose terminals have visited at least two base station clusters are filtered through the actual occupancy time of the base station cluster corresponding to the terminal where the user in the set of users to be filtered can accurately filter out the high-speed rail users whose corresponding communication data on the target high-speed rail line can be used for the high-speed rail simulated road test, and the high-speed rail users associated with the target high-speed rail line can be accurately simulated on the target high-speed rail line through the communication data corresponding to the terminal where the target high-speed rail user is located, thereby effectively improving the accuracy of the simulated road test data of the target high-speed rail line. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A schematic flow chart of a user identification method in one embodiment;

[0044] Figure 2 A schematic diagram of multiple target areas on a target high-speed railway line in one embodiment;

[0045] Figure 3 A schematic diagram of base station cluster identifiers corresponding to a base station cluster on a target high-speed railway line in one embodiment;

[0046] Figure 4A schematic diagram of an initial window corresponding to a terminal where a user is located in one embodiment;

[0047] Figure 5 This is a schematic diagram of a first sliding window corresponding to a terminal where a user is located in one embodiment;

[0048] Figure 6 This is a schematic diagram of a second sliding window corresponding to a terminal where a user is located in one embodiment;

[0049] Figure 7 This is a schematic diagram of a target time window corresponding to a terminal where a user is located in one embodiment;

[0050] Figure 8 A schematic diagram of two target time windows corresponding to a terminal where a user is located in one embodiment;

[0051] Figure 9 This is a schematic diagram of a target time window corresponding to a terminal where another user is located in one embodiment;

[0052] Figure 10 A schematic diagram of a predicted occupancy period of a base station cluster and an actual occupancy period of a base station cluster in one embodiment;

[0053] Figure 11 This is a schematic diagram of equivalent time intervals corresponding to multiple base station clusters on a target high-speed railway line in one embodiment;

[0054] Figure 12 is a flow chart of a user identification method in another embodiment;

[0055] Figure 13 is a flow chart of another user identification method in one embodiment;

[0056] Figure 14 is a structural block diagram of a user identification device in one embodiment;

[0057] Figure 15 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0059] In one embodiment, Figure 1As shown, a user identification method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0060] Step S110 , determining multiple base station clusters on the target high-speed rail line.

[0061] The base station cluster is a collection of base station cells within the target area.

[0062] The target area range is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius.

[0063] The preset distance is a constant set based on prior knowledge.

[0064] Among them, the target high-speed rail line is the high-speed rail line to be subjected to simulated road testing.

[0065] In the specific implementation, since the target high-speed rail line, including the stations, will be covered by 4G (4th generation mobile communication technology) / 5G (5th Generation Mobile Communication Technology, fifth generation mobile communication technology) base stations along the way, the terminal can use at least one sampling point on the target high-speed rail line as the center of the circle and a preset distance as the radius to determine multiple target area ranges, so that the set of base station cells within each target area can be regarded as a base station cluster, and can be regarded as a "large base station", so that the terminal can determine multiple base station clusters on the target high-speed rail line.

[0066] In actual applications, the terminal can randomly select sampling points on the target high-speed rail line to determine multiple base station clusters on the target high-speed rail line.

[0067] To facilitate understanding by those skilled in the art, Figure 2 A schematic diagram of multiple target areas on a target high-speed rail line is provided. Figure 2 As shown, the set of all base station cell combinations within each circular range is considered a base station cluster. To conserve resources, the number of base station cells in the multiple base station clusters ultimately determined by the terminal can be no less than 30% and no more than 70% of the total number of base station cells along the target high-speed rail line.

[0068] Step S120 : acquiring base station cluster access data of each base station cluster, and determining a user set to be screened based on the base station cluster access data.

[0069] The terminal where the user in the to-be-screened user set is located has visited at least two base station clusters.

[0070] The base station cluster access data of each base station cluster includes a mapping relationship between a terminal identifier and a base station cluster identifier, which is used to represent a mapping relationship between a terminal identifier corresponding to a user terminal and a base station cluster identifier corresponding to a visited base station cluster.

[0071] The base station cluster access data of each base station cluster can be determined based on the historical XDR (External Data Representation) signaling data of each base station cluster within the same preset time period (wherein the time period can be in units of days, weeks, months, etc.).

[0072] The historical XDR signaling data includes a base station cluster identifier, a terminal identifier corresponding to the user's terminal, and a base station cluster access time point corresponding to when the user's terminal accesses the base station cluster.

[0073] The base station cluster identifier can be obtained by sequentially identifying multiple base station clusters determined by the terminal according to the line sequence of the target high-speed railway line (regardless of the line direction). Figure 3 As shown in a schematic diagram of base station cluster identification corresponding to a base station cluster on a target high-speed rail line, according to the order of multiple base station clusters on the target high-speed rail line, the base station cluster identification corresponding to each base station cluster is B1, B2, B3, B4, and B5.

[0074] Among them, the mapping relationship between the above-mentioned terminal identifier and the base station cluster identifier is obtained by the terminal by performing user deduplication on the acquired historical XDR signaling data, so that when the user's terminal visits the same base station cluster multiple times, only the mapping relationship between the terminal identifier corresponding to the user's terminal and the base station cluster identifier corresponding to the visited base station cluster is retained.

[0075] In a specific implementation, the terminal can obtain the base station cluster access data of each base station cluster. Based on the base station cluster access data, among the users whose terminals have visited the base station clusters on the target high-speed rail line, the users whose terminals have visited at least two base station clusters are determined to form a user set to be screened.

[0076] Specifically, since the base station cluster access data of each base station cluster includes a mapping relationship between the terminal identifier corresponding to the user's terminal and the base station cluster identifier corresponding to the visited base station cluster, the terminal can use a preset counting method (for example, a map-reduce method) to count the number of visits of the user's terminal to different base station clusters. For example, if the user's terminal ua has visited base station clusters B1, B3, and B5, its count is 3. Therefore, the terminal can delete the users whose terminals have only visited one base station cluster among the users whose terminals have visited the base station clusters on the target high-speed rail line based on the counting result corresponding to the user's terminal, and obtain a set of users to be screened.

[0077] For example, the mapping relationship between terminal identifiers (assuming that the terminal identifiers corresponding to the user's terminal include u1, u2, and u3) and base station cluster identifiers is shown in Table 1 (assuming that the base station clusters on the target high-speed rail line include base station cluster B1, base station cluster B2, and base station cluster B3). As shown in Table 1, since user terminal u3 has only visited base station cluster B1, the terminal needs to delete the user corresponding to user terminal u3 from the users whose terminals have visited the base station cluster on the target high-speed rail line.

[0078] Table 1 Mapping relationship between terminal identifier and base station cluster identifier

[0079]

[0080] Step S130 , filtering the user set to be filtered to obtain a filtered user set.

[0081] The actual occupancy time of the base station cluster corresponding to the terminal where the user in the filtered user set is located meets the preset first filtering condition.

[0082] The actual occupancy time of the base station cluster corresponding to the user's terminal is the actual occupancy time corresponding to when the user's terminal accesses the base station cluster.

[0083] In a specific implementation, the terminal can filter the users in the user set to be filtered according to the actual occupancy time of the base station cluster corresponding to the terminal where the user is located, and filter out the users whose actual occupancy time of the base station cluster corresponding to the terminal meets the preset first filtering condition to form the filtered user set.

[0084] Step S140: Identify users in the filtered user set as target high-speed rail users associated with the target high-speed rail line.

[0085] Among them, the communication data corresponding to the terminal where the target high-speed rail user is located is used to perform simulated road testing on the target high-speed rail line.

[0086] In a specific implementation, the terminal can identify the users in the filtered user set as target high-speed rail users associated with the target high-speed rail line, and the communication data corresponding to the terminal where the target high-speed rail user is located can be used to perform a simulated road test on the target high-speed rail line, so as to obtain the simulated road test data corresponding to the target high-speed rail line based on the communication data corresponding to the terminal where the target high-speed rail user is located.

[0087] In the above user identification method, multiple base station clusters on the target high-speed rail line are determined; the base station cluster is a collection of base station cells within the target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius; the base station cells of the target high-speed rail line can be effectively planned to have signal coverage; the base station cluster access data of each base station cluster is obtained, and the user set to be screened is determined based on the base station cluster access data; the terminal where the user in the user set to be screened is located has visited at least two base station clusters; since the terminal where the high-speed rail user is located has usually visited at least two base station clusters on the target high-speed rail line, and even if the high-speed rail user has only visited one base station cluster, the communication data of the high-speed rail user cannot provide a reference for the high-speed rail simulation road test data. Therefore, the base station cluster access data corresponding to the base station cluster on the target high-speed rail line is used to determine the user whose terminal has visited at least two base station clusters, and among the users whose terminals have visited the base station cluster on the target high-speed rail line, the corresponding communication data that cannot be used for the high-speed rail simulation road test can be accurately excluded.

[0088] By filtering the set of users to be filtered, a filtered user set is obtained; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the filtered user set is located meets the preset first filtering condition; the user in the filtered user set is identified as a target high-speed rail user associated with the target high-speed rail line; the communication data corresponding to the terminal where the target high-speed rail user is located is used to simulate a road test on the target high-speed rail line. In this way, since the actual occupancy time of the base station cluster corresponding to the base station cluster on the target high-speed rail line for the terminal where the high-speed rail user is located meets certain conditions, the users whose terminals have visited at least two base station clusters are filtered through the actual occupancy time of the base station cluster corresponding to the terminal where the user in the set of users to be filtered can accurately filter out the high-speed rail users whose corresponding communication data on the target high-speed rail line can be used for the high-speed rail simulated road test, and the high-speed rail users associated with the target high-speed rail line can be accurately simulated on the target high-speed rail line through the communication data corresponding to the terminal where the target high-speed rail user is located, thereby effectively improving the accuracy of the simulated road test data of the target high-speed rail line.

[0089] In one embodiment, the base station cluster access data includes the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; the user set to be screened is screened to obtain the screened user set, including: determining the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened according to the access base station cluster time point corresponding to the terminal where the user is located in the user set to be screened; determining the predicted occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened; the predicted occupancy time of the base station cluster is matched with the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line; based on the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened and the corresponding predicted occupancy time of the base station cluster, the users in the user set to be screened are screened to obtain the screened user set.

[0090] The base station cluster access data of each base station cluster also includes the corresponding access time point of the base station cluster when the terminal of the user in the to-be-screened user set accesses the base station cluster.

[0091] The predicted occupancy time of the base station cluster is the corresponding predicted occupancy time when the user's terminal accesses the base station cluster.

[0092] In a specific implementation, when the terminal is filtering the user set to be filtered and obtaining the filtered user set, the terminal can determine the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set is located based on the access time point of the base station cluster corresponding to the terminal where the user in the user set is located. In addition, the terminal can also determine the predicted occupancy time of the base station cluster corresponding to the terminal where the user in the user set is located, and the predicted occupancy time of the base station cluster matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line; in this way, the terminal can filter the users in the user set to be filtered based on the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set is located, and the corresponding predicted occupancy time of the base station cluster, to obtain the filtered user set.

[0093] The technical solution of this embodiment is that the base station cluster access data includes the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; by determining the actual base station cluster occupancy time corresponding to the terminal where the user in the user set to be screened according to the access base station cluster time point corresponding to the terminal where the user in the user set to be screened; determining the predicted base station cluster occupancy time corresponding to the terminal where the user in the user set to be screened; the predicted base station cluster occupancy time matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line; based on the actual base station cluster occupancy time corresponding to the terminal where the user in the user set to be screened and the corresponding predicted base station cluster occupancy time, the users in the user set to be screened are screened to obtain a screened user set. In this way, the users in the user set to be screened are screened by using the actual base station cluster occupancy time corresponding to the terminal where the user is located and the predicted base station cluster occupancy time corresponding to the terminal where the user is located, which matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line, and the target high-speed rail users whose corresponding communication data on the target high-speed rail line can be used for high-speed rail simulation road testing can be accurately screened out, so as to effectively improve the accuracy of the simulation road test data of the target high-speed rail line through the communication data corresponding to the terminal where the target high-speed rail user is located.

[0094] In one embodiment, the actual occupation time of the base station cluster includes the actual occupation period of the base station cluster; the predicted occupation time of the base station cluster includes the predicted occupation period of the base station cluster; according to the actual occupation time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located, and the corresponding predicted occupation time of the base station cluster, the users in the user set to be screened are screened to obtain a screened user set, including: determining the total actual occupation time of the base station cluster corresponding to the terminal where the user in the user set to be screened according to the actual occupation period of the base station cluster; determining the base station cluster occupation overlap period corresponding to the terminal where the user in the user set to be screened according to the actual occupation period of the base station cluster and the corresponding predicted occupation period of the base station cluster; determining the total base station cluster occupation overlap time corresponding to the terminal where the user in the user set to be screened according to the base station cluster occupation overlap time; screening out users whose target duration ratio corresponding to their terminals is greater than a preset ratio threshold in the user set to be screened, to obtain a screened user set; the target duration ratio is the ratio between the total base station cluster occupation overlap time and the corresponding total base station cluster actual occupation time.

[0095] In which, in the process of determining the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located based on the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located, for any base station cluster, the base station cluster data of any base station cluster includes the access base station cluster time point corresponding to when the terminal where the user in the user set to be screened accesses the any base station cluster. The terminal can use the access base station cluster time point corresponding to when the terminal where the user is located initially accesses the any base station cluster as the first access base station cluster time point corresponding to the terminal where the user is located accessing the any base station cluster.

[0096] The terminal can determine, among the access base station cluster time points corresponding to when the terminal where the user in the user set to be screened is located accesses any base station cluster, an access base station cluster time point whose time interval with the first access base station cluster time point meets the preset time interval condition, as the target access base station cluster time point. The target access base station cluster time point is the candidate access base station cluster time point with the largest time interval with the first access base station cluster time point. The candidate access base station cluster time point is the access base station cluster time point corresponding to when the terminal where the user in the user set to be screened is located accesses any base station cluster, and the time interval with the first access base station cluster time point is less than the predicted occupancy time. The predicted occupancy time is determined based on the preset distance corresponding to the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line.

[0097] In this way, for any base station cluster, the terminal can determine the actual occupancy period of the base station cluster corresponding to the user's terminal in any base station cluster based on the time point at which the user's terminal accessed the target base station cluster and the time point at which the user's terminal first accessed the base station cluster. Thus, the terminal can determine the actual occupancy period of the base station cluster corresponding to the user's terminal in each base station cluster on the target high-speed rail line.

[0098] In actual applications, when the terminal determines the target access base station cluster time point whose time interval with the first access base station cluster time point meets the preset time interval condition, the terminal can use a sliding window-based deduplication method to find the first access base station cluster time point corresponding to the initial access to the any base station cluster by the user terminal under any base station cluster in chronological order, and then traverse the subsequent access base station cluster time points corresponding to the user terminal under any base station cluster in sequence. As long as Tg>Ta is satisfied, the access base station cluster time point is included in the window, and the window is then slid back in the same way. Wherein, Tg refers to the time interval between the access base station cluster time point currently to be determined and the current initial access base station cluster time point (i.e., the first access base station cluster time point) of the sliding window. Wherein, Ta is the predicted occupancy time, which is used to characterize the most likely occupancy time of a high-speed rail user on the target high-speed rail line when accessing the base station cluster. Wherein, Ta = (v + 2*r*a) / 2*r. Wherein, v is the high-speed rail speed, usually 250 km / h, a is a constant, and r is the preset radius corresponding to the target area range.

[0099] To facilitate understanding by those skilled in the art, when any base station cluster is a B1 base station cluster and the terminal identifier corresponding to the user's terminal is u1, Figure 4 A schematic diagram of the initial window corresponding to the user terminal u1 is provided. Figure 4As shown, the base station cluster access data of the B1 base station cluster includes the access base station cluster time points corresponding to the user's terminal u1 when accessing the B1 base station cluster in the set of users to be screened, which are respectively represented by (u1, t1), (u1, t2), (u1, t3), and (u1, t4) in chronological order. Among them, t1 is the first access base station cluster time point corresponding to the user's terminal u1 when initially accessing the B1 base station cluster. According to (u1, t1), the initial window corresponding to the user terminal u1 under the B1 base station cluster can be determined.

[0100] Then, the terminal can determine whether the t2 access base station cluster time point belongs to this initial window. When (t2 - t1) < Ta, the time (u1, t2) will be included in the initial window to obtain the first slid window. For the convenience of those skilled in the art to understand, Figure 5 A schematic diagram of the first slid window is provided.

[0101] Then, the terminal can determine whether the t3 access base station cluster time point belongs to this first slid window. When (t3 - t1) < Ta, the time (u1, t3) will be included in the first slid window to obtain the second slid window. For the convenience of those skilled in the art to understand, Figure 6 A schematic diagram of the second slid window is provided.

[0102] Then, the terminal can determine whether the t4 access base station cluster time point belongs to this second slid window. When (t4 - t1) > Ta, the second slid window is closed and will not be extended. At this time, it enters the duplicate removal stage within the window. The head and tail two access base station cluster time points in the second slid window are retained, and other access base station cluster time points are removed to obtain the de-duplicated second slid window, which is used as the target time window corresponding to the user's terminal u1 under the B1 base station cluster to determine the actual occupied period of the base station cluster corresponding to the user's terminal u1 under the B1 base station cluster. For the convenience of those skilled in the art to understand, Figure 7 A schematic diagram of the target time window is provided.

[0103] Among them, under the B1 base station cluster, the storage format of the actual occupied period of the base station cluster corresponding to the user's terminal u1 is {B1: (u1, t1, t3)}. Among them, the t2 access base station cluster time point is de-duplicated. Among them, when the t3 access base station cluster time point is the first access base station cluster time point corresponding to the user's terminal u1 under the B1 base station cluster as t1, the t3 access base station cluster time point is the target access base station cluster time point corresponding to the user's terminal u1.

[0104] Continuing with the previous example, assume that the corresponding base station cluster access time points when user terminal u1 accesses base station cluster B1 also include t5, t6, and t7, represented in chronological order by (u1, t5), (u1, t6), and (u1, t7), respectively. Since base station cluster access time point t4 is not included in the second sliding window, the terminal can use base station cluster access time point t4 as the first base station cluster access time point corresponding to the new initial window. Repeat the above steps to obtain another target time window corresponding to user terminal u1 in base station cluster B1, thereby determining the actual occupied time period of another base station cluster corresponding to user terminal u1 in base station cluster B1. The format for storing the actual occupied time period of another base station cluster is {B1: (u1, t4, t7)}, where base station cluster access time points t5 and t6 are deduplicated. The base station cluster access time point t7 is used as the target base station cluster access time point corresponding to user terminal u1 in base station cluster B1 when the first base station cluster access time point corresponding to user terminal u1 is t4.

[0105] To facilitate understanding by those skilled in the art, Figure 8 A schematic diagram of two target time windows (including a first target time window and a second target time window) corresponding to a user terminal u1 in a B1 base station cluster is provided.

[0106] Similarly, the terminal can obtain the actual occupied time period of the base station cluster corresponding to the terminals of other users under the B1 base station cluster (such as user terminal u2, user terminal u3, and other terminals of users who have visited the B1 base station cluster), as well as the actual occupied time period of the base station cluster corresponding to the terminal of the user under other base station clusters (such as B2 base station cluster, B3 base station cluster, and other base station clusters).

[0107] To facilitate understanding by those skilled in the art, Figure 9 A schematic diagram of target time windows corresponding to terminals of other users in a B1 base station cluster is provided.

[0108] In addition, in the process of determining the predicted occupancy time of the base station cluster corresponding to the terminal where the user in the set of users to be screened is located, due to the significant characteristic behavior of high-speed rail users, the terminal where the user in the set of users to be screened is located, the base station cluster is accessed in sequence according to the order of the base station clusters. For example, if the user terminal u1 has visited the B1, B2, and B3 base station clusters, the corresponding base station cluster access order is B1-B2-B3 or B3-B2-B1. If the user terminal ua has visited the B1, B3, and B5 base station clusters, the corresponding base station cluster access order is B1-B3-B5 or B5-B3-B1. For any user in the set of users to be screened, the terminal can obtain the actual occupancy time period of the base station cluster corresponding to the time when the user terminal accesses the first base station cluster, as the actual occupancy time period of the first base station cluster; wherein the first base station cluster is the first base station cluster visited by the user. After obtaining the high-speed rail travel time corresponding to the target high-speed rail line between base station clusters, the terminal can determine the predicted occupancy time period of the base station cluster corresponding to when the terminal of any user visits the base station cluster based on the predicted occupancy time of the base station cluster, the actual occupancy time period of the first base station cluster corresponding to any user, and the high-speed rail travel time.

[0109] For example, if the base station cluster access order corresponding to the user terminal u1 is B1-B2-B3, the terminal can obtain the actual occupancy period of the first base station cluster corresponding to the user terminal u1 accessing the B1 base station cluster, represented by {B1: (u1, tb, te)}. Among them, tb is the earliest access time point of the base station cluster corresponding to the actual occupancy period of the first base station cluster, and te is the latest access time point of the base station cluster corresponding to the actual occupancy period of the first base station cluster. The terminal can perform time compensation for the latest access time point te corresponding to the user terminal u1. Among them, the formula of the time compensation value (represented by t) is as follows:

[0110]

[0111] In this way, the terminal can determine the predicted base station cluster occupancy period corresponding to when the terminal where any user is located accesses the next base station cluster according to the formulas (te+Tn-t) and (te+Tn+Ta+t), according to the base station cluster access order corresponding to the terminal where any user is located. Among them, Tn is the high-speed rail travel time before reaching the next base station cluster, that is, the high-speed rail travel time between two adjacent base station clusters in the corresponding base station cluster access order. For example, if the next base station cluster of the user terminal u1 is the B2 base station cluster, then Tn is the high-speed rail travel time between the B1 base station cluster and the B2 base station cluster, which can be represented by T12. In this way, through the above formula, the predicted base station cluster occupancy period corresponding to when the terminal where any user is located accesses the base station cluster can be determined, so that the predicted base station cluster occupancy period can be matched with the preset distance corresponding to the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line, and then the predicted base station cluster occupancy period can be made to conform to the high-speed rail travel characteristics.

[0112] To facilitate understanding by those skilled in the art, following the above example, Figure 10 A diagram is provided showing the predicted occupancy period and actual occupancy period of a user terminal u1 with a corresponding BS cluster access order of B1-B2-B3. The solid line represents the actual occupancy period, and the dotted line represents the predicted occupancy period.

[0113] In this way, when the terminal screens the users in the user set to be screened based on the actual occupation time of the base station cluster corresponding to the terminal where the user is located and the corresponding predicted occupation time of the base station cluster, and obtains the screened user set, the terminal can determine the total actual occupation time of the base station cluster corresponding to the terminal where the user is located based on the actual occupation time period of the base station cluster corresponding to the terminal where the user is located in the user set; at the same time, the terminal can determine the overlapping occupation time period of the base station cluster corresponding to the terminal where the user is located based on the actual occupation time period of the base station cluster corresponding to the terminal where the user is located and the corresponding predicted occupation time period of the base station cluster, so as to obtain the total overlapping occupation time of the base station cluster corresponding to the terminal where the user is located, and determine the target duration ratio corresponding to the terminal where the user is located based on the ratio between the total overlapping occupation time of the base station cluster corresponding to the terminal where the user is located and the actual total occupation time of the base station cluster, and determine the user whose target duration ratio corresponding to the terminal where the user is located is greater than the preset ratio threshold. The user is the user whose actual occupation time of the base station cluster corresponding to the terminal meets the preset first screening condition. Therefore, the terminal can screen out the users whose target duration ratio corresponding to the terminal is greater than the preset ratio threshold in the user set to obtain the screened user set.

[0114] The technical solution of this embodiment is that the actual occupation time of the base station cluster includes the actual occupation period of the base station cluster; the predicted occupation time of the base station cluster includes the predicted occupation period of the base station cluster; the total actual occupation time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located is determined according to the actual occupation period of the base station cluster; the base station cluster occupation overlap period corresponding to the terminal where the user in the user set to be screened is determined according to the actual occupation period of the base station cluster and the corresponding predicted occupation period of the base station cluster; the total overlapping occupation time of the base station cluster corresponding to the terminal where the user in the user set to be screened is determined according to the base station cluster occupation overlap period; users whose target duration ratio corresponding to their terminals is greater than a preset ratio threshold are screened out from the user set to be screened; the screened user set is obtained by screening out users whose target duration ratio corresponding to their terminals is greater than a preset ratio threshold; the target duration ratio is the ratio between the total overlapping occupation time of the base station cluster and the total actual occupation time of the corresponding base station cluster.

[0115] In this way, since the predicted occupancy period of the base station cluster corresponding to the user's terminal matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line, the predicted occupancy period of the base station cluster corresponding to the user's terminal when accessing the base station cluster conforms to the high-speed rail travel characteristics corresponding to the target high-speed rail line. Therefore, the total overlapping occupancy time of the base station cluster corresponding to the user's terminal is determined through the predicted occupancy period of the base station cluster corresponding to the user's terminal and the corresponding actual occupancy period of the base station cluster. According to the ratio between the total overlapping occupancy time of the base station cluster corresponding to the terminal and the actual total occupancy time of the corresponding base station cluster, the users in the user set to be screened are screened, and then users whose actual occupancy period of the base station cluster corresponding to their terminal is closer to the corresponding predicted occupancy period of the base station cluster can be screened out. As users whose actual occupancy period of the base station cluster corresponding to their terminal is more consistent with the high-speed rail travel characteristics, the filtered user set is formed. When the users in the filtered user set are identified as target high-speed rail users, the accuracy of identifying the target high-speed rail users can be improved, so that the simulated road test data of the target high-speed rail line obtained according to the communication data corresponding to the terminal of the target high-speed rail user is more accurate.

[0116] In one embodiment, the method also includes: determining a user set to be clustered in the user set to be screened; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be clustered is located does not meet the preset first screening condition; clustering the users in the user set to be clustered to obtain a clustered user set; the base station cluster access characteristics corresponding to the terminal where the user in the clustered user set is located meet the preset second screening condition; and identifying the users in the user set after clustering as target high-speed rail users.

[0117] In a specific implementation, after the terminal determines the filtered user set from the set of users to be filtered, it can also determine the set of users to be clustered from the set of users to be filtered. The actual occupancy time of the base station cluster corresponding to the terminals containing the users in the set of users to be clustered does not meet the preset first filtering condition. Then, due to the consistency of high-speed rail user characteristics, the terminal can cluster the users in the set of users to be clustered to obtain a clustered user set. The base station cluster access characteristics corresponding to the terminals containing the users in the clustered user set meet the preset second filtering condition. In this way, the terminal can identify the users in the clustered user set as target high-speed rail users.

[0118] The technical solution of this embodiment is as follows: in the set of users to be screened, a set of users to be clustered is determined; the actual occupancy time of the base station cluster corresponding to the terminal where the users in the set of users to be clustered is located does not meet a preset first screening condition; the users in the set of users to be clustered are clustered to obtain a clustered user set; the access characteristics of the base station cluster corresponding to the terminal where the users in the set of users to be clustered are located meet a preset second screening condition; and the users in the set of users to be clustered are identified as target high-speed rail users. In this way, due to the consistency of high-speed rail user characteristics, for the set of users to be clustered whose actual occupancy time of the base station cluster corresponding to the terminal does not meet the preset first filtering condition, the users in the set of users to be clustered can also be clustered according to the base station cluster access characteristics corresponding to the user's terminal, and users whose base station cluster access characteristics corresponding to the terminal meet the preset second filtering condition are screened out to obtain the clustered user set. Then, when the users in the clustered user set are identified as target high-speed rail users, users with similar base station cluster access characteristics corresponding to the terminal can be classified as target high-speed rail users, which effectively improves the identification accuracy of the target high-speed rail users and avoids missing the target high-speed rail users in the set of users to be clustered, and realizes the comprehensive identification of the target high-speed rail users among the users whose terminals have visited the base station cluster on the target high-speed rail line, so that the simulated road test data of the target high-speed rail line obtained according to the communication data corresponding to the terminal where the target high-speed rail user is located is more accurate.

[0119] In one embodiment, the base station cluster access data includes the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; clustering processing is performed on the users in the clustered user set to obtain a clustered user set, including: obtaining equal time intervals corresponding to multiple base station clusters; the equal time intervals are obtained by dividing the total time period occupied by the target base station cluster corresponding to the target high-speed rail line into equal values; the total time period occupied by the target base station cluster is determined based on the earliest access base station cluster time point and the latest access base station cluster time point among the access base station cluster time points; clustering processing is performed on the users in the clustered user set according to the equal time intervals where the access base station cluster time point corresponding to the terminal where the user in the user set to be clustered is located to obtain multiple user cluster clusters; a user cluster cluster in which the number of users is greater than a preset number threshold among the multiple user cluster clusters is used as a target user cluster cluster; and the users in the target user cluster cluster are added to the clustered user set.

[0120] In a specific implementation, when a terminal clusters users in a clustered user set to obtain a clustered user set, the terminal can first determine the earliest access time point and the latest access time point of the base station cluster among the access time points of the base station cluster corresponding to the terminals of the users in the to-be-screened user set, so as to obtain the total occupancy period of the target base station cluster corresponding to the target high-speed rail line. This allows the total occupancy period of the target base station cluster to be divided into equal parts to obtain equal time intervals corresponding to multiple base station clusters on the target high-speed rail line. The equal time interval is an equal time interval based on each base station cluster on the target high-speed rail line.

[0121] For the convenience of those skilled in the art, Figure 11 A schematic diagram of equivalent time intervals corresponding to multiple base station clusters on a target high-speed rail line is provided. Figure 11 As shown, multiple base station clusters (including base station cluster B1, base station cluster B2, and base station cluster B3) on the target high-speed rail line correspond to five equal-value time intervals. The number of equal-value time intervals is not specifically limited here.

[0122] In this way, the access time points of the base station clusters corresponding to the terminals of the users in the user set to be clustered will fall into each equivalent time interval in chronological order. For example, if the terminal u1 of the user set to be clustered accesses the base station clusters according to the corresponding B1-B2-B3 base station cluster access order, if the access time points corresponding to each base station cluster fall into equivalent time interval 1, equivalent time interval 3, and equivalent time interval 5, respectively, they are represented by (B1, 1), (B2, 3), and (B3, 5), respectively. Then, the data storage format for the equivalent time interval in which the access time points of the base station clusters corresponding to the user terminal u1 fall can be {u1:[1, 3, 5]}.

[0123] Then, the terminal may cluster the users in the user set to be clustered based on the equivalent time intervals of the access time points of the base station clusters corresponding to the terminals of the users in the user set to be clustered, thereby obtaining multiple user clusters. The terminal may determine that, among the multiple user clusters, the base station cluster access characteristics corresponding to the terminals of the users in the user cluster with a number of users greater than a preset number threshold satisfy a preset second screening condition, and select the user cluster with a number of users greater than the preset number threshold as the target user cluster, thereby adding the users in the target user cluster to the clustered user set.

[0124] The technical solution of this embodiment is to obtain equivalent time intervals corresponding to multiple base station clusters; the equivalent time intervals are obtained by dividing the total occupied time period of the target base station cluster corresponding to the target high-speed rail line into equal values; the total occupied time period of the target base station cluster is determined based on the earliest access base station cluster time point and the latest access base station cluster time point among the access base station cluster time points; cluster the users in the user set to be clustered according to the equivalent time intervals where the access base station cluster time points corresponding to the terminals where the users are located are located, and obtain multiple user clustering clusters; a user clustering cluster in which the number of users is greater than a preset number threshold among the multiple user clustering clusters is used as a target user clustering cluster; and the users in the target user clustering cluster are added to the clustered user set.

[0125] In this way, clustering is performed on the users in the user set to be clustered according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the users in the user set to be clustered are located. This can accurately realize clustering of the users in the user set to be clustered according to the base station cluster access characteristics corresponding to the terminal where they are located, so that users with highly consistent base station cluster access characteristics corresponding to their terminals are added to the same user clustering cluster to obtain multiple user clustering clusters. Since high-speed rail users have feature consistency, the more users in the user clustering cluster, that is, the more users whose terminals in the user set to be clustered meet the base station cluster access characteristics corresponding to the user clustering cluster, the more it meets the condition that high-speed rail users have feature consistency. Therefore, users in the user clustering cluster whose number of users in the user clustering cluster is greater than the preset number threshold are added to the clustered user set, so that the users in the clustered user set are identified as target high-speed rail users. This can improve the identification accuracy of the target high-speed rail users, so that the simulated road test data of the target high-speed rail line obtained according to the communication data corresponding to the terminal where the target high-speed rail user is located is more accurate.

[0126] In one embodiment, clustering is performed on users in the user set to be clustered according to the equivalent time intervals of the access base station cluster time points corresponding to the terminals of the users in the user set to be clustered, to obtain multiple user clustering clusters, including: determining the access feature similarity between the terminals of the users in the user set to be clustered according to the equivalent time intervals of the access base station cluster time points corresponding to the terminals of the users in the user set to be clustered; in the user set to be clustered, adding users whose access feature similarity corresponding to their terminals is greater than a preset similarity threshold to the same user clustering cluster.

[0127] The access feature similarity can be represented by cosine similarity.

[0128] In a specific implementation, the terminal clusters the users in the user set to be clustered according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the users in the user set to be clustered are located, and in the process of obtaining multiple user clusters, the terminal can determine the cosine similarity between the terminals where the users in the user set to be clustered are located according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the users in the user set to be clustered are located, as the access feature similarity between the terminals where the users in the user set to be clustered are located.

[0129] For example, continuing with the previous example, if the user terminal u2 is included in the user set to be clustered, when accessing the base station cluster according to the corresponding B1-B2-B3 base station cluster access order, if the access base station cluster time points corresponding to each base station cluster fall in the equivalent time interval 1, the equivalent time interval 3 and the equivalent time interval 4 respectively, then the equivalent time interval of the access base station cluster time point corresponding to the user terminal u2 is {u2:[1, 3, 4]}; and the equivalent time interval of the access base station cluster time point corresponding to the user terminal u1 is {u1:[1, 3, 5]}.

[0130] Then the access feature similarity between the user terminal u2 and the user terminal u1 is expressed as cosine similarity (ω 12 ) represents, then

[0131]

[0132] Among them, b1, b2, b3, b4, and b5 correspond to the equivalent time interval 1, the equivalent time interval 2, the equivalent time interval 3, the equivalent time interval 4, and the equivalent time interval 5, respectively.

[0133] In this way, the terminal can add users whose access feature similarity corresponding to their terminals is greater than a preset similarity threshold to the same user cluster in the set of users to be clustered. Specifically, the terminal can select any two users in the set of users to be clustered, determine the access feature similarity between the terminals of the arbitrary two users, and if the access feature similarity is greater than the preset similarity threshold, add the arbitrary two users to the same user cluster. Then, using any one of the arbitrary two users as the target matching user, calculate the access feature similarity between the terminals of other users in the set of users to be clustered, excluding the arbitrary two users, and the terminal of the target matching user, so as to improve the computational efficiency when determining the access feature similarity between the terminals of users in the set of users to be clustered.

[0134] The technical solution of this embodiment determines the access feature similarity between the terminals of the users in the user set to be clustered based on the equivalent time intervals of the time points at which the terminals of the users in the user set to be clustered access the base station cluster; in the user set to be clustered, users whose terminals have access feature similarity greater than a preset similarity threshold are added to the same user cluster. In this way, the access feature similarity between the terminals of the users in the user set to be clustered can be accurately determined based on the equivalent time intervals of the time points at which the terminals of the users in the user set to be clustered access the base station cluster, thereby accurately classifying users with high access feature similarity into one category based on the access feature similarity, and further accurately clustering the user set to be clustered into multiple user clusters.

[0135] In another embodiment, Figure 12 As shown, a user identification method is provided, which is described by taking the method applied to a terminal as an example, and includes the following steps:

[0136] Step S1210: Determine multiple base station clusters on the target high-speed rail line.

[0137] Step S1220 : Acquire base station cluster access data of each base station cluster, and determine a user set to be screened based on the base station cluster access data.

[0138] Step S1230 , determining the actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the time point at which the terminal where the user in the to-be-screened user set is located accesses the base station cluster.

[0139] Step S1240 : determining the predicted occupancy time of the base station cluster corresponding to the terminal where the user in the set of users to be screened is located.

[0140] Step S1250 , filtering the users in the user set to be filtered according to the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be filtered is located and the corresponding predicted occupancy time of the base station cluster to obtain a filtered user set.

[0141] Step S1260: Determine a user set to be clustered from the user set to be screened.

[0142] Step S1270 , clustering the users in the to-be-clustered user set to obtain a clustered user set.

[0143] Step S1280: Identify the users in the filtered user set and the users in the clustered user set as target high-speed rail users associated with the target high-speed rail line.

[0144] Thus, this user identification method improves the accuracy of identifying target high-speed rail users compared to traditional methods for identifying high-speed rail users through simulated drive testing, thereby increasing the accuracy of simulated drive test data for the target high-speed rail line and further enhancing the effectiveness of simulated drive testing for the target high-speed rail line. Furthermore, due to its computational simplicity, this user identification method can complete a complete search of users who have visited multiple base station clusters on the target high-speed rail line in a very short time. Due to its high efficiency in identifying target high-speed rail users, this method can be used for real-time target high-speed rail user identification on the target high-speed rail line.

[0145] It should be noted that the specific definition of the above steps can refer to the specific definition of a user identification method above.

[0146] To facilitate understanding by those skilled in the art, Figure 13 A flowchart of another user identification method is provided. When a user's terminal has at least two corresponding base station clusters with actual occupancy periods, the terminal can use a preset first screening condition and a preset second screening condition based on the actual occupancy period of each base station cluster to determine whether the user is a target high-speed rail user. It should be noted that the specific definitions of the steps in the above method can be found in the specific definitions of a user identification method above and are not further elaborated here.

[0147] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0148] Based on the same inventive concept, embodiments of the present application also provide a user identification device for implementing one of the user identification methods mentioned above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more user identification device embodiments provided below can be found in the above-mentioned limitations of a user identification method and will not be repeated here.

[0149] In one embodiment, Figure 14As shown, a user identification device is provided, including: a base station cluster determination module 1410, a user determination module 1420, a screening module 1430 and an identification module 1440, wherein:

[0150] The base station cluster determination module 1410 is used to determine multiple base station clusters on the target high-speed rail line; the base station cluster is a collection of base station cells within the target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius.

[0151] The user determination module 1420 is configured to obtain base station cluster access data of each base station cluster and determine a user set to be screened based on the base station cluster access data; the terminal where the user in the user set to be screened is located has visited at least two base station clusters.

[0152] The screening module 1430 is configured to screen the user set to be screened to obtain a screened user set; the actual occupancy time of the base station cluster corresponding to the terminal where the user is located in the screened user set meets a preset first screening condition.

[0153] The identification module 1440 is used to identify the users in the filtered user set as target high-speed rail users associated with the target high-speed rail line; the communication data corresponding to the terminal where the target high-speed rail user is located is used to perform simulated road testing on the target high-speed rail line.

[0154] In one embodiment, the base station cluster access data includes the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; the screening module 1430 is specifically used to determine the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located based on the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; determine the predicted occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened; the predicted occupancy time of the base station cluster matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line; according to the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened, and the corresponding predicted occupancy time of the base station cluster, the users in the user set to be screened are screened to obtain the screened user set.

[0155] In one embodiment, the actual occupancy time of the base station cluster includes the actual occupancy period of the base station cluster; the predicted occupancy time of the base station cluster includes the predicted occupancy period of the base station cluster; the screening module 1430 is specifically used to determine the total actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located based on the actual occupancy period of the base station cluster; determine the base station cluster occupancy overlap period corresponding to the terminal where the user in the user set to be screened is located based on the actual occupancy period of the base station cluster and the corresponding predicted occupancy period of the base station cluster; determine the total overlapping occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located based on the base station cluster occupancy overlap period; screen out users in the user set to be screened whose corresponding target time ratio of their terminals is greater than a preset ratio threshold to obtain the screened user set; the target time ratio is the ratio between the total overlapping occupancy time of the base station cluster and the total actual occupancy time of the corresponding base station cluster.

[0156] In one embodiment, the device also includes: a clustering module, used to determine the user set to be clustered in the user set to be screened; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be clustered is located does not meet the preset first screening condition; clustering processing is performed on the users in the user set to be clustered to obtain a clustered user set; the base station cluster access characteristics corresponding to the terminal where the user in the clustered user set are located meet the preset second screening condition; and the users in the clustered user set are identified as the target high-speed rail users.

[0157] In one embodiment, the base station cluster access data includes the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; the clustering module is specifically used to obtain the equivalent time intervals corresponding to the multiple base station clusters; the equivalent time intervals are obtained by dividing the total occupancy period of the target base station cluster corresponding to the target high-speed rail line into equal values; the total occupancy period of the target base station cluster is determined based on the earliest access base station cluster time point and the latest access base station cluster time point among the access base station cluster time points; according to the equivalent time intervals where the access base station cluster time point corresponding to the terminal where the user in the user set to be clustered is located, the users in the user set to be clustered are clustered to obtain multiple user cluster clusters; the user cluster cluster in which the number of users is greater than a preset number threshold among the multiple user cluster clusters is used as the target user cluster cluster; the users in the target user cluster cluster are added to the clustered user set.

[0158] In one embodiment, the clustering module is specifically used to determine the access feature similarity between the terminals of users in the user set to be clustered based on the equivalent time interval of the access base station cluster time point corresponding to the terminals of the users in the user set to be clustered; in the user set to be clustered, users whose access feature similarity corresponding to their terminals is greater than a preset similarity threshold are added to the same user clustering cluster.

[0159] Each module in the above-mentioned user identification device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0160] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 15 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store base station cluster access data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a user identification method is implemented.

[0161] Those skilled in the art will understand that Figure 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0162] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0163] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0164] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0165] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0166] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0167] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0168] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A user identification method, characterized in that: The method comprises: Determine multiple base station clusters on a target high-speed rail line; the base station cluster is a collection of base station cells within a target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius; Obtaining base station cluster access data for each of the base station clusters, and determining a user set to be screened based on the base station cluster access data; wherein a terminal of a user in the user set to be screened has accessed at least two of the base station clusters; the base station cluster access data includes a mapping relationship between a terminal identifier and a base station cluster identifier, and the mapping relationship between the terminal identifier and the base station cluster identifier is determined by deduplicating the users; Filtering the user set to be filtered to obtain a filtered user set; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the filtered user set is located meets a preset first filtering condition; Identifying users in the filtered user set as target high-speed rail users associated with the target high-speed rail line; using communication data corresponding to a terminal where the target high-speed rail user is located to perform a simulated road test on the target high-speed rail line; The base station cluster access data includes a time point at which a terminal of a user in the to-be-screened user set accesses the base station cluster; and filtering the to-be-screened user set to obtain a screened user set includes: Determine the actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the access time point of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located; Determine a predicted occupancy time of a base station cluster corresponding to a terminal where a user in the set of users to be screened is located; the predicted occupancy time of the base station cluster matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line; The users in the user set to be screened are screened according to the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located and the corresponding predicted occupancy time of the base station cluster to obtain the screened user set.

2. The method according to claim 1, characterized in that The determining, based on the access time point of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located, the actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located includes: For any base station cluster, the base station cluster access data of the any base station cluster includes a base station cluster access time point corresponding to when a terminal of a user in the to-be-screened user set accesses the any base station cluster, and the base station cluster access time point corresponding to when the terminal of the user initially accesses the any base station cluster is used as the first base station cluster access time point corresponding to the user terminal accessing the any base station cluster; Among the access base station cluster time points corresponding to when the terminal of a user in the user set to be screened accesses any base station cluster, determine the access base station cluster time point whose time interval with the first access base station cluster time point meets the preset time interval condition, and use it as the target access base station cluster time point; the target access base station cluster time point is the candidate access base station cluster time point with the largest time interval between it and the first access base station cluster time point; the candidate access base station cluster time point is the access base station cluster time point corresponding to when the terminal of a user in the user set to be screened accesses any base station cluster, and the time interval between it and the first access base station cluster time point is less than the predicted occupancy time; the predicted occupancy time is determined based on the preset distance corresponding to the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line.

3. The method according to claim 2, characterized in that The actual occupied time of the base station cluster includes the actual occupied time period of the base station cluster; the predicted occupied time of the base station cluster includes the predicted occupied time period of the base station cluster; The filtering of users in the to-be-filtered user set according to the actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-filtered user set is located and the corresponding predicted occupancy time of the base station cluster to obtain the filtered user set includes: Determine the total actual occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the actual occupancy time period of the base station cluster; Determine, based on the actual occupied period of the base station cluster and the corresponding predicted occupied period of the base station cluster, the overlapping occupied period of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located; Determine the total overlapping occupancy time of the base station cluster corresponding to the terminal where the user in the to-be-screened user set is located according to the overlapping occupancy time period of the base station cluster; From the set of users to be screened, users whose terminal's corresponding target duration ratio is greater than a preset ratio threshold are screened out to obtain the screened user set; the target duration ratio is the ratio between the total overlapping duration of the base station cluster occupancy and the total actual occupancy duration of the corresponding base station cluster.

4. The method according to claim 1, wherein The method further comprises: Determining a user set to be clustered from the set of users to be screened; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the set of users to be clustered is not satisfied with the preset first screening condition; Performing clustering processing on the users in the to-be-clustered user set to obtain a clustered user set; base station cluster access characteristics corresponding to terminals where users in the clustered user set are located meet a preset second screening condition; Identify the users in the clustered user set as the target high-speed rail users.

5. The method according to claim 4, characterized in that The base station cluster access data includes the access time point of the base station cluster corresponding to the terminal where the user in the user set to be screened is located; the clustering process of the users in the user set to be clustered to obtain the clustered user set includes: Obtaining equivalent time intervals corresponding to the multiple base station clusters; the equivalent time intervals are obtained by dividing the total occupied time period of the target base station cluster corresponding to the target high-speed rail line into equal values; the total occupied time period of the target base station cluster is determined based on the earliest access base station cluster time point and the latest access base station cluster time point among the access base station cluster time points; Clustering the users in the user set to be clustered according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the users in the user set to be clustered are located, to obtain multiple user clusters; Taking a user cluster with a user number greater than a preset number threshold among the multiple user clusters as a target user cluster; Add users in the target user cluster to the clustered user set.

6. The method according to claim 5, characterized in that The clustering of users in the user set to be clustered is performed according to the equivalent time interval of the access base station cluster time point corresponding to the terminal where the users in the user set to be clustered are located, to obtain multiple user clusters, including: Determining access feature similarities between terminals where users in the set of users to be clustered are located, based on equivalent time intervals where time points at which terminals where users in the set of users to be clustered access the base station cluster belong; In the set of users to be clustered, users whose terminals have access feature similarities greater than a preset similarity threshold are added to the same user cluster.

7. A user identification device, characterized in that: The device comprises: A base station cluster determination module is configured to determine multiple base station clusters on a target high-speed rail line; the base station cluster is a collection of base station cells within a target area; the target area is an area with at least one sampling point on the target high-speed rail line as the center and a preset distance as the radius; A user determination module is configured to obtain base station cluster access data of each of the base station clusters and determine a user set to be screened based on the base station cluster access data; a terminal containing a user in the user set to be screened has accessed at least two of the base station clusters; the base station cluster access data includes a mapping relationship between a terminal identifier and a base station cluster identifier, and the mapping relationship between the terminal identifier and the base station cluster identifier is determined by deduplicating the users; A screening module is configured to screen the user set to be screened to obtain a screened user set; the actual occupancy time of the base station cluster corresponding to the terminal where the user in the screened user set is located satisfies a preset first screening condition; An identification module is configured to identify users in the filtered user set as target high-speed rail users associated with the target high-speed rail line; and communication data corresponding to the terminal where the target high-speed rail user is located is used to perform a simulated road test on the target high-speed rail line; The base station cluster access data includes the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; the screening module is also used to determine the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened is located based on the access base station cluster time point corresponding to the terminal where the user in the user set to be screened is located; determine the predicted occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened; the predicted occupancy time of the base station cluster matches the target area range and the high-speed rail travel speed corresponding to the target high-speed rail line; according to the actual occupancy time of the base station cluster corresponding to the terminal where the user in the user set to be screened, and the corresponding predicted occupancy time of the base station cluster, the users in the user set to be screened are screened to obtain the screened user set.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • User identification method, device and equipment and computer readable storage medium

    CN111314947A