A high-iron user identification method and device, and a storage medium
By identifying the main control cell for high-speed rail and building a model using the LightGBM algorithm, high-speed rail users are accurately identified, solving the problem of inaccurate identification in existing technologies, optimizing network quality, and improving user experience.
Patent Information
- Application Number
- CN202310739330.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-06-20
AI Technical Summary
Existing methods for identifying high-speed rail users are inefficient and inaccurate, resulting in insufficient network quality optimization and impacting user experience.
By identifying the main control cell for high-speed rail, a first model is constructed using the LightGBM algorithm. The confidence level is calculated based on the fingerprint features of the main control cell for high-speed rail, and serving cells with high confidence levels are selected. An initial model is then built for training to identify high-speed rail users.
It enables accurate identification of high-speed rail users, discovers potential coverage quality issues, optimizes network quality, improves user experience, and creates a higher-quality high-speed rail network.
Smart Images

Figure CN116567672B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication, and in particular to a high-speed rail user identification method and device and storage medium. BACKGROUND
[0002] In recent years, high-speed rail has gradually become the first choice for people to travel long distances and commute. As the number of people taking high-speed rail increases, the demand for online services on high-speed rail by high-speed rail users is also increasing. Against this background, it is particularly important for operators to establish high-quality networks to meet the online service needs of high-speed rail users on high-speed rail. Therefore, it is necessary to accurately identify high-speed rail users, and then optimize high-speed rail network quality and improve user perception. The current method for identifying high-speed rail users still cannot efficiently and accurately identify high-speed rail users. SUMMARY
[0003] The present application provides a high-speed rail user identification method, device and storage medium, which can identify high-speed rail users.
[0004] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0005] In a first aspect, the present application provides a high-speed rail user identification method, which comprises: determining a high-speed rail master cell; wherein the high-speed rail master cell is used to provide services for a high-speed rail user group; determining a first model according to the fingerprint characteristics of the high-speed rail master cell; wherein the first model is used to determine the confidence of the high-speed rail master cell; determining a first target service cell according to the confidence of the high-speed rail master cell; the first target service cell is a service cell in the high-speed rail master cell whose confidence is greater than a pre-set confidence threshold; determining a second model according to the first target service cell; wherein the second model is used to identify high-speed rail users; and identifying high-speed rail users in the to-be-identified users according to the second model.
[0006] In a possible implementation, the determination of the high-speed rail master cell specifically comprises: obtaining a second target service cell; wherein the second target service cell is a service cell whose switching times in a pre-set time period on a high-speed rail line are less than a pre-set number of times; determining a high-speed rail user group from a user group corresponding to the second target service cell according to a long-interval speed algorithm; and determining the high-speed rail master cell according to a first pre-set feature condition and the high-speed rail user group.
[0007] In a possible implementation, the first pre-set feature condition comprises one or more of the following: the moving route has periodicity, the distance from the high-speed rail line is less than a first pre-set distance threshold, and the ECI of the corresponding service cell in the pre-set time period is the same.
[0008] In a possible implementation, the acquiring the second target serving cell specifically includes: determining a third target serving cell; wherein the third target serving cell has a shortest distance to the high-speed railway line less than a second preset distance threshold; and acquiring S1-MME interface data of users in the third target serving cell.
[0009] The S1-MME interface data contains an ECI of the third target serving cell; the S1-MME interface data is time-sequenced; and the second target serving cell is acquired from the third target serving cell according to the sequenced S1-MME interface data.
[0010] In a possible implementation, the first model and the second model are constructed according to a LightGBM algorithm. In a possible implementation, the second model is determined according to the first target serving cell, specifically including: constructing an initial model according to the LightGBM algorithm; determining whether an identification accuracy of the initial model is greater than or equal to an accuracy threshold according to a K-fold cross-validation algorithm; and in a case where the identification accuracy of the initial model is greater than or equal to the accuracy threshold, determining the initial model as the second model.
[0011] In a second aspect, the present application provides a high-speed rail user identification device, which includes: a processing unit; the processing unit is configured to determine a high-speed rail master cell; wherein the high-speed rail master cell is configured to provide services for a high-speed rail user group; the processing unit is further configured to determine a first model according to a fingerprint feature of the high-speed rail master cell; wherein the first model is configured to determine a confidence degree of the high-speed rail master cell; the processing unit is further configured to determine a first target serving cell according to the confidence degree of each cell; the first target serving cell is a serving cell in the high-speed rail master cell with a confidence degree greater than a preset confidence degree threshold; the processing unit is further configured to determine a second model according to the first target serving cell; wherein the second model is configured to identify high-speed rail users; and the processing unit is further configured to identify high-speed rail users in a user to be identified according to the second model.
[0012] In a possible implementation, the device further includes: an acquisition unit; the acquisition unit is configured to acquire a second target serving cell; wherein the second target serving cell is a serving cell with a number of handovers in a preset time period less than a preset number on a high-speed railway line; the processing unit is further configured to determine a high-speed rail user group from a user group corresponding to the second target serving cell according to a long-interval speed algorithm; and the processing unit is further configured to determine a high-speed rail master cell according to a first preset feature condition and the high-speed rail user group.
[0013] In a possible implementation, the first preset feature condition includes one or more of the following: a mobile route has periodicity, a distance to the high-speed railway line is less than a first preset distance threshold, and an ECI of a corresponding serving cell in a preset time period is the same.
[0014] In a possible implementation, the acquiring the second target serving cell specifically includes: the processing unit is further configured to determine a third target serving cell; the third target serving cell has a shortest distance to the high-speed railway line less than a second preset distance threshold; the acquiring unit is further configured to acquire S1-MME interface data of the user in the third target serving cell; the S1-MME interface data contains an ECI of the third target serving cell; the processing unit is further configured to perform time sorting on the S1-MME interface data; and the processing unit is further configured to acquire the second serving cell from the third target serving cell according to the sorted S1-MME interface data.
[0015] In a possible implementation, the processing unit is further configured to construct the first model and the second model according to a LightGBM algorithm. In a possible implementation, the processing unit is further configured to determine the second model according to the first target serving cell, specifically including: the processing unit is further configured to construct an initial model according to the LightGBM algorithm; the processing unit is further configured to determine whether an identification accuracy of the initial model is greater than or equal to an accuracy threshold according to a K-fold cross-validation algorithm; and the processing unit is further configured to determine the initial model as the second model in a case where the identification accuracy of the initial model is greater than or equal to the accuracy threshold.
[0016] In a third aspect, the present application provides a high-speed rail user identification device, which includes a processor and a communication interface; the communication interface is coupled with the processor, and the processor is configured to run a computer program or instruction to implement the high-speed rail user identification method described in the first aspect and any possible implementation of the first aspect.
[0017] In a fourth aspect, the present application provides a computer readable storage medium, which stores an instruction, and when the instruction runs on a terminal, causes the terminal to perform the high-speed rail user identification method described in the first aspect and any possible implementation of the first aspect.
[0018] In a fifth aspect, the present application provides a computer program product containing an instruction, and when the computer program product runs on a high-speed rail user identification device, causes the high-speed rail user identification device to perform the high-speed rail user identification method described in the first aspect and any possible implementation of the first aspect.
[0019] In a sixth aspect, the present application provides a chip, which includes a processor and a communication interface; the communication interface is coupled with the processor, and the processor is configured to run a computer program or instruction to implement the high-speed rail user identification method described in the first aspect and any possible implementation of the first aspect.
[0020] Specifically, the chip provided in the embodiments of the present application further includes a memory configured to store the computer program or instruction.
[0021] Based on the above technical solution, the high-speed rail user identification method provided by the embodiment of the application first determines a high-speed rail master cell, then determines a first model according to the fingerprint feature of the high-speed rail master cell and the LightGBM algorithm, calculates the confidence of all high-speed rail master cells according to the first model, and then determines a service cell with a confidence greater than a preset confidence threshold as a first target service cell, then constructs an initial model according to the LightGBM algorithm and the first target service cell, and trains the initial model according to the features of the first target service cell, and when the identification accuracy of the initial model is greater than or equal to an accuracy threshold, the initial model is determined as a second model, so as to identify the high-speed rail users in the to-be-identified users according to the second model. Therefore, the application can accurately identify high-speed rail users, thereby discovering potential coverage quality problems, optimizing network quality, improving user perception, and creating a better network. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 A service cell neighbor relation schematic diagram provided by the embodiment of the application;
[0023] Figure 2 A high-speed rail base station and high-speed rail distance relation schematic diagram provided by the embodiment of the application;
[0024] Figure 3 A histogram algorithm forming a histogram flowchart provided by the embodiment of the application;
[0025] Figure 4 A histogram difference optimization flowchart provided by the embodiment of the application;
[0026] Figure 5 A comparison schematic diagram of different growth modes provided by the embodiment of the application;
[0027] Figure 6 A model construction flowchart according to the LightGBM algorithm provided by the embodiment of the application;
[0028] Figure 7 An architecture schematic diagram of a high-speed rail identification system provided by the embodiment of the application;
[0029] Figure 8 A hardware structure schematic diagram of a high-speed rail user identification device provided by the embodiment of the application;
[0030] Figure 9 A hardware structure schematic diagram of another high-speed rail user identification device provided by the embodiment of the application;
[0031] Figure 10 A flowchart of a high-speed rail user identification method provided by the embodiment of the application;
[0032] Figure 11 A schematic diagram of the confidence of part of the high-speed rail master cell provided by the embodiment of the present application;
[0033] Figure 12 A schematic diagram of the prediction accuracy of the initial model according to K-fold cross-validation provided by the embodiment of the present application;
[0034] Figure 13 A schematic diagram of the driving trajectory of a certain identified high-speed rail user on a certain high-speed rail line provided by the embodiment of the present application;
[0035] Figure 14 A schematic diagram of the driving trajectory and MR data comparison of a certain identified high-speed rail user on a certain high-speed rail line provided by the embodiment of the present application;
[0036] Figure 15 A flowchart of another high-speed rail user identification method provided by the embodiment of the present application;
[0037] Figure 16 A schematic diagram of a certain high-speed rail line service cell provided by the embodiment of the present application;
[0038] Figure 17 A schematic diagram of a certain identified high-speed rail user according to a long interval algorithm provided by the embodiment of the present application;
[0039] Figure 18 A flowchart of another high-speed rail user identification method provided by the embodiment of the present application;
[0040] Figure 19 A schematic diagram of a certain high-speed rail line service cell distribution in A area provided by the embodiment of the present application;
[0041] Figure 20 A schematic diagram of S1-MME interface data of a certain user in a third target service cell provided by the embodiment of the present application;
[0042] Figure 21 A schematic diagram of S1-MME interface data of a certain user in a third target service cell provided by the embodiment of the present application;
[0043] Figure 22 A schematic diagram of a certain user switching between service cells provided by the embodiment of the present application;
[0044] Figure 23 A schematic diagram of a high-speed rail user identification device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0045] The high-speed rail user identification method and device provided by the embodiments of the present application and the storage medium are described in detail below with reference to the drawings.
[0046] The term "and / or" in this document merely describes an association relationship of associated objects, and indicates that three relationships can exist, for example, A and / or B can represent three cases of existence of A alone, existence of A and B simultaneously, and existence of B alone.
[0047] The terms "first" and "second" and the like in the description of the present application and the drawings are used to distinguish different objects or different treatments of the same object, and are not used to describe a specific order of the objects.
[0048] In addition, the terms "include" and "have" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include other steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0049] It should be noted that in the embodiments of the present application, the words "exemplary" or "for example" are used to represent an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplary" or "for example" are intended to present the relevant concept in a specific manner.
[0050] In order to facilitate understanding of the technical solutions of the present application, the technical terms related to the present application are introduced as follows:
[0051] 1. High-speed rail signal coverage characteristics.
[0052] In a long term evolution (LTE) wireless network, the coverage of high-speed rail signals mainly has the following characteristics:
[0053] (1) Large wireless signal penetration loss.
[0054] Because the carriages of high-speed trains adopt a fully enclosed structure, the wireless signal penetration loss is large. Table 1 below shows the wireless signal penetration loss of various high-speed train models:
[0055] Table 1 Wireless signal penetration loss of various high-speed train models
[0056]
[0057] (2) Doppler shift.
[0058] Doppler shift refers to the signal wave changes with the transmitter and receiver relative position changes, the faster the relative movement, the greater the frequency shift. Frequency shift will cause sub-carrier interference, thereby affecting the cell switching and selection. Table 2 below shows the Doppler shift under different frequency bands:
[0059] Table 2 Doppler shift under different frequency bands
[0060]
[0061]
[0062] Reasonable selection of base station site can reduce the impact of Doppler shift, in addition to the adaptive frequency offset correction algorithm developed by each equipment manufacturer, can also be corrected for frequency shift, to improve the baseband performance demodulation.
[0063] (3) Frequent switching.
[0064] High-speed rail has the characteristics of high-speed movement, in the high-speed moving scene, high-speed rail through the cell time is very short, so the user terminal will be between the service cell switching. Service cell switching is too frequent, will make the drop rate increases, affect the user perception. For the problem of frequent switching of cells, the current solution is to merge multiple service cells, merging multiple adjacent service cells into a logical service cell. Thus in the high-speed moving scene, the user terminal will not switch within a logical cell, which will reduce the number of switching cells and improve the success rate of switching.
[0065] In the high-speed moving scene, the switching between the service cell value needs a larger overlap coverage range, so the spacing of the service cell needs to be re-planned.
[0066] 2, high-speed rail signal coverage strategy.
[0067] (1) Special network coverage.
[0068] High-speed rail coverage is divided into public network networking and special network networking. Public network networking uses existing or newly built sites and uses the same frequency as the surrounding base station, thereby covering high-speed rail users and surrounding users at the same time. Special network networking is composed of high-speed rail sites, only serving high-speed rail users, and high-speed rail surrounding users are still served by public networks. Among them, the special network frequency can be the same as the public network, or it can be different from the public network.
[0069] High-speed rail LTE networks provide signal coverage using a dedicated network. A chain-like design of adjacent cells is employed along the route, eliminating handoffs with the public network. This ensures excellent network continuity and improves communication quality for users traveling at high speeds. Due to frequency resource limitations, using a dedicated, off-frequency network solution—with the private network and public network each using 10 MHz—would halve the peak rate of LTE on high-speed rail, reducing spectrum utilization and impacting user experience. Therefore, a co-frequency dedicated network is preferred for LTE networks.
[0070] (2)Switching strategy.
[0071] When the high-speed rail LTE network adopts the same frequency private network, it is necessary to make good handover between the private network and the public network along the high-speed rail line. Figure 1 As shown in the figure, the private network cells along the high-speed rail line need to configure neighboring relationships with the two cells before and after them, but not with the surrounding public network cells.
[0072] (3) Station spacing setting.
[0073] For example, Figure 2 As shown in the figure, the inter-station spacing of high-speed rail base stations is related to factors such as the distance between the base station and the rails and the cell coverage radius. When the base station is close to the rails, the incident angle becomes smaller, and the wireless signal suffers greater loss when penetrating the vehicle body, and the corresponding inter-station spacing becomes larger. When the base station is far away from the rails, the wireless signal attenuation will be too large, and by the time it reaches the high-speed train, it will no longer be able to meet the terminal's call needs.
[0074] High-speed rail technology continues to evolve, and wireless signal coverage scenarios are constantly changing. Therefore, targeted solutions are needed to address wireless signal coverage issues and enhance the user experience in high-speed rail scenarios, thereby earning a positive reputation for operators. To accurately reflect the actual wireless signal coverage and service perception of high-speed rail users, operators use network management data, mapreduce (MR) data, and external data representation (XDR) data to analyze the location of high-speed rail users. Currently, algorithms for identifying high-speed rail users use three rule models: user signaling occurs on the high-speed rail dedicated network, user location trajectories match the high-speed rail route, and user movement speed exceeds a specific threshold.
[0075] The current rule-based algorithm for identifying high-speed rail users has the following problems: identifying high-speed rail users by matching multiple consecutive cells in the switching chain can easily lead to missed identification of high-speed rail users; calculating the user's movement speed by the line distance between cells that appear successively in the user signaling and the time difference between the line mappings has a large error, which can easily lead to errors in identifying high-speed rail users; when calculating the user's movement distance, the passenger's waiting time and the time spent in the station are not taken into account, resulting in the calculated movement speed being lower than the actual speed, causing missed identification of high-speed rail users.
[0076] Currently, the traditional high-speed rail recognition method is to use an XDR data-based recognition method, which is as follows:
[0077] 3. Introduction to the principle of the traditional algorithm.
[0078] (1) DPI principle and XDR data acquisition.
[0079] The deep packet inspection (DPI) system detects and analyzes the traffic and packet content of the network key interface, filters and controls the traffic according to the policy, realizes the collection of the signaling plane and the user plane, and thus filters and collects the information generated by the user's online behavior.
[0080] The DPI system is divided into three layers of architecture, namely the collection layer, the decoding layer and the reference layer. Among them, the collection layer and the decoding layer are responsible for data collection, traffic analysis and log synthesis, and are generally stored in the database of the decoding layer in the form of call detail record (CDR) and transaction detail record (TDR) records. The application layer mainly completes the calculation, arrangement, statistics, reasonable organization and storage of CDR and TDR record data, and performs presentation.
[0081] (2) XDR data mining.
[0082] After data cleaning, standardization and warehousing operations are performed on the collected XDR data, the following four types of data can be obtained: S1-mobility management entity (MME) interface data, hypertext transfer protocol (HTTP), video (video) and domain name system (DNS), among which S1-MME is the data of the user signaling plane, HTTP is the user webpage data, video is the user video watching data, and DNS is the user online path data. For high-speed rail user recognition, S1-MME data needs to be analyzed and mined.
[0083] S1 interface is the communication interface between LTE base station and evolved packet core (EPC), and S1 interface is divided into two interfaces, namely S1-MME and S1-U. Among them, S1-MME is used for the control plane, and S1-U is used for the user plane.
[0084] The S1-MME interface data in the XDR data is mainly the content of the user online signaling surface. The international mobile equipment identity (IMEI), user occupation cell start time, user occupation end time, user occupation cell condition and the like can be obtained from the S1-MME interface data, and are used for high-speed rail user identification.
[0085] 4. High-speed rail user behavior characteristics and identification algorithm model.
[0086] The longitude and latitude of the high-speed rail service cell can be obtained from the basic information of the high-speed rail service cell, so that the distance between the high-speed rail service cells can be calculated. From the information of the S1-MME interface data, the list of user occupation service cells and the start and end time of user occupation service cells can be obtained, so that the speed of the user in a certain interval can be calculated. Then, after analyzing and processing the above information by an algorithm, the high-speed rail user can be identified.
[0087] (1) According to the large granularity model, the high-speed rail user with long time span is identified.
[0088] The high-speed rail user with long time span is analyzed every 30 minutes, and the start and end cells, cell number, distance, time length, speed, direction and the like of the high-speed rail user are recorded. Table 3 below shows the large granularity model.
[0089] Table 3 Large Granularity Model
[0090]
[0091] (2) According to the small granularity model, the high-speed rail user with short time span is identified.
[0092] The high-speed rail user with short time span is analyzed every 10 minutes, and the start and end cells, cell number, distance, time length, speed, direction and the like of the high-speed rail user are recorded. Table 4 below shows the small granularity model.
[0093] Table 4 Small Granularity Model
[0094]
[0095] (3) According to the fusion model, the high-speed rail user is identified.
[0096] The fusion model can obtain the identification result of the high-speed rail user. Every hour, the users identified by the large granularity model and the small granularity model in the last hour are merged to obtain the fusion model, and the start and end cells, cell number, distance, time length, speed, direction and the like of the high-speed rail user are recorded for analysis by the upper layer application.
[0097] From the traditional high-speed rail user identification method, it can be seen that the traditional high-speed rail user identification method mainly obtains the list of user occupied service cells and the start and end time of the occupied service cells from the S1-MME interface data, calculates the speed of the user in each interval, and based on different time span high-speed rail users, uses different granularity models. Therefore, after the algorithm analyzes and processes these basic information, the high-speed rail user is identified.
[0098] 5、LightGBM algorithm.
[0099] Gradient boosting decision tree (GBDT) is a long-lasting model in machine learning, which mainly uses weak classifiers to iteratively train the optimal model. This model has the advantages of good training effect and not easy to overfit. In 2017, Microsoft proposed the LightGBM algorithm, which is an improved algorithm based on GBDT. Compared with GBDT, it can more effectively process massive data. LightGBM algorithm mainly includes the following features: Histogram algorithm, leaf-wise leaf growth strategy with depth limit, gradient-based one-side sampling (GOSS), exclusive feature bundling (EFB), support for category features, efficient parallelism and cache hit rate optimization.
[0100] (1) Histogram algorithm.
[0101] The histogram algorithm is to "bin" the feature values of the original data, divide the data into different discrete areas, and then traverse the discrete data to find the optimal partition point. Among them, each "bin" of the feature value division has two meanings. One is the number of samples in each "bin"; the other is the gradient sum of samples in each "bin" (the square mean of the first-order gradient sum is equivalent to the mean square loss).
[0102] From the above introduction of the histogram algorithm, it can be seen that the model processed by the histogram algorithm can reduce the complexity of the model. And after "binning" according to the feature value of the original data, only discrete values are saved, so the memory occupancy rate is greatly reduced. Secondly, the histogram algorithm uses bin instead of the original data, which is equivalent to increasing the regularization ratio. This will discard more detailed features, and similar data may be divided into the same bin, so the number of bins affects the regularization degree, and the fewer the bins, the lower the risk of overfitting.
[0103] For example, the process of forming a histogram by the histogram algorithm is as follows: Figure 3 shown.
[0104] In the LightGBM algorithm, the histogram algorithm also includes a histogram difference optimization, that is, after the LightGBM algorithm obtains the histogram of a leaf, it can obtain the histogram of its brother leaves at a very low cost by histogram difference.
[0105] For example, Figure 4 As shown in the figure, after obtaining the histogram of a leaf and the histogram of its parent node, the histogram of the sibling leaves of the leaf can also be obtained. In this way, the speed of the LightGBM algorithm can be further optimized.
[0106] (2) Leaf-wise leaf growth strategy with depth restriction.
[0107] Both GBDT and the extreme gradient boosting (XGBoost) model use a layer-wise leaf-wise splitting strategy for their leaf growth strategies. This method targets every node in the same layer during the split, meaning that each iteration requires traversing the entire dataset. While this layer-wise leaf-wise splitting approach allows for parallel processing of leaves in each layer and controls model complexity, since each iteration requires traversing the entire dataset, it results in unnecessary searches and splits, consuming more memory and increasing computational costs.
[0108] The LightGBM algorithm improves on the leaf-wise growth strategy by adopting a leaf-wise splitting approach. Specifically, only the leaf with the highest splitting gain is split each time. This leaf-wise splitting approach reduces error and accelerates learning. However, because it doesn't split other leaves, the split results are less refined. Furthermore, splitting only one leaf per layer increases the tree depth, causing model overfitting. Therefore, the LightGBM algorithm limits the tree depth during leaf-wise growth to avoid overfitting.
[0109] For example, the growth mode of layer-wise splitting and the growth mode of leaf-wise splitting are as follows: Figure 5 shown.
[0110] (3)GOSS.
[0111] In GBDT, each sample has a different gradient value, and the gradient of the sample can reflect the contribution degree to the model. The greater the gradient of the sample, the more information gain the sample contributes to the model, and the smaller the gradient of the sample, the better the sample performs in the model.
[0112] The LightGBM algorithm introduces GOSS, which is based on the idea of reducing samples. The gradient size information of the sample is used as the weight of the importance of the sample. All samples with large gradients are retained, and samples with small gradients are randomly sampled in proportion. In order to not change the data distribution of the samples, a constant is introduced to balance the samples with small gradients when calculating the gain. In this way, the number of samples can be reduced without changing the original data distribution, and the training speed of the model can be improved.
[0113] (4) EFB.
[0114] High-dimensional data is usually very sparse, and there is mutual exclusivity between features. Exemplarily, several features generated after one-hot encoding will not be 0 at the same time. Such data has a certain impact on the effect and running speed of the model. The EFB can solve the sparsity problem of high-dimensional data. If two features are not completely exclusive, an index can be used to measure the degree of non-exclusivity of the features. When the index value is small, we can choose to bundle two features that are not completely exclusive without affecting the final accuracy.
[0115] Exemplarily, as shown in Table 3, assume that feature 1, feature 2, and feature 3 are mutually exclusive sparse features. Through the EFB algorithm, the three features are bundled into a new dense feature, and then the new feature replaces the original three features, thereby reducing the feature dimension without losing information, avoiding unnecessary 0 value calculation, and improving the speed of the gradient boosting algorithm. Table 5 shows that the three features are bundled into a new dense feature through the EFB algorithm:
[0116] Table 5 Three features bundled into a new dense feature
[0117] # Feature 1 Feature 2 Feature 3 New Feature 1 0 2 0 2 2 0 0 0 0 3 0 0 0 0 4 0 0 1 1 5 3 0 0 3
[0118] From the above description of the LightGBM algorithm, it can be seen that the LightGBM algorithm is a highly optimized GBDT algorithm, and at the same time, it can also be regarded as an optimization algorithm of XGboost.
[0119] Exemplarily, the LightGBM algorithm can be expressed by the following formula 1:
[0120] LightGBM = XGboost + Histogram + GOSS + EFB Formula 1
[0121] (5) LightGBM algorithm parameters.
[0122] LightGBM algorithm parameters are complex, and can be roughly divided into core parameters, learning control parameters, IO parameters, target parameters, and metric parameters. Generally, the core parameters, learning control parameters, and metric parameters need to be adjusted.
[0123] For example, Table 6 below shows the default values and interpretations of commonly used important parameters:
[0124] Table 6 Default values and interpretations of commonly used important parameters
[0125]
[0126] For example, as shown in Table 6, the default values and interpretations of commonly used important parameters are as follows: Figure 6 Figure 6 are the specific steps for building a model according to the LightGBM algorithm.
[0127] The above introduces the technical terms involved in the present application.
[0128] At present, high-speed rail has gradually become the first choice for people to travel long distances. High-speed rail users' demand for online is also increasing, so the quality of mobile communication networks on high-speed rail will also affect the brand reputation of operators. Therefore, it is particularly important for operators to establish a high-quality network to meet the online needs of high-speed rail users on high-speed rail. In order to establish a high-quality network to meet the online needs of high-speed rail users, operators and related units have analyzed the positioning of high-speed rail users according to network management data, measurement report (MR) data, XDR data, etc. and on this basis, analyzed the network coverage, service performance and capacity of high-speed rail, and thus proposed corresponding network construction and optimization schemes.
[0129] The traditional high-speed rail user identification method has the following problems:
[0130] (1) Identifying high-speed rail users by matching multiple consecutive service cells of the handover chain is easy to cause missed identification of high-speed rail users;
[0131] (2) Calculate the user's moving speed by the distance between the service cells where the user's signaling appears and the time difference of line mapping, which has a large error;
[0132] (3) The calculation does not consider the user's waiting time and in-station stay time.
[0133] Therefore, this will cause missed identification of some high-speed rail users, so that high-speed rail users cannot be accurately identified.
[0134] To solve the above problems, the high-speed rail user identification method provided in the application first determines the high-speed rail master cell through preliminary screening, then determines a first model according to the fingerprint features of the high-speed rail master cell and a LightGBM algorithm, calculates the confidence of all high-speed rail master cells according to the first model, and then determines the service cell with a confidence greater than a preset confidence threshold as a first target service cell; then an initial model is constructed according to the LightGBM algorithm and the first target service cell, and the initial model is trained according to the features of the first target service cell; when the identification accuracy of the initial model is greater than or equal to an accuracy threshold, the initial model is determined as a second model, so as to identify the high-speed rail user in the user to be identified according to the second model. Thus, the application performs preliminary screening on the service cell, and performs re-screening on the service cell after the preliminary screening through the first model to obtain the service cell with a high confidence as the training object of the initial model, so as to determine the second model. Thus, the application can accurately identify the high-speed rail user through the second model, so as to find potential coverage quality problems, optimize network quality, improve user perception, and create a better high-speed rail network. Before the high-speed rail user identification method of the embodiment of the application is described in detail, the implementation environment and application scenario of the embodiment of the application are introduced.
[0135] As shown in the example of FIG. 7, Figure 7 As shown in the example of FIG. 7,
[0136] Optionally, the base station 701 can be a base station (base transceiver station, BTS) in a global system for mobile communication (GSM), a code division multiple access (CDMA), a base station (node B) in wideband code division multiple access (WCDMA), a base station (eNB) in internet of things (IoT) or narrowband-internet of things (NB-IoT), a base station in a future 5th generation mobile communication technology (5G) mobile communication network or a future evolved public land mobile network (PLMN), and the embodiment of the application does not make any limitation thereto.
[0137] It should be understood that,Figure 7 In the actual application, the number of base stations 701 can be set according to actual conditions, and the present application does not make a specific limitation. For example, the number of base stations 701 is specifically set according to the length of the route of the high-speed rail 702.
[0138] It should be noted that in the embodiment of the present application, the coverage range of the base station 701 is a service cell, wherein the high-speed rail user on the high-speed rail 702 can access the base station 701 in the service cell, and the high-speed rail user can perform online service after accessing the base station 701.
[0139] Optionally, the high-speed rail 702 is a railway with high design standard and a maximum speed of 200 km / h or more. The design standard of high-speed rail is different in different periods, and the definition of high-speed rail is also constantly updated. In the embodiment of the present application, no specific limitation is made.
[0140] As shown in Figure 8 Fig. 1 is a schematic diagram of a hardware structure of a high-speed rail user identification device provided in an embodiment of the present application. The high-speed rail user identification device includes a processor 81, a memory 82, a communication interface 83, and a bus 84. The processor 81, the memory 82, and the communication interface 83 can be connected through the bus 84.
[0141] The processor 81 is the control center of the high-speed rail user identification device, which can be one processor or a general term of multiple processing elements. For example, the processor 81 can be a general central processing unit (CPU), or other general-purpose processors, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0142] As an embodiment, the processor 81 can include one or more CPUs, such as the CPU 0 and the CPU 1 shown in Figure 8
[0143] The memory 82 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.
[0144] In a possible implementation, the memory 82 can exist independently of the processor 81, and the memory 82 can be connected to the processor 81 through the bus 84, for storing instructions or program codes. When the processor 81 invokes and executes the instructions or program codes stored in the memory 82, the high-speed rail user identification method provided in the embodiments of the present application can be implemented.
[0145] In another possible implementation, the memory 82 can also be integrated with the processor 81.
[0146] The communication interface 83 is configured to connect the high-speed rail user identification apparatus to other devices through a communication network. The communication network can be an Ethernet, a wireless access network, a wireless local area network (WLAN), or the like. The communication interface 83 can include a receiving unit configured to receive data, and a sending unit configured to send data.
[0147] The bus 84 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For the convenience of representation, Figure 8 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.
[0148] Figure 9 Another hardware structure of the high-speed rail user identification apparatus in the embodiments of the present application is shown. As shown in the figure, Figure 9 The high-speed rail user identification apparatus can include a processor 91 and a communication interface 92. The processor 91 is coupled to the communication interface 92.
[0149] The functions of the processor 91 can refer to the description of the processor 81. In addition, the processor 91 also has a storage function, which can function as the memory 82.
[0150] The communication interface 92 is configured to provide data for the processor 91. The communication interface 92 can be an internal interface of the high-speed rail user identification apparatus, or an external interface (equivalent to the communication interface 83) of the high-speed rail user identification apparatus.
[0151] It should be noted that, Figure 8 the structure shown in the figure does not constitute a limitation on the high-speed rail user identification apparatus, and Figure 9 except for the structure shown in the figure, other structures can also be used. Figure 8(or Figure 9 The high-speed rail user identification device can include more or fewer components than shown, or combine certain components, or arrange different components, in addition to the components shown in the figure.
[0152] The high-speed rail user identification method provided in the present application will be described in detail below in conjunction with the accompanying drawings of the specification:
[0153] Exemplarily, as shown in the figure, Figure 10 Figure 10 is a flowchart of a high-speed rail user identification method provided in the present application, which includes the following steps:
[0154] S1001, the high-speed rail user identification device determines a high-speed rail master cell.
[0155] The high-speed rail master cell is used to provide services for a high-speed rail user group.
[0156] In one possible implementation, determining the high-speed rail master cell specifically includes the following steps: obtaining a second target service cell; wherein the second target service cell is a service cell with a switching number less than a preset number within a preset time period on a high-speed rail line; determining the high-speed rail user group from a user group corresponding to the second target service cell according to a long interval speed algorithm and a first preset characteristic condition; and determining the high-speed rail master cell according to the high-speed rail user group. It should be noted that the process of specifically determining the high-speed rail master cell by the high-speed rail user identification device can be referred to S1501-S1503 described below, which will not be described here.
[0157] S1002, the high-speed rail user identification device determines a first model according to the fingerprint characteristics of the high-speed rail master cell.
[0158] The first model is used to determine the confidence of the high-speed rail master cell.
[0159] Optionally, the first model is constructed according to a LightGBM algorithm, which is used to extract and train the fingerprint characteristics of the high-speed rail master cell, and determine the confidence of the high-speed rail master cell according to the fingerprint characteristics of the high-speed rail master cell. The introduction and principle of the LightGBM algorithm are described in the fifth part of the technical term introduction in the foregoing, which will not be described here.
[0160] It should be noted that the confidence is also called reliability, or confidence level, or confidence coefficient. When estimating the total parameter by sampling, the conclusion is always uncertain due to the randomness of the sample, therefore, a probability-based statement method is needed, that is, the interval estimation method in mathematical statistics. That is, the corresponding probability of the estimated value and the total parameter within a certain allowable error range is how large, which is called confidence. In this embodiment, the confidence is the weight value of the high-speed rail master cell parameter.
[0161] Optionally, the fingerprint feature of the high-speed rail master cell is a feature such as a start time of user occupation of the cell, an end time of user occupation of the cell, and signal quality when the user occupies the cell.
[0162] S1003, the high-speed rail user identification device determines a first target service cell according to the confidence of the high-speed rail master cell.
[0163] It should be noted that the first target service cell is a service cell in the high-speed rail master cell whose confidence is greater than a preset confidence threshold.
[0164] Optionally, in actual application, the preset confidence threshold can be set according to actual needs, and the present application does not make specific limitations thereto.
[0165] Exemplarily, Figure 11 The confidence of the high-speed rail master cell after K (K=4) verification, and the last column is the average value of the confidence.
[0166] The preset confidence threshold is 0.9999. Specifically, the feature cell whose confidence is greater than the preset confidence threshold can be used as the first target service cell, so that the high-speed rail master cell can be simplified and the computing performance can be improved.
[0167] S1004, the high-speed rail user identification device determines a second model according to the first target service cell.
[0168] The second model is used for identifying high-speed rail users.
[0169] Optionally, the second model is constructed according to the LightGBM algorithm. For the introduction and principle of the LightGBM algorithm, please refer to the fifth part of the technical term introduction in the foregoing, which will not be repeated here.
[0170] It should be noted that the second model is determined according to the first target service cell, which specifically includes the following three steps:
[0171] (1) An initial model is constructed according to the LightGBM algorithm. The LightGBM algorithm is used for feature extraction and training of the first target service cell. For the introduction and principle of the LightGBM algorithm, please refer to the fifth part of the technical term introduction in the foregoing, which will not be repeated here.
[0172] (2) According to the K-fold cross-validation algorithm, it is judged whether the identification accuracy of the initial model is greater than or equal to an accuracy threshold;
[0173] (3) In a case where the identification accuracy of the initial model is greater than or equal to the accuracy threshold, the initial model is determined as the second model. Optionally, although the first model can also achieve high-precision identification, the first model is abnormally large in size because it is trained using the fingerprints of all high-speed rail master cells. Therefore, the initial model is constructed using the LightGBM algorithm for the second time, and the first target service cell is used for training, so as to obtain the second model. In this way, the computing performance can be optimized, and the identification accuracy will not be reduced.
[0174] Exemplarily, as shown in FIG. 6, Figure 12 Figure 12 The prediction accuracy of the initial model according to K-fold (n = 4) cross-validation is shown. The positive samples are high-speed rail users determined according to the long interval algorithm, and the negative samples are other users. It can be seen that the identification accuracy reaches 98.61%, and accurate high-speed rail user identification can be achieved.
[0175] S1005, the high-speed rail user identification device identifies the high-speed rail user in the to-be-identified user according to the second model.
[0176] Exemplarily, the identified high-speed rail user is tested on a certain high-speed rail line according to an actual scene, as shown in FIG. 7, Figure 13-14 Figure 13 is the driving trajectory of the user on a certain high-speed rail line, Figure 14 is the driving trajectory and MR data of the high-speed rail user on a certain high-speed rail line. It can be seen from Figure 14 that the MR data and the driving trajectory of the high-speed rail user are highly coincident, thereby indicating that the high-speed rail user identification method proposed in the present application can accurately identify high-speed rail users.
[0177] Based on the technical solution, the high-speed rail user identification method provided in the application first determines a high-speed rail master cell, and then determines a first model according to the fingerprint feature of the high-speed rail master cell and a LightGBM algorithm; then the confidence of all high-speed rail master cells is calculated according to the first model, and then the service cell with a confidence greater than a preset confidence threshold is determined as a first target service cell; then an initial model is constructed according to the LightGBM algorithm and the first target service cell, and the initial model is trained according to the feature of the first target service cell; when the identification accuracy of the initial model is greater than or equal to an accuracy threshold, the initial model is determined as a second model, so as to identify the high-speed rail user in the to-be-identified user according to the second model. Thus, after the service cell is preliminarily screened, the service cell after the preliminary screening is further screened by the first model, and the service cell with a high confidence is obtained as the training object of the initial model, so as to determine the second model. Thus, the high-speed rail user can be accurately identified by the second model, so as to find potential coverage quality problems, optimize network quality, improve user perception, and create a better high-speed rail network.
[0178] Exemplarily, in combination with Figure 10 As Figure 15 shown, the step 1001 can be implemented by the following S1501-S1503:
[0179] S1501, the high-speed rail user identification device acquires a second target service cell.
[0180] The second target service cell is a service cell with a switching number less than a preset number in a preset period on a high-speed rail line.
[0181] It should be noted that the preset period and the preset number can be set according to actual needs, and the application does not make specific limitations thereto.
[0182] It should be noted that the specific way in which the high-speed rail user identification device acquires the second target service cell can be referred to S1801-S1804 below, which will not be described here.
[0183] S1502, the high-speed rail user identification device determines a high-speed rail user group from a user group corresponding to the second target service cell according to a long-interval speed algorithm.
[0184] In one possible implementation, a plurality of users in the user group corresponding to the second target service cell have the same motion feature and a moving speed of 200km / h or more at one time or more times, and are determined as the high-speed rail user group.
[0185] Exemplarily, in combination with Figure 16 The long-interval algorithm is specifically described.
[0186] Optionally, assuming that the coverage range of each service cell is 5 kilometers, first take the switching time of the user from A to B and D to E as the time difference, and then take the distance between the median of AB and the intersection M of the high-speed rail line and the median of DE and the intersection N of the high-speed rail line as MN, and calculate the speed of the user corresponding to the second target service cell by the formula "speed = distance / time difference".
[0187] In a possible implementation, the high-speed rail user group may be determined from the user group corresponding to the second target serving cell by using a long interval speed algorithm and a first preset characteristic condition.
[0188] For example, Figure 17 As shown, this is the calculation result of a random sampling of identified high-speed rail users. It can be seen that the speed difference is small. The high-speed rail user group can be effectively determined from the user group corresponding to the second target service cell through the long interval speed algorithm and the first preset characteristic condition.
[0189] It should be noted that there must be an error between the distance between the above MNs and the actual high-speed rail route. However, when the distance between AB and DE is far enough, the error can be ignored, so the negative impact on this scheme can also be ignored.
[0190] For example, assuming that the distance between AB and DE is 5000 meters, and the distance between AB, DE and the high-speed railway line is 100 meters, the error calculation process is:
[0191]
[0192] From this we can see that the error is less than 0.1% and can be ignored.
[0193] S1503. The high-speed rail user identification device determines the high-speed rail master control cell according to the first preset characteristic condition and the high-speed rail user group.
[0194] It should be noted that the first preset characteristic condition includes one or more of the following: the moving route is periodic, the distance from the high-speed rail route is less than the first preset distance threshold, and the cell unique identifier ECI of the corresponding service cell within the preset time period is the same.
[0195] In one possible implementation, combined with big data analysis and other methods, a service cell that meets the first preset characteristic condition can be obtained from the service cells passed by the high-speed rail user group, and determined as the high-speed rail master cell.
[0196] Based on the above technical solution, the high-speed rail user identification method provided by this application can screen the service cells. First, the service cells with a switching frequency less than the preset frequency within the preset time period are obtained and determined as the second target service cells. Then, based on the long interval speed algorithm and the first preset characteristic condition, the high-speed rail user group is determined from the user group corresponding to the second target service cell. Thus, the high-speed rail master cell is determined based on the high-speed rail user group. Therefore, after the preliminary screening of the service cells, this application uses the first model to screen the service cells after the preliminary screening again, and obtains the service cells with the highest confidence as the training objects of the initial model, thereby determining the second model. Therefore, this application can accurately identify high-speed rail users through the second model, thereby discovering potential coverage quality problems, optimizing network quality, improving user perception, and creating a higher-quality high-speed rail network.
[0197] For example, in combination Figure 15 ,like Figure 18 As shown, the above S1501 can be specifically implemented through the following S1801-S1804:
[0198] S1801. The high-speed rail user identification device obtains a third target serving cell.
[0199] The shortest distance between the third service cell and the high-speed rail route is less than a second preset distance threshold;
[0200] It should be noted that the second preset distance threshold can be set according to actual needs, and this application does not impose any specific restrictions on this.
[0201] Optionally, the railway layer in the GIS layer is used to mark all cells along the high-speed railway as third target service cells.
[0202] For example, taking region A as an example, Figure 19 As shown, Figure 19 This is a distribution map of service cells along the high-speed rail lines in Region A. Region A currently has six high-speed rail lines: High-speed Rail Line 1, High-speed Rail Line 2, High-speed Rail Line 3, High-speed Rail Line 4, High-speed Rail Line 5, and High-speed Rail Line 6. Extract all cells within N kilometers (N can be 0 to 10) along the high-speed rail lines as the third target service cells.
[0203] S1802. The high-speed rail user identification device obtains S1-MME interface data of users in the third target service cell.
[0204] Among them, the S1-MME interface data includes the ECI of the third target serving cell.
[0205] For example, Figure 20 It is the S1-MME interface data of a user in the third target serving cell.
[0206] S1803, the high-speed user identification device time-sequences the S1-MME interface data.
[0207] Exemplarily, as Figure 21 is the result of time-sequencing the S1-MME interface data of the user in Figure 20 .
[0208] S1804, the high-speed user identification device acquires the second target service cell from the third target service cell according to the sequenced S1-MME interface data.
[0209] Among them, the second target service cell is a service cell that does not frequently switch in the third target service cell.
[0210] Exemplarily, in combination with Figure 21 , from the time result, there is a user frequently switching service cells, for example, at 08:08:08:23 7548 occupies cell 18349057, switches to cell 18131459 after 2 seconds, and switches back to cell 18349057 after 17 seconds. This situation only occurs in a static scenario.
[0211] Exemplarily, as Figure 22 is shown, Figure 22 is the switching process diagram between three cells of the user. As can be seen from the diagram, the user frequently switches between service cell A and service cell B, and finally switches to service cell C. In actual application, the user will frequently switch between multiple service cells, and to obtain the second target service cell, an algorithm needs to be used to exclude the frequently switched cells from the third target service cell.
[0212] Exemplarily, the algorithm used is:
[0213] (1) Frequently switching within n cells in 10 seconds (n currently supports less than 5), only taking the last time of the appearance of the service cell as the sequence time of the service cell, i.e. taking leaving the cell as the basis for judgment, the time slice value may be lost, taking entering the cell or leaving the cell has little effect on the final calculation speed, for explanation, refer to the speed algorithm part.
[0214] (2) Frequently switching within n cells for more than 30 seconds (n currently supports less than 5), and there is a neighboring cell relationship between the cells, then it is judged that the user is likely to be static, at this time the speed of the user is marked as 0.
[0215] Based on the technical scheme, the high-speed rail user identification method provided by the application can preliminarily screen the service cells, determine the service cells with the shortest distance to the high-speed rail line less than the second preset distance threshold as the third target service cells, then acquire the S1-MME interface data of the third target service cells, and perform time sorting on the S1-MME interface data, so as to acquire the service cells with the switching times less than the preset times in the preset time period from the third target service cells according to the algorithm, and determine the service cells as the second target service cells. Thus, the application can further screen the service cells after the preliminary screening by the first model, obtain the service cells with high confidence as the training objects of the initial model, and determine the second model, so that the application can accurately identify the high-speed rail users by the second model, thereby discovering potential coverage quality problems, optimizing the network quality, improving the user perception, and creating a better high-speed rail network.
[0216] Exemplarily, as shown in Figure 23 a structure schematic diagram of a high-speed rail user identification device provided by an embodiment of the application, the device comprises a processing unit 2301 and an acquisition unit 2302.
[0217] Optionally, the processing unit is configured to determine a high-speed rail master cell; wherein the high-speed rail master cell is configured to provide services for a high-speed rail user group.
[0218] Optionally, the processing unit 2301 is further configured to determine a first model according to a fingerprint feature of the high-speed rail master cell; wherein the first model is configured to determine a confidence degree of the high-speed rail master cell.
[0219] Optionally, the processing unit 2301 is further configured to determine a first target service cell according to the confidence degree of the high-speed rail master cell.
[0220] Optionally, the processing unit 2301 is further configured to determine a second model according to the first target service cell.
[0221] Optionally, the processing unit 2301 is further configured to identify the high-speed rail users in the to-be-identified users according to the second model.
[0222] Optionally, the acquisition unit 2302 is configured to acquire the second target service cell.
[0223] Optionally, the processing unit 2301 is further configured to determine a high-speed rail user group from a user group corresponding to the second target service cell according to a long-interval speed algorithm and a first preset feature condition.
[0224] Optionally, the processing unit 2301 is further configured to determine the high-speed rail master cell according to the high-speed rail user group.
[0225] Optionally, the processing unit 2301 is further configured to determine the third target serving cell.
[0226] Optionally, the obtaining unit 2302 is further configured to obtain S1-MME interface data of the user in the third target serving cell.
[0227] Optionally, the processing unit 2301 is further configured to time-sort the S1-MME interface data.
[0228] Optionally, the processing unit 2301 is further configured to obtain the second target serving cell from the third serving cell according to the sorted S1-MME interface data.
[0229] Optionally, the processing unit 2301 is further configured to construct the first model and the second model according to a LightGBM algorithm.
[0230] Optionally, the processing unit 2301 is further configured to construct the initial model according to the LightGBM algorithm.
[0231] Optionally, the processing unit 2301 is further configured to determine, according to a K-fold cross-validation algorithm, whether the recognition accuracy of the initial model is greater than or equal to an accuracy threshold.
[0232] Optionally, the processing unit 2301 is further configured to determine the initial model as the second model in a case where the recognition accuracy of the initial model is greater than or equal to the accuracy threshold.
[0233] In addition, Figure 23 The technical effects of the high-speed rail user identification device can refer to the technical effects of the high-speed rail user identification method of the above-mentioned embodiments, which will not be repeated here.
[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0235] The embodiment of the application provides a computer program product containing instructions, when the computer program product runs on a computer, so that the computer executes the high-speed rail user identification method in the above method embodiments.
[0236] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the high-speed rail user identification method in the method flow shown in the method embodiment.
[0237] The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a register, a hard disk, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing, or any other medium from which a computer can read instructions. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Of course, the storage medium can be a part of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). In the embodiment of the present application, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device.
[0238] Since the high-speed rail user identification device, the computer readable storage medium, and the computer program product in the embodiment of the present application can be applied to the above method, the technical effects that can be obtained are also referable to the above method embodiment, and the embodiment of the present application will not be described here.
[0239] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited to this. Any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A high-speed rail user identification method, characterized in that: The method comprises: Determine a high-speed rail master cell; wherein the high-speed rail master cell is used to provide services for the high-speed rail user group; Determine a first model based on the fingerprint features of the high-speed rail master cell and the LightGBM algorithm; wherein the first model is used to determine the confidence of the high-speed rail master cell; the fingerprint features include: the start time of the user occupying the cell, the end time of the user occupying the cell, and the signal quality when the user occupies the cell; Determining a first target serving cell according to the confidence of the high-speed rail master cell; the first target serving cell is a serving cell in the high-speed rail master cell whose confidence is greater than a preset confidence threshold; Constructing an initial model based on the fingerprint features of the first target serving cell and the LightGBM algorithm; the LightGBM algorithm is used to extract and train fingerprint features of the first target serving cell; Determine whether the recognition accuracy of the initial model is greater than or equal to the accuracy threshold according to the K-fold cross validation algorithm; When the recognition accuracy of the initial model is greater than or equal to the accuracy threshold, the initial model is determined as the second model; wherein the second model is used to identify high-speed rail users; The high-speed rail users among the users to be identified are identified according to the second model.
2. The method according to claim 1, characterized in that The determining of the high-speed rail master control cell specifically includes: Obtain a second target serving cell; wherein the second target serving cell is a serving cell with a switching frequency less than a preset frequency within a preset period of time on the high-speed railway line; determining the high-speed rail user group from the user groups corresponding to the second target serving cell according to a long interval speed algorithm; The high-speed rail master cell is determined according to the first preset characteristic condition and the high-speed rail user group.
3. The method according to claim 2, characterized in that The first preset characteristic condition includes one or more of the following: the moving route is periodic, the distance from the high-speed rail route is less than a first preset distance threshold, and the cell unique identifier ECI of the corresponding service cell within a preset time period is the same.
4. The method according to claim 3, characterized in that The acquiring of the second target serving cell specifically includes: Determining a third target serving cell; wherein the shortest distance between the third target serving cell and the high-speed rail route is less than a second preset distance threshold; Acquire S1-mobility management entity (MME) interface data of a user in the third target serving cell; the S1-MME interface data includes the ECI of the third target serving cell; Time sorting of the S1-MME interface data; The second target serving cell is obtained from the third target serving cell according to the sorted S1-MME interface data.
5. A high-speed rail user identification device, characterized in that: The device comprises: a processing unit; The processing unit is configured to determine a high-speed rail master cell; wherein the high-speed rail master cell is configured to provide services for a high-speed rail user group; The processing unit is further configured to determine a first model based on the fingerprint features of the high-speed rail master cell and the LightGBM algorithm; wherein the first model is used to determine the confidence of the high-speed rail master cell; the fingerprint features include: the start time of the user occupying the cell, the end time of the user occupying the cell, and the signal quality when the user occupies the cell; The processing unit is further configured to determine a first target serving cell according to the confidence of the high-speed rail master cell; the first target serving cell is a serving cell in the high-speed rail master cell whose confidence is greater than a preset confidence threshold; The processing unit is further configured to construct an initial model based on the fingerprint features of the first target serving cell and the LightGBM algorithm; the LightGBM algorithm is configured to extract and train fingerprint features of the first target serving cell; The processing unit is further configured to determine whether the recognition accuracy of the initial model is greater than or equal to an accuracy threshold according to a K-fold cross validation algorithm; The processing unit is further configured to determine the initial model as a second model when the recognition accuracy of the initial model is greater than or equal to the accuracy threshold; wherein the second model is used to identify high-speed rail users; The processing unit is further configured to identify the high-speed rail user among the users to be identified according to the second model.
6. The device according to claim 5, characterized in that The device further includes: an acquisition unit; The acquisition unit is configured to acquire a second target serving cell; wherein the second target serving cell is a serving cell whose number of handovers within a preset time period on the high-speed railway line is less than a preset number; The processing unit is further configured to determine the high-speed rail user group from the user group corresponding to the second target serving cell according to a long interval speed algorithm; The processing unit is further used to determine the high-speed rail master cell according to the first preset characteristic condition and the high-speed rail user group.
7. The device according to claim 6, characterized in that The first preset characteristic condition includes one or more of the following: the moving route is periodic, the distance from the high-speed rail route is less than a first preset distance threshold, and the ECI of the corresponding service cell within a preset time period is the same.
8. The device according to claim 7, characterized in that The acquiring the second target serving cell specifically includes: The processing unit is further configured to determine a third target serving cell; wherein the shortest distance between the third target serving cell and the high-speed rail route is less than a second preset distance threshold; The acquiring unit is further configured to acquire S1-MME interface data of a user in the third target serving cell; the S1-MME interface data includes the ECI of the third target serving cell; The processing unit is further configured to perform time sorting on the S1-MME interface data; The processing unit is further configured to obtain the second target serving cell from the third target serving cell according to the sorted S1-MME interface data.
9. A high-speed rail user identification device, characterized in that: include: A processor and a communication interface; the communication interface is coupled to the processor, and the processor is used to run a computer program or instruction to implement the high-speed rail user identification method as described in any one of claims 1-4.
10. A computer-readable storage medium storing instructions, characterized in that: When a computer executes the instruction, the computer executes the high-speed rail user identification method described in any one of claims 1 to 4 above.
Citation Information
Patent Citations
Wide-area high-speed rail base station identification method and device, server and storage medium
CN112333689A
User identification method and device, computer equipment and storage medium
CN116033468A