Center position determination method and device and electronic equipment

By using deep packet inspection signaling data and cluster analysis from operators, the regional center can be accurately located, solving the problem of inaccurate location determination of the regional center in existing technologies. This provides high-precision, low-cost data support and promotes the dynamic identification and evaluation of urban functional areas.

CN121310067APending Publication Date: 2026-01-09CHINA TELECOM CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511377394.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently and accurately determine the location of regional centers, making it impossible to provide high-precision, low-cost data support for public services.

Method used

Based on the deep packet inspection signaling data of the operator, the characteristic data of the serving cell is determined, and the regional center point is identified through cluster analysis, including K-means clustering and DBSCAN density clustering. Combined with geographic information system technology, it is converted into a structured address.

Benefits of technology

It enables efficient and accurate determination of regional center locations, providing high-precision and low-cost data support for public services, and timely identification of emerging business hotspots and residential areas, supporting urban planning and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121310067A_ABST
    Figure CN121310067A_ABST
Patent Text Reader

Abstract

The invention discloses a central position determination method and apparatus, and an electronic device. The method comprises the following steps: on the basis of deep packet inspection signaling data of an operator, determining feature data of a service cell, the deep packet inspection signaling data at least comprising a user identifier, a service cell identifier and service occurrence time, and the feature data at least comprising stay durations of different users in different service cells; based on the feature data, performing clustering processing on the service cells to obtain a plurality of first clustering results; performing clustering processing on the physical position of the service cell in each first clustering result to obtain a plurality of second clustering results; and determining the central point of the second clustering result according to the mean value of the second clustering result, and determining the position information of the central point. The technical problem that high-precision and low-cost data support cannot be provided for public services due to the fact that the center position of the area cannot be efficiently and accurately determined in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of big data processing, and in particular, to a center position determination method and device and electronic equipment. BACKGROUND

[0002] With the vigorous development of digital economy and the trend of consumption upgrading, the spatial function of the city is evolving at an unprecedented speed. Emerging consumption places, entertainment centers, and non-traditional service areas are constantly emerging, which has a profound impact on the city's lifestyle and economic pattern. In order to adapt to this change, it is urgently needed to have a tool that can accurately and dynamically identify and evaluate the city's functional areas, so as to facilitate scientific planning and resource allocation and promote the high-quality development of service consumption.

[0003] Current city functional area identification mainly relies on two traditional methods: one is field research, which makes empirical judgments through direct observation of indicators such as pedestrian flow and commercial facility density; the other is data integration, which combines POI (Point of Interest) and land planning data for spatial analysis. Although these two methods are effective to some extent, they have the following obvious shortcomings when dealing with the rapid changes in the digital economy era:

[0004] 1. Field research is time-consuming and costly, and it is difficult to update continuously. Static POI data is convenient, but it lacks dynamic and flexibility, and cannot capture the immediate changes of emerging, characteristic or non-traditional areas, resulting in information lag and incomplete area identification.

[0005] 2. Related methods divide functional areas based on fixed administrative divisions or pre-set commercial district boundaries. This one-size-fits-all approach ignores the true continuity of functional areas and cannot accurately represent the actual structure and function distribution of urban space, especially in large commercial districts, multi-commercial district clusters or functional overlapping areas, which may cause boundary conflicts and misclassification.

[0006] In summary, related technologies cannot efficiently and accurately determine the center position of the area, which leads to the inability to provide high-precision and low-cost data support for public services.

[0007] To address the above problems, no effective solutions have been proposed so far. SUMMARY

[0008] The present application provides a center position determination method, device and electronic equipment to at least solve the technical problem of being unable to provide high-precision and low-cost data support for public services due to the inability of related technologies to efficiently and accurately determine the center position of the area.

[0009] According to an aspect of the present application, a method for determining a central position is provided, comprising: determining characteristic data of a serving cell based on deep packet inspection signaling data of an operator, wherein the deep packet inspection signaling data at least includes user identification, serving cell identification, and service occurrence time, and the characteristic data at least includes the staying duration of different users in different serving cells; performing clustering processing on the serving cells based on the characteristic data to obtain a plurality of first clustering results; performing clustering processing on the physical positions of the serving cells in each first clustering result to obtain a plurality of second clustering results; determining a center point of each second clustering result according to the mean value of the second clustering result, and determining the position information of the center point.

[0010] Optionally, the determining of the characteristic data of the serving cell based on the deep packet inspection signaling data of the operator comprises: determining the staying duration of different users in different serving cells based on the deep packet inspection signaling data, and constructing a first serving cell portrait according to the staying duration of different users in different serving cells; determining an operation performance index of the users in different preset time periods in the first serving cell portrait, wherein the operation performance index at least includes user activity, per capita staying duration, people flow fluctuation rate, and traffic volume; constructing a second serving cell portrait according to the operation performance index; determining the connection change rate and the number change rate of the users in a first target time period in the second serving cell portrait, wherein the first target time period includes a holiday; adding the connection change rate, the number change rate, and the position information of the serving cell to the second serving cell portrait to obtain the characteristic data of the serving cell.

[0011] Optionally, based on the feature data, the serving cells are clustered to obtain a plurality of first clustering results, including: determining a first clustering cluster satisfying a first condition in the serving cells, wherein the first condition includes that a per capita stay time in a first preset time period, a second preset time period and a third preset time period is greater than a first preset threshold, and a human flow fluctuation rate is less than a second preset threshold, a latest time in the first preset time period is earlier than an earliest time in the second preset time period, and a latest time in the second preset time period is earlier than an earliest time in the third preset time period; determining a second clustering cluster satisfying a second condition in the serving cells, wherein the second condition includes that the per capita stay time is less than a third preset threshold, and a number change rate in a first target time period is greater than a fourth preset threshold and a human flow fluctuation rate in the first target time period is greater than a fifth preset threshold; determining a third clustering cluster satisfying a third condition in the serving cells, wherein the third condition includes that a user activity in the first preset time period is less than a sixth preset threshold, a traffic volume in the second preset time period is less than a seventh preset threshold, and a number change rate in the first target time period is greater than an eighth preset threshold, wherein the eighth preset threshold is greater than the fourth preset threshold; determining a fourth clustering cluster satisfying a fourth condition in the serving cells, wherein the fourth condition includes that the user activity in the first preset time period is greater than a ninth preset threshold, the number change rate in the first target time period is less than a tenth preset threshold, and a human flow fluctuation rate in a second target time period is greater than an eleventh preset threshold, wherein the second target time period includes a weekday; and determining the first clustering cluster, the second clustering cluster, the third clustering cluster and the fourth clustering cluster as the plurality of first clustering results.

[0012] Optionally, the physical locations of the serving cells in each first clustering result are clustered to obtain a plurality of second clustering results, including: clustering the physical locations of the serving cells in each first clustering result based on a neighborhood radius and a minimum point number to obtain a plurality of initial second clustering results; determining a target area where the initial second clustering results are located; obtaining a preset type of geographical area entity in the target area, determining a first number of the geographical area entity, and determining a second number of the serving cells in the initial second clustering results; calculating a target difference value between the first number and the second number; in a case where the target difference value is not in a preset interval, adjusting the neighborhood radius and / or the minimum point number, and re-clustering the physical locations of the serving cells in each first clustering result until the target difference value is in the preset interval to obtain the plurality of second clustering results.

[0013] Optionally, the second clustering result comprises a discrete set of latitude and longitude points; after obtaining the plurality of second clustering results, the method further comprises: identifying extreme latitude and longitude points in the discrete set of latitude and longitude points in the second clustering result; determining the extreme latitude and longitude points as initial boundary points, establishing a polar coordinate system with the lowest latitude point as a reference; performing polar angle sorting processing on the remaining latitude and longitude points other than the initial boundary points in the polar coordinate system to obtain an ordered point sequence; sequentially traversing each latitude and longitude point in the ordered point sequence, removing the top point that causes concavity in each latitude and longitude point until the latitude and longitude point remains convex; after completing the traversal, obtaining a convex hull boundary point set; sequentially connecting the latitude and longitude coordinates of the points in the convex hull boundary point set according to the polar angle order to obtain a minimum convex polygon geographic area; determining a vertex coordinate sequence composed of the vertex coordinates of the minimum convex polygon geographic area as a geographic boundary description of the second clustering result.

[0014] Optionally, the position information of the center point comprises: obtaining latitude and longitude information of the center point in the network management system, and constructing an interface request address based on the latitude and longitude information; sending the interface request address to a map service providing object, and receiving standardized format response data returned by the map service providing object; processing the standardized format response data using a recursive parsing algorithm to extract address components layer by layer to obtain address components at each level; combining and splicing the address components at each level according to a preset address formatting rule from large to small level to obtain position information of the center point described by a structured physical address.

[0015] Optionally, after determining the feature data of the service cell, the method further comprises: for target feature data of the same dimension in the feature data, calculating a mean value and a standard deviation, determining an upper limit and a lower limit of the abnormal value judgment of the target feature data based on the mean value and three times the standard deviation calculation; identifying abnormal data points in the target feature data that exceed the upper limit and the lower limit of the abnormal value judgment, and replacing and filling the abnormal data points using the upper limit value and / or the lower limit value corresponding to the target feature data to obtain first feature data; calculating a Pearson correlation coefficient between feature data of any two dimensions in the first feature data, and constructing a correlation coefficient matrix based on the Pearson correlation coefficient between the feature data of any two dimensions; identifying a feature pair in the correlation coefficient matrix that is greater than or equal to a preset correlation threshold, and deleting the feature data of one dimension in the feature pair to obtain second feature data; converting the data scale in the second feature data to a preset standard range.

[0016] According to still another aspect of the present application, a center position determination apparatus is also provided, comprising: a first determination module configured to determine feature data of service cells based on deep packet inspection signaling data of an operator, wherein the deep packet inspection signaling data at least comprises user identity, service cell identity, and service occurrence time, and the feature data at least comprises stay duration of different users in different service cells; a generation module configured to perform clustering processing on the service cells based on the feature data to obtain a plurality of first clustering results; a clustering module configured to perform clustering processing on physical positions of the service cells in each first clustering result to obtain a plurality of second clustering results; and a second determination module configured to determine a center point of each second clustering result according to a mean value of the second clustering result, and determine position information of the center point.

[0017] According to still another aspect of the present application, a non-volatile storage medium is also provided, comprising a stored program, wherein the program, when executed, controls a device in which the storage medium is located to perform the above center position determination method.

[0018] According to still another aspect of the present application, an electronic device is also provided, comprising a memory and a processor, wherein the processor is configured to execute a program stored in the memory, and the program, when executed, performs the above center position determination method.

[0019] According to still another aspect of the present application, a computer program is also provided, wherein the computer program, when executed by a processor, implements the above center position determination method.

[0020] According to still another aspect of the present application, a computer program product is also provided, comprising a non-volatile computer readable storage medium, wherein the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the above center position determination method.

[0021] In the present application, feature data of service cells is determined based on deep packet inspection signaling data of an operator, wherein the deep packet inspection signaling data at least comprises user identity, service cell identity, and service occurrence time, and the feature data at least comprises stay duration of different users in different service cells; clustering processing is performed on the service cells based on the feature data to obtain a plurality of first clustering results; clustering processing is performed on physical positions of the service cells in each first clustering result to obtain a plurality of second clustering results; and a center point of each second clustering result is determined according to a mean value of the second clustering result, and position information of the center point is determined, thereby achieving the purpose of efficiently and accurately determining regional center positions, and thereby achieving the technical effect of providing high-precision and low-cost data support for public services, and further solving the technical problem that high-precision and low-cost data support for public services cannot be provided due to the inability of related technologies to efficiently and accurately determine regional center positions. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0023] Figure 1 This is a flowchart of a method for determining the center position according to an embodiment of this application;

[0024] Figure 2 This is a schematic diagram of a clustering scenario corresponding to a first cluster, a second cluster, a third cluster, and a fourth cluster according to an embodiment of this application;

[0025] Figure 3 This is a schematic diagram illustrating the use of the elbow method to determine the value of K according to an embodiment of this application;

[0026] Figure 4 This is a schematic diagram illustrating the determination of the K value using the profile coefficient method according to an embodiment of this application;

[0027] Figure 5 This is a schematic diagram of density clustering in a clustering scenario according to an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of a second clustering result according to an embodiment of this application;

[0029] Figure 7 This is a schematic diagram illustrating a geographical boundary description according to an embodiment of this application;

[0030] Figure 8 This is a schematic diagram of the Pearson correlation coefficient between feature data of any two dimensions according to an embodiment of this application;

[0031] Figure 9 This is a flowchart of another method for determining the center position according to an embodiment of this application;

[0032] Figure 10 A structural diagram of a device for determining a center position according to an embodiment of this application;

[0033] Figure 11 This is a hardware structure block diagram of a computer terminal according to an embodiment of the present application of a method for determining the center position. Detailed Implementation

[0034] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.

[0035] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] According to the embodiments of the present application, a method embodiment of a center position determination method is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0037] Figure 1 is a flowchart of a center position determination method according to an embodiment of the present application, as shown in Figure 1 The method comprises the following steps:

[0038] Step S102, determining feature data of the service cell based on the deep packet inspection signaling data of the operator, wherein the deep packet inspection signaling data at least includes: user identification, service cell identification, service occurrence time, and the feature data at least includes: the stay duration of different users in different service cells.

[0039] The core of step S102 is to extract feature data of the serving cell based on the operator's Deep Packet Inspection (DPI) signaling data to reveal user behavior patterns and functional attributes within its coverage area. DPI signaling data is detailed data recording user network behavior within the operator's network, including key information such as user identifier, serving cell identifier, and service occurrence time. By analyzing this data, feature data of the serving cell can be constructed, one of the most representative features being the dwell time of different users in different serving cells. This feature data reflects the activity level of the area covered by the serving cell and users' usage habits in that area, serving as an important basis for identifying the area's function. For example, a longer dwell time may indicate that the area covered by the serving cell is office or residential, while a shorter dwell time may suggest commercial, entertainment, or transportation hub attributes.

[0040] Step S104: Based on the feature data, cluster the serving cells to obtain multiple first clustering results.

[0041] In step S104, the K-means clustering algorithm can be used to cluster the serving cells based on feature data. Specifically, an initial number of clusters K needs to be determined first, which can be estimated through prior analysis or using techniques such as the elbow rule or silhouette coefficient. After selecting a value for K, the algorithm randomly selects or initializes K cluster centers according to a certain strategy. Then, the feature data of each serving cell is compared with these K centers, and each serving cell is assigned to the nearest center. This distance can be Euclidean distance, Manhattan distance, or other measures suitable for the feature data type. After assignment, the algorithm recalculates the center of each cluster, updating the center position by calculating the mean of the feature vectors of all serving cells within the cluster. Then, the serving cells are compared with the updated centers again, and reassigned. This process iterates repeatedly until the assignment of all serving cells no longer changes or reaches a preset convergence criterion. At this point, K-means clustering forms a stable clustering result, i.e., multiple first clustering results.

[0042] Step S106: Perform clustering processing on the physical location of the serving cell in each first clustering result to obtain multiple second clustering results.

[0043] In step S106, the DBSCAN density clustering algorithm can be used for secondary clustering to discover geographically proximate clusters with consistent behavioral patterns. The DBSCAN algorithm, by defining the neighborhood radius eps and the minimum number of points MinPts, can identify areas whose "density" meets specific criteria, even if these areas have irregular shapes. The clustering in step S106, based on the latitude and longitude coordinates of the serving cell, combined with the clustering results from step S104, can more accurately depict the spatial extent of the functional area.

[0044] In step S108, the center point of the second clustering result is determined according to the mean value of the second clustering result, and the position information of the center point is determined.

[0045] In step S108, the center point of each cluster, that is, the "core" position of the region, is determined by calculating the mean value of the second clustering result. Specifically, by taking the longitude and latitude coordinates of the service cell as input, using the K-Means clustering algorithm with K=1, the unique clustering center point obtained after the algorithm converges is actually the geometric center of the service cell position in the region, that is, the center point. Although the center point may not completely correspond to the hotspot or the point with the most dense flow in the region, as a representative of the clustering result, it can best reflect the collective position characteristics of all service cells in the region. Then, by using the reverse geocoding API provided by the map service provider, the center point coordinates are converted into structured physical address information, thereby realizing intelligent and refined mining from the original signaling data to the specific regional functional scene, and providing valuable spatial decision basis for public services, commercial layout, traffic planning, etc.

[0046] The above steps S102 to S108 form a closed-loop information transformation process from abstract extraction of data to clustering analysis and then to specific positioning in geographic space. Not only the real-time and comprehensive nature of the operator DPI signaling data is effectively utilized, but also the limitations of related research methods are overcome through hierarchical clustering strategies, realizing dynamic and refined identification and evaluation of urban functional areas. This high-precision and low-cost identification method can not only discover emerging commercial hotspots and residential areas in a timely manner, but also provide scientific basis for the optimized allocation of public services.

[0047] The steps shown in FIG. 8 will be exemplarily described and explained as follows. Figure 1

[0048] According to some optional embodiments of the present application, based on the deep packet inspection signaling data of the operator, the feature data of the service cell can be determined by the following method: based on the deep packet inspection signaling data, the residence duration of different users in different service cells is determined, and a first service cell portrait is constructed according to the residence duration of different users in different service cells; in the first service cell portrait, the operation efficiency indicators of the users in different preset time periods are determined, wherein the operation efficiency indicators at least include: user activity, average residence duration, traffic fluctuation rate and traffic volume; according to the operation efficiency indicators, a second service cell portrait is constructed; in the second service cell portrait, the connection change rate and the number change rate of the users in the first target period are determined, wherein the first target period includes: holidays; the connection change rate, the number change rate and the location information of the service cell are added to the second service cell portrait to obtain the feature data of the service cell.​

[0049] Before determining the residence duration of different users in different service cells based on the deep packet inspection signaling data, the following steps can be performed: operating the DPI signaling data in the form of SQL statements through SparkSQL technology. In this process, the outliers, duplicates and invalid records in the original signaling data are removed, ensuring the purity of the data set and the accuracy of subsequent analysis. At the same time, through the aggregation function and window function of SparkSQL, the residence duration and service times of each user in different service cells can be conveniently calculated. Three core dimensions in the DPI signaling data are extracted: user identification, service cell and service occurrence time, wherein the user identification allows tracking the behavior trajectory of a specific user, the service cell identification is the basis for locating the specific area of user activity, and the service occurrence time is used to analyze the user's usage pattern and the active period of the service cell.

[0050] In this embodiment, the single residence duration and service times of each user in each service cell can be counted on a daily basis to form the first service cell portrait T1. The residence duration calculation criterion is as follows: according to the service time sorting, the first data of the user accessing the service cell is retained for the continuous multiple records of the same user in the same service cell, and the residence duration of the user in the service cell is obtained by subtracting the access time of the user switching to the next different service cell. The service times calculation criterion is as follows: the number of DPI records during the service cell is recorded as the service times of the user during the period.

[0051] Exemplarily, the first service cell portrait is shown in the following table:

[0052] User identification Service cell coding Access time Residence duration (s) Service times 460110 10098437_1024 20250620073015 1253 374 460113 10098437_1025 20250620124530 18 2 460112 10098437_1027 20250620183045 5422 566 460118 10098437_1028 20250620213000 72 12 460119 10098437_1029 20250620081520 315 13 460119 10098439_1027 20250620174510 423 21 460117 10098439_1028 20250620110555 677 148 460116 10098439_1029 20250620220030 363 76 460111 10098440_1024 20250620093000 271 28 460119 10098440_1025 20250620140015 901 153 460117 10098440_1027 20250620205020 484 37 460115 10098440_1028 20250620161040 1248 94 ... ... ... ... ...

[0053] Further, based on the first service cell image T1, the residence time and the service number fields are identified by the IQR method to identify outliers, and the outliers are removed. Then, according to the midnight 0:00-6:00, daytime 9:00-17:00, and night 19:00-0:00 time periods, each service cell is counted by taking "day" as the account period. Through aggregation calculation, the following labels are constructed for each service cell to form the second service cell image T2. Among them, the daytime activity calculation caliber is: (09:00-17:00 user number) / total user number of the day; the night activity calculation caliber is: (19:00-00:00 user number) / total user number of the day; the midnight activity calculation caliber is: (0:00-6:00 active user number) / total user number of the day; the daytime per capita residence time calculation caliber is: take the 50% percentile value (median, reduce the difference of the long tail distribution of residence time) of 09:00-17:00 user residence time; the night per capita residence time calculation caliber is: take the 50% percentile value of 19:00-00:00 user residence time; the midnight per capita residence time calculation caliber is: take the 50% percentile value of 0:00-6:00 user residence time; the daytime per capita connection number calculation caliber is: (9:00-17:00 service number) / (09:00-17:00 user number); the night per capita connection number calculation caliber is: (19:00-00:00 service number) / (19:00-00:00 user number); the midnight per capita connection number calculation caliber is: (0:00-6:00 service number) / (0:00-6:00 user number); the traffic fluctuation rate calculation caliber is: (max(hourly traffic)-min(hourly traffic)) / avg(hourly traffic); the daily connection number calculation caliber is: sum(service number); the daily user number calculation caliber is: the number of users whose residence time is more than 15 minutes.

[0054] Exemplarily, the second service cell image is shown in the following table:

[0055]

[0056]

[0057] Further, in the second service cell image, the connection change rate and the number change rate of the user in the first target period are determined by the following method: weekend connection change rate: (weekend average connection number-workday average connection number) / workday average connection number; weekend personnel change rate: (weekend connection user number-workday connection user number) / workday connection user number. Finally, the latitude and longitude information of the service cell is obtained by associating the work parameter data, wherein the work parameter data refers to the base station engineering parameter in the mobile communication network, including the physical location of the base station, the antenna direction angle, the downtilt angle, the transmission power and other information.

[0058] The connection change rate, the number change rate, and the location information (latitude and longitude information) of the serving cell are added to the second serving cell image, and thus the characteristic data of the serving cell can be obtained.

[0059] Exemplarily, the characteristic data of the serving cell is shown in the following table.

[0060]

[0061]

[0062] According to some optional embodiments of the present application, based on the characteristic data, the serving cell is clustered to obtain a plurality of first clustering results, which can be achieved by the following method: in the serving cell, a first clustering cluster satisfying a first condition is determined, wherein the first condition includes that the average stay time per person in a first preset time period, a second preset time period, and a third preset time period is greater than a first preset threshold, and the traffic fluctuation rate is less than a second preset threshold, the latest time in the first preset time period is earlier than the earliest time in the second preset time period, and the latest time in the second preset time period is earlier than the earliest time in the third preset time period; in the serving cell, a second clustering cluster satisfying a second condition is determined, wherein the second condition includes that the average stay time per person is less than a third preset threshold, and the number change rate in a first target time period is greater than a fourth preset threshold and the traffic fluctuation rate in the first target time period is greater than a fifth preset threshold; in the serving cell, a third clustering cluster satisfying a third condition is determined, wherein the third condition includes that the user activity in the first preset time period is less than a sixth preset threshold, the traffic volume in the second preset time period is less than a seventh preset threshold, and the number change rate in the first target time period is greater than an eighth preset threshold, wherein the eighth preset threshold is greater than the fourth preset threshold; in the serving cell, a fourth clustering cluster satisfying a fourth condition is determined, wherein the fourth condition includes that the user activity in the first preset time period is greater than a ninth preset threshold, the number change rate in the first target time period is less than a tenth preset threshold, and the traffic fluctuation rate in a second target time period is greater than an eleventh preset threshold, wherein the second target time period includes weekdays; and the first clustering cluster, the second clustering cluster, the third clustering cluster, and the fourth clustering cluster are determined as the plurality of first clustering results.

[0063] It can be understood that the first preset time period is used to represent a time period in the daytime, the second preset time period is used to represent a time period in the night other than the early morning, and the third preset time period is used to represent a time period in the early morning of the night.

[0064] The first clustering cluster, the second clustering cluster, the third clustering cluster, and the fourth clustering cluster correspond to clustering scenario 1, clustering scenario 2, clustering scenario 3, and clustering scenario 4, respectively. Figure 2is a schematic diagram of a first clustering scenario, a second clustering scenario, a third clustering scenario, and a fourth clustering scenario respectively corresponding to a first clustering cluster, a second clustering cluster, a third clustering cluster, and a fourth clustering cluster according to an embodiment of the present application, wherein Figure 2 It can be known that, for the clustering scenario 1, the average stay time of people in the day, night, and late night is relatively high, and the flow fluctuation rate is relatively low, which is more in line with some areas where the number of people changes little, and the average stay time of people is relatively long for 24 hours, such as school, hospital, industrial park, and the like. For the clustering scenario 2, the average stay time of people is relatively low, and the personnel change rate and the flow fluctuation rate are relatively high on the weekend, and the business is relatively active in the day and at night, which is more in line with some areas with large floating population and active users at night, such as transportation hubs, business districts, night markets, and the like. For the clustering scenario 3, the daytime activity is low, the nighttime business volume is the least, and the personnel change rate is the highest on the weekend, which is more in line with the characteristics of residential areas. For the clustering scenario 4, the daytime is the most active, the personnel change rate is the lowest on the weekend, and the flow fluctuation rate is high on weekdays, which is more in line with the characteristics of some office and business areas.

[0065] In this embodiment, the K-Means clustering algorithm is used for in-depth clustering analysis of the service cells, and the service cells are divided into clusters with similar attributes based on the behavior characteristics of the service cells. The K value, i.e., the number of clusters, can be determined by comprehensively applying the elbow method and the silhouette coefficient. Figure 3 is a schematic diagram for confirming the K value by using the elbow method, wherein the elbow method selects a point where the SSE (Sum of Squared Error) trend significantly flattens as the K value by observing the SSE trend under different K values. Figure 4 is a schematic diagram for confirming the K value by using the silhouette coefficient method, which measures the closeness of data points in their own clusters and the separation degree from other clusters to evaluate the clustering effect, and can find the K value when the silhouette coefficient is the largest.

[0066] The feature of the clustering scenario 1 is that the average stay time of people in the day, night, and late night is relatively high, and the flow fluctuation rate is relatively low, which indicates that the flow of the area is relatively stable and the stay time is relatively long, and it is some areas with single function and small personnel mobility, such as schools, hospitals, industrial parks, and the like. Such areas have stable personnel stay for most of the day, and there is no significant flow change due to time period, which is suitable for developing services and activities that require long stay or continuous operation.

[0067] Cluster scenario 2 presents a lower average stay time, a higher weekend personnel change rate and a higher flow fluctuation rate, and is active during the day and at night. This mode is consistent with the transportation hub, business district, night market and other areas with large mobile population and frequent night activities. In these areas, people move quickly or stay for a short time, but during certain time periods such as weekends and nights, there will be a flow peak due to shopping, dining, entertainment and other activities, which is a concentrated embodiment of urban vitality and business opportunities.

[0068] Cluster scenario 3 is characterized by low daytime activity, minimal night traffic, and the highest weekend personnel change rate, which is highly consistent with residential areas. Residential areas may be quiet during the day on weekdays as most residents go out to work or school, but the flow and activity will increase significantly during weekends and nights as residents return and rest.

[0069] Cluster scenario 4 shows the highest daytime activity, the lowest weekend personnel change rate, and a higher flow fluctuation rate on weekdays. This is consistent with the characteristics of office business districts, which are crowded during the day on weekdays due to a large number of office and business activities, but are relatively quiet during weekends and nights. In addition, the high flow fluctuation rate on weekdays reflects significant flow changes during morning and evening rush hours, which is of great reference value for traffic planning and business site selection.

[0070] Preferably, taking the first cluster that meets the first condition as an example, the step can be implemented by the following method: obtaining the average stay time data and flow fluctuation rate data of each service cell in the first, second, and third preset time periods; constructing a multi-dimensional feature vector containing three time period stay time dimensions and flow fluctuation rate dimensions for each service cell based on the average stay time data and flow fluctuation rate data; standardizing the multi-dimensional feature vector to generate a standardized feature vector; performing clustering analysis on the standardized feature vector using a K-means clustering algorithm to generate a plurality of cluster clusters containing cluster centers; based on the cluster centers, filtering a target cluster from the plurality of cluster clusters that simultaneously satisfies the three time period stay time dimensions being higher than the overall average and the flow fluctuation rate dimension being lower than the overall average; outputting the service cell set contained in the target cluster to obtain the first cluster.

[0071] In some optional embodiments of the present application, the physical locations of the serving cells in each first clustering result are clustered to obtain a plurality of second clustering results, which can be achieved by the following method: based on a neighborhood radius and a minimum point number, the physical locations of the serving cells in each first clustering result are clustered to obtain a plurality of initial second clustering results; a target area where the initial second clustering results are located is determined; a preset type of geographical area entity in the target area is obtained, a first number of the geographical area entities is determined, and a second number of the serving cells in the initial second clustering results is determined; a target difference value between the first number and the second number is calculated; in a case where the target difference value is not in a preset interval, the neighborhood radius and / or the minimum point number are adjusted, and the physical locations of the serving cells in each first clustering result are clustered again until the target difference value is in the preset interval, thereby obtaining a plurality of second clustering results.

[0072] It should be noted that a series of first clustering results are obtained, which represent clusters of serving cells with similar behavior characteristics. In order to further deepen the analysis, the focus is shifted to the geographical space, and the physical locations of the serving cells are subjected to secondary clustering processing to identify more specific geographical area entities, such as commercial districts, residential areas, or night economy areas, etc.

[0073] In the present embodiment, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) density clustering algorithm is used for secondary clustering processing, and the core parameters are neighborhood radius (eps) and minimum point number (MinPts). Among them, the neighborhood radius (eps) is used to define the distance within which a point is considered to be part of its neighbor point, and the minimum point number (MinPts) is used to define the minimum number of points required for a region to be identified as a "dense region". By adjusting these two parameters, the granularity and density of the clustering results can be controlled, and contiguous areas composed of multiple serving cells with similar behavior characteristics and close physical locations can be identified. When the DBSCAN algorithm is initially run, based on the preset neighborhood radius and minimum point number, the physical locations of the serving cells in each first clustering result are subjected to density clustering, thereby obtaining a plurality of initial second clustering results, i.e., the boundaries of the preliminary delineated geographical area entities.

[0074] Then focus on the target area where these initial second clustering results are located, which refers to the geographical space with specific functional attributes identified after the initial clustering. Further explore the preset type of geographical area entities within the target area, such as known business districts, residential areas, or nightlife active areas, etc., to determine the "first quantity" of these geographical area entities, i.e., the number of known and clearly defined entities. At the same time, the "second quantity" of service cells in the initial second clustering results is also counted, i.e., the number of entities automatically identified by the algorithm. By comparing the first quantity and the second quantity, calculate the "target difference" between the two, i.e., the gap between the quantities, to evaluate the accuracy and reasonableness of the clustering results.

[0075] If the target difference is not within the preset allowable interval, it means that the number of geographical area entities identified by the algorithm deviates significantly from the actual situation, which may be caused by improper selection of neighborhood radius or minimum point number. In this case, take response strategies to adjust these two parameters and re-execute the DBSCAN algorithm until the target difference falls within the preset reasonable interval. The basis for setting the preset interval is the understanding of a specific geographical area, for example, in an area known to have multiple contiguous business districts, if the number of entities identified by the algorithm is much less than expected, the neighborhood radius needs to be increased to capture a wider range of areas; on the contrary, if the number of entities is too large, the minimum point number needs to be increased to filter out too dispersed points, ensuring that the boundaries of the identified entities are more coherent and compact.

[0076] The above process needs to be iterated and optimized repeatedly until the ideal clustering effect is achieved. By adjusting the parameters and continuously refining the clustering results, the boundaries of geographical area entities can be more accurately defined, while ensuring that the service cells within the entities have high spatial and behavioral consistency.

[0077] That is, a set of parameter values can be set first, the DBSCAN algorithm is run, and then the number and size distribution of the identified areas are observed to see if they meet the business expectations. Business expectations can be based on point of interest (POI) data on the map, urban planning information, or other prior knowledge. For example, Figure 5In the density clustering of the shown clustering scenario 2, it is expected to identify areas with large people flow similar to traditional commercial districts, shopping malls, transportation hubs, night markets, etc. If the number of areas identified by DBSCAN is too many or too few, or the size distribution of the areas does not match the expectations, the values of eps and MinPts will be adjusted accordingly, and the algorithm will be run again until the clustering results can cover both known hotspots (such as Guan Yin Bridge, Times Square) and some emerging but not yet fully identified areas (such as the D7 area east of Dashiba, Hongqi Hegou, and the popular area near Min'an Avenue). These emerging areas are often caused by emerging market trends, social media effects, or urban development planning, which enrich the functional diversity of the city but also pose challenges to related urban management and business planning.

[0078] The above-mentioned secondary clustering based on DBSCAN not only overcomes the limitations of K-Means algorithm in handling irregular-shaped or differently-sized entities, but also effectively identifies emerging geographic area entities in the city that have not been fully captured by traditional methods, providing more accurate, timely, and comprehensive data support for urban planning, business analysis, and public services. The final second clustering results will be more close to the functional layout of the real-world city, providing a strong basis for subsequent refined management and decision-making, such as Figure 6 as shown.

[0079] As some optional embodiments of the present application, the second clustering result includes a set of discrete latitude and longitude points. After obtaining multiple second clustering results, the following steps can be performed: identifying the latitude and longitude extreme points in the set of discrete latitude and longitude points in the second clustering result; determining the latitude and longitude extreme points as initial boundary points, and establishing a polar coordinate system with the lowest latitude point as the reference; performing polar angle sorting processing on the remaining latitude and longitude points in the polar coordinate system except the initial boundary points to obtain an ordered point sequence; sequentially traversing each latitude and longitude point in the ordered point sequence, removing the top-of-stack point that causes concavity in each latitude and longitude point until the latitude and longitude point remains convex; after completing the traversal, obtaining a convex hull boundary point set; connecting the latitude and longitude coordinates of each point in the convex hull boundary point set in polar angle order to obtain a minimum convex polygon geographic area; and determining the vertex coordinate sequence composed of the vertex coordinates of the minimum convex polygon geographic area as the geographic boundary description of the second clustering result.

[0080] In this embodiment, after completing the secondary clustering of DBSCAN, the second clustering result is obtained, which is a set of dense clusters of multiple service cells in geographical space, and each cluster represents a geographic area with specific functional attributes. In order to more accurately describe the boundaries of these geographic areas, a convex hull algorithm is used to extract and determine the geographic boundaries.

[0081] Firstly, the "latitude and longitude extreme points" are identified from the discrete set of latitude and longitude points. The latitude and longitude extreme points include the northernmost point, the southernmost point, the easternmost point and the westernmost point in the set, i.e. the points that reach the maximum or minimum value in the latitude and longitude dimensions. The latitude and longitude extreme points provide a preliminary outline of the geographical area entity distribution range, and these points constitute the outermost boundary of the geographical area, laying the foundation for subsequent boundary refinement and description.

[0082] Then the lowest latitude point (i.e. the southernmost point of the geographical area entity) among these latitude and longitude extreme points is determined as the origin of the polar coordinate system, thereby establishing a polar coordinate system. In the polar coordinate system, the position of each point is determined by two parameters: the polar radius (the radial distance of the point from the origin) and the polar angle (the angle between the line connecting the point to the origin and the north direction). By taking the lowest latitude point as the reference of the polar coordinate system, it is ensured that the subsequent polar angle sorting and boundary point extraction process can be performed around this point.

[0083] After the polar coordinate system is established, the remaining latitude and longitude points other than the initial boundary points are subjected to polar angle sorting processing, and all non-boundary points are sorted according to their angle (i.e. polar angle) with the reference point to form an ordered point sequence. The purpose of the polar angle sorting processing is to follow a certain direction when constructing the boundary, ensuring the continuity and logical order of the boundary points, so as to more accurately depict the boundary shape of the area entity.

[0084] Each latitude and longitude point in the ordered point sequence is traversed in turn, and the following steps are performed: removing the top-of-stack point that causes concavity in each latitude and longitude point. Here, "stack" is a data structure that follows the principle of first-in last-out. In each iteration of the algorithm, it is checked whether the current point (the point in the ordered point sequence) and its preceding and following points (i.e. the adjacent points in the stack) form a triangle that maintains convexity. If the addition of the current point causes the boundary to have a concave shape, i.e. the interior angle of the triangle is greater than 180 degrees, then this point is considered as the "top-of-stack point" and is removed from the boundary point set.

[0085] After the traversal of the entire ordered point sequence and the checking and removal of the top-of-stack points are completed, the convex hull boundary point set composed of boundary points is obtained. The concept of convex hull refers to the smallest convex polygon that can contain all points in a set of points in a plane. The points in the convex hull boundary point set are the boundary points of the smallest convex polygon, which constitute the convex boundary of the geographical area entity.

[0086] The points in the convex hull boundary point set are connected in order of their polar angle in the polar coordinate system to form a minimum convex polygon geographic area. The polygon geographic area not only contains all the service cell points in the original cluster, but also accurately describes the boundary of the geographic area entity in the simplest convex shape. The sequence of the vertex coordinates of the minimum convex polygon geographic area (i.e., the longitude and latitude coordinates of the boundary points arranged in order of the polar angle) is determined as the geographic boundary description of the second clustering result, as shown in Figure 7 The geographic boundary description means obtaining a structured and ordered data set for representing and storing the boundary information of the geographic area.

[0087] In some optional embodiments of the present application, the position information of the center point can be determined by the following method: obtaining the longitude and latitude information of the center point in the network management system, and constructing an interface request address based on the longitude and latitude information; sending the interface request address to a map service providing object and receiving standardized format response data returned by the map service providing object; processing the standardized format response data using a recursive parsing algorithm to extract address components layer by layer to obtain address components at each level; and combining and splicing the address components at each level according to a preset address formatting rule in descending order of level to obtain the position information of the center point represented by a structured physical address description.

[0088] In this embodiment, the longitude and latitude information of the center point is first extracted from the network management system. The network management system includes detailed engineering parameters of the base station, including accurate longitude and latitude coordinates. Based on the obtained longitude and latitude information, a specific interface request address is constructed, which includes key parameters required by the map service providing object, such as longitude and latitude coordinates, request format, and response type, etc. The construction of the interface request address needs to follow the API specification of the map service providing object to ensure the effectiveness of the request and the standardization of the response data.

[0089] Then the constructed interface request address is sent to the map service providing object, and the standardized format response data is waited to be returned. These response data include structured address information such as country, province, city, district, street, and house number, which are organized according to certain hierarchical relationship and accurately describe the physical location of the center point.

[0090] After receiving the response data, a recursive parsing algorithm is used to process the response data. The recursive parsing algorithm starts from the outermost address component and gradually goes deeper to smaller levels until the address components at the finest granularity are extracted. This process is similar to peeling an onion, and each layer of parsing reveals more detailed information about the location of the center point. Through recursive parsing, all levels of address information can be accurately extracted and saved, providing a basis for subsequent address combination.

[0091] Finally, according to the preset address formatting rules, the address components obtained through recursive parsing are combined and spliced in descending order of level to generate a complete and structured physical address. The address formatting rules include the order of address components, the use of separators, and the standardization of address format to ensure that the final generated physical address meets the address writing specifications. Through the above steps, not only the precise position of the center point is determined, but also it is converted into a more intuitive and practical structured physical address description, greatly improving the readability and usability of the data, and providing more accurate location information support for urban planning, business decision-making and public service optimization.

[0092] As some optional embodiments of the present application, after determining the feature data of the serving cell, the following steps can also be performed: for the target feature data of the same dimension in the feature data, the mean and standard deviation are calculated, and the upper limit and lower limit of the abnormal value judgment of the target feature data are determined based on the mean and three times the standard deviation calculation; identify the abnormal data points in the target feature data that exceed the upper limit and lower limit of the abnormal value judgment, and replace and fill the abnormal data points using the upper limit and / or lower limit value corresponding to the target feature data to obtain the first feature data; calculate the Pearson correlation coefficient between the feature data of any two dimensions in the first feature data, and construct a correlation coefficient matrix based on the Pearson correlation coefficient between the feature data of any two dimensions; identify the feature pairs in the correlation coefficient matrix that are greater than or equal to a preset correlation threshold, and delete the feature data of one dimension in the feature pair to obtain the second feature data; convert the data scale in the second feature data to a preset standard range.

[0093] In this embodiment, the target feature data of the same dimension is first statistically analyzed, including calculating the mean and standard deviation of each dimension feature. Among them, the mean represents the central tendency of the feature value, and the standard deviation reflects the distribution range of the data points around the mean. Based on three times the standard deviation, the upper and lower limits of the abnormal value judgment are calculated. The setting principle of this upper and lower limit is that any feature value that exceeds the range of mean plus three times the standard deviation or mean minus three times the standard deviation will be regarded as an abnormal value. This method, i.e. the three-sigma principle, based on the normal distribution theory, can effectively identify those data points that significantly deviate from the average level.

[0094] Then, the target feature dataset is subjected to outlier identification, and all data points exceeding the upper and lower limits of the outlier judgment are marked as outliers. These outlier data points are caused by measurement errors, data entry errors or extreme events, and if not handled, they will interfere with the subsequent clustering analysis results. In order to correct the outliers, the upper limit value and / or the lower limit value are used to replace the outlier data points to ensure the continuity and reasonableness of the dataset, and a first feature dataset is obtained. The replaced data points are not the true values of the original records, but they are closer to the normal range of the feature, avoiding the influence of extreme values on model training.

[0095] After processing the outliers, feature dimension reduction is performed to reduce the redundancy of the dataset and avoid overfitting problems. Specifically, the Pearson correlation coefficient is used to measure the linear correlation between features. The Pearson correlation coefficient is a value between -1 and 1, which is used to represent the correlation between two variables, a positive value indicates a positive correlation, a negative value indicates a negative correlation, and a value close to 0 indicates a low correlation between the two variables. The Pearson correlation coefficient between any two dimensional feature data in the first feature dataset can be calculated, and a correlation coefficient matrix as shown in Figure 8 is constructed, which intuitively shows the correlation between all features.

[0096] Based on the correlation coefficient matrix, identify feature pairs with a correlation coefficient greater than or equal to a preset correlation threshold. In order to reduce the dimension of the dataset and avoid the influence of high correlation between features on the model, one dimension in the feature pair is selected for deletion, and the other is retained. The deletion decision can be based on the importance of the feature, the distribution characteristics of the dataset. After this dimension reduction process, a second feature dataset is obtained, which reduces the feature dimension while maintaining the integrity of the data information and simplifies the subsequent analysis process.

[0097] Finally, in order to ensure that all feature data is comparable in subsequent clustering analysis and machine learning models, the second feature dataset is subjected to data scale conversion, i.e. data standardization, to convert the data scale to a preset standard range, such as the common 0-1 range or -1 to 1 range. Data standardization removes the influence of the data unit of the feature and scales the value range of the feature, making the comparison between different features more fair and avoiding the disproportionate influence of a feature with a large value range on the model results. This process can be completed using the Z-Score standardization method or the min-max scaling method.

[0098] Figure 9 is a flowchart of another method for determining a center position according to an embodiment of the present application, as shown in Figure 9 , the method comprises the following steps:

[0099] 1. Obtain operator DPI data: DPI signaling data includes user behavior information in the network, such as which base station is connected, connection time, data traffic, etc., which is the basis for subsequent base station portrait construction.

[0100] 2. Base station portrait label processing: Convert raw DPI data into multi-dimensional base station portrait labels, including but not limited to: average stay time, average connection times, weekend personnel change rate, etc.

[0101] 3. Data processing: The data after base station portrait label processing needs subsequent processing, including outlier processing, data normalization, and data dimension reduction. Outlier processing aims to eliminate or correct data that deviates from the normal range to avoid affecting subsequent analysis. Data normalization standardizes features of different scales to a unified range, making it easier for clustering algorithms to process. Data dimension reduction removes redundant features, reduces analysis dimensions, and improves algorithm efficiency.

[0102] 4. Multi-layer clustering algorithm strategy: This is the core part of the technical solution, which includes the following main steps:

[0103] Elbow method analysis and silhouette coefficient analysis: used to determine the best K value in the K-Means algorithm, i.e. the number of clustering clusters. The elbow method analyzes the total squared error (SSE) change under different K values to find the best K value turning point; the silhouette coefficient evaluates the tightness of data points in their clusters and the separation between clusters, and selects the best K value to obtain the most reasonable clustering result. Combined with the analysis results of the elbow method and the silhouette coefficient, a best K value is confirmed for subsequent K-Means clustering.

[0104] K-Means algorithm: After determining the best K value, apply the K-Means algorithm to the full base station portrait label data to identify clusters of base stations with similar attributes, each cluster representing a preliminary classified functional area, such as commercial district, residential area, night economy area, etc.

[0105] Density clustering algorithm (e.g. DBSCAN): secondary clustering of each cluster produced by the K-Means algorithm aims to identify geographically continuous and behaviorally consistent geographic area entities, further refining the division of functional areas.

[0106] K=1 K-Means algorithm: within the geographic area entities identified by DBSCAN, use the K=1 K-Means clustering algorithm to locate the center point of each area. This method is equivalent to calculating the geometric center of all service cell location points in the area, resulting in a representative point coordinate.

[0107] 5. Multi-level clustering results: output the region classification with clear geographical boundaries and functional attributes obtained after multi-level clustering processing, and the center point coordinates of each region.

[0108] The above steps adopt a multi-level linkage clustering strategy of K-Means, DBSCAN and K-Means (K = 1), forming a progressive analysis process from behavior analysis to geographical space exploration to regional center positioning. First, the K-Means clustering algorithm identifies service cells with similar behavior characteristics based on the multi-dimensional labels of the base station image, and preliminarily divides clusters such as business circles, residential areas or night economy zones. Subsequently, the DBSCAN density clustering algorithm further focuses on service cells with similar geographical locations and consistent behavior patterns, and mines specific geographical space region entities such as specific commercial street blocks, residential communities or night economy active zones. Finally, the K = 1 K-Means clustering algorithm is applied within each geographical space region to locate the "center point" of the region, i.e. the most representative geographical location. This multi-level clustering analysis strategy from coarse to fine not only improves the accuracy and efficiency of identification, but also captures more complex and subtle regional functional features from the dimensions of behavior and geography.

[0109] Figure 10 Figure 1 is a structural diagram of a center position determination device according to an embodiment of the present application, as shown in the figure, the device comprises: Figure 10

[0110] The first determination module 1002 is configured to determine feature data of the service cell based on the deep packet inspection signaling data of the operator, wherein the deep packet inspection signaling data at least includes: user identifier, service cell identifier, service occurrence time, and the feature data at least includes: the stay duration of different users in different service cells.

[0111] The generation module 1004 is configured to perform clustering processing on the service cell based on the feature data to obtain a plurality of first clustering results.

[0112] The clustering module 1006 is configured to perform clustering processing on the physical location of the service cell in each first clustering result to obtain a plurality of second clustering results.

[0113] The second determination module 1008 is configured to determine the center point of the second clustering result according to the mean value of the second clustering result, and determine the position information of the center point.

[0114] ​Optionally, the feature data of the service cell is determined based on the deep packet inspection signaling data of the operator, specifically including the following steps: determining the residence duration of different users in different service cells based on the deep packet inspection signaling data, and constructing a first service cell portrait according to the residence duration of different users in different service cells; determining the operation performance indicators of the users in different preset time periods in the first service cell portrait, wherein the operation performance indicators at least include: user activity, average residence duration, crowd fluctuation rate, and traffic volume; constructing a second service cell portrait according to the operation performance indicators; determining the connection change rate and the number change rate of the users in the first target period in the second service cell portrait, wherein the first target period includes: holidays; adding the connection change rate, the number change rate, and the location information of the service cell to the second service cell portrait to obtain the feature data of the service cell.

[0115] Optionally, the service cells are clustered based on the feature data to obtain a plurality of first clustering results, specifically including the following steps: determining a first clustering cluster in the service cells that meets a first condition, wherein the first condition includes: the average residence duration in the first preset time period, the second preset time period, and the third preset time period is greater than a first preset threshold, and the crowd fluctuation rate is less than a second preset threshold, the latest time in the first preset time period is earlier than the earliest time in the second preset time period, and the latest time in the second preset time period is earlier than the earliest time in the third preset time period; determining a second clustering cluster in the service cells that meets a second condition, wherein the second condition includes: the average residence duration is less than a third preset threshold, and the number change rate in the first target period is greater than a fourth preset threshold and the crowd fluctuation rate in the first target period is greater than a fifth preset threshold; determining a third clustering cluster in the service cells that meets a third condition, wherein the third condition includes: the user activity in the first preset time period is less than a sixth preset threshold, the traffic volume in the second preset time period is less than a seventh preset threshold, and the number change rate in the first target period is greater than an eighth preset threshold, wherein the eighth preset threshold is greater than the fourth preset threshold; determining a fourth clustering cluster in the service cells that meets a fourth condition, wherein the fourth condition includes: the user activity in the first preset time period is greater than a ninth preset threshold, the number change rate in the first target period is less than a tenth preset threshold, and the crowd fluctuation rate in the second target period is greater than an eleventh preset threshold, wherein the second target period includes: weekdays; and determining the first clustering cluster, the second clustering cluster, the third clustering cluster, and the fourth clustering cluster as the plurality of first clustering results.

[0116] Optionally, the physical positions of the serving cells in each first clustering result are clustered to obtain a plurality of second clustering results, specifically including the following steps: based on the neighborhood radius and the minimum point number, the physical positions of the serving cells in each first clustering result are clustered to obtain a plurality of initial second clustering results; a target area where the initial second clustering result is located is determined; a geographic area entity of a preset type in the target area is obtained, a first number of the geographic area entity is determined, and a second number of the serving cells in the initial second clustering result is determined; a target difference value between the first number and the second number is calculated; in the case that the target difference value is not in a preset interval, the neighborhood radius and / or the minimum point number are adjusted, and the physical positions of the serving cells in each first clustering result are clustered again until the target difference value is in the preset interval, thereby obtaining a plurality of second clustering results.

[0117] Optionally, the second clustering result includes a discrete latitude and longitude point set; after obtaining the plurality of second clustering results, the following steps can be further performed: identifying latitude and longitude extreme points in the discrete latitude and longitude point set in the second clustering result; determining the latitude and longitude extreme points as initial boundary points, and establishing a polar coordinate system based on the lowest latitude point; performing polar angle sorting processing on the remaining latitude and longitude points other than the initial boundary points in the polar coordinate system to obtain an ordered point sequence; sequentially traversing each latitude and longitude point in the ordered point sequence, and removing the top point that causes concavity in each latitude and longitude point until the latitude and longitude point remains convex; after completing the traversal, a convex hull boundary point set is obtained; connecting the latitude and longitude coordinates of the points in the convex hull boundary point set in turn according to the polar angle order to obtain a minimum convex polygon geographic area; a vertex coordinate sequence composed of the vertex coordinates of the minimum convex polygon geographic area is determined as a geographic boundary description of the second clustering result.

[0118] Optionally, the position information of the center point is determined, specifically including the following steps: obtaining the latitude and longitude information of the center point in the network management system, and based on the latitude and longitude information, constructing an interface request address; sending the interface request address to a map service providing object, and receiving standardized format response data returned by the map service providing object; processing the standardized format response data using a recursive parsing algorithm to extract address components layer by layer to obtain address components at different levels; according to a preset address formatting rule, combining and splicing the address components at different levels in descending order of level to obtain the position information of the center point described by a structured physical address.

[0119] Optionally, after the characteristic data of the serving cell is determined, the following steps can be further performed: for target characteristic data of the same dimension in the characteristic data, calculating a mean value and a standard deviation, determining an upper limit of abnormal value judgment and a lower limit of abnormal value judgment based on the mean value and three times the standard deviation, identifying abnormal data points in the target characteristic data that exceed the upper limit of abnormal value judgment and the lower limit of abnormal value judgment, and replacing and filling the abnormal data points using the upper limit value and / or the lower limit value corresponding to the target characteristic data to obtain first characteristic data; calculating a Pearson correlation coefficient between characteristic data of any two dimensions in the first characteristic data, and constructing a correlation coefficient matrix based on the Pearson correlation coefficient between the characteristic data of any two dimensions; identifying a feature pair in the correlation coefficient matrix that is greater than or equal to a preset correlation threshold, and deleting the characteristic data of one dimension in the feature pair to obtain second characteristic data; and converting the data scale in the second characteristic data to a preset standard range.

[0120] It should be noted that the above Figure 10 Each module in the above

[0121] It should be noted that the preferred embodiments of the above Figure 10 embodiments can be referred to the related description of the above Figure 1 embodiments, which will not be repeated here.

[0122] Figure 11 A hardware structure block diagram of a computer terminal for implementing the method of determining the central position is shown. As Figure 11 shown, the computer terminal 110 can include one or more (shown in the figure as 1102a, 1102b, …, 1102n) processors 1102 (the processor 1102 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1104 for storing data, and a transmission module 1106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 11 The structure shown is only schematic, and does not limit the structure of the above electronic device. For example, the computer terminal 110 can include more or fewer components than those shown in Figure 11 or have a different configuration than that shown in Figure 11 .

[0123] It should be noted that the one or more processors 1102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any of the other elements of the computer terminal 110. As referred to in embodiments of the present application, the data processing circuitry functions as a processor to control, for example, selection of the variable resistance terminal path connected to the interface.

[0124] The memory 1104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the center position determination method in embodiments of the present application. The processor 1102 can execute various functional applications and data processing, i.e., implement the center position determination method described above, by running the software programs and modules stored in the memory 1104. The memory 1104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 1104 can further include a memory disposed remotely with respect to the processor 1102, which can be connected to the computer terminal 110 through a network. Examples of the network can include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0125] The transmission module 1106 is configured to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 110. In one example, the transmission module 1106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to be able to communicate with the Internet. In one example, the transmission module 1106 can be a radio frequency (RF) module, which is configured to communicate with the Internet in a wireless manner.

[0126] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computer terminal 110.

[0127] It should be noted that in some optional embodiments, the above-mentioned Figure 11 The computer terminal shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or combinations of both hardware and software elements. It should be noted that in some embodiments, the functions of the computer terminal described above can be combined in a single module or implemented in a distributed manner over multiple modules. Figure 11Only one of the specific examples, and is intended to show the type of components that can be present in the above computer terminal.

[0128] It should be noted that, Figure 11 The computer terminal shown is used to execute Figure 1 The center position determination method shown, so the relevant explanation of the execution method of the above command also applies to the electronic device, hereinafter will not be described.

[0129] The embodiments of the present application also provide a non-volatile storage medium, the non-volatile storage medium comprises a stored program, wherein the program runs to control the device where the storage medium executes the above center position determination method.

[0130] The non-volatile storage medium executes the program of the following functions: determining feature data of a service cell based on operator deep packet inspection signaling data, wherein the deep packet inspection signaling data at least includes: user identification, service cell identification, service occurrence time, and the feature data at least includes: the stay duration of different users in different service cells; performing clustering processing on the service cells based on the feature data to obtain a plurality of first clustering results; performing clustering processing on the physical positions of the service cells in each first clustering result to obtain a plurality of second clustering results; determining the center point of the second clustering result according to the mean value of the second clustering result, and determining the position information of the center point.

[0131] The embodiments of the present application also provide an electronic device, comprising: a memory and a processor, the processor is used to run the program stored in the memory, wherein the program runs to execute the above center position determination method.

[0132] The processor is used to run the program of the following functions: determining feature data of a service cell based on operator deep packet inspection signaling data, wherein the deep packet inspection signaling data at least includes: user identification, service cell identification, service occurrence time, and the feature data at least includes: the stay duration of different users in different service cells; performing clustering processing on the service cells based on the feature data to obtain a plurality of first clustering results; performing clustering processing on the physical positions of the service cells in each first clustering result to obtain a plurality of second clustering results; determining the center point of the second clustering result according to the mean value of the second clustering result, and determining the position information of the center point.

[0133] The above sequence number of the embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments.

[0134] In the above embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0135] In the above embodiments of the present application, the collected information is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with relevant laws, regulations and standards, necessary protection measures are taken, the public order and good customs are not violated, and corresponding operation portals are provided for the user to select authorization or refusal.

[0136] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.

[0137] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place or distributed to multiple units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0138] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0139] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part that essentially contributes to the related art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk and various program code storage media.

[0140] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A method for determining the center position, characterized in that, include: Based on the deep packet inspection signaling data of the operator, the characteristic data of the serving cell is determined. The deep packet inspection signaling data includes at least: user identifier, serving cell identifier, and service occurrence time. The characteristic data includes at least: the dwell time of different users in different serving cells. Based on the feature data, the serving cells are clustered to obtain multiple first clustering results; The physical locations of the serving cells in each of the first clustering results are clustered to obtain multiple second clustering results; The center point of the second clustering result is determined based on the mean of the second clustering result, and the location information of the center point is determined.

2. The method according to claim 1, characterized in that, Based on the operator's deep packet inspection signaling data, characteristic data of the serving cell are determined, including: Based on the deep packet inspection signaling data, the dwell time of different users in different serving cells is determined, and a first serving cell profile is constructed based on the dwell time of different users in different serving cells. In the first service cell profile, the operational efficiency indicators of users in different preset time periods are determined, wherein the operational efficiency indicators include at least: user activity, average dwell time per person, traffic fluctuation rate, and business volume; Based on the aforementioned operational efficiency indicators, construct a profile of the second service community; In the second service cell profile, the connection change rate and the number of users change rate during the first target time period are determined, wherein the first target time period includes: holidays; The connection change rate, the number of users change rate, and the location information of the serving cell are added to the second serving cell profile to obtain the feature data of the serving cell.

3. The method according to claim 2, characterized in that, Based on the aforementioned feature data, clustering is performed on the serving cells to obtain multiple first clustering results, including: In the service cell, a first cluster that meets the first condition is determined, wherein the first condition includes: the average stay time per person in the first preset time period, the second preset time period, and the third preset time period is greater than the first preset threshold, and the fluctuation rate of the flow of people is less than the second preset threshold, the latest time in the first preset time period is earlier than the earliest time in the second preset time period, and the latest time in the second preset time period is earlier than the earliest time in the third preset time period. In the service cell, a second cluster that meets the second condition is determined, wherein the second condition includes: the average stay time per person is less than a third preset threshold, and the change rate of the number of people in the first target time period is greater than a fourth preset threshold and the fluctuation rate of the flow of people in the first target time period is greater than a fifth preset threshold. In the service cell, a third cluster that meets the third condition is determined, wherein the third condition includes: the user activity level in the first preset time period is less than the sixth preset threshold, the service volume in the second preset time period is less than the seventh preset threshold, and the number of users change rate in the first target time period is greater than the eighth preset threshold, wherein the eighth preset threshold is greater than the fourth preset threshold. In the service cell, a fourth cluster that meets the fourth condition is determined, wherein the fourth condition includes: the user activity level is greater than the ninth preset threshold during the first preset time period, the number of people change rate is less than the tenth preset threshold during the first target time period, and the traffic fluctuation rate is greater than the eleventh preset threshold during the second target time period, wherein the second target time period includes: weekdays; The first cluster, the second cluster, the third cluster, and the fourth cluster are identified as the plurality of first clustering results.

4. The method according to claim 1, characterized in that, Clustering is performed on the physical locations of the serving cells in each of the first clustering results to obtain multiple second clustering results, including: Based on the neighborhood radius and the minimum number of points, the physical location of the serving cell in each of the first clustering results is clustered to obtain multiple initial second clustering results; Determine the target region where the initial second clustering result is located; Obtain geographic region entities of a preset type in the target region, determine the first number of the geographic region entities, and determine the second number of serving cells in the initial second clustering result; Calculate the target difference between the first quantity and the second quantity; If the target difference is not within the preset range, adjust the neighborhood radius and / or the minimum number of points, and re-cluster the physical locations of the serving cells in each of the first clustering results until the target difference is within the preset range, thus obtaining the plurality of second clustering results.

5. The method according to claim 1, characterized in that, The second clustering result includes: a discrete set of latitude and longitude points; after obtaining multiple second clustering results, the method further includes: Identify the latitude and longitude extreme points in the discrete latitude and longitude point set in the second clustering result; The extreme latitude and longitude points are determined as the initial boundary points, and a polar coordinate system is established with the lowest latitude point as the reference. In the polar coordinate system, the remaining latitude and longitude points other than the initial boundary point are sorted by polar angle to obtain an ordered point sequence. Iterate through each latitude and longitude point in the ordered point sequence, removing stack vertices that cause concavity at each latitude and longitude point until the latitude and longitude point remains convex; after completing the traversal, obtain the set of convex hull boundary points; Connect the latitude and longitude coordinates of each point in the convex hull boundary point set in order of polar angle to obtain the smallest convex polygon geographical region. The vertex coordinate sequence formed by the vertex coordinates of the minimum convex polygon geographic region is determined as the geographic boundary description of the second clustering result.

6. The method according to claim 1, characterized in that, Determining the location information of the center point includes: Obtain the latitude and longitude information of the center point in the network management system, and construct the interface request address based on the latitude and longitude information; Send the interface request address to the map service provider object and receive the standardized format response data returned by the map service provider object; The standardized format response data is processed using a recursive parsing algorithm, and address components are extracted layer by layer to obtain address components at each level. According to the preset address formatting rules, the address components at each level are combined and spliced ​​in descending order to obtain the location information of the center point represented by the structured physical address description.

7. The method according to claim 1, characterized in that, After determining the characteristic data of the serving cell, the method further includes: For target feature data of the same dimension in the feature data, calculate the mean and standard deviation, and calculate and determine the upper limit and lower limit of outlier judgment for the target feature data based on the mean and three times the standard deviation; Identify abnormal data points in the target feature data that exceed the upper limit and lower limit of the outlier judgment, and replace and fill the abnormal data points with the upper limit and / or lower limit corresponding to the target feature data to obtain the first feature data; Calculate the Pearson correlation coefficient between any two dimensions of feature data in the first feature data, and construct a correlation coefficient matrix based on the Pearson correlation coefficient between any two dimensions of feature data; Identify feature pairs in the correlation coefficient matrix that are greater than or equal to a preset correlation threshold, and delete feature data of one dimension in the feature pair to obtain second feature data; The data scale in the second feature data is converted to a preset standard range.

8. A device for determining the center position, characterized in that, include: The first determining module is used to determine the characteristic data of the serving cell based on the deep packet inspection signaling data of the operator, wherein the deep packet inspection signaling data includes at least: user identifier, serving cell identifier, and service occurrence time, and the characteristic data includes at least: the dwell time of different users in different serving cells; The generation module is used to perform clustering processing on the serving cells based on the feature data to obtain multiple first clustering results; The clustering module is used to cluster the physical locations of the serving cells in each of the first clustering results to obtain multiple second clustering results; The second determining module is used to determine the center point of the second clustering result based on the mean of the second clustering result, and to determine the location information of the center point.

9. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the non-volatile storage medium to perform the method for determining the center location as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the method for determining the center position as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for determining the center position as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Power transmission line lightning stroke fault identification method and lightning detection device

    CN121476841A