Method and apparatus for identifying a business cluster

By utilizing decision tree models and data supplementation techniques, combined with scenario, engineering parameters, performance, and MDT data, commercial cluster areas are identified, solving the problem of low accuracy caused by reliance on human labor in existing technologies, and achieving higher identification accuracy and objectivity.

CN116975660BActive Publication Date: 2026-06-26CHINA MOBILE GRP GUANGDONG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE GRP GUANGDONG CO LTD
Filing Date
2022-04-20
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

The identification of commercial clusters in existing technologies relies on human resources, resulting in low accuracy and a large influence of subjective factors. In particular, the identification results are inaccurate when the optimization personnel are unfamiliar with the area.

Method used

By acquiring scene data, engineering parameter data, performance data, and MDT data of the area to be identified, prediction is made using a target decision tree model, missing scene data is supplemented, scoring is calculated and density clustering is performed, and commercial value is evaluated by combining voice data and traffic data to identify commercial cluster areas.

Benefits of technology

This significantly reduces subjective factors in the process of identifying commercial clusters, improves the accuracy and credibility of the identification results, and ensures more accurate identification of commercial clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975660B_ABST
    Figure CN116975660B_ABST
Patent Text Reader

Abstract

The application provides a kind of identification method and device for commercial aggregation area, comprising obtaining the scene data, performance data, MDT data and work parameter data of the area to be identified after preprocessing, inputting the scene data, performance data, MDT data and work parameter data into target decision tree model to obtain predicted scene data, predicted performance data, predicted MDT data and predicted work parameter data with data label as commercial nature;Based on predicted work parameter data, the predicted scene data is supplemented with scene missing data to obtain target scene data;The comprehensive total score data is obtained by scoring calculation on predicted performance data and predicted MDT data;The density clustering grouping result is obtained by density clustering grouping on target scene data and comprehensive total score data, and the commercial aggregation area is identified from the area to be identified according to the density clustering grouping result, so that the identification result is more convincing, and the identified commercial aggregation area is more accurate through comprehensive scoring combined with scene data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to an identification method and apparatus for commercial cluster areas. Background Technology

[0002] Currently, the construction of 5G networks is progressing steadily. However, due to the high initial costs of 5G network construction, priority is being given to applying 5G technology in commercially concentrated areas to make rational use of resources and maximize commercial value. However, the identification of commercially concentrated areas still requires a significant amount of manpower, relying heavily on the experience and subjective perception of the personnel. Therefore, the identification of concentrated areas is greatly influenced by the subjective opinions of the personnel. Furthermore, when the personnel are unfamiliar with certain areas, the accuracy of the identification of concentrated areas will decrease significantly. Summary of the Invention

[0003] This invention provides a method and apparatus for identifying commercial clusters, which addresses the shortcomings of existing technologies that still require a large number of human resources to identify commercial clusters, resulting in low accuracy when personnel are unfamiliar with certain areas. This invention improves the accuracy of commercial cluster identification.

[0004] This invention provides a method for identifying commercial clusters, including:

[0005] The scene data, engineering parameter data, performance data, and MDT data of the region to be identified are obtained after preprocessing. The scene data, engineering parameter data, performance data, and MDT data are then sequentially input into the target decision tree model to obtain prediction scene data, prediction engineering parameter data, prediction performance data, and prediction MDT data, all of which are commercially labeled.

[0006] Based on the predicted working parameter data, the missing scene data of the predicted scene data is supplemented to obtain the target scene data;

[0007] The prediction performance data and the prediction MDT data are scored and calculated to obtain the comprehensive total score data;

[0008] Density clustering is performed on the target scene data and the comprehensive total score data to obtain density clustering results. Based on the density clustering results, commercial cluster areas are identified from the areas to be identified.

[0009] According to a method for identifying commercial cluster areas provided by the present invention, after identifying the commercial cluster area from the area to be identified based on the density clustering grouping result, the method further includes:

[0010] Acquire voice data, traffic data, and SMS data from the aforementioned commercial cluster area;

[0011] The commercial value of the commercial cluster area is scored based on the voice data, traffic data, and SMS data to obtain the quantitative value of the commercial cluster area.

[0012] According to the present invention, a method for identifying commercial cluster areas, wherein the step of supplementing the predicted scene data with missing scene data based on the predicted working parameter data to obtain target scene data specifically includes:

[0013] Generate first Geometry data corresponding to the predicted engineering parameter data, generate second Geometry data corresponding to the predicted scene data, and supplement the first Geometry data with missing scene data based on the second Geometry data to obtain target Geometry data;

[0014] The target Geometry data is subjected to density clustering to obtain several cluster data groups. Under the condition of a preset scenario generation threshold setting, the several cluster data groups are aggregated to obtain the aggregated target Geometry data.

[0015] The predicted working parameter data is rasterized to obtain first raster data, and the predicted scene data is rasterized to obtain second raster data. The second raster data is supplemented based on the first raster data to obtain target raster data.

[0016] Based on the target raster data and the target Geometry data, the target scene data is obtained.

[0017] According to the identification method for commercial cluster areas provided by the present invention, the step of scoring and calculating the predicted performance data and the predicted MDT data to obtain a comprehensive total score specifically includes:

[0018] Obtain the first linear scoring factor of the predicted performance data and the second linear scoring factor of the predicted MDT data;

[0019] The comprehensive score of the prediction performance data is calculated based on the first linear scoring factor, and the comprehensive score of the prediction MDT data is calculated based on the second linear scoring factor.

[0020] Obtain the first weight of the predicted performance data and the second weight of the predicted MDT data, and perform a weighted calculation on the comprehensive score of the predicted performance data and the comprehensive score of the predicted MDT data based on the first weight and the second weight to obtain the comprehensive total score data.

[0021] According to the identification method for commercial cluster areas provided by the present invention, the step of calculating the comprehensive score of the predicted performance data based on the first linear scoring factor and calculating the comprehensive score of the predicted MDT data based on the second linear scoring factor specifically includes:

[0022] The index scores of each performance index in the predicted performance data are calculated, and the comprehensive score of the predicted performance data is calculated based on the first linear scoring factor and the index scores of each index. The performance index data of the predicted performance data includes average daily 4G traffic, maximum number of RRC connections during busy hours, and VoLTE voice traffic.

[0023] The index scores of each MDT index data in the predicted MDT data are calculated, and the comprehensive score of the predicted MDT data is calculated based on the second linear scoring factor and the index scores of each MDT index data. The MDT index data of the predicted MDT data includes the total number of cells and the total number of samples.

[0024] According to the present invention, a method for identifying commercial cluster areas, before sequentially inputting the scene data, the engineering parameter data, the performance data, and the MDT data into the target decision tree model, further includes:

[0025] Decision tree nodes are constructed based on the data attributes of the training data to generate an initial decision tree model. The training data includes training scenario data, training parameter data, training performance data, and training MDT data.

[0026] The test data is input into the initial decision tree model to obtain the predicted labels of the test data. The test data includes test scenario data, test parameter data, test performance data and test MDT data with artificial labels.

[0027] The initial decision tree model is iteratively trained based on the predicted labels and the artificial labels until the target decision tree model is obtained.

[0028] According to the identification method for commercial cluster areas provided by the present invention, the construction of decision tree nodes based on data attributes of training data specifically includes:

[0029] Determine the data attributes of the training data, calculate the weighted average Gini value of each data attribute, and select the first data attribute with the smallest weighted average Gini value as the root node of the decision tree;

[0030] Select the second data attribute with the smallest weighted average Gini value besides the first data attribute, and select the second data attribute as an internal node located at the next level below the root node;

[0031] Using the second data attribute as the first data attribute, the step of selecting the second data attribute with the smallest weighted average Gini value (other than the first data attribute) is repeated until all data attributes are selected.

[0032] According to the present invention, a method for identifying commercial cluster areas includes performing density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and identifying commercial cluster areas from the area to be identified based on the density clustering grouping results. Specifically, this includes:

[0033] Density clustering is performed on the target scene data and the comprehensive total score data to obtain the first density clustering result;

[0034] Based on the first density clustering grouping results, target data is selected from the target scene data and the comprehensive total score data;

[0035] The target data is subjected to density clustering to obtain a second density clustering result. Based on the second density clustering result, commercial cluster areas are identified from the area to be identified.

[0036] According to the present invention, a method for identifying commercial cluster areas, before acquiring the preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified, further includes:

[0037] The initial scene data of the region to be identified is obtained, and the initial scene data is standardized and filtered sequentially to obtain preprocessed scene data;

[0038] Obtain the initial engineering parameter data of the region to be identified, and standardize the initial engineering parameter data to obtain the preprocessed engineering parameter data;

[0039] The initial performance data of the region to be identified is obtained, and the initial performance data is sequentially rasterized, denoised, and filtered to obtain the preprocessed performance data.

[0040] The initial MDT data of the region to be identified is obtained, and the initial MDT data is sequentially rasterized, denoised, and filtered to obtain the preprocessed MDT data.

[0041] The present invention also provides an identification device for commercial cluster areas, comprising:

[0042] The acquisition unit is used to acquire scene data, engineering parameter data, performance data and MDT data of the area to be identified after preprocessing, and input the scene data, engineering parameter data, performance data and MDT data into the target decision tree model in sequence to obtain prediction scene data, prediction engineering parameter data, prediction performance data and prediction MDT data with commercial labels.

[0043] The supplementary unit is used to supplement the missing scene data of the predicted scene data based on the predicted working parameter data to obtain the target scene data;

[0044] The scoring unit is used to score and calculate the prediction performance data and the prediction MDT data to obtain the comprehensive total score data.

[0045] The identification unit is used to perform density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and to identify commercial clustering areas from the area to be identified based on the density clustering grouping results.

[0046] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the identification method for commercial cluster areas as described above.

[0047] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the identification method for commercial cluster areas as described above.

[0048] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the identification method for commercial cluster areas as described above.

[0049] The present invention provides a method and apparatus for identifying commercial cluster areas. This method acquires preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified. These data are then sequentially input into a target decision tree model to obtain predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercial in nature. Next, missing scene data is supplemented based on the predicted engineering parameter data to obtain the target scene data. Finally, the predicted performance data and predicted MDT data are scored to obtain a comprehensive total score. A density function is then applied between the target scene data and the comprehensive total score data. Clustering is performed to obtain density clustering results. Based on these density clustering results, commercial clusters are identified from the areas to be identified. Then, a target decision tree model is used to filter commercially relevant prediction scenario data, prediction parameter data, prediction performance data, and prediction MDT data, significantly reducing subjective factors in the process of identifying commercial clusters and making the identification results more convincing. The prediction scenario data is supplemented by prediction parameter data to make the prediction scenario data more complete. Finally, the prediction performance data and prediction MDT data are comprehensively scored and combined with the target scenario data to ensure that the final identified commercial clusters are more accurate. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0051] Figure 1 This is one of the flowcharts illustrating the identification method for commercial cluster areas provided by the present invention;

[0052] Figure 2 This is the second flowchart illustrating the identification method for commercial cluster areas provided by the present invention;

[0053] Figure 3 This is the third flowchart illustrating the identification method for commercial cluster areas provided by the present invention;

[0054] Figure 4 This is the fourth flowchart illustrating the identification method for commercial cluster areas provided by the present invention;

[0055] Figure 5 This is the fifth flowchart illustrating the identification method for commercial cluster areas provided by the present invention;

[0056] Figure 6This is the sixth flowchart illustrating the identification method for commercial cluster areas provided by the present invention;

[0057] Figure 7 This is a schematic diagram of the structure of the identification device for commercial cluster areas provided by the present invention;

[0058] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0060] The following is combined Figures 1-6 This invention describes a method for identifying commercial cluster areas.

[0061] Figure 1 This is one of the flowcharts illustrating the identification method for commercial cluster areas provided by the present invention, such as... Figure 1 As shown, the method includes:

[0062] Step 100: Obtain the preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified, and input the scene data, engineering parameter data, performance data, and MDT data into the target decision tree model in sequence to obtain prediction scene data, prediction engineering parameter data, prediction performance data, and prediction MDT data with commercial labels.

[0063] It should be noted that, in this invention, when acquiring data of the area to be identified, the area to be identified can be divided into units such as each city, street, community, or building, and then scene data, engineering parameter data, performance data, and MDT data of the area to be identified can be collected. In addition, data can also be collected based on the regional division of the area to be identified in the electronic map. In other words, the scene data, engineering parameter data, performance data, and MDT data in this invention include data from several sets of sampling points. For ease of explanation, this invention will be explained using a set of scene data, engineering parameter data, performance data, and MDT data.

[0064] Specifically, scene data refers to scene border information and scene tags. Scene tags include information such as scene name and scene type. Scene border information is mainly used to determine the outline of the commercial scene area, while scene tags are mainly used to determine the commercial attributes of the scene. For example, if the tag has keywords such as "shopping mall", "supermarket", "hotel" or "shopping center", then the scene can be identified as a commercial scene.

[0065] It should be noted that, since there is no unified standard for the format and standard of scene border information and scene labels, some scene data may be incorrect or omitted during the collection and integration process. Therefore, in order to solve this problem, the present invention can also use the enterprise directory to assist in the identification of scene labels. The enterprise directory refers to the directory of names of retail stores or buildings registered in the database. For example, if the enterprise directory contains "enterprise a", "enterprise b", etc., when the extracted scene information contains "enterprise a" or "enterprise b", it is identified as a scene label such as "shopping mall", "supermarket" or "shopping center".

[0066] The engineering parameter data includes 2G, 4G, and 5G engineering parameter data. Among them, the engineering parameter data includes base station longitude, access latitude, base station name, cell identifier, serving cell offset, cell name, cell radius, cell activation status, cell instance status, etc. Because the engineering parameter data contains more detailed scene attribute information, it can be used to supplement the missing parts of the scene data, making the identification of commercial cluster areas more accurate and complete.

[0067] Furthermore, it should be noted that the performance data in this invention mainly includes three indicators: average daily 4G traffic, maximum number of RRC connections during busy hours, and VoLTE voice traffic. MDT (Minimization Drive Test) refers to the Minimization Drive Test technique, which mainly obtains relevant parameters needed for network optimization through measurement reports reported by terminals. In this invention, MDT data mainly includes two indicators: total number of cells and total number of samples. The target decision tree model refers to a pre-trained decision tree model that can be used to classify data. The training data can include data with artificial labels, such as scenario data, engineering parameter data, performance data, and MDT data. Thus, the target decision tree model is constructed using data with artificial labels, which will not be elaborated further here.

[0068] Since there is no unified standard for the format and standard of scene data, engineering parameter data, performance data, and MDT data, in order to improve the model processing efficiency, this invention requires data preprocessing before inputting the data into the target decision tree model, thereby obtaining data with a unified format. Specifically, before obtaining the preprocessed scene data, engineering parameter data, performance data, and MDT data of the region to be identified, this invention also includes:

[0069] The initial scene data of the region to be identified is obtained, and the initial scene data is standardized and filtered sequentially to obtain preprocessed scene data;

[0070] Obtain the initial engineering parameter data of the region to be identified, and standardize the initial engineering parameter data to obtain the preprocessed engineering parameter data;

[0071] The initial performance data of the region to be identified is obtained, and the initial performance data is sequentially rasterized, denoised, and filtered to obtain the preprocessed performance data.

[0072] The initial MDT data of the region to be identified is obtained, and the initial MDT data is sequentially rasterized, denoised, and filtered to obtain the preprocessed MDT data.

[0073] In this invention, initial scene data, initial engineering parameter data, initial performance data, and initial MDT data all refer to unprocessed data. After obtaining the initial scene data, the initial scene data is standardized to make the fields of the initial scene data uniform in format. Then, error correction and record cleaning are performed to filter out duplicate or abnormal data and retain the scene data with valid format standard.

[0074] Similarly, the initial parameter data is standardized to transform diverse parameter data into standardized parameter data. For the initial performance data and initial MDT data, the initial performance data and initial MDT data are rasterized, and then the rasterized data is denoised. The head MDT data and tail MDT data are filtered out to reduce the impact of extreme data on the recognition results.

[0075] Therefore, this invention preprocesses the data to filter out duplicate and special data in advance, thereby significantly reducing the amount of data processing and further reducing the error of the recognition results.

[0076] Step 200: Based on the predicted working parameter data, supplement the predicted scene data with missing scene data to obtain the target scene data;

[0077] In practical applications, it is easy to understand that initial scene data is extracted from electronic maps. Electronic maps are constructed by collecting vector data and other supplementary data from various locations over a certain period of time. If the electronic map is not updated in a timely manner or if there are omissions in data collection, the initial scene data extracted from the electronic map will have data gaps. This will lead to the prediction scene data after classification based on the target decision tree model also having data gaps. Since the engineering parameter data contains more comprehensive scene attribute information, this invention uses the prediction engineering parameter data to supplement the missing scene data in the prediction scene data, thereby obtaining more comprehensive target scene data, which makes the identification of commercial cluster areas more accurate and complete.

[0078] Specifically, in this step, the predicted engineering parameter data is compared with the predicted scene data to obtain the engineering parameter data and scene data contained at each sampling point. For sampling points that only contain engineering parameter data or where there are missing information points in the scene data, the missing information in the scene data of the sampling point is supplemented by the information recorded in the engineering parameter data.

[0079] Step 300: Calculate the scores for the prediction performance data and the prediction MDT data to obtain the overall total score data;

[0080] Specifically, this step quantifies the value of the predictive performance data and predictive MDT data by utilizing the average daily 4G traffic, maximum number of RRC connections during busy hours, and VoLTE voice traffic from the predictive performance data, and the total number of cells and total number of samples from the predictive MDT data, in order to calculate the comprehensive total score. In this invention, the average daily 4G traffic is calculated in GB, and the maximum number of RRC connections during busy hours is the number of devices currently connected in the cell.

[0081] In practical applications, based on MDT technology, as long as the user terminal has GPS enabled and supports MDT functionality, the terminal can automatically report MDT data containing user location information to the base station. In this invention, the data labels output by the target decision tree model are all commercially relevant predicted MDT data. When the MDT data is collected in IDLE state, the total number of cells and the total number of samples can be calculated based on the RSCP and RSRQ of the serving cell, the RSCP and RSRQ of neighboring cells, location information, and other data in the MDT data. When the MDT data is collected in IDLE, CELL-PCH, and URA-PCH states, the total number of cells and the total number of samples can be calculated based on the RSCP and Ec / No of the serving cell, the RSCP and Ec / No of neighboring cells, location information, and other data in the MDT data. It should be noted that the above method for calculating the total number of cells and the total number of samples in the predicted MDT data is one implementation method proposed in this invention. Other methods can also be used to obtain the total number of cells and the total number of samples in the predicted MDT data, which will not be elaborated here.

[0082] Step 400: Perform density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and identify the commercial clustering area from the area to be identified based on the density clustering grouping results.

[0083] Specifically, density clustering is a density-based clustering method. The main idea of ​​density-based clustering is to find high-density regions that are separated from low-density regions. For commercial clusters, due to the large flow of people, the data density in commercial clusters is also high. Therefore, density clustering can identify commercial clusters with higher credibility. In this invention, the concave-convex envelope density algorithm can be used for density clustering grouping. In addition, other density clustering algorithms can also be used for density clustering grouping. This invention does not limit the use of density clustering algorithms.

[0084] In this invention, when performing density clustering on the data, the cluster centers are first determined, and then the cluster groups centered around each cluster center are determined, thus obtaining multiple cluster groups. Finally, the cluster group that conforms to the clustering characteristics of commercial clusters is selected from the multiple cluster groups, thereby identifying the commercial cluster area from the area to be identified.

[0085] Furthermore, it should be noted that the target scene data also needs to be rasterized in this invention to facilitate density clustering between the target scene data and the overall score data. It should be noted that in some application scenarios, the target scene data may be in vector format, such as vector or bitmap scene data. In this case, it is necessary to first rasterize the vector or bitmap scene data. Simply put, the raster data obtained after rasterization is a data organization method that represents the distribution of spatial features or phenomena in the form of a two-dimensional matrix. Each matrix unit is called a raster cell, and each data point in the raster represents the attribute data of the feature or phenomenon. In other words, raster data has the characteristics of explicit attributes and implicit location. Therefore, using rasterized target scene data and overall score data for density clustering in this invention can make the clustering results more accurate.

[0086] It is easy to understand that for commercial clusters, due to the large volume of business and people, the amount of performance data and MDT data in commercial clusters will be larger than that in other types of areas. Therefore, when quantifying the value of the data to obtain the comprehensive total score, the comprehensive total score of commercial clusters will also be larger than that of other types of areas. Therefore, this invention uses the method of combining target scene data with comprehensive total score data to determine commercial clusters, making the identification results of commercial clusters more accurate.

[0087] In this invention, scene data, engineering parameter data, performance data, and MDT data of the region to be identified are preprocessed and then sequentially input into the target decision tree model. This yields predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercially relevant. Next, missing scene data is supplemented based on the predicted engineering parameter data to obtain the target scene data. Finally, the predicted performance data and predicted MDT data are scored to obtain a comprehensive total score. The supplemented predicted scene data and the comprehensive total score data are then subjected to density clustering to obtain a density clustering group. The density clustering results are used to identify commercial clusters from the areas to be identified. Then, a target decision tree model is used to filter commercially relevant prediction scenario data, prediction parameter data, prediction performance data, and prediction MDT data, significantly reducing subjective factors in the identification of commercial clusters and further improving the accuracy of commercial cluster identification. Prediction parameter data is used to supplement the prediction scenario data, making the target scenario data more complete. Finally, the prediction performance data and prediction MDT data are comprehensively scored, combined with the supplemented target scenario data, thus ensuring a more accurate identification of commercial clusters.

[0088] Alternatively, in another embodiment of the present invention, reference is made to... Figure 2 , Figure 2 This is a second flowchart illustrating the identification method for commercial cluster areas provided by the present invention, as shown below. Figure 2 As shown: The step of supplementing the missing scene data in the predicted scene data based on the predicted working parameter data to obtain the target scene data specifically includes:

[0089] Step 2001: Generate the first Geometry data corresponding to the predicted engineering parameter data, and generate the second Geometry data corresponding to the predicted scene data. Then, based on the second Geometry data, supplement the first Geometry data with missing scene data to obtain the target Geometry data.

[0090] In this step, the Geometry data can be obtained by processing the target scene data using Geometry processing software, and there are no restrictions on this.

[0091] It should be noted that since scene data includes scene border information and scene labels, and scene border information includes the latitude and longitude data of the border vertices, in order to make the missing scene data more comprehensive, this invention converts the latitude and longitude data of the border vertices into Geometry data, so that when supplementing missing scene data, the missing border vertices and other scene data can also be supplemented.

[0092] Specifically, when supplementing missing scene data, the first Geometry data can be compared with the second Geometry data. In practical applications, when there are no missing scene data, the second Geometry data and the first Geometry data at the same sampling point should be consistent. However, when there are missing scene data, the second Geometry data and the first Geometry data at the same sampling point will be inconsistent. For example, there may be extra second Geometry data or extra first Geometry data. When there is extra first Geometry data, the extra first Geometry data is added to the second Geometry data, thereby making the scene data more comprehensive in terms of Geometry.

[0093] Step 2002: Perform density clustering on the target Geometry data to obtain several clustered data groups, and aggregate the several clustered data groups under the condition of a preset scenario generation threshold setting to obtain aggregated target Geometry data;

[0094] In this step, when performing density clustering grouping on the target Geometry data, since the target Geometry data contains multiple types of data, such as point data, line data and area data, the present invention performs density clustering grouping on each type of data to obtain multiple clusters of point data, line data and area data. Then, the clusters of each type are aggregated by generating a threshold through a preset scenario.

[0095] The preset scene generation threshold settings in this invention are shown in Table 2 below:

[0096] Table 2: Initial Settings for Scene Generation Threshold

[0097]

[0098]

[0099] As shown in Table 2, after density clustering the target geometry data, at least two clusters of point data and two clusters of areal data are obtained. After obtaining the density clustering results, for the point data clusters, they are first transformed into areal data clusters based on the point data expansion radius and the point data clustering radius. Then, the multiple areal data clusters are aggregated based on the areal data expansion radius and the areal data clustering radius, thereby generating more complete target geometry data.

[0100] Step 2003: Rasterize the predicted working parameter data to obtain first raster data, and rasterize the predicted scene data to obtain second raster data, and supplement the second raster data based on the first raster data to obtain target raster data;

[0101] It should be noted that raster data is a data organization method that represents the distribution of spatial features or phenomena in the form of a two-dimensional matrix. Each matrix unit is called a raster cell, and each cell contains an information value. Each data point in the raster represents the attribute data of the feature or phenomenon. Geometry data structure, on the other hand, uses points, lines, and surfaces to represent the real world. Therefore, in this invention, in addition to using Geometry data to supplement the scene boundary information of the scene data, it is also necessary to use raster data to supplement the scene attribute data of the scene data.

[0102] Specifically, when supplementing scene attribute data in scene data, the first raster data can be compared with the second raster data, and then the second raster data can be supplemented based on the comparison result to obtain the target raster data. The comparison and supplementation process of raster data in this step is the same as the comparison process of Geometry data in step 2001, and will not be described again here.

[0103] Step 2004: Based on the target raster data and the target Geometry data, obtain the target scene data.

[0104] Specifically, in this step, the target raster data can be supplemented based on the target Geometry data. For example, the target Geometry data can be rasterized to obtain the third raster data. Then, the third raster data is compared with the target raster data, and the missing raster data in the target raster data is supplemented based on the comparison results. In this way, the integrity of the scene data is further improved through secondary supplementation.

[0105] In this invention, missing data is supplemented by combining raster data and Geometry data, making the supplemented scene data more comprehensive, thereby making the subsequent identification results of commercial cluster areas more accurate.

[0106] Alternatively, in another embodiment of the present invention, reference is made to... Figure 3 , Figure 3 This is the third flowchart illustrating the identification method for commercial cluster areas provided by the present invention, as shown below. Figure 3 As shown: The process of scoring and calculating the predictive performance data and the predictive MDT data to obtain the comprehensive total score data specifically includes:

[0107] Step 3001: Obtain the first linear scoring factor of the prediction performance data and the second linear scoring factor of the prediction MDT data;

[0108] Specifically, in this invention, a linear scoring factor is used to calculate the comprehensive score of each data source. The calculation formula for the linear scoring factor in this invention is as follows:

[0109]

[0110] The predictive performance data includes three indicators: average daily 4G traffic, maximum number of RRC connections during busy hours, and VoLTE voice traffic. The predictive MDT data includes two indicators: total number of cells and total number of samples. The value in the linear scoring factor calculation formula... max The value refers to the maximum value of multiple metrics selected from the data source. min It refers to the minimum value of an indicator among multiple indicator data selected from the data source.

[0111] Specifically, for the first linear scoring factor of the predictive performance data, the index value of each indicator data in the predictive performance data is calculated, and based on the value of the indicator data in the predictive performance data... max and value min The first linear scoring factor is calculated. For the second linear scoring factor in the predicted MDT data, the index values ​​(values) for each indicator based on the total number of cells and the total number of samples are calculated. Then, based on the values ​​from the total number of cells and the total number of samples... max and value min The second linear scoring factor was calculated.

[0112] Step 3002: Calculate the comprehensive score of the prediction performance data based on the first linear scoring factor, and calculate the comprehensive score of the prediction MDT data based on the second linear scoring factor;

[0113] Since both the prediction performance data and the prediction MDT data contain multiple indicator data, this invention first calculates the indicator score of each indicator data based on the linear scoring factor, and then performs weighted calculations on the multiple indicator data of the prediction performance data and the prediction MDT data to obtain their respective comprehensive scores.

[0114] Specifically, in another embodiment, the step of calculating the comprehensive score of the prediction performance data based on the first linear scoring factor and calculating the comprehensive score of the prediction MDT data based on the second linear scoring factor specifically includes:

[0115] The index scores of each performance index in the predicted performance data are calculated, and the comprehensive score of the predicted performance data is calculated based on the first linear scoring factor and the index scores of each index. The performance index data of the predicted performance data includes average daily 4G traffic, maximum number of RRC connections during busy hours, and VoLTE voice traffic.

[0116] The index scores of each MDT index data in the predicted MDT data are calculated, and the comprehensive score of the predicted MDT data is calculated based on the second linear scoring factor and the index scores of each MDT index data. The MDT index data of the predicted MDT data includes the total number of cells and the total number of samples.

[0117] Specifically, the present invention uses the following formula to calculate the index scores of each index data:

[0118]

[0119] Here, "value" refers to the quantitative value of each indicator data in the data source. In this invention, when the indicator score of the indicator data exceeds 100, the indicator score is set to 100, thereby avoiding extreme data from affecting the accuracy of identifying commercial cluster areas.

[0120] Furthermore, it should be noted that since each data source contains multiple indicator data, and the importance of each indicator data varies in each data source, i.e., the degree of influence of each indicator data on each data source is different, this invention uses the following formula to calculate the comprehensive score of each data source in order to reasonably calculate the comprehensive score:

[0121]

[0122] in,

[0123] Among them, score x The scores for each indicator data, weight x The indicator weights refer to the weights of each indicator data. card(x) represents the cardinality of the indicator data from each data source. For example, when the data source is performance data, x = {1, 2, 3}, then card(x) = 3.

[0124] It should also be noted that the indicator weight represents the importance of the indicator in the data source. In this invention, the indicator weight of each indicator data is determined according to the importance of each indicator data in each data source. For example, for performance data, the indicators ranked from important to unimportant are, in order, daily average 4G traffic, maximum number of RRC connections during busy time, and VoLTE voice traffic. The indicator weights of each indicator data can be set to 0.8, 0.5, and 0.3 respectively.

[0125] Step 3003: Obtain the first weight of the predicted performance data and the second weight of the predicted MDT data, and perform a weighted calculation on the comprehensive score of the predicted performance data and the comprehensive score of the predicted MDT data based on the first weight and the second weight to obtain the comprehensive total score data.

[0126] Specifically, the following calculation formula is used in this invention for weighted calculation to obtain the overall total score:

[0127]

[0128] Among them, Score total The total score is calculated using the Score method. pm Score is a comprehensive score based on performance data sources. mdt For MDT data source comprehensive score, weight pm Weight is the data source weight for performance data. mdtThe data source weights for MDT data.

[0129] In addition, in order to prevent extreme data from adversely affecting the final result, and at the same time to significantly reduce the amount of computation and improve the efficiency of the algorithm, the data processing thresholds shown in Table 3 below are used to constrain the calculation process when calculating the overall score.

[0130] Table 3 Data Processing Thresholds

[0131]

[0132] In this invention, a comprehensive total score is obtained by weighting the first linear scoring factor and various indicator data related to the prediction performance data, and the second linear scoring factor and various indicator data related to the prediction MDT data. The comprehensive total score is then used to assist the scene data in identifying commercial cluster areas, reflecting the degree of commercial clustering in each area in a quantitative way, so that the finally identified commercial cluster areas are more in line with the real results.

[0133] Alternatively, in another embodiment of the present invention, reference is made to... Figure 4 , Figure 4 This is the fourth flowchart illustrating the identification method for commercial cluster areas provided by the present invention, as shown below. Figure 4 As shown: The step of performing density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and identifying the commercial clustering area from the area to be identified based on the density clustering grouping results, specifically includes: Step 4001, performing density clustering grouping on the target scene data and the comprehensive total score data to obtain the first density clustering grouping result;

[0134] In this step, it is easy to understand that when acquiring data of the area to be identified in this invention, the area to be identified can be divided into units such as each city, street, community, or building, and then scene data, engineering parameter data, performance data, and MDT data of the area to be identified can be collected. In addition, data can also be collected based on the regional division of the area to be identified in the electronic map. In other words, the scene data, engineering parameter data, performance data, and MDT data in this invention include data at several sets of sampling points. Therefore, the target scene data and the comprehensive total score data are data at several sets of sampling points. Since the density of commercial cluster areas is relatively high, the sampling points in commercial cluster areas are relatively dense. Therefore, this invention adopts density grouping, first dividing multiple data into multiple cluster groups to obtain the first density clustering grouping result.

[0135] In some application scenarios, the concave-convex envelope density algorithm can be used to separate some isolated data points in the target scene data and the comprehensive total score data. In addition, other density clustering algorithms can be used for density clustering grouping, which will not be elaborated here.

[0136] Step 4002: Based on the first density clustering grouping results, select target data from the target scene data and the comprehensive total score data;

[0137] In this invention, after obtaining the density clustering results, the clustered data is filtered. Specifically, target data for identifying commercial cluster areas is selected from all data participating in the density clustering according to the following rules:

[0138] The first method is to select "scene clustering data" that contains the center point of "overall score data".

[0139] The second method is to select the "overall score data" that includes the center point of the "scene clustering data".

[0140] The third type: "Comprehensive Total Score Data" which includes the central point of "Scenario Supplementary Data".

[0141] The fourth type: "Scenario Supplementary Data" which includes the central point of "Overall Score Data".

[0142] It should be noted that, in this invention, after identifying target scene data and target engineering parameter data with commercial nature through the target decision tree model, since there may be some missing data in the target scene data, the target engineering parameter data will be used to supplement the target scene data. Therefore, the scene clustering data in this invention refers to the original scene data in the target scene data, while the scene supplement data refers to the scene supplement data in the target engineering parameter data used to supplement the target scene data.

[0143] Step 4003: Perform density clustering on the target data to obtain a second density clustering result, and identify the commercial cluster area from the area to be identified based on the second density clustering result.

[0144] In this step, the concave-convex envelope density algorithm can be used to perform density clustering on the target data again. Specifically, the parameters in the concave-convex envelope density algorithm can be adjusted to make the second density clustering grouping result more accurate than the first density clustering grouping result.

[0145] The rule for selecting data for identifying commercial cluster areas based on the second density clustering grouping results in this invention is the same as the rule for selecting data for identifying commercial cluster areas based on the first density clustering grouping results in step 4002, and will not be repeated here.

[0146] Furthermore, due to the larger business volume and pedestrian flow in commercial clusters, the amount of performance data and MDT data in these areas will be greater than in other types of areas. Therefore, when quantifying the value of this data to obtain a comprehensive score, the comprehensive score for commercial clusters will also be higher than that for other types of areas. Thus, when selecting data, setting a comprehensive score threshold can make the selected data from commercial clusters more accurate. For example, the following rules can be used to select data for identifying commercial clusters:

[0147] The first method is to select "scenario clustering data" that contains the center point of "comprehensive total score data that is greater than the threshold of comprehensive total score data".

[0148] The second method is to select "comprehensive total score data that is greater than the threshold of comprehensive total score data" that contains the center point of "scene clustering data".

[0149] The third type: "Comprehensive total score data that is greater than the threshold of the comprehensive total score data" includes the center point of "supplementary scenario data".

[0150] The fourth type: "Scenario Supplementary Data" includes the center point of "Comprehensive Total Score Data that is greater than the threshold of the Comprehensive Total Score Data".

[0151] In this invention, by combining target scene data and comprehensive total score data to identify commercial cluster areas, the identification results are made more convincing. Furthermore, a second density clustering grouping is performed on the target scene data and comprehensive total score data to make the grouping results more accurate, thereby improving the accuracy of commercial cluster area identification.

[0152] Alternatively, in another embodiment of the present invention, reference is made to... Figure 5 , Figure 5 This is the fifth flowchart illustrating the identification method for commercial cluster areas provided by the present invention, as shown below. Figure 5 As shown: Before sequentially inputting the scene data, the engineering parameter data, the performance data, and the MDT data into the target decision tree model, the method further includes:

[0153] Step 5001: Construct decision tree nodes based on the data attributes of the training data to generate an initial decision tree model. The training data includes training scenario data, training parameter data, training performance data, and training MDT data.

[0154] Specifically, data attributes refer to data attributes that can distinguish commercial clusters. Each data attribute is then used as a node in the decision tree to construct an initial decision tree model. The data attributes in this invention are shown in Table 4 below:

[0155] Table 4: Data Attributes and Types

[0156]

[0157]

[0158] In practical applications, scene data, including scene border information, scene labels, scene names, and scene types, is extracted from electronic maps. This extracted scene data is then standardized, conforming the initial scene data fields to a uniform format. Error correction and record cleaning follow, filtering out duplicate or abnormal data and retaining only valid, standardized scene data. Similarly, the engineering parameter data is standardized, transforming diverse parameters into standardized data. Performance data and MDT data are rasterized, resulting in preprocessed scene data, engineering parameter data, performance data, and MDT data.

[0159] To improve the accuracy of the constructed decision tree model, this invention randomly selects 70% of the data from the preprocessed scene data, engineering parameter data, performance data, and MDT data as training scene data, training engineering parameter data, training performance data, and training MDT data. When selecting 70% of the data, 70% of the scene data, 70% of the engineering parameter data, 70% of the performance data, and 70% of the MDT data can be selected separately. Alternatively, 70% of the data can be randomly selected from all the data as training data.

[0160] In another embodiment, since there is more than one data attribute, the decision tree can generate a variety of decision tree models based on different root node and internal node selection schemes. The selection of the root node and internal nodes determines the accuracy of the decision tree model's prediction results. Therefore, in order to improve the accuracy of the decision tree model's prediction results, the construction of decision tree nodes based on the data attributes of the training data in this invention specifically includes:

[0161] Determine the data attributes of the training data, calculate the purity of each data attribute, and select the data attribute with the highest purity as the root node of the decision tree;

[0162] Determine the data attributes of the training data, calculate the weighted average Gini value of each data attribute, and select the first data attribute with the smallest weighted average Gini value as the root node of the decision tree;

[0163] Select the second data attribute with the smallest weighted average Gini value besides the first data attribute, and select the second data attribute as an internal node located at the next level below the root node;

[0164] Using the second data attribute as the first data attribute, the step of selecting the second data attribute with the smallest weighted average Gini value (other than the first data attribute) is repeated until all data attributes are selected.

[0165] For the input training data, the weighted average Gini value of each data attribute is calculated as the purity of the decision tree node. The data attribute with the highest purity is selected as the root node of the decision tree. Then, when selecting internal nodes, the purity of the remaining data attributes is calculated, and the data attributes with high purity are selected as the internal nodes that are closer to the root node. This process is repeated until all nodes on the decision tree are constructed.

[0166] Specifically, the algorithm for the weighted average Gini value in this invention is as follows:

[0167]

[0168] Furthermore, the algorithm for calculating the Gini value is as follows:

[0169]

[0170] Where P(j) represents the probability that the sample belongs to a commercial cluster area.

[0171] For example, if the input data for a specific region includes "whether it is a well-known brand," "whether it is a high-traffic warning cell," and "call volume," the weighted average Gini values ​​for these three data attributes are calculated to be 1, 1 / 2, and 0, respectively. If the calculated weighted average Gini value is small, it indicates the lowest uncertainty and highest purity, making it suitable as a classification node in a decision tree. Conversely, if the calculated weighted average Gini value is large, it indicates the highest uncertainty and lowest purity, making it unsuitable as a classification node in a decision tree. Therefore, in this invention, "whether it is a well-known brand" with the highest purity is selected as the root node, and "whether it is a high-traffic warning cell" with the second highest purity is selected as an internal node.

[0172] Step 5002: Input the test data into the initial decision tree model to obtain the predicted labels of the test data. The test data includes test scenario data, test parameter data, test performance data, and test MDT data with artificial labels.

[0173] Specifically, after randomly selecting 70% of the preprocessed scene data, parameter data, performance data, and MDT data as training scene data, training parameter data, training performance data, and training MDT data, the remaining 30% of the data is used as test scene data, test parameter data, test performance data, and test MDT data. Furthermore, based on manual identification, the preprocessed scene data, parameter data, performance data, and MDT data are manually labeled to obtain scene data, parameter data, performance data, and MDT data with manual labels. The accuracy of the current initial decision model is then verified through manual labeling.

[0174] In this context, "artificial labels" refer to labels indicating commercial nature. In practical applications, data with commercial nature can be labeled with "1" to represent its commercial nature, while data without commercial nature can be labeled with "0" to represent its non-commercial nature. Therefore, the predicted labels for the output test data also consist of "0" and "1".

[0175] Step 5003: Iteratively train the initial decision tree model based on the predicted label and the artificial label until the target decision tree model is obtained.

[0176] After obtaining the actual human labels and the predicted labels output by the model, the predicted labels output by the model are compared with the actual human labels to calculate the accuracy of the model's predictions. If the accuracy is lower than the set accuracy, it indicates that the current model has not met the target. In this case, it is necessary to select data from other regions and divide them into training data and test data to continue updating the initial decision tree model until a target decision tree model with an accuracy higher than the set accuracy is obtained.

[0177] In this invention, decision tree nodes are constructed based on the data attributes of training data to generate an initial decision tree model. Then, test data is input into the initial decision tree model to obtain predicted labels for the test data. Finally, the initial decision tree model is iteratively trained based on the predicted labels and artificial labels until the target decision tree model is obtained. Thus, the nodes of the decision tree are constructed based on the data attributes of training data of the same data type, making the classification results of various types of data in the target decision tree more consistent with the real results, reducing subjective factors in the process of identifying commercial cluster areas. Through a certain amount of training and verification, a high-accuracy target decision tree model is generated, further improving the accuracy of commercial cluster area identification.

[0178] Alternatively, in another embodiment of the present invention, reference is made to... Figure 6 , Figure 6 This is the sixth flowchart illustrating the identification method for commercial cluster areas provided by the present invention, as shown below. Figure 6 As shown: After identifying the commercial cluster area from the area to be identified based on the density clustering grouping results, the method further includes:

[0179] Step 6001: Obtain voice data, traffic data, and SMS data of the commercial cluster area;

[0180] Specifically, the voice data includes the total voice volume and the unit price of the voice service under the commercial cluster area; the traffic data includes the total traffic volume and the unit price of the traffic service under the commercial cluster area; and the SMS data includes the total SMS volume and the unit price of the SMS service under the commercial cluster area.

[0181] Step 6002: Based on the voice data, traffic data, and SMS data, score the commercial value of the commercial cluster area to obtain the quantitative value of the commercial cluster area.

[0182] Specifically, the total revenue for each business is first calculated based on the business data volume and unit price of each business. The formula for calculating the total revenue for each business in this invention is as follows:

[0183] Income totaltraffic =Price traffic *Traffic total

[0184] Income totalflow =Price flow *Flow total

[0185] Income totalsms =Price sms *Sms total

[0186] Among them, Income totaltraffic For total revenue from voice services, Income totalflow Income represents the total revenue from traffic-related services. totalsms For total revenue from SMS services, Traffic total For the total amount of speech, Flow total For total traffic, Sms total For total SMS volume, Price traffic Price is the unit price for voice services. flow Price is the unit price per unit of traffic. sms This refers to the unit price for SMS messages.

[0187] Next, after calculating the total revenue for each business segment, a weighted average is calculated based on the weight of each business segment to obtain the quantitative value of the commercial cluster area. Specifically, the formula for calculating the quantitative value of the commercial cluster area is as follows:

[0188] value = ∑Traffic total *Weight traffic +flow total *Weight flow

[0189] +Sms total *Weight sms

[0190] Among them, Weight traffic Weight flow and Weight sms These are the weights of each business segment. In this invention, the weight of each business segment is determined by calculating its proportion in the total revenue of the three business segments. Specifically, the formula for calculating the weight of each business segment is as follows:

[0191] Income total =Income totaltraffic +Income totalflow +Income totalsms

[0192]

[0193]

[0194]

[0195] Among them, Income total For total revenue, Weight Traffic Weight for voice services flow Weight is the weight for traffic business. sms Weighting of SMS services.

[0196] In this invention, after identifying commercial clusters, the commercial value of these clusters is quantified. This allows for the improvement and planning of business activities within these clusters based on their commercial value, thereby maximizing their overall value.

[0197] The identification device for commercial clusters provided by the present invention is described below. The identification device for commercial clusters described below and the identification method for commercial clusters described above can be referred to in correspondence.

[0198] refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of the identification device for commercial cluster areas provided by the present invention, as shown below. Figure 7 As shown, the identification device for commercial cluster areas includes: an acquisition unit 710, used to acquire preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified, and sequentially input the scene data, engineering parameter data, performance data, and MDT data into a target decision tree model to obtain predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercial; a supplementation unit 720, used to supplement the predicted scene data with missing scene data based on the predicted engineering parameter data to obtain target scene data; a scoring unit 730, used to calculate a score for the predicted performance data and the predicted MDT data to obtain a comprehensive total score; and an identification unit 740, used to perform density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and identify the commercial cluster area from the area to be identified based on the density clustering grouping results.

[0199] The identification device for commercial cluster areas provided by the present invention acquires preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified, and sequentially inputs these data into a target decision tree model to obtain predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercial in nature. Then, based on the predicted engineering parameter data, missing scene data is supplemented to obtain target scene data. Finally, the predicted performance data and predicted MDT data are scored to obtain a comprehensive total score. A density function is then applied between the target scene data and the comprehensive total score data. Clustering is performed to obtain density clustering results. Based on these density clustering results, commercial clusters are identified from the areas to be identified. Then, a target decision tree model is used to filter commercially relevant prediction scenario data, prediction parameter data, prediction performance data, and prediction MDT data, significantly reducing subjective factors in the process of identifying commercial clusters and making the identification results more convincing. The prediction scenario data is supplemented by prediction parameter data to make the prediction scenario data more complete. Finally, the prediction performance data and prediction MDT data are comprehensively scored and combined with the target scenario data to ensure that the final identified commercial clusters are more accurate.

[0200] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logic instructions in the memory 830 to execute a method for identifying commercial cluster areas. This method includes: acquiring preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified; sequentially inputting the scene data, engineering parameter data, performance data, and MDT data into a target decision tree model to obtain predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercial; supplementing the predicted scene data with missing scene data based on the predicted engineering parameter data to obtain target scene data; calculating a score for the predicted performance data and the predicted MDT data to obtain a comprehensive total score; performing density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results; and identifying the commercial cluster area from the area to be identified based on the density clustering grouping results.

[0201] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0202] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the identification method for commercial cluster areas provided by the above methods. The method includes: acquiring preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified; sequentially inputting the scene data, engineering parameter data, performance data, and MDT data into a target decision tree model to obtain predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercial; supplementing the predicted scene data with missing scene data based on the predicted engineering parameter data to obtain target scene data; calculating a score for the predicted performance data and the predicted MDT data to obtain a comprehensive total score; performing density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results; and identifying the commercial cluster area from the area to be identified based on the density clustering grouping results.

[0203] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the identification method for commercial cluster areas provided by the above methods. The method includes acquiring preprocessed scene data, engineering parameter data, performance data, and MDT data of the area to be identified; sequentially inputting the scene data, engineering parameter data, performance data, and MDT data into a target decision tree model to obtain predicted scene data, predicted engineering parameter data, predicted performance data, and predicted MDT data, all labeled as commercial in nature; supplementing the predicted scene data with missing scene data based on the predicted engineering parameter data to obtain target scene data; calculating a score for the predicted performance data and the predicted MDT data to obtain a comprehensive total score; performing density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results; and identifying the commercial cluster area from the area to be identified based on the density clustering grouping results.

[0204] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0205] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying a business cluster area, characterized by, include: The scene data, engineering parameter data, performance data, and MDT data of the region to be identified are obtained after preprocessing. The scene data, engineering parameter data, performance data, and MDT data are then sequentially input into the target decision tree model to obtain prediction scene data, prediction engineering parameter data, prediction performance data, and prediction MDT data, all of which are commercially labeled. Based on the predicted working parameter data, the missing scene data of the predicted scene data is supplemented to obtain the target scene data; The prediction performance data and the prediction MDT data are scored and calculated to obtain the comprehensive total score data; Density clustering is performed on the target scene data and the comprehensive total score data to obtain density clustering results. Based on the density clustering results, commercial cluster areas are identified from the areas to be identified. The step of supplementing the missing scene data in the predicted scene data based on the predicted working parameter data to obtain the target scene data specifically includes: Generate first Geometry data corresponding to the predicted engineering parameter data, and generate second Geometry data corresponding to the predicted scene data; When there is redundant first Geometry data at the same location compared to the second Geometry data, the redundant first Geometry data is added to the second Geometry data to obtain the target Geometry data. The point data in the target Geometry data is transformed into area data based on the expansion radius of the point data, and multiple area data are aggregated based on the expansion radius of the area data to obtain the aggregated target Geometry data; The predicted working parameter data is rasterized to obtain first raster data, the predicted scene data is rasterized to obtain second raster data, and the second raster data is supplemented based on the first raster data to obtain target raster data. Based on the target raster data and the aggregated target Geometry data, the target scene data is obtained.

2. The business cluster area-oriented recognition method according to claim 1, characterized by, After identifying the commercial cluster area from the area to be identified based on the density clustering grouping results, the method further includes: Acquire voice data, traffic data, and SMS data from the aforementioned commercial cluster area; The commercial value of the commercial cluster area is scored based on the voice data, traffic data, and SMS data to obtain the quantitative value of the commercial cluster area.

3. The business cluster area-oriented recognition method according to claim 1, characterized by, The step of scoring and calculating the comprehensive total score based on the prediction performance data and the prediction MDT data specifically includes: Obtain the first linear scoring factor of the predicted performance data and the second linear scoring factor of the predicted MDT data; The comprehensive score of the prediction performance data is calculated based on the first linear scoring factor, and the comprehensive score of the prediction MDT data is calculated based on the second linear scoring factor. Obtain the first weight of the predicted performance data and the second weight of the predicted MDT data, and perform a weighted calculation on the comprehensive score of the predicted performance data and the comprehensive score of the predicted MDT data based on the first weight and the second weight to obtain the comprehensive total score data.

4. The business-oriented cluster area facing recognition method according to claim 3, characterized by, The calculation of the comprehensive score of the prediction performance data based on the first linear scoring factor, and the calculation of the comprehensive score of the prediction MDT data based on the second linear scoring factor, specifically includes: The index scores of each performance index in the predicted performance data are calculated, and the comprehensive score of the predicted performance data is calculated based on the first linear scoring factor and the index scores of each performance index. The performance index data of the predicted performance data includes average daily 4G traffic, maximum number of RRC connections during busy hours, and VoLTE voice traffic. The index scores of each MDT index data in the predicted MDT data are calculated, and the comprehensive score of the predicted MDT data is calculated based on the second linear scoring factor and the index scores of each MDT index data. The MDT index data of the predicted MDT data includes the total number of cells and the total number of samples.

5. The identification method for commercial cluster areas according to claim 1, characterized in that, Before sequentially inputting the scene data, the engineering parameter data, the performance data, and the MDT data into the target decision tree model, the method further includes: Decision tree nodes are constructed based on the data attributes of the training data to generate an initial decision tree model. The training data includes training scenario data, training parameter data, training performance data, and training MDT data. The test data is input into the initial decision tree model to obtain the predicted labels of the test data. The test data includes test scenario data, test parameter data, test performance data and test MDT data with artificial labels. The initial decision tree model is iteratively trained based on the predicted labels and the artificial labels until the target decision tree model is obtained.

6. The identification method for commercial cluster areas according to claim 5, characterized in that, The decision tree nodes are constructed based on the data attributes of the training data, specifically including: Determine the data attributes of the training data, calculate the weighted average Gini value of each data attribute, and select the first data attribute with the smallest weighted average Gini value as the root node of the decision tree; Select the second data attribute with the smallest weighted average Gini value besides the first data attribute, and select the second data attribute as an internal node located at the next level below the root node; Using the second data attribute as the first data attribute, the step of selecting the second data attribute with the smallest weighted average Gini value (other than the first data attribute) is repeated until all data attributes are selected.

7. The identification method for commercial cluster areas according to claim 1, characterized in that, The step of performing density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and identifying commercial clustering areas from the area to be identified based on the density clustering grouping results, specifically includes: Density clustering is performed on the target scene data and the comprehensive total score data to obtain the first density clustering result; Based on the first density clustering grouping results, target data is selected from the target scene data and the comprehensive total score data; The target data is subjected to density clustering to obtain a second density clustering result. Based on the second density clustering result, commercial cluster areas are identified from the area to be identified.

8. The identification method for commercial cluster areas according to any one of claims 1-7, characterized in that, Before acquiring the preprocessed scene data, engineering parameter data, performance data, and MDT data of the region to be identified, the process also includes: The initial scene data of the region to be identified is obtained, and the initial scene data is standardized and filtered sequentially to obtain preprocessed scene data; Obtain the initial engineering parameter data of the region to be identified, and standardize the initial engineering parameter data to obtain the preprocessed engineering parameter data; The initial performance data of the region to be identified is obtained, and the initial performance data is sequentially rasterized, denoised, and filtered to obtain the preprocessed performance data. The initial MDT data of the region to be identified is obtained, and the initial MDT data is sequentially rasterized, denoised, and filtered to obtain the preprocessed MDT data.

9. An identification device for commercial cluster areas, characterized in that, include: The acquisition unit is used to acquire scene data, engineering parameter data, performance data and MDT data of the area to be identified after preprocessing, and input the scene data, engineering parameter data, performance data and MDT data into the target decision tree model in sequence to obtain prediction scene data, prediction engineering parameter data, prediction performance data and prediction MDT data with commercial labels. The supplementary unit is used to supplement the missing scene data of the predicted scene data based on the predicted working parameter data to obtain the target scene data; The scoring unit is used to score and calculate the prediction performance data and the prediction MDT data to obtain the comprehensive total score data. The identification unit is used to perform density clustering grouping on the target scene data and the comprehensive total score data to obtain density clustering grouping results, and to identify commercial clustering areas from the area to be identified based on the density clustering grouping results; The step of supplementing the missing scene data in the predicted scene data based on the predicted working parameter data to obtain the target scene data specifically includes: Generate first Geometry data corresponding to the predicted engineering parameter data, and generate second Geometry data corresponding to the predicted scene data; When there is redundant first Geometry data at the same location compared to the second Geometry data, the redundant first Geometry data is added to the second Geometry data to obtain the target Geometry data. The point data in the target Geometry data is transformed into area data based on the expansion radius of the point data, and multiple area data are aggregated based on the expansion radius of the area data to obtain the aggregated target Geometry data; The predicted working parameter data is rasterized to obtain first raster data, the predicted scene data is rasterized to obtain second raster data, and the second raster data is supplemented based on the first raster data to obtain target raster data. Based on the target raster data and the aggregated target Geometry data, the target scene data is obtained.

Citation Information

Patent Citations

  • Urban business function zoning method and device based on multi-factor spatial clustering and medium

    CN113379269A

  • Hotspot area positioning method and device, equipment and storage medium

    CN113392338A