Cell identification method and apparatus, storage medium, and computer program product
Patent Information
- Application Number
- PCT/CN2026/083773
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-16
- Publication Date
- 2026-09-24
Smart Images

Figure CN2026083773_24092026_PF_FP_ABST
Abstract
Description
A cell identification method, device, storage medium, and computer program product
[0001] Cross-references to related applications
[0002] This disclosure is based on and claims priority to Chinese Patent Application No. 202510329791.0, filed on March 19, 2025, the entire contents of which are incorporated herein by way of inclusion. Technical Field
[0003] This disclosure relates to the field of wireless technology, and in particular to a cell identification method, apparatus, storage medium, and computer program product. Background Technology
[0004] With the increasing number of 5G users on the existing network, the opportunities for dormancy will continue to decrease. Power-saving strategies based solely on user load forecasting are no longer sufficient to meet the energy-saving needs of the current network. More refined energy-saving technologies are urgently needed, such as setting differentiated energy-saving thresholds based on the energy-saving requirements of different coverage scenarios of wireless network cells. Current clustering algorithms can divide multiple cells into different clusters. For example, cells under the same coverage scenario can be divided into the same cluster, thereby allowing different energy-saving thresholds to be set according to different clusters.
[0005] However, current clustering algorithms generally use the default hyperparameter values when dividing multiple cells into clusters. This will cause the clustering algorithm to fail to output the optimal clustering effect, thus reducing the accuracy of the clustering algorithm in clustering cells. Summary of the Invention
[0006] This disclosure provides a cell identification method, apparatus, storage medium, and computer program product that can improve the accuracy of cell clustering.
[0007] The technical solution of this disclosure embodiment is implemented as follows:
[0008] In a first aspect, embodiments of this disclosure provide a cell identification method, the method comprising:
[0009] Obtain the first feature matrix corresponding to the first data; wherein the first data includes sample data of N cells, and the first feature matrix includes name features and location features corresponding to the N cells, where N is a positive integer;
[0010] The first preset algorithm is optimized based on the first feature matrix to obtain the optimized first preset algorithm; wherein, the optimized first preset algorithm includes a first parameter, which is configured to represent a first number of neighboring cells associated with each cell, and the first number is a variable value;
[0011] The optimized first preset algorithm is used to identify the clusters corresponding to the N cells.
[0012] Secondly, embodiments of this disclosure provide a cell identification device, which includes: an acquisition unit, an optimization unit, and an identification unit; wherein,
[0013] The acquisition unit is configured to acquire a first feature matrix corresponding to the first data; wherein the first data includes sample data of N cells, and the first feature matrix includes name features and location features corresponding to the N cells, where N is a positive integer;
[0014] The optimization unit is configured to optimize the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm; wherein the optimized first preset algorithm includes a first parameter, the first parameter is configured to characterize a first number of neighboring cells associated with each cell, and the first number is a variable value;
[0015] The identification unit is configured to identify the clusters corresponding to the N cells based on the optimized first preset algorithm.
[0016] Thirdly, embodiments of this disclosure provide a cell identification device, the cell identification device comprising: a processor and a memory; wherein,
[0017] The memory is configured to store computer programs that can run on the processor;
[0018] The processor is configured to execute the cell identification method as described above when running the computer program.
[0019] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer program code, which, when executed by a computer, implements the cell identification method described above.
[0020] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the cell identification method described above.
[0021] This disclosure provides a cell identification method, apparatus, storage medium, and computer program product. The method includes: acquiring a first feature matrix corresponding to first data; wherein the first data includes sample data of N cells, and the first feature matrix includes name features and location features corresponding to the N cells, where N is a positive integer; optimizing a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm; wherein the optimized first preset algorithm includes a first parameter, which is configured to represent a first number of neighboring cells associated with each cell, and the first number is a variable value; and identifying clusters corresponding to the N cells based on the optimized first preset algorithm. Therefore, it can be seen that a first feature matrix can be obtained first, which includes the name features and location features of the corresponding cells. Then, the first preset algorithm can be assisted in parameter tuning based on the feature matrix, so that the first parameters of the optimized first preset algorithm are variable. That is, the embodiments of this disclosure can dynamically obtain similar cell samples based on the name features and location features of each cell, so that the number of neighboring cells associated with each cell is variable. This ensures that the optimized first preset algorithm is more in line with the actual scene distribution characteristics when identifying clusters of N cells, thereby improving the accuracy of cell clustering. Attached Figure Description
[0022] Figure 1 is a schematic diagram of the cell identification method proposed in an embodiment of this disclosure;
[0023] Figure 2 is a schematic diagram of the first relationship network structure proposed in an embodiment of this disclosure;
[0024] Figure 3 is a schematic diagram of the target result acquisition queue proposed in an embodiment of this disclosure;
[0025] Figure 4 is a schematic diagram of the cluster division process proposed in the embodiments of this disclosure;
[0026] Figure 5 is a schematic diagram of the composition structure of the cell identification device proposed in the embodiment of this disclosure;
[0027] Figure 6 is a schematic diagram of the composition structure of the cell identification device proposed in the embodiments of this disclosure. Embodiments of the present invention
[0028] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the relevant applications and not for limiting the applications. Furthermore, it should be noted that, for ease of description, only the parts related to the relevant applications are shown in the accompanying drawings.
[0029] With the growth of 5G users in the live network, the dormancy opportunity will continue to decrease, and only using the power saving strategy of unified configuration based on user load prediction cannot meet the energy saving needs of the live network. The live network urgently needs more refined energy saving technology. Differentiated energy saving thresholds are set according to the energy saving needs of different coverage scenarios of the wireless network cell, and finally the potential of energy saving of the live network is fully tapped under the premise of guaranteeing the network quality and user experience of the live network, further improving the power saving amount of the live network.
[0030] Because the coverage scenarios in the work parameter data are set relatively widely, only a large category of scenarios can be differentiated for energy saving strategy, and the refined energy saving needs of the actual subdivided scenarios in the live network cannot be met. For example, the work parameter data sets the coverage scenario of all colleges and universities as "college and university", but in the live network, different energy saving strategies are often needed according to the actual energy saving needs of each college and university (such as Tsinghua University, Peking University, Renmin University, etc.) at different times to achieve refined energy saving of subdivided scenarios.
[0031] The current clustering algorithm can divide multiple cells into different clusters. For example, cells in the same coverage scenario can be divided into the same cluster, so that different energy saving thresholds can be set according to different clusters.
[0032] However, when the current clustering algorithm divides multiple cells into clusters, the default hyperparameter value of the clustering algorithm is generally used, which will result in the clustering algorithm being unable to output the optimal clustering effect, thereby causing the accuracy of the clustering algorithm for cell clustering to decrease.
[0033] To solve the problem of the decline in the accuracy of cell clustering of the current clustering algorithm, the embodiment of the present disclosure provides a cell identification method, device, storage medium and computer program product, which comprises: obtaining a first feature matrix corresponding to first data; wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and location features corresponding to the N cells, and N is a positive integer; performing optimization processing on a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm; wherein the optimized first preset algorithm comprises a first parameter, and the first parameter is used to represent a first number of neighbor cells associated with each cell, and the first number is a variable value; and the optimized first preset algorithm is used to identify and process clustering clusters corresponding to the N cells. As can be seen, the first feature matrix can be obtained first, the feature matrix comprises name features and location features corresponding to the cells, and then the first preset algorithm can be parameterized based on the feature matrix, so that the first parameter included in the optimized first preset algorithm is variable, i.e., the embodiment of the present disclosure can dynamically obtain similar cell samples based on the name features and location features of each cell, so that the number of neighbor cells associated with each cell is variable, ensuring that the optimized first preset algorithm is more in line with the actual scene distribution characteristics when identifying and processing the clustering clusters of the N cells, thereby improving the accuracy of cell clustering.
[0034] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure.
[0035] The embodiment of the present disclosure provides a cell identification method, and FIG. 1 is a schematic diagram of the cell identification method provided by the embodiment of the present disclosure, as shown in FIG. 1, the cell identification method can comprise the following operations:
[0036] Operation 101, a first feature matrix corresponding to first data is obtained; wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and location features corresponding to the N cells, and N is a positive integer.
[0037] In the embodiment of the present disclosure, the cell identification device can obtain a first feature matrix corresponding to first data.
[0038] It should be noted that in the embodiment of the present disclosure, the first data can comprise sample data of N cells, for example, it can be full-amount cell sample data under the coverage scene category in the engineering parameter data, and the present disclosure does not make specific limitation on the number of cell sample data included in the first data.
[0039] It should be noted that in the embodiments of the present disclosure, the first feature matrix can include name features and position features corresponding to the N cells, the name features can include cell Chinese name feature values, and the position features can include longitude and latitude features of the cells, and the present disclosure does not make specific limitation on the number and types of features included in the first feature matrix.
[0040] For example, in the embodiments of the present disclosure, the first feature matrix can be as shown in Table 1.
[0041] Table 1
[0042]
[0043] Optionally, in the embodiments of the present disclosure, when the cell identification apparatus obtains the first feature matrix corresponding to the first data, the cell identification apparatus can first perform vectorization processing on the name information of the N cells to obtain a second feature matrix; wherein the second feature matrix includes name features of the N cells; and then the first feature matrix can be determined based on the second feature matrix and the position information corresponding to the N cells.
[0044] It should be noted that in the embodiments of the present disclosure, the position information can include longitude and latitude information of the cells, and the present disclosure does not make specific limitation on the types and number of information included in the position information.
[0045] For example, in the embodiments of the present disclosure, the cell identification apparatus can use natural language processing (NLP) technology to vectorize the Chinese names of the full-sample (i.e., the name information of the N cells) and generate an N M feature matrix (i.e., the second feature matrix) of the full-sample; wherein N represents the number of samples of the selected cells, and M represents the dimension of the output vector; and then the N (M+2) feature matrix (i.e., the first feature matrix) of the full-sample based on the second feature matrix and in combination with the position information of the cells. In order to ensure that features of different dimensions have consistent scales, the longitude and latitude data need to be normalized, and the finally generated first feature matrix is as shown in Table 1.
[0046] That is, in the embodiments of the present disclosure, the cell identification apparatus can construct the first feature matrix based on the name features and the position features of the cells at the same time, so that the first feature matrix can include the name information and the longitude and latitude information of each cell at the same time.
[0047] Operation 102, optimizing the first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm; wherein the optimized first preset algorithm includes a first parameter, and the first parameter is configured to represent a first number of neighbor cells associated with each cell, and the first number is a variable value.
[0048] In the embodiment of the present disclosure, after obtaining the first feature matrix corresponding to the first data, the cell identification device can optimize the first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm.
[0049] It should be noted that in the embodiment of the present disclosure, the first preset algorithm can be an OPTICS density algorithm, and the present disclosure does not make specific limitations on the type of the first preset algorithm.
[0050] It should be noted that in the embodiment of the present disclosure, the optimized first preset algorithm can include a first parameter, or can include other parameters, and the present disclosure does not make specific limitations on the number and type of parameters included in the optimized first preset algorithm.
[0051] It should be noted that in the embodiment of the present disclosure, the first parameter can represent the core hyperparameter min_samples (i.e., the minimum number of points in the neighborhood) of the OPTICS density clustering algorithm, and the first parameter can be used to represent the first number of neighbor cells associated with each cell. The first number is a variable value, that is, the number of neighbor cells associated with each cell in the embodiment of the present disclosure is variable, and is not a uniform min_samples value. The present disclosure does not make specific limitations on the size of the first parameter.
[0052] It should be noted that in the embodiment of the present disclosure, when the cell identification device optimizes the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm, the cell identification device can optimize the initial parameters in the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm. The initial parameters include a first initial parameter and a second initial parameter. The first initial parameter is used to represent the second number of neighbor cells associated with each cell. The second number is a fixed value. The second initial parameter is used to divide the clustering clusters corresponding to the N cells.
[0053] It should be noted that in the embodiment of the present disclosure, the first initial parameter can be the initial min_samples (i.e., the initial minimum number of points in the neighborhood) of the OPTICS density clustering algorithm, and the first initial parameter represents the second number of neighbor cells associated with each cell. The second number is a fixed value, that is, the number of neighbor cells associated with each cell is fixed.
[0054] It should be noted that in the embodiment of the present disclosure, the second initial parameter can be the (i.e., the initial neighborhood radius), and the present disclosure does not make specific limitations on the size of the second initial parameter.
[0055] Furthermore, in the embodiments of this disclosure, when the cell identification device optimizes the initial parameters in the first preset algorithm based on the first feature matrix, it can perform dimensionality reduction on the first feature matrix to obtain a dimensionality-reduced first feature matrix; then, it can generate a first relationship network corresponding to N cells based on the dimensionality-reduced first feature matrix; wherein, the first relationship network includes a set of triangular networks formed by each of the N cells and the target cell, the number of triangular networks contained in the triangular network set is different, and the target cell includes cells associated with the cell; then, the first initial parameters can be optimized based on the first relationship network to obtain the first parameters.
[0056] For example, in embodiments of this disclosure, the cell identification device can utilize Principal Component Analysis (PCA) dimensionality reduction algorithm to reduce N... The (M+2) cell clustering feature matrix (i.e., the first feature matrix) is reduced to 2 dimensions, resulting in N. The first feature matrix (i.e., the first feature matrix after dimensionality reduction) is used to generate a first relationship network corresponding to N cells. Figure 2 is a schematic diagram of the first relationship network structure proposed in this embodiment. As shown in Figure 2, the first relationship network can contain a triangular network set formed by each cell in the N cells and the target cell. The number of triangular networks contained in the triangular network set is different, that is, the number of neighbor points of each sample can be different. As shown in Figure 2, the number of neighbor points of sample 1 can be 5, sample 2 can be 10, and sample 3 can be 4. Therefore, the min_samples hyperparameter values of samples 1, 2, and 3 can be set to 5, 10, and 4, respectively. The target cell can contain cells associated with a certain cell, that is, each point represents 1 cell sample, and the points connected to it represent cells strongly associated with it. For example, the name information and latitude and longitude information of each cell can be obtained based on the first feature matrix, and then the cells strongly associated with each cell sample can be obtained based on the name information and latitude and longitude information of each cell.
[0057] It should be noted that, in the embodiments of this disclosure, after the cell identification device generates a first relationship network corresponding to N cells based on the first feature matrix after dimensionality reduction, it can optimize the first initial parameters based on the first relationship network to obtain the first parameters.
[0058] That is, in the embodiment of the present disclosure, the cell identification device can generate a first relationship network corresponding to the N cells based on the name information and the longitude and latitude information of each cell, that is, a set of triangular networks formed by each cell in the N cells and the target cell can be obtained, so that the number of strongly associated cells of each cell sample can be obtained based on different sets of triangular networks, and the value of the number is a variable value, and then the first initial parameter can be optimized based on the first relationship network, that is, the initial min_samples (that is, the initial minimum number of neighbors) of the OPTICS density clustering algorithm can be optimized, so that the fixed number of neighbor cells associated with each cell is optimized to a variable number, thereby avoiding the use of a uniform numerical value for the number of neighbor cells of all cell samples, and thus the accuracy of the subsequent clustering results can be improved.
[0059] It should be noted that in the embodiment of the present disclosure, the optimized first preset algorithm can further include a second parameter, and the second parameter is used to divide the clustering clusters corresponding to the N cells, and the first parameter can be a parameter obtained by optimizing the first initial parameter, and the present disclosure does not limit the size of the second parameter.
[0060] It should be noted that in the embodiment of the present disclosure, when the cell identification device optimizes the initial parameter in the first preset algorithm based on the first feature matrix, the second initial parameter can be optimized based on the second preset algorithm and the first feature matrix to generate the second parameter; wherein the second preset algorithm is configured to search for an optimal second parameter, and the second preset algorithm includes a preset search condition; the preset search condition can include that the proportion of noise samples is minimum and the number of clustering clusters is greater than a first preset threshold, and the noise samples include cells not belonging to any clustering cluster.
[0061] It should be noted that in the embodiment of the present disclosure, the second preset algorithm can include a Bayesian optimization algorithm, and the present disclosure does not limit the type of the second preset algorithm.
[0062] It should be noted that in the embodiment of the present disclosure, the first preset threshold can be 1, and the present disclosure does not limit the size of the first preset threshold.
[0063] It should be noted that in the embodiment of the present disclosure, the preset search condition can be as shown in the following formula (1), and the preset search condition can include that the proportion of noise samples is minimum and the number of clustering clusters is greater than a first preset threshold, for example, the number of clustering clusters is greater than 1, thereby avoiding the case that all samples are classified into the same scene clustering cluster, losing the significance of the scene subdivision algorithm, and minimizing the proportion of noise samples can avoid the case that the division standard of the scene clustering cluster is too fine when the second parameter is small, thereby causing too many cells not to belong to any scene and become noise samples.
[0064] (1)
[0065] wherein, denotes the noise sample ratio, N is the total number of samples, and k is the number of clustering clusters.
[0066] That is, in the embodiment of the present disclosure, the cell identification device can pre-set the search condition of the Bayesian optimization algorithm, that is, the noise sample ratio is minimum and the number of clustering clusters is greater than the first preset threshold, so that the proportion of samples misidentified as noise can be effectively reduced, and the clustering effect is ensured. In the embodiment of the present disclosure, the second initial parameter (i.e. the initial neighborhood radius) of the OPTICS density clustering algorithm can be optimized based on the search condition of the Bayesian optimization algorithm to obtain the second parameter. For example, the initial neighborhood radius of the N cells in the first feature matrix can be optimized according to the search condition to obtain the optimal neighborhood radius (i.e. the second parameter) that satisfies the above search condition.
[0067] Operation 103, identifying and processing the clustering clusters corresponding to the N cells based on the optimized first preset algorithm.
[0068] In the embodiment of the present disclosure, after the cell identification device optimizes the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm, the clustering clusters corresponding to the N cells can be identified and processed based on the optimized first preset algorithm.
[0069] It should be noted that in the embodiment of the present disclosure, when the cell identification device identifies and processes the clustering clusters corresponding to the N cells based on the optimized first preset algorithm, the first cell can be obtained in the first data; wherein the first cell represents a core cell; the core cell includes a first number of cells greater than a second preset threshold; then the first cell can be added to the initial result queue, and the P first neighbor cells associated with the first cell can be obtained based on the first parameter; wherein P is a positive integer; further, the P first neighbor cells can be added to the ordered queue, and the target result queue can be obtained based on the ordered queue; so that the clustering clusters corresponding to the N cells can be determined based on the second parameter and the target result queue.
[0070] It should be noted that in the embodiment of the present disclosure, the second preset threshold can be set to any positive integer according to demand, and the size of the second preset threshold is not limited in the present disclosure.
[0071] For example, in the embodiment of the present disclosure, if the first number of neighbor cells associated with a certain cell is greater than the second preset threshold, the cell can be considered as a core cell, and the judgment rule of the core cell is not limited in the present disclosure.
[0072] It should be noted that, in the embodiments of this disclosure, after the cell identification device adds P first neighbor cells to the ordered queue, it can determine the first distance between each first neighbor cell and the first cell; then it can sort the first neighbor cells based on the first distance and the preset sorting order.
[0073] It should be noted that in the embodiments of this disclosure, the preset sorting order can be an ascending order, and this disclosure does not specifically limit the setting of the preset sorting order.
[0074] For example, in an embodiment of this disclosure, the cell identification device can sort the first distances in ascending order, thereby sorting the first neighboring cells in an ordered queue in ascending order.
[0075] It should be noted that, in the embodiments of this disclosure, when the cell identification device determines the first distance between each first neighbor cell and the first cell, it can determine the cosine similarity between each first neighbor cell and the first cell; wherein, the cosine similarity is configured to measure the similarity between the first neighbor cell and the first cell; then, it can determine the second distance between the first base station corresponding to the first neighbor cell and the second base station corresponding to the first cell; wherein, the second distance is obtained based on the latitude and longitude information of the first base station and the latitude and longitude information of the second base station; and then, the cosine similarity and the second distance can be subjected to a first preset calculation to obtain the first distance.
[0076] For example, in an embodiment of this disclosure, the cell identification device can calculate the cosine similarity between the first neighbor cell and the first cell using the following formula (2).
[0077] (2)
[0078] Here, A and B represent the feature vectors of two community names. The value range of Cosine Similarity(A,B) is [-1,1], which represents the angular difference between the two vectors. A higher cosine similarity indicates that the semantic similarity between the names is stronger, and vice versa.
[0079] It should be noted that, in the embodiments of this disclosure, in practical applications, in order to convert cosine similarity into a distance quantity, the following methods can be used: The semantic "distance" between two community names can be calculated using the following formula (3).
[0080] (3)
[0081] in, denotes the semantic distance between two cell names.
[0082] It should be noted that in the embodiments of the present disclosure, when the cell identifying apparatus determines the second distance between the first base station corresponding to the first neighbor cell and the second base station corresponding to the first cell, the second distance can be calculated by the following formula (4).
[0083] (4)
[0084] wherein, , and , denote the latitude and longitude of the two base stations respectively, denotes the second distance.
[0085] It should be noted that in the embodiments of the present disclosure, after the cell identifying apparatus obtains the cosine similarity between the first neighbor cell and the first cell and the second distance between the first base station corresponding to the first neighbor cell and the second base station corresponding to the first cell, the cell identifying apparatus can perform the first preset operation on the cosine similarity and the second distance to obtain the first distance.
[0086] It should be noted that in the embodiments of the present disclosure, when the cell identifying apparatus performs the first preset operation on the cosine similarity and the second distance to obtain the first distance, the first distance can be calculated by the following formula (5).
[0087] (5)
[0088] wherein, denotes the first distance, and are weighting coefficients, denotes the semantic distance between two cell names, denotes the second distance.
[0089] It should be noted that in the embodiments of the present disclosure, as shown in the above formula (5), and are weighting coefficients, which are used to balance the importance of the semantic information of the cell name and the geographic information, and the selection of the weighting coefficients directly affects the clustering effect. The weighting coefficients can be adjusted according to the scene requirements and the data distribution characteristics, and the present disclosure does not make specific limitations on the size of the weighting coefficients.
[0090] That is, in the embodiments of the present disclosure, the distance metric manner can be defined based on a weighted combination of the cosine similarity of the cell name and the Euclidean distance of the base station latitude and longitude, to obtain the first distance between each first neighbor cell and the first cell, and sequentially sorted, so as to sufficiently reflect the similarity of the cell in the semantic space and the base station in the geographical space, so that the clustering result is more in line with the actual scene distribution characteristics.
[0091] It should be noted that in the embodiments of the present disclosure, the cell identification device can obtain the target result queue based on the ordered queue after obtaining P first neighbor cells associated with the first cell based on the first parameter, adding the P first neighbor cells to the ordered queue, and obtaining the first distance between each first neighbor cell and the first cell, and sorting according to the size of the first distance.
[0092] Further, in the embodiments of the present disclosure, when the cell identification device obtains the target result queue based on the ordered queue, it can be judged whether the ordered queue is an empty queue, and in the case that the ordered queue is not an empty queue, the cells in the ordered queue can be sequentially added to the initial result queue; or, in the case that the ordered queue is an empty queue, other cells in the first data can be added to the initial result queue to obtain the target result queue; wherein the other cells at least contain other core cells except the first cell.
[0093] Exemplarily, in the embodiment of the present disclosure, FIG. 3 is a schematic diagram of obtaining a target result queue according to the embodiment of the present disclosure. As shown in FIG. 3, the flow of obtaining the target result queue can include the following operations: 201. judging whether the first data (i.e., all sample sets D) is empty; 202. in the case that the first data is not empty, a first cell (i.e., a core object P) can be obtained from the first data; 203. the first cell can be added to an initial result queue (i.e., a result queue O); 204. based on a first parameter, P first neighbor cells (i.e., a number of neighbor points N) associated with the first cell are obtained; 205. the P first neighbor cells are added to an ordered queue Q and sorted according to the size of a first distance (i.e., a reachable distance); 206. judging whether the ordered queue is empty; 207. in the case that the ordered queue is not empty, the cells in the ordered queue can be added to the initial result queue one by one, for example, a first object q in the ordered queue Q can be taken out first; 208. the object q is added to the initial result queue; 209. judging whether the object q is a core point (i.e., a core cell); 210. if the object q is a core point, then based on the first parameter again, neighbor cells (i.e., a number of neighbor points N) associated with the object q are obtained; 211. the neighbor cells associated with the object q are added to the ordered queue Q and sorted according to the first distance; 212. judging whether the ordered queue is empty, if not, then a first object S in the Q is taken out again; and the above operations 207-212 are repeated until the ordered queue is empty; 213. in the case that the ordered queue Q is empty, it can be judged whether the first data is empty, in the case that the first data is not empty, another core object is taken out again and the above operations 202-212 are repeated until the first data is empty, then the flow ends, so as to obtain the target result queue.
[0094] It should be noted that in the embodiment of the present disclosure, after the cell identification device obtains the target result queue, the cell identification device can determine the clustering clusters corresponding to the N cells based on the second parameter and the target result queue.
[0095] It should be noted that in the embodiment of the present disclosure, when the cell identification device determines the clustering clusters corresponding to the N cells based on the second parameter and the target result queue, the cell identification device can perform division processing on the clustering clusters corresponding to the cells in the target result queue based on the second parameter to obtain M clustering clusters; wherein M is greater than a first preset threshold, and M is a positive integer.
[0096] For example, in the embodiments of the present disclosure, FIG. 4 is a schematic diagram of a clustering cluster division process proposed by the embodiments of the present disclosure. As shown in FIG. 4, when the cell identification device divides the clustering cluster corresponding to the cells in the target result queue based on the second parameter, the following operations can be included: 301. The cell identification device obtains the target result queue and determines the neighborhood radius (i.e., the second parameter); 302. Determine whether the target result queue O is empty; 303. In the case where the target result queue O is not empty, take out the head element P; 304. Determine whether the reachable distance (i.e., the first distance) of the element P is greater than the neighborhood radius (i.e., the second parameter); 305. If it is less than or equal to the neighborhood radius, the element P is added to the current clustering cluster; 306. If it is greater than the neighborhood radius, determine whether the core distance of the element P is less than or equal to the neighborhood radius; 307. If it is greater than the neighborhood radius, it is considered that the element P is a noise point; 308. If it is less than or equal to the neighborhood radius, the element P is added to a new clustering cluster; 309. Determine whether the target result queue O is empty again, and repeat the above operations 303~operations 309 in the case where the target result queue O is not empty; 310. Until it is determined that the target result queue O is empty, record the clustering cluster result and calculate the noise sample proportion; 311. Determine whether the optimization iteration is ended; 312. In the case where the optimization iteration is ended, output the clustering cluster corresponding to the optimal neighborhood radius as the final clustering result (M clustering clusters).
[0097] That is, in the embodiments of the present disclosure, the cell identification device can pre-set the search condition of the Bayesian optimization algorithm, so that the second initial parameter (i.e., the initial neighborhood radius) of the OPTICS density clustering algorithm can be optimized based on the search condition of the Bayesian optimization algorithm to obtain the optimal second parameter. When the cell identification device divides the clustering cluster corresponding to the cells in the target result queue based on the second parameter, the proportion of samples misidentified as noise can be effectively reduced, while the clustering effect is ensured, that is, the noise sample proportion is minimized and the number of clustering clusters is greater than the first preset threshold, so that the accuracy of the clustering result can be improved.
[0098] Further, in the embodiments of the present disclosure, after the cell identification device divides the clustering cluster corresponding to the cells in the target result queue based on the second parameter to obtain M clustering clusters, the M clustering clusters can be respectively subjected to semantic identification to obtain the semantic information corresponding to each of the M clustering clusters. Then, the keyword frequency statistics processing can be performed on the semantic information, and the keywords can be sorted based on the statistical result to obtain the actual scene information corresponding to the M clustering clusters.
[0099] For example, in the embodiments of the present disclosure, when the cell identification apparatus respectively performs semantic identification on the M clustering clusters to obtain semantic information corresponding to each of the M clustering clusters, the Chinese identification of the subdivision scene can be performed on each clustering cluster (for example, cluster_1) as shown in Table 2, so as to identify the actual subdivision scene represented by each clustering cluster, and implement a customized power saving strategy for each subdivision scene.
[0100] Table 2
[0101]
[0102] It should be noted that in the embodiments of the present disclosure, after the cell identification apparatus respectively performs semantic identification on the M clustering clusters to obtain semantic information corresponding to each of the M clustering clusters, the keyword in the semantic information can be subjected to word frequency statistical processing, and the keywords can be sorted based on the statistical results to obtain actual scene information corresponding to the M clustering clusters as shown in Table 3.
[0103] Table 3
[0104]
[0105] That is, in the embodiments of the present disclosure, after the cell identification apparatus divides the clustering cluster corresponding to the cell to obtain M clustering clusters, the M clustering clusters can be further respectively subjected to semantic identification to obtain semantic information corresponding to each of the M clustering clusters, and the keyword in the semantic information can be subjected to word frequency statistical processing, that is, the subdivision scene name of the clustering cluster can be automatically identified in the embodiments of the present disclosure, so as to realize further subdivision identification of the coverage scene (university, core business district, station, etc.), that is, the embodiments of the present disclosure can further accurately subdivide and identify the scene to which the cell belongs, so as to support a fine energy saving solution based on the subdivision scene.
[0106] In summary, the cell identification apparatus can construct a first feature matrix based on the name features and the location features of the cells at the same time, so that the first feature matrix can include the name information and the latitude and longitude information of each cell at the same time, and then the first feature matrix can be processed for dimension reduction to obtain a first feature matrix after dimension reduction; then the first relationship network corresponding to the N cells can be generated based on the first feature matrix after dimension reduction; wherein the first relationship network contains a triangular net set formed by each cell in the N cells and the target cell, and the number of triangular nets contained in the triangular net set is different; the target cell contains a cell associated with the cell; further, the first initial parameter can be optimized based on the first relationship network to obtain the first parameter; that is, the name information and the latitude and longitude information of each cell can be used to generate the first relationship network corresponding to the N cells, and then the triangular net set formed by each cell in the N cells and the target cell can be obtained, so that the number of cells strongly associated with each cell sample can be obtained based on different triangular net sets, and the value of the number is a variable value, and then the first initial parameter can be optimized based on the first relationship network, that is, the initial min_samples (i.e. the initial neighborhood minimum point number) of the OPTICS density clustering algorithm can be optimized, so that the fixed number of neighbor cells associated with each cell is optimized to a variable number, so that the number of neighbor cells of all cell samples can be avoided. Uniform numerical value, so as to improve the accuracy of the subsequent clustering result; the cell identification apparatus can also pre-set the search condition of the Bayesian optimization algorithm, that is, the noise sample ratio is the smallest and the number of clustering clusters is greater than the first preset threshold, so as to effectively reduce the proportion of samples misidentified as noise, while ensuring the clustering effect, and in the embodiment of the present disclosure, the second initial parameter (i.e. the initial neighborhood radius) of the OPTICS density clustering algorithm can be optimized based on the search condition of the Bayesian optimization algorithm to obtain the second parameter, for example, the initial neighborhood radius of the N cells in the first feature matrix can be optimized according to the search condition to obtain the optimal neighborhood radius (i.e. the second parameter) that satisfies the above search condition; further, when the cell identification apparatus divides the clustering cluster corresponding to the cell based on the second parameter, it can effectively reduce the proportion of samples misidentified as noise, while ensuring the clustering effect, that is, the noise sample ratio is the smallest and the number of clustering clusters is greater than the first preset threshold, so as to improve the accuracy of the clustering result.
[0107] The embodiment of the present disclosure provides a cell identification method, which comprises: obtaining a first feature matrix corresponding to first data; wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and position features corresponding to the N cells, N is a positive integer; performing optimization processing on a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm; wherein the optimized first preset algorithm comprises a first parameter configured to represent a first number of neighbor cells associated with each cell, and the first number is a variable value; performing identification processing on clustering clusters corresponding to the N cells based on the optimized first preset algorithm. As can be seen, the first feature matrix can be obtained first, and the feature matrix comprises name features and position features corresponding to the cells. Then, the first preset algorithm can be parameterized based on the feature matrix, so that the first parameter included in the optimized first preset algorithm is variable, that is, the present embodiment can dynamically obtain similar cell samples based on the name features and position features of each cell, so that the number of neighbor cells associated with each cell is variable, ensuring that the optimized first preset algorithm is more in line with the actual scene distribution characteristics when identifying the clustering clusters of the N cells, thereby improving the accuracy of cell clustering.
[0108] Based on the above embodiment, another embodiment of the present disclosure provides a cell identification method, which can comprise the following contents. First, the cell identification device can construct a cell name and latitude-longitude joint feature matrix (i.e. the first feature matrix), then the prior information of the cell feature matrix (i.e. the first feature matrix) can be used to assist the OPTICS density clustering algorithm (i.e. the first preset algorithm) in hyperparameter optimization, and the sample similarity measurement method is optimized, and the clustering cluster is subdivided by the automatic identification of the scene name based on the word frequency statistics of the Chinese name of the cell, to realize the further subdivision and identification of the coverage scene (university, core business circle, station, etc.), to support the fine energy-saving solution based on the subdivided scene.
[0109] It should be noted that in the embodiment of the present disclosure, when the cell identification device constructs the cell clustering feature matrix (i.e. the first feature matrix), it can select the full-amount cell sample data (i.e. the first data) under the coverage scene category that needs to be subdivided from the engineering parameter data, then the NLP technology can be used to vectorize the Chinese name of the full-amount sample to generate N M feature matrix (i.e. the second feature matrix); wherein N is the sample quantity under the selected scene, M is the output vector dimension of the Chinese name vectorization model, and on this basis, the latitude and longitude information of the sample is combined to form the N The (M+2) feature matrix (i.e., the first feature matrix) needs to be normalized to ensure that features of different dimensions have consistent scales. An example of the finally generated cell clustering feature matrix is shown in Table 1 above.
[0110] That is, in the embodiments of the present disclosure, the cell identification device can construct the first feature matrix based on the name feature and the location feature of the cell at the same time, so that the first feature matrix can include the name information and the latitude and longitude information of each cell at the same time.
[0111] It should be noted that, in the embodiments of the present disclosure, the OPTICS density clustering algorithm (i.e., the first preset algorithm) can be improved and optimized from two aspects of super parameter selection and sample similarity measurement in combination with the characteristics of the clustering scene, so that the clustering result is more in line with the actual scene sample distribution characteristics, and the accuracy of the clustering result is improved; wherein, regarding the optimization of the super parameter selection of the OPTICS density algorithm, the two core super parameters (initial parameters) of the OPTICS density clustering algorithm are min_samples (neighborhood minimum point number) and neighborhood radius), the selection of different super parameters will have a decisive influence on the clustering effect, in order to ensure that the selection of the super parameter is in line with the characteristics of the corresponding clustering scene to achieve the most accurate clustering effect, the prior information of the cell feature matrix (i.e., the first feature matrix) can be used to assist the selection of the optimal super parameter in the embodiments of the present disclosure, which can include the following contents, (1) variable min_samples (neighborhood minimum point number) super parameter selection, the OPTICS density clustering algorithm sets a uniform min_samples value (i.e., the first initial parameter) for all samples by default, the embodiments of the present disclosure can use the prior information of the clustering scene characteristics (i.e., the first feature matrix) to propose a super parameter optimization scheme of a single sample variable neighborhood minimum point number based on a Delaunay triangular net neighbor relationship (i.e., the first relationship net) to obtain the first parameter, the specific operation is as follows:
[0112] Operation 1: N (M+2) cell clustering feature matrix (i.e., the first feature matrix) is reduced to 2 dimensions by using the PCA dimension reduction algorithm, to obtain an N 2 feature matrix (i.e., the first feature matrix after dimension reduction); Operation 2: N 2The feature matrix generates a Delaunay triangulation (i.e., a first relationship network), as shown in FIG. 2, where each point represents a cell sample, and the points connected thereto represent the strongly associated cells thereof. According to the neighbor relationship of the triangulation, a variable min_samples value is set for each cell sample according to the number of neighbor points. As can be seen from FIG. 2, the number of neighbor points of sample No. 1 is 5, the number of neighbor points of sample No. 2 is 10, and the number of neighbor points of sample No. 3 is 4. Therefore, in the clustering algorithm (i.e., the first preset algorithm after optimization) of the embodiments of the present disclosure, the min_samples hyperparameter values of samples No. 1, 2, and 3 can be set to 5, 10, and 4, respectively. Since the neighbor points of the sample points in the Delaunay triangulation represent the samples closest to the feature, it can be ensured that in the clustering algorithm, samples with similar cell name and latitude and longitude joint information can be clustered into one class.
[0113] It should be noted that in the embodiments of the present disclosure, the first relationship network can include a triangulation set formed by each cell in the N cells and a target cell, the number of triangulations included in the triangulation set is different, and the target cell includes a cell associated with the cell. Further, the first initial parameter can be optimized based on the first relationship network to obtain the first parameter.
[0114] That is, in the embodiments of the present disclosure, the cell recognition device can generate a first relationship network corresponding to the N cells based on the name information and the latitude and longitude information of each cell, that is, a triangulation set formed by each cell in the N cells and a target cell can be obtained, so that the number of strongly associated cells of each cell sample can be obtained based on different triangulation sets, and the value of the number is a variable value. Further, the first initial parameter can be optimized based on the first relationship network, that is, the initial min_samples (i.e., the initial minimum number of neighbors) of the OPTICS density clustering algorithm can be optimized, so that the fixed number of neighbor cells associated with each cell is optimized to a variable number. In this way, the accuracy of the subsequent clustering result can be improved.
[0115] It should be noted that in the embodiments of the present disclosure, the cell recognition device can also optimize another core hyperparameter (neighbor radius) (i.e., a second initial parameter) of the OPTICS density clustering algorithm. The hyperparameter of the OPTICS algorithm controls the maximum distance threshold between samples, and the value thereof directly affects the proportion of noise samples (i.e., the proportion of cells not divided into any one sub-scene clustering cluster) and the number of sub-scene clustering clusters. In a smaller case, the division standard of the scene clustering cluster is too fine, so that too many cells are not classified into any scene and become noise samples. When the value is When too large, the division criterion of the scene clustering cluster is too broad, leading to unclear cluster boundary, and even all samples are classified into the same scene clustering cluster, thereby losing the significance of the scene subdivision algorithm. To effectively reduce the proportion of samples misidentified as noise while ensuring the clustering effect, the Bayesian optimization algorithm (i.e., the second preset algorithm) is used to search for the optimal value (i.e., the second parameter), and the optimization objective is to minimize the noise sample proportion and the number of clustering clusters needs to be greater than 1 (i.e., at least two clustering clusters exist) to avoid the case that all samples are divided into the same cluster. The optimal value can be expressed as a solution satisfying the above formula (1).
[0116] It should be noted that in the embodiments of the present disclosure, the cell identification apparatus can perform optimization processing on the second initial parameter based on the second preset algorithm and the first feature matrix to generate the second parameter; the second preset algorithm is configured to search for the optimal second parameter, and the second preset algorithm includes a preset search condition; the preset search condition can include that the noise sample proportion is minimum and the number of clustering clusters is greater than a first preset threshold, and the noise samples include cells not classified into any clustering cluster.
[0117] It should be noted that in the embodiments of the present disclosure, the second preset algorithm can include a Bayesian optimization algorithm, and the type of the second preset algorithm is not limited in the present disclosure.
[0118] It should be noted that in the embodiments of the present disclosure, the first preset threshold can be 1, and the size of the first preset threshold is not limited in the present disclosure.
[0119] It should be noted that in the embodiments of the present disclosure, the preset search condition can be as shown in the above formula (1), the preset search condition can include that the noise sample proportion is minimum and the number of clustering clusters is greater than the first preset threshold, for example, the number of clustering clusters is greater than 1, thereby avoiding the case that all samples are classified into the same scene clustering cluster, losing the significance of the scene subdivision algorithm, and minimizing the noise sample proportion can avoid that when the second parameter is small, the division criterion of the scene clustering cluster is too fine, thereby leading to too many cells not being classified into any scene and becoming noise samples.
[0120] That is, in the embodiments of the present disclosure, the cell identification device can pre-set the search condition of the Bayesian optimization algorithm, that is, the noise sample ratio is the smallest and the number of clustering clusters is greater than the first preset threshold, so that the proportion of samples misidentified as noise can be effectively reduced, while ensuring the clustering effect, and the second initial parameter (that is, the initial neighborhood radius) of the OPTICS density clustering algorithm can be optimized based on the search condition of the Bayesian optimization algorithm in the embodiments of the present disclosure to obtain the second parameter. For example, the initial neighborhood radius of the N cells in the first feature matrix can be optimized according to the search condition to obtain the optimal neighborhood radius (that is, the second parameter) that satisfies the above search condition.
[0121] For example, in the embodiments of the present disclosure, when the cell identification device optimizes the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm, for example, the initial parameters in the first preset algorithm can be optimized based on the first feature matrix to obtain the optimized first preset algorithm; wherein the initial parameters include a first initial parameter and a second initial parameter, the first initial parameter is used to represent the second number of neighbor cells associated with each cell, and the second number is a fixed value, and the second initial parameter is used to divide the clustering clusters corresponding to the N cells; so that the optimized first preset algorithm can include the first parameter and the second parameter, and then the clustering clusters corresponding to the N cells can be identified based on the optimized first preset algorithm.
[0122] It should be noted that in the embodiments of the present disclosure, the optimization of the OPTICS density algorithm sample similarity measurement method can include the following contents, in order to enable the OPTICS algorithm to accurately measure the similarity between base stations, the combination of the semantic information of the cell name and the geographic location information of the base station is considered; in some embodiments, the distance measurement method is defined by calculating the weighted combination of the cosine similarity of the cell name and the Euclidean distance of the base station longitude and latitude, which can fully reflect the similarity of the cell in the semantic space and the base station in the geographic space, so that the clustering result is more consistent with the actual scene distribution characteristics. In the embodiments of the present disclosure, the cell name is used as a kind of semantic information, which can be vectorized by a pre-trained Word2Vec model for each character in the name, and converted into a high-dimensional vector, and the cosine similarity is used to measure the similarity between two cell name vectors to capture the semantic proximity of the name. The cosine similarity can be calculated by the above formula (2), and a higher cosine similarity indicates that the semantic similarity between the names is stronger, and vice versa. In practical applications, in order to convert the cosine similarity into a distance, the following formula can be used to represent the "distance" between two cell names in semantics, The first distance D can be calculated by formula (5) above. The geographic location information of the base station includes latitude and longitude, which reflects the position of the base station in the geographic space. The Euclidean distance of the latitude and longitude (i.e., the second distance) is used to measure the spatial distance between cells, which can be calculated by formula (4) above. By combining the cosine similarity distance of the cell name and the Euclidean distance of the geographic space, the clustering result is more in line with the actual scenario. The final total distance D (i.e., the first distance) can be calculated by formula (5) above.
[0123] It should be noted that in the embodiments of the present disclosure, as shown in FIG. 3, D represents a set to be clustered; Q represents an ordered queue, elements are sorted by reachable distance, and the smallest reachable distance is at the head of the queue; O represents a result queue, an ordered queue of the point set of the final output result; the process of obtaining the target result queue can include the following operations: 201. Determine whether the first data (i.e., all sample sets D) is empty; 202. In the case that the first data is not empty, a first cell (i.e., the core object P) can be obtained from the first data; 203. The first cell can be added to the initial result queue (i.e., the result queue O); 204. Based on the first parameter, P first neighbor cells (i.e., the number of neighbor points N) associated with the first cell are obtained; 205. P first neighbor cells are added to the ordered queue Q and sorted by the size of the first distance (i.e., the reachable distance); 206. Determine whether the ordered queue is empty; 207. In the case that the ordered queue is not empty, the cells in the ordered queue can be added to the initial result queue one by one, for example, we can first take out the first object q in the ordered queue Q; 208. The object q is added to the initial result queue; 209. Determine whether the object q is a core point (i.e., a core cell); 210. If the object q is a core point, then based on the first parameter, the neighbor cells (i.e., the number of neighbor points N) associated with the object q are obtained again; 211. The neighbor cells associated with the object q are added to the ordered queue Q and sorted by the first distance; 212. Determine whether the ordered queue is empty, if not, then the first object S in Q is taken out again; and the above operations 207~212 are repeated until the ordered queue is empty; 213. In the case that the ordered queue Q is empty, it can be determined whether the first data is empty, and in the case that it is not empty, another core object is taken out again and the above operations 202~212 are repeated until the first data is empty, then the process ends. In this way, the target result queue can be obtained.
[0124] It should be noted that in the embodiments of the present disclosure, after obtaining the target result queue, the cell identification device can use the Bayesian optimization algorithm to search for the optimal a value of the neighborhood radius, so that the noise sample ratio is minimized and the number of clustering clusters is greater than 1 (that is, there are at least two clustering clusters to avoid all samples being divided into the same cluster); and then the clustering cluster corresponding to the N cells can be determined based on the optimal the neighborhood radius (that is, the second parameter) and the target result queue.
[0125] It should be noted that in the embodiments of the present disclosure, when the cell identification apparatus determines the clustering cluster corresponding to the N cells based on the second parameter and the target result queue, the cell identification apparatus can perform division processing on the clustering cluster corresponding to the cells in the target result queue based on the second parameter to obtain M clustering clusters; wherein M is greater than the first preset threshold, and M is a positive integer.
[0126] For example, in the embodiments of the present disclosure, as shown in FIG. 4, when the cell identification apparatus performs division processing on the clustering cluster corresponding to the cells in the target result queue based on the second parameter, the following operations can be included: 301. The cell identification apparatus obtains the target result queue and determines the neighborhood radius (that is, the second parameter); 302. Determine whether the target result queue O is empty; 303. In the case that the target result queue O is not empty, take out the head element P; 304. Determine whether the reachable distance (that is, the first distance) of the element P is greater than the neighborhood radius (that is, the second parameter); 305. If it is less than or equal to the neighborhood radius, the element P is added to the current clustering cluster; 306. If it is greater than the neighborhood radius, determine whether the core distance of the element P is less than or equal to the neighborhood radius; 307. If it is greater than the neighborhood radius, it is considered that the element P is a noise point; 308. If it is less than or equal to the neighborhood radius, the element P is added to a new clustering cluster; 309. Determine whether the target result queue O is empty again, and repeat the above operations 303-operations 309 in the case that the target result queue O is not empty; 310. Until it is determined that the target result queue O is empty, record the result of the clustering cluster and calculate the noise sample ratio; 311. Determine whether the optimization iteration is ended; 312. In the case that the optimization iteration is ended, output the clustering cluster corresponding to the optimal neighborhood radius as the final clustering result (M clustering clusters).
[0127] That is, in the embodiments of the present disclosure, the cell identification apparatus can pre-set the search condition of the Bayesian optimization algorithm, so that the second initial parameter (that is, the initial neighborhood radius) of the OPTICS density clustering algorithm can be optimized based on the search condition of the Bayesian optimization algorithm to obtain the optimal second parameter, so that when the cell identification apparatus performs division processing on the clustering cluster corresponding to the cells in the target result queue based on the second parameter, the proportion of samples misidentified as noise can be effectively reduced, and the clustering effect can be ensured, that is, the noise sample ratio can be minimized and the number of clustering clusters is greater than the first preset threshold, so that the accuracy of the clustering result can be improved.
[0128] Further, in the embodiments of the present disclosure, after the cell identification device outputs the sub-scene clustering cluster by improving the OPTICS density clustering algorithm (i.e., the first preset algorithm after optimization), the Chinese identification (i.e., semantic identification) of the sub-scene needs to be performed on each clustering cluster, so as to enable the mobile operator user to identify the actual sub-scene represented by each clustering cluster, and implement a customized power saving strategy for each sub-scene.
[0129] For example, in the embodiments of the present disclosure, in order to identify the Chinese scene name of each clustering cluster, the full amount of cell Chinese names in each clustering cluster can be first summarized and segmented, the word frequency is counted and sorted from high to low, and the user can identify and mark the Chinese scene name of the clustering cluster through the TOP N word frequency, for example, as shown in the following table 2, which is the cluster_1 clustering cluster result of a certain province 5G work reference university coverage scene, and table 3 is the cell Chinese name segmentation and word frequency statistics result, according to the TOP 5 word frequency, the clustering cluster can be marked as a certain medical college sub-scene under the university coverage scene.
[0130] That is, in the embodiments of the present disclosure, after the cell identification device divides the clustering cluster corresponding to the cell and obtains M clustering clusters, the M clustering clusters can be further subjected to semantic identification, to obtain the semantic information corresponding to each of the M clustering clusters, and the key words in the semantic information are subjected to word frequency statistical processing, that is, in the embodiments of the present disclosure, the sub-scene name of the clustering cluster can be automatically identified, to realize further sub-scene identification for the coverage scene (university, core business district, station, etc.), that is, the embodiments of the present disclosure can further accurately subdivide and identify the scene to which the cell belongs, to support the fine energy saving solution based on the sub-scene.
[0131] In summary, the cell identification apparatus can construct a first feature matrix based on the name features and the location features of the cells at the same time, so that the first feature matrix can include the name information and the latitude and longitude information of each cell at the same time, and then the first feature matrix can be processed for dimension reduction to obtain a first feature matrix after dimension reduction; then the first relationship network corresponding to the N cells can be generated based on the first feature matrix after dimension reduction; wherein the first relationship network contains a triangular net set formed by each cell in the N cells and the target cell, and the number of triangular nets contained in the triangular net set is different; the target cell contains a cell associated with the cell; further, the first initial parameter can be optimized based on the first relationship network to obtain the first parameter; that is, the name information and the latitude and longitude information of each cell can be used to generate the first relationship network corresponding to the N cells, and then the triangular net set formed by each cell in the N cells and the target cell can be obtained, so that the number of strongly associated cells of each cell sample can be obtained based on different triangular net sets, and the value of the number is a variable value, and then the first initial parameter can be optimized based on the first relationship network, that is, the initial min_samples (i.e. the initial neighborhood minimum point number) of the OPTICS density clustering algorithm can be optimized, so that the fixed number of neighbor cells associated with each cell is optimized to a variable number, so as to avoid that the number of neighbor cells of all cell samples adopts a uniform value, thereby improving the accuracy of the subsequent clustering result; the cell identification apparatus can also pre-set the search condition of the Bayesian optimization algorithm, that is, the noise sample ratio is the smallest and the number of clustering clusters is greater than the first preset threshold, so as to effectively reduce the proportion of samples misidentified as noise and ensure the clustering effect; in the embodiment of the present disclosure, the second initial parameter (i.e. the initial neighborhood radius) of the OPTICS density clustering algorithm can be optimized based on the search condition of the Bayesian optimization algorithm to obtain the second parameter, for example, the initial neighborhood radius of the N cells in the first feature matrix can be optimized according to the search condition to obtain the optimal neighborhood radius (i.e. the second parameter) that satisfies the above search condition; further, when the cell identification apparatus divides the clustering clusters corresponding to the cells based on the second parameter, it can effectively reduce the proportion of samples misidentified as noise and ensure the clustering effect, that is, the noise sample ratio is the smallest and the number of clustering clusters is greater than the first preset threshold, so as to improve the accuracy of the clustering result.
[0132] The embodiment of the present disclosure provides a cell identification method, which comprises: obtaining a first feature matrix corresponding to first data; wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and position features corresponding to the N cells, and N is a positive integer; performing optimization processing on a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm; wherein the optimized first preset algorithm comprises a first parameter, and the first parameter is used to represent a first number of neighbor cells associated with each cell, and the first number is a variable value; and performing identification processing on clustering clusters corresponding to the N cells based on the optimized first preset algorithm. Therefore, the first feature matrix can be obtained first, the feature matrix comprises name features and position features corresponding to the cells, then the first preset algorithm can be parameter-optimized based on the feature matrix, so that the first parameter included in the optimized first preset algorithm is variable, that is, the embodiment of the present disclosure can dynamically obtain similar cell samples based on the name features and position features of each cell, so that the number of neighbor cells associated with each cell is variable, and it is ensured that the optimized first preset algorithm is more in line with the actual scene distribution characteristics when identifying the clustering clusters of the N cells, thereby improving the accuracy of cell clustering.
[0133] Based on the above embodiment, the embodiment of the present disclosure provides a cell identification device, and FIG. 5 is a schematic diagram of the composition structure of the cell identification device one, as shown in FIG. 5, the device 10 comprises: an acquisition unit 11, an optimization unit 12 and an identification unit 13; wherein,
[0134] The acquisition unit 11 is configured to obtain a first feature matrix corresponding to first data; wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and position features corresponding to the N cells, and N is a positive integer;
[0135] The optimization unit 12 is configured to perform optimization processing on a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm; wherein the optimized first preset algorithm comprises a first parameter, and the first parameter is used to represent a first number of neighbor cells associated with each cell, and the first number is a variable value;
[0136] The identification unit 13 is configured to perform identification processing on clustering clusters corresponding to the N cells based on the optimized first preset algorithm.
[0137] In the embodiments of the present disclosure, further, as shown in FIG. 6, the cell identification device 10 can further include a processor 14, a memory 15 storing executable instructions of the processor 14, further, the device 10 can further include a communication interface 16, and a bus 17 for connecting the processor 14, the memory 15 and the communication interface 16.
[0138] In the embodiments of the present disclosure, the processor 14 can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, or a microprocessor. It can be understood that, for different devices, the electronic device for realizing the functions of the processor can also be other devices, and the embodiments of the present disclosure are not limited specifically. The device 10 can further include a memory 15, which can be connected with the processor 14, wherein the memory 15 is used for storing executable program codes, the program codes including computer operation instructions, and the memory 15 can include a high-speed RAM memory and can also include a non-volatile memory, for example, at least two disk memories.
[0139] In the embodiments of the present disclosure, the bus 17 is used for connecting the communication interface 16, the processor 14 and the memory 15 and the mutual communication among these devices.
[0140] In the embodiments of the present disclosure, the memory 15 is used for storing instructions and data.
[0141] Further, in the embodiments of the present disclosure, the processor 14 is configured to obtain a first feature matrix corresponding to first data, wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and location features corresponding to the N cells, and N is a positive integer; perform optimization processing on a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm, wherein the optimized first preset algorithm comprises a first parameter, and the first parameter is used to represent a first number of neighbor cells associated with each cell, and the first number is a variable value; and perform identification processing on clustering clusters corresponding to the N cells based on the optimized first preset algorithm.
[0142] In actual applications, the memory 15 can be a volatile memory, such as a Random-Access Memory (RAM), or a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a Hard Disk Drive (HDD) or a Solid-State Drive (SSD), or a combination of the above types of memories, and provides instructions and data to the processor 14.
[0143] The embodiments of the present disclosure provide a cell identification device, which obtains a first feature matrix corresponding to first data, wherein the first data comprises sample data of N cells, the first feature matrix comprises name features and location features corresponding to the N cells, and N is a positive integer; performs optimization processing on a first preset algorithm based on the first feature matrix to obtain an optimized first preset algorithm, wherein the optimized first preset algorithm comprises a first parameter, and the first parameter is used to represent a first number of neighbor cells associated with each cell, and the first number is a variable value; and performs identification processing on clustering clusters corresponding to the N cells based on the optimized first preset algorithm. As can be seen, the first feature matrix can be obtained first, the feature matrix comprises name features and location features corresponding to the cells, and then the first preset algorithm can be parameterized based on the feature matrix, so that the first parameter included in the optimized first preset algorithm is variable, that is, the embodiments of the present disclosure can dynamically obtain similar cell samples based on the name features and location features of each cell, so that the number of neighbor cells associated with each cell is variable, and it is ensured that the optimized first preset algorithm is more consistent with the actual scene distribution characteristics when identifying the clustering clusters of the N cells, thereby improving the accuracy of cell clustering.
[0144] This disclosure provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the cell identification method as described above.
[0145] The program instructions corresponding to a cell identification method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to a cell identification method in the storage media are read or executed by an electronic device, the following operations are included:
[0146] Obtain the first feature matrix corresponding to the first data; wherein the first data includes sample data of N cells, and the first feature matrix includes name features and location features corresponding to the N cells, where N is a positive integer;
[0147] The first preset algorithm is optimized based on the first feature matrix to obtain the optimized first preset algorithm; wherein, the optimized first preset algorithm includes a first parameter, which is used to characterize the first number of neighboring cells associated with each cell, and the first number is a variable value;
[0148] The optimized first preset algorithm is used to identify the clusters corresponding to the N cells.
[0149] This disclosure also provides a computer program product, including a computer program that can be executed by the processor 14 of the cell identification device 10 to perform the operations described in any of the foregoing methods.
[0150] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0151] This disclosure is described with reference to implementation flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations thereof, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowcharts and / or one or more blocks of the block diagrams.
[0152] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0153] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide operations for implementing the functions specified in one or more flowcharts and / or one or more blocks in a block diagram.
[0154] The above description is merely a preferred embodiment of this disclosure and is not intended to limit the scope of protection of this disclosure.
Claims
1. A cell identification method, wherein, The method includes: Obtain the first feature matrix corresponding to the first data; wherein the first data includes sample data of N cells, and the first feature matrix includes name features and location features corresponding to the N cells, where N is a positive integer; The first preset algorithm is optimized based on the first feature matrix to obtain the optimized first preset algorithm; wherein, the optimized first preset algorithm includes a first parameter, which is configured to represent a first number of neighboring cells associated with each cell, and the first number is a variable value; The optimized first preset algorithm is used to identify the clusters corresponding to the N cells.
2. The method according to claim 1, wherein, The step of obtaining the first feature matrix corresponding to the first data includes: The name information of the N cells is vectorized to obtain a second feature matrix; wherein the second feature matrix includes the name features of the N cells; The first feature matrix is determined based on the second feature matrix and the location information corresponding to the N cells.
3. The method according to claim 1, wherein, The optimization of the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm includes: Based on the first feature matrix, the initial parameters in the first preset algorithm are optimized to obtain the optimized first preset algorithm; The initial parameters include a first initial parameter and a second initial parameter. The first initial parameter is configured to represent a second number of neighboring cells associated with each cell, and the second number is a fixed value. The second initial parameter is configured to divide the clusters corresponding to the N cells.
4. The method according to claim 3, wherein, The optimization of the initial parameters in the first preset algorithm based on the first feature matrix includes: The first feature matrix is reduced in dimension to obtain the reduced first feature matrix. Based on the first feature matrix after dimensionality reduction, a first relationship network is generated corresponding to the N cells; wherein, the first relationship network includes a set of triangular networks formed by each of the N cells and the target cell, the number of triangular networks in the set of triangular networks is different, and the target cell includes cells associated with the cell. The first initial parameters are optimized based on the first relationship network to obtain the first parameters.
5. The method according to claim 3, wherein, The optimized first preset algorithm further includes a second parameter, which is configured to divide the clusters corresponding to the N cells. The optimization of the initial parameters in the first preset algorithm based on the first feature matrix includes: The second initial parameters are optimized based on the second preset algorithm and the first feature matrix to generate the second parameters; wherein the second preset algorithm is configured to search for the optimal second parameters, and the second preset algorithm includes preset search conditions; The preset search conditions include having the smallest proportion of noise samples and having a number of clusters greater than a first preset threshold. The noise samples include cells that have not been assigned to any cluster.
6. The method according to claim 1, wherein, The process of identifying clusters corresponding to the N cells based on the optimized first preset algorithm includes: A first cell is obtained from the first data; wherein the first cell represents a core cell; the core cell includes cells whose first number is greater than a second preset threshold. Add the first cell to the initial result queue, and obtain P first neighbor cells associated with the first cell based on the first parameter; where P is a positive integer; Add the P first neighbor cells to an ordered queue, and obtain the target result queue based on the ordered queue; Based on the second parameter, the target result queue determines the clusters corresponding to the N cells.
7. The method according to claim 6, wherein, After adding the P first neighbor cells to the ordered queue, the method further includes: Determine the first distance between each of the first neighboring cells and the first cell; The first neighboring cells are sorted based on the first distance and a preset sorting order.
8. The method according to claim 7, wherein, Determining the first distance between each of the first neighboring cells and the first cell includes: For each of the first neighboring cells, a cosine similarity is determined between the first neighboring cell and the first cell; wherein the cosine similarity is configured to measure the similarity between the first neighboring cell and the first cell. Determine a second distance between a first base station corresponding to the first neighboring cell and a second base station corresponding to the first cell; wherein the second distance is obtained based on the latitude and longitude information of the first base station and the latitude and longitude information of the second base station; The first distance is obtained by performing a first preset operation on the cosine similarity and the second distance.
9. The method according to claim 6, wherein, The step of obtaining the target result queue based on the ordered queue includes: Determine whether the ordered queue is empty. If the ordered queue is not empty, add the cells in the ordered queue to the initial result queue in sequence. or, If the ordered queue is empty, other cells in the first data are added to the initial result queue to obtain the target result queue; wherein, the other cells include at least other core cells besides the first cell.
10. The method according to claim 6, wherein, The step of determining the clusters corresponding to the N cells based on the second parameter and the target result queue includes: Based on the second parameter, the clusters corresponding to the cells in the target result queue are divided to obtain M clusters; wherein, M is greater than a first preset threshold and M is a positive integer.
11. The method according to claim 10, wherein, The method further includes: Semantic labeling is performed on each of the M clusters to obtain the semantic information corresponding to each of the M clusters; The keywords in the semantic information are subjected to word frequency statistics processing, and the keywords are sorted based on the statistical results to obtain the actual scene information corresponding to the M clusters.
12. The method according to claim 3, wherein, The first preset algorithm is the OPTICS density algorithm; and / or The first initial parameter is the minimum number of initial neighborhood points in the OPTICS density clustering algorithm; and / or The second initial parameter is the initial neighborhood radius of the OPTICS density clustering algorithm.
13. The method according to claim 4, wherein, The step of generating the first relationship network corresponding to the N cells based on the dimensionality-reduced first feature matrix includes: The first relationship network corresponding to the N cells is generated based on the name information and latitude and longitude information of each cell.
14. A cell identification device, wherein, The cell identification device includes: an acquisition unit, an optimization unit, and an identification unit; wherein... The acquisition unit is configured to acquire a first feature matrix corresponding to the first data; wherein the first data includes sample data of N cells, and the first feature matrix includes name features and location features corresponding to the N cells, where N is a positive integer; The optimization unit is configured to optimize the first preset algorithm based on the first feature matrix to obtain the optimized first preset algorithm; wherein the optimized first preset algorithm includes a first parameter, the first parameter is configured to characterize a first number of neighboring cells associated with each cell, and the first number is a variable value; The identification unit is configured to identify the clusters corresponding to the N cells based on the optimized first preset algorithm.
15. A cell identification device, wherein, The cell identification device includes: a processor and a memory; wherein... The memory is configured to store computer programs that can run on the processor; The processor is configured to perform the method according to any one of claims 1-13 when running the computer program.
16. A computer-readable storage medium, wherein, The storage medium stores computer program code, which, when executed by a computer, performs the method according to any one of claims 1-13.
17. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method according to any one of claims 1-13.