Clustering method, device and server

By dividing the target area into multiple meshes and calculating the first mesh density of each mesh, and using a preset clustering algorithm for clustering calculation, the problem of low matching between the clustering results and the actual requirements is solved, and clustering efficiency and calculation complexity are improved.

CN114266317BActive Publication Date: 2025-06-24AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202111633500.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-06-24
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

In the actual process of achieving commercial site selection through clustering algorithms, the existence of obstacles such as rivers, buildings, and public facilities leads to a low degree of matching between clustering results and actual needs.

Method used

By dividing the target area into multiple meshes, the first mesh density of each mesh is calculated, and clustering calculations are performed using a preset clustering algorithm, multiple target clusters are obtained, and each target cluster includes at least one mesh.

Benefits of technology

The matching degree between clustering results and actual requirements is improved, clustering efficiency and computational complexity are improved, and algorithm complexity is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266317B_ABST
    Figure CN114266317B_ABST
Patent Text Reader

Abstract

The present application provides a clustering method, apparatus, and server. The method includes: The server can divide a target area into multiple grids, and data points and obstacles within the target area are distributed in the grids. The server can traverse all the grids within the target area and calculate the first grid density of these grids one by one. For obstacle grids, the server can use a corresponding density calculator to calculate their first grid density. The server can use a preset clustering algorithm to cluster these grids according to the first grid density of each grid in the target area. The server can obtain multiple target clusters by clustering, and each target cluster can include at least one grid. Each target cluster can include a clustering center. The method of the present application improves the matching degree between the clustering result and the actual demand, and improves the clustering efficiency and computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to a clustering method, apparatus, and server. Background Art

[0002] In data analysis, clustering algorithms are widely used to explore or organize data. For example, clustering algorithms have extensive applications in fields such as commercial site selection based on user location information, image segmentation, and recommendation algorithms.

[0003] In the prior art, the focus of clustering algorithms is concentrated on how to handle large-scale data processing. For example, in the application scenario of commercial site selection based on user location information, clustering algorithms usually start from a large range such as a city or a region and consider the commercial site selection and layout as a whole.

[0004] However, in the actual process of implementing commercial site selection through clustering algorithms, the existence of obstacles such as rivers, buildings, and public facilities will lead to a problem that the clustering results have a low matching degree with the actual needs. Summary of the Invention

[0005] This application provides a clustering method, apparatus, and server to solve the problem that in the actual process of implementing commercial site selection through clustering algorithms, the existence of obstacles such as rivers, buildings, and public facilities will lead to a problem that the clustering results have a low matching degree with the actual needs.

[0006] In a first aspect, this application provides a clustering method, including:

[0007] Dividing a target area into a plurality of grids, where data points and obstacles in the target area are distributed in the grids;

[0008] Determining a first grid density of each grid according to the grid information, data point information, and obstacle information in each grid, where the grid information includes grid length and grid width;

[0009] Performing clustering calculation on all grids in the target area according to the first grid density of each grid and a preset clustering algorithm to obtain at least one target cluster, where each target cluster includes at least one grid.

[0010] In a second aspect, this application provides a clustering apparatus, including:

[0011] An acquisition module, configured to divide a target area into a plurality of grids, where data points and obstacles in the target area are distributed in the grids;

[0012] A processing module, configured to determine a first grid density of each grid according to the grid information, the data point information, and the obstacle information within each grid, where the grid information includes the grid length and the grid width; perform clustering calculations on all grids within the target area according to the first grid density of each grid and a preset clustering algorithm to obtain at least one target cluster, and each target cluster includes at least one grid.

[0013] In a third aspect, the present application provides a server, including: a memory and a processor;

[0014] The memory is used to store a computer program; the processor is configured to execute the clustering method in the first aspect and any possible design of the first aspect according to the computer program stored in the memory.

[0015] In a fourth aspect, the present application provides a readable storage medium, in which a computer program is stored. When at least one processor of the server executes the computer program, the server executes the clustering method in the first aspect and any possible design of the first aspect.

[0016] In a fifth aspect, the present application provides a computer program product, where the computer program product includes a computer program. When at least one processor of the server executes the computer program, the server executes the clustering method in the first aspect and any possible design of the first aspect.

[0017] The clustering method provided by the present application divides the target area into multiple grids, and the data points and obstacles within the target area are distributed in the grids; traverses all grids within the target area and calculates the first grid density of these grids one by one, where the obstacle grids use the corresponding density calculator to calculate their first grid density; uses a preset clustering algorithm to cluster these grids according to the first grid density of each grid in the target area, and the clustering obtains multiple target clusters. Each target cluster may include at least one grid, and each target cluster may include a clustering center, thereby improving the matching degree of the clustering result with the actual demand, and improving the clustering efficiency and computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic diagram of the clustering result when there is no obstacle;

[0020] Figure 2 Schematic diagram of data distribution in the presence of an obstacle;

[0021] Figure 3 Schematic diagram of clustering results in the presence of an obstacle;

[0022] Figure 4 Schematic diagram of a partitioning-based obstacle clustering method;

[0023] Figure 5 Schematic diagram of the overall framework of a clustering device capable of handling obstacles provided by an embodiment of the present application;

[0024] Figure 6 Schematic diagram of the process of an application scenario of a clustering method provided by an embodiment of the present application;

[0025] Figure 7 Flowchart of a clustering method provided by an embodiment of the present application;

[0026] Figure 8 Flowchart of an obstacle stationary clusterer provided by an embodiment of the present application;

[0027] Figure 9 Schematic diagram of an obstacle grid provided by an embodiment of the present application;

[0028] Figure 10 Flowchart of a clustering method provided by an embodiment of the present application;

[0029] Figure 11 Flowchart of an obstacle addition clusterer provided by an embodiment of the present application;

[0030] Figure 12 Flowchart of an obstacle reduction clusterer provided by an embodiment of the present application;

[0031] Figure 13 Schematic diagram of the structure of a clustering device provided by an embodiment of the present application;

[0032] Figure 14 Schematic diagram of the hardware structure of a server provided by an embodiment of the present application. Detailed implementation manners

[0033] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0034] The terms "first", "second", "third", "fourth", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances. Furthermore, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context indicates otherwise.

[0035] In data analysis, clustering algorithms are widely used to explore or organize data. For example, clustering algorithms have extensive applications in fields such as commercial site selection based on user location information, image segmentation, and recommendation algorithms. Clustering algorithms are mainly implemented based on statistical models and neural network models. Currently, the focus of the application of clustering algorithms in the industry mainly concentrates on how to solve large-scale data processing and how to adapt to various types of applications. For example, in commercial site selection based on user location information, the prior art can analyze commercial site selection through user location information. In clustering scenarios such as commercial site selection, the target area is usually a relatively large area. The distribution of data points (user locations) in the target area is usually scattered. Therefore, in order to improve the clustering efficiency, the target area is gridded, and the standardized and normalized grid density is used as a sample in the clustering algorithm, which can usually effectively improve the clustering efficiency. Among them, a grid refers to a two-dimensional space in the target area, and the server can divide the target area into multiple grids of the same size according to the preset grid length and grid width. Grid density refers to the ratio of the number of data points in a grid to the area of the grid.

[0036] For example, as Figure 1 shown, each small circle can correspond to a data point (user location), or each small circle can also correspond to a sample (grid density). The server can cluster based on the data points in the target area to obtain two clustering centers. These two clustering centers are the two commercial site selections obtained from the clustering analysis. However, in practice, the target space may be an obstacle space with obstacles distributed in the space. The obstacles in this obstacle space can be entities such as rivers and buildings. The prior art does not consider the problem that there may be obstacles between the user location and the clustering center during the clustering process of user locations. As Figure 2 shown, the horizontal line can represent an obstacle such as a river in the target area. The existence of this obstacle divides the target area into upper and lower regions, and users in these two regions usually cannot directly cross this obstacle to move. Therefore, when this obstacle exists, the clustering result of the server can be as Figure 3 shown. Obviously, the existence of this obstacle will have a very large impact on the clustering result.

[0037] Currently, there are not many solutions proposed for the clustering scenario in the presence of obstacles. Considering that in the actual scenario, the main impact of the presence of obstacles on users lies in the distance. That is, when there is an obstacle such as a river between a user and a destination, the user usually needs to detour to a bridge to cross the obstacle and then can reach the destination. Therefore, the server can use the distance that the data point bypasses the obstacle to reach the clustering center as the distance from the data point to the clustering center, thereby adding the impact of the obstacle on the clustering result to the clustering calculation. For example, as Figure 4 shown, when the server has selected two clustering centers C A and C B , the server can calculate the distances from the data point p to the clustering center C A and the clustering center C B respectively. Among them, there is no obstacle between the data point p and the clustering center C A . Therefore, the distance dis(p, C A ) from the data point p to the clustering center C A is the straight-line distance between the data point p and the clustering center C A . Among them, there is an obstacle between the data point p and the clustering center C B . Therefore, the distance dis(p, C B ) from the data point p to the clustering center C B is not the straight-line distance. Among them, the obstacle between the data point p and the clustering center C B can include the endpoint D. dis(p, C B ) can be the sum of the distance from the data point p to the endpoint D and the distance from the endpoint D to the clustering center C B . The server can determine the partition of the data point p according to the calculated dis(p, C A ) and dis(p, C B ). Obviously, although using this method can meet the clustering calculation in the presence of obstacles, the complexity of the obstacle will have a greater impact on the calculation amount. For example, when the shape of the obstacle is complex, using the above two-stage calculation will not be able to obtain the distance between the obstacle between the data point p and the clustering center C B . Or, when there are multiple obstacles, the path from the data point p to the clustering center C B will become complex. It can be seen that the above method has the problem of low calculation efficiency.

[0038] In addition, when Figure 4 the midpoint p corresponds to a data point, using the above method can achieve the clustering of the data points in the target area. However, when Figure 4When the point p in [[]] corresponds to a sample, the grid density corresponding to the sample p cannot achieve clustering through distance. When the target area is large enough and there are enough users in the target area, the method of clustering for each data point is obviously inefficient. The method of using grids for clustering will cause problems with low matching degree to the actual requirements because obstacles cannot be calculated.

[0039] In view of the above problems, this application proposes a clustering method. The clustering method proposed in this application aims to establish an accurate, efficient, and concise clustering method that can be used in the presence of obstacles. The use of this method can effectively expand the application scenarios of clustering algorithms. For example, in the clustering analysis and calculation of the population distribution or traffic congestion situation, the clustering algorithm of this application can process data in the presence of obstacles such as buildings and roads.

[0040] To improve the processing efficiency of the clustering algorithm, this application uses the method of grid density. In this application, the server first needs to perform grid processing on the target area and divide the target area into multiple grids. Based on these grids, this application abstracts obstacles into two-dimensional graphics. For example, when the obstacle is a building, the obstacle can be abstracted as a polygon. Another example is that when the obstacle is a river, the obstacle can be abstracted as a line segment.

[0041] In the target area, the two-dimensional image of the obstacle will be cut by the grids and distributed in different grids. These grids with obstacles can be called obstacle grids. For obstacle grids, the calculation of their grid density will be affected by the obstacles. For example, when the obstacle is a polygon, the server can use the grid area after removing the obstacle area as the available area to calculate the grid density of the obstacle grid. Another example is that when the obstacle is linear, the server can divide the obstacle grid into multiple sub-grids according to the obstacles in the obstacle grid and calculate the sub-grid density of each sub-grid. The above processing of obstacles can ignore the shape of the obstacles, avoiding the need to frequently calculate the distance from data points to obstacles during the clustering process through distance, and at the same time avoiding the problem of increased distance calculation complexity caused by the complex shape of the obstacles. In addition, by combining obstacles with grids, the grid density can be calculated more accurately. Furthermore, the server can use these grid densities to execute the clustering algorithm to obtain multiple target clusters.

[0042] This application uses a clustering algorithm based on grid density to solve the problems of high computational complexity and low clustering efficiency in calculating clustering results based on distance metrics in the prior art. At the same time, this application also abstracts obstacles and adds them to the grid to make up for the problem of low matching degree between the clustering result and the actual demand in the existing grid density-based clustering algorithm. The clustering method of this application greatly reduces the algorithm complexity, improves the clustering efficiency, and improves the clustering accuracy.

[0043] In addition, this application can also use the already calculated clustering result as input information to calculate the situation of obstacle changes and update the target clusters obtained by clustering. This application can also achieve the effect of quickly updating the target clusters in both cases of increasing and decreasing obstacles.

[0044] The technical solution of this application will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0045] Figure 5 The overall framework diagram of a clustering device capable of handling obstacles provided by an embodiment of this application is shown. As shown in the figure, the entire device can be mainly divided into four units: a data model construction unit, a density calculation unit, a clustering calculation unit, and a visualization display unit. Among them, the density calculation unit can include two parts: a proportional density calculator and a direction proportional density calculator. The proportional density calculator is used to calculate the grid density of obstacle grids with polygon obstacles. The proportional density calculator is used to calculate the grid density of obstacle grids with linear obstacles. The clustering calculation unit includes three parts: an obstacle static clusterer, an obstacle addition clusterer, and an obstacle reduction clusterer.

[0046] This clustering device can be a virtual device and is set in the server. Each unit and actuator in this clustering device can actually be code modules or functions stored in this server.

[0047] The original data input into this clustering device can be the user information of each user in the target area and the obstacle information of this target area directly obtained by the server from other devices. When this clustering device obtains this original data, the execution process of this clustering device can be as Figure 6 shown.

[0048] After the original data is input into the clustering device, the data model construction unit can construct a data model through the data model construction unit. The data model construction unit can model the disordered and unstructured data to obtain ordered, structured, and processable data for subsequent clustering methods. For example, during the construction process, the data model construction unit can standardize the original data. Specifically, the standardization process can coordinate the location information in the user information of the original data. The coordinate-transformed user information can be displayed in the form of data points in the two-dimensional plane of the target area. The server can also abstractly process the obstacles based on the obstacle information in the target area. The abstract processing can include abstracting the obstacles in the target area into lines or polygons. When the obstacle is abstracted into a line, the abstracted obstacle information can include the starting coordinates and terminal coordinates of the line. When the obstacle is abstracted into a polygon, the abstracted obstacle information can include the vertex coordinates of the polygon. The server can also grid the target area through the data model construction unit. The server can divide the data points and obstacles in the target area into each grid according to the grid division. The server can determine the first data point information and the first obstacle information in each grid.

[0049] The server can calculate the grid density of each grid through the density calculation unit. The server can determine the density calculator used to calculate the grid density by judging whether the obstacle in the grid is a polygon. For example, when the obstacle in the obstacle grid is a polygon, the server can use a proportional density calculator to calculate the first grid density of the obstacle grid. When the grid is an obstacle grid with a linear obstacle or a grid without an obstacle, the server can use a directional proportional density calculator to calculate the first grid density of the grid.

[0050] The server can determine the type of the obstacle according to the obstacle information. The type can include stationary obstacles, increasing obstacles, and decreasing obstacles. For different obstacle types, the server can correspond to different clusterers in the clustering calculation unit. For example, when the obstacle type is a stationary obstacle, the server can use an obstacle stationary clusterer to cluster the grid density. When the obstacle type is an increasing obstacle, the server can use an obstacle increasing clusterer to cluster the grid density. When the obstacle type is a decreasing obstacle, the server can use an obstacle decreasing clusterer to cluster the grid density.

[0051] The server can continuously iterate the operations in the density calculation unit and the clustering calculation unit until all the grids in the target area are processed. When the server finishes processing the target area, the server can obtain the clustering result. The clustering result can include at least one target cluster. Each target cluster can include at least one grid. Each target cluster can include a clustering center. The server can convert information such as data points, obstacle abstraction results, grid density, target clusters, and clustering centers in the target area into visual information through the visualization display unit and display it on the monitor corresponding to the server. Users can perform subsequent analysis and decision-making based on this visual information.

[0052] In this application, taking the server as the execution entity, the clustering method of the following embodiments is executed. Specifically, the execution entity can be the hardware device of the server, or the software application in the server that implements the following embodiments, or the computer-readable storage medium installed with the software application that implements the following embodiments, or the code of the software application that implements the following embodiments.

[0053] Figure 7 The flowchart of a clustering method provided by an embodiment of this application is shown. As Figure 7 shown, taking the server as the execution entity, the method of this embodiment can include the following steps:

[0054] S101. Divide the target area into multiple grids, and the data points and obstacles in the target area are distributed within the grids.

[0055] In this embodiment, the server can obtain the original data. The original data can be the user information of each user in the target area and the obstacle information of the target area directly obtained by the server from other devices. The server can coordinate according to the location information in the user information in the original data through the data model construction unit. The server can convert the user information of each user into data point information in the target area. The server can also abstract the obstacle information. For example, for obstacles with a certain area such as buildings, the server can model them as two-dimensional plane polygons according to the obstacle information in the original data and use the vertices of the polygon in the target area as the obstacle information of the obstacle. Another example is that for linear obstacles such as roads and rivers, the server can model them as two-dimensional line segments according to the obstacle information in the original data and use the endpoints of the line segment as the obstacle information of the obstacle.

[0056] The server can also determine grid information based on the data point information. The server can divide the target area into multiple grids according to the grid information. The grid information can include the grid width and the grid height. For example, when the target area is a 30m×30m area and both the grid width and the grid height are 1m, the server can divide the target area into 900 1m×1m grids.

[0057] In one example, the specific steps for the server to implement the grid division of the target area may include:

[0058] Step 1: Determine grid information based on the data point information of all data points in the target area and the step size mapping table. The grid information includes the grid length and the grid width. The step size mapping table is used to indicate the mapping relationship from the second data point information to the grid length and the grid width.

[0059] In this step, the server can determine the grid information according to the dispersion degree of each data point in the target area and the step size mapping table. For example, the server can use a two-dimensional discrete distribution function to calculate the dispersion degree of the data points in the target area based on the data point information of each data point in the target area. The server can determine the grid information corresponding to the dispersion degree according to the dispersion degree and the step size mapping table. The step size mapping table may include the grid information corresponding to each discrete value. For example, when the discrete value is between 10 - 20, the server can determine that both the grid length and the grid width are 1 meter according to the step size mapping table. Or, when the discrete value is between 10 - 20, the server can determine that the target area will be divided into 30×30 grids according to the step size mapping table. When the target area is a 90m×60 area, the grid length is 3m and the grid width is 2m.

[0060] Alternatively, the server can also determine the grid information according to the density of the data points in the target area and the step size mapping table. For example, the server can determine the density of the data points in the target area by calculating the ratio of the number of data points in the target area to the area of the target area. The server can determine the grid information corresponding to the density of the data points in the target area according to the density of the data points in the target area and the step size mapping table.

[0061] Alternatively, the server can also determine the grid information according to the number of data points in the target area and the step size mapping table. In this application, there is no specific limitation on the determination of the grid information.

[0062] Step 2: Divide the target area into multiple grids according to the grid length and the grid width.

[0063] In this step, after the target area is divided into multiple grids, the data points and obstacles in the target area are correspondingly divided into each grid. There may be no data points in some grids. When there are no data points in a grid, the data point information of this grid is empty. Some grids may include at least one data point. There may be no obstacles in some grids. When there are no obstacles in a grid, the obstacle information of this grid is empty. Some grids may include obstacles, and such grids are obstacle grids.

[0064] S102. Determine the first grid density of each grid according to the grid information, the data point information in the grid, and the obstacle information in the grid, where the grid information includes the grid length and the grid width.

[0065] In this embodiment, the server can traverse all the grids in the target area and calculate the first grid density of these grids one by one. The implementation of this step is equivalent to Figure 5 the density calculation unit in the clustering device shown. For grids in different situations, the server can use different density calculators. Different density calculators can correspond to different density calculation methods and density calculation formulas. In the target area, the grids can be mainly divided into two categories: obstacle grids with obstacles and ordinary grids without obstacles. For obstacle grids, they can also be divided into two categories: polygon obstacle grids and linear obstacle grids according to the shape of the obstacles.

[0066] In one example, for an ordinary grid, the server can determine the first grid density of this grid by calculating the ratio of the number of data points in the grid to the area of the grid. Among them, the number of data points in the grid can be determined by counting the data points in the grid. Among them, the area of the grid can be determined according to the product of the grid length and the grid width of the grid.

[0067] In another example, assume that the grid to be calculated currently is the first grid. When the first grid includes a polygon obstacle, the server can use Figure 5 and Figure 6 the proportional density calculator shown in to calculate the first grid density of the first grid. Among them, the main execution steps of the first grid density of the first grid may include:

[0068] Step 1. Determine the available area of the first grid according to the grid information of the first grid and the obstacle information in the first grid.

[0069] In this step, the server can determine the obstacle area of the polygonal obstacles within the first grid based on the obstacle information in the first grid. The server can also determine the grid area of the first grid according to the grid information of the first grid. Since the places where there are obstacles in the first grid are generally considered as places that cannot be used by users, the server can determine the available area of the first grid based on the difference between the grid area and the obstacle area.

[0070] Step 2: Determine the number of data points in the first grid according to the data point information in the first grid.

[0071] In this step, the server can count the number of data points within the first grid to obtain the number of data points in the first grid.

[0072] Step 3: Determine the first grid density of the first grid according to the number of data points in the first grid and the available area of the first grid.

[0073] In this step, the server can calculate the ratio of the number of data points to the available area. This ratio is the first grid density of the first grid.

[0074] In another example, assume that the grid to be calculated currently is the second grid. When the second grid includes linear obstacles, the server can use Figure 5 and Figure 6 the directional proportional density calculator shown in to calculate the first grid density. Among them, the main execution steps of the first grid density of the second grid can include:

[0075] Step 1: Divide the grid into at least one sub-grid according to the obstacle information in the second grid and the grid information of the second grid, and determine the sub-grid area of each sub-grid.

[0076] In this step, the server can determine whether the linear obstacles in the second grid divide the second grid according to the obstacle information in the second grid. When the linear obstacle divides the second grid, the second grid is divided into multiple sub-grids. For example, when the linear obstacle in the second grid connects the midpoint of the left boundary of the second grid to the midpoint of the right boundary of the second grid, the linear obstacle divides the second grid into two sub-grids. Another example is that when the linear obstacle in the second grid connects the midpoint of the left boundary of the second grid to a certain point inside the second grid, although there is an obstacle in the second grid, the obstacle does not divide the second grid. The second grid can be regarded as a sub-grid. When the second grid is not divided into multiple sub-grids by the obstacle but is regarded as a sub-grid, the calculation of the first grid density of the second grid is the same as the calculation of the first grid density of an ordinary grid. Therefore, in Figure 6Among them, the grids of the server's non-polygonal obstacles are all calculated using this directional proportional density calculator.

[0077] After the server completes the division of the sub-grids, it can also calculate the area of each sub-grid. Since this obstacle is a linear obstacle, it can be considered that this linear obstacle does not occupy the area of the grid. For example, when the linear obstacle in the second grid connects from the midpoint of the left boundary of the second grid to the midpoint of the right boundary of the second grid, the areas of the two sub-grids are respectively 1 / 2 of the area of the second grid. Another example is when the linear obstacle in the second grid connects from the midpoint of the left boundary of the second grid to a certain point inside the second grid, the area of one sub-grid is the area of the second grid.

[0078] Step 2: Determine the number of sub-grid data points in each sub-grid according to the data point information of the second grid and the sub-grids of the second grid.

[0079] In this step, the server can obtain the data point information according to the second grid, and after determining that the second grid is divided into at least one sub-grid, obtain the data point information in each sub-grid. The server can count the number of data points in each sub-grid to obtain the number of sub-grid data points. For example, when the second grid is divided into one sub-grid, the number of sub-grid data points is the number of data points of the second grid. Another example is when the second grid is a linear obstacle connecting from the midpoint of the left boundary of the second grid to the midpoint of the right boundary of the second grid and is divided into two sub-grids, the server can respectively count the number of sub-grid data points in the two sub-grids.

[0080] Step 3: Determine the sub-grid density of each sub-grid according to the sub-grid area and the number of sub-grid data points of each sub-grid.

[0081] In this step, the server can calculate the ratio of the sub-grid area of each sub-grid in the second grid to the number of sub-grid data points of the sub-grid one by one. The server can use this ratio as the sub-grid density of the sub-grid.

[0082] S103: According to the first grid density of each grid and a preset clustering algorithm, perform clustering calculations on all grids in the target area to obtain at least one target cluster, and each target cluster includes at least one grid.

[0083] In this embodiment, the server can use a preset clustering algorithm to cluster these grids according to the first grid density of each grid in the target area. The server can cluster to obtain multiple target clusters, and each target cluster can include at least one grid. Each target cluster can include a clustering center. The above steps are equivalent to Figure 5 the process executed by the clustering calculation unit shown.

[0084] In one example, the clustering process is equivalent to Figure 5 the obstacle static clusterer in the clustering device shown. The execution process of the obstacle static clusterer can be as Figure 8 shown. Combining with the Figure 8 execution process shown, the specific process for the server to perform clustering can include the following steps:

[0085] Step 1: Divide all the grids in the target area into dense grids and sparse grids according to the first grid density of each grid and the grid density threshold.

[0086] In this step, the server can obtain the first grid density of each grid according to the above steps. The grid density threshold can be stored in the server. The server can compare the first grid density of each grid in the target area with the grid density threshold. When the first grid density of a grid is greater than or equal to the grid density threshold, the grid is a dense grid. When the first grid density of a grid is less than the grid density threshold, the grid is a sparse grid.

[0087] Step 2: Traverse all the grids in the target area and establish a connection relationship for two adjacent dense grids.

[0088] In this step, the server can traverse the grids in the target area. When the server obtains a grid, the server first determines whether the grid is a dense grid. When the grid is a sparse grid, the server can obtain the next grid. When the grid is a dense grid, the server can obtain the adjacent grids in the cycle of the grid. Among them, a grid can include four adjacent grids: up, down, left, and right. The server can determine whether the adjacent grids of the grid are dense grids. When there is a dense grid among the adjacent grids, the server can establish a connection relationship for the grid and the adjacent grid.

[0089] In one implementation, the server can mark the connection relationship of each grid in the grid. For example, when the third grid is the currently traversed grid. The server can determine that among the four adjacent grids around the third grid, only the fourth grid located above the third grid is a dense grid. The server can record the connection relationship between the third grid and the fourth grid in the third grid.

[0090] In another implementation, the server can traverse the grids in the target area in the form of a linked list, and establish the connection relationship by adding the grids with connection relationships to the linked list. For example, when the server sequentially traverses to the third grid and determines that the third grid is a dense grid. The server can start the formation of the target cluster from the third grid. The server can determine that among the four adjacent grids around the third grid, the fourth grid above the third grid and the fifth grid on the right side of the third grid are dense grids. The server can point the pointer of the third grid to the fourth grid and the fifth grid. When the server traverses to the fourth grid, the server can connect the fourth grid to the sixth grid on the right side of the fourth grid. When the server traverses to the fifth grid, the server can connect the fifth grid to the sixth grid above the fifth grid. The connection between the fourth grid and the fifth grid and the sixth grid can ensure that any two adjacent dense grids are connected together. The setting of this connection relationship can better update the target cluster and improve the target cluster update efficiency when the obstacles in the grid change.

[0091] During this traversal process, the server can use breadth - first traversal or depth - first traversal to complete the establishment of the connection relationships of each grid in a target cluster. When the server completes the generation of a target cluster, the server can continue to traverse backward from the third grid according to the preset order. When the subsequent grid is a grid that has been added to the target cluster or is a sparse grid, the server can continue to read the next grid. When the server reads a dense grid that has not been added to the target cluster, the server can start establishing the connection relationships of the grids in the new target cluster from this dense grid.

[0092] Step 3: Determine that multiple grids with connection relationships form a target cluster according to the connection relationships of the grids in the target area.

[0093] In this step, after the server completes the establishment of the above - mentioned connection relationship, the server can determine at least one target cluster in the target area according to this connection relationship. The server can display the target cluster and the clustering center of the target cluster through the visualization display unit, so that the user can perform subsequent analysis and decision - making according to the clustering result through the visualization display unit.

[0094] In one example, for obstacle grids, the clustering process is not exactly the same as that of ordinary grids. When the grids in the target area include the first grid with a polygon obstacle and the second grid with a linear obstacle, the specific process of the server performing clustering can include the following steps:

[0095] Step 1: Determine whether the first grid density of the first grid and the sub-grid densities of each sub-grid in the second grid are greater than or equal to the grid density threshold. When the first grid density of the first grid is greater than or equal to the grid density threshold, the first grid is a dense grid. When the sub-grid density of each sub-grid in the second grid is greater than or equal to the grid density threshold, the sub-grid is a dense grid.

[0096] Step 2: When the first grid or the sub-grid is a dense grid, determine the adjacent grids of the first grid or the sub-grid based on the grids in the target area and the first grid or the sub-grid.

[0097] In this step, the first grid may include a polygonal obstacle. The boundary of the polygonal obstacle may partially coincide with the boundary of the first grid. At the partial boundary where the first grid coincides with the obstacle, the first grid does not actually contact the adjacent grid. Therefore, the adjacent grids corresponding to the boundaries of the first grid that do not coincide with the obstacle are the adjacent grids of the first grid. For example, the obstacle of the first grid may include Figure 9 the shaded part in (a). The boundary of the obstacle of the first grid completely coincides with the left boundary of the first grid. The boundary of the obstacle of the first grid also coincides with the left half of the upper and lower boundaries of the first grid. Therefore, the boundaries of the first grid that do not coincide with the obstacle include a part of the upper boundary, a part of the lower boundary, and the entire right boundary of the first grid. The grids above, below, and to the right of the first grid that contact these three boundaries are the adjacent grids having a neighbor relationship with the first grid.

[0098] The second grid may include at least one sub-grid. When the second grid includes one sub-grid, the sub-grid may include four adjacent grids above, below, to the left, and to the right of the first grid. When the second grid includes multiple sub-grids, the linear obstacle may divide the second grid into multiple sub-grids. For example, as shown by the curve in the grid of Figure 9 (b), the second grid may be divided into a first sub-grid and a second sub-grid. Among them, the boundary where the first sub-grid overlaps with the second grid may include the upper boundary, a part of the left boundary, and a part of the right boundary. Therefore, the adjacent grids corresponding to the first sub-grid may include three adjacent grids above, to the left, and to the right of the second grid. The boundary where the second sub-grid overlaps with the second grid may include the lower boundary, a part of the left boundary, and a part of the right boundary. Therefore, the adjacent grids corresponding to the second sub-grid may include three adjacent grids below, to the left, and to the right of the second grid.

[0099] Step 3: When the adjacent grid is a dense grid, establish a connection relationship between the first grid or the sub-grid and the adjacent grid.

[0100] In this step, the server can determine whether the adjacent grids of the first grid or the sub-grid are dense grids. When the adjacent grids of the first grid or the sub-grid are dense grids, the server can establish a connection relationship between the first grid or the sub-grid and the adjacent grids. It should be noted that, among the multiple sub-grids of a second grid, even if all the multiple sub-grids are dense grids, no connection relationship is established among the multiple sub-grids.

[0101] For the clustering method provided in this application, the server can divide the target area into multiple grids, and the data points and obstacles in the target area are distributed within the grids. The server can traverse all the grids in the target area and calculate the first grid density of these grids one by one. For obstacle grids, the server can use the corresponding density calculator to calculate their first grid density. The server can use a preset clustering algorithm to cluster these grids according to the first grid density of each grid in the target area. The server can cluster to obtain multiple target clusters, and each target cluster can include at least one grid. Each target cluster can include a clustering center. In this application, by using different density calculators to calculate the first grid density of obstacle grids, the influence of obstacles on the clustering result is increased during clustering, and the matching degree between the clustering result and the actual requirements is improved. Moreover, in this application, the grid information can also be determined according to the distribution of data points in the target area, realizing the method of clustering using grids, and improving the clustering efficiency and computational complexity.

[0102] Figure 10 The flowchart of a clustering method provided by an embodiment of this application is shown. In Figures 7 to 9 Based on the shown embodiment, after the server completes the clustering of each grid in the target area when the obstacles are in a static state, in this embodiment, the server can also update the clustering result when the obstacles change dynamically, as Figure 10 shown. Taking the server as the execution subject, the method of this embodiment can include the following steps:

[0103] S201. Obtain the changed grids where the obstacles in the target area have changed, and the obstacle information within the changed grids.

[0104] In this embodiment, the server can periodically obtain obstacle information within a target area. This period can be determined according to actual needs. For example, this period can be 1 day, 1 hour, etc. The server can compare the obstacle information in the current period with the obstacle information in the previous period to determine the changed obstacle information. The server can determine the changed grids corresponding to this changed obstacle information. Also, the server can obtain the changed obstacle information in these changed grids. Among them, a changed grid can be a grid that was originally an ordinary grid and an obstacle appears in this grid after the change. A changed grid can also be a grid that was originally an obstacle grid and the information such as the size and position of the obstacle in this obstacle grid has changed after the change. A changed grid can also include a grid that was originally an obstacle grid and the obstacle in this grid disappears after the change.

[0105] S202. Determine the second grid density of the changed grid according to the data point information and the obstacle information within each changed grid.

[0106] In this embodiment, the server can recalculate the second grid density of each changed grid. Among them, the calculation method of the second grid density of this changed grid is the same as Figure 7 the calculation method of the first density of the grid in the shown embodiment, which will not be elaborated in this embodiment.

[0107] S203. Determine the adjacent grids of the changed grid.

[0108] In this embodiment, the server can re - determine the adjacent grids of these changed grids. Or, the server can re - determine the adjacent grids of the sub - grids of these changed grids. The specific method is the same as Figure 7 the method for determining adjacent grids in the shown embodiment, which will not be elaborated in this embodiment.

[0109] S204. Update the connection relationship between the changed grid and the adjacent grids according to the second grid density of the changed grid, the first grid density of the adjacent grids, and a preset clustering algorithm.

[0110] In this embodiment, the server can traverse each changed grid. When the changed grid changes from a dense grid to a sparse grid, the server can delete the connection relationship between this changed grid and its adjacent grids. When the changed grid changes from a sparse grid to a dense grid, the server can add the connection relationship between this dense grid and its adjacent grids. When the grid type of the changed grid does not change, the server does not modify the connection relationship of this changed grid.

[0111] This change in the connection relationship is equivalent to Figure 5 the operations corresponding to the obstacle addition clusterer and the obstacle reduction clusterer in the shown clustering device. Among them, the specific process of the obstacle addition clusterer can be as Figure 11As shown. The specific process of the obstacle reduction clustering device can be as follows Figure 12 As shown. The obstacle update clustering device specifically may include the following steps:

[0112] Step 1: The server saves the changed grids in the target area that change due to the increase of obstacles.

[0113] Step 2: When there are data points in the changed grid, the server calculates the second grid density of the changed grid according to the obstacle information after the change of the changed grid.

[0114] Step 3: The server compares the second grid density of the changed grid with the grid density threshold to determine whether the second grid belongs to a dense grid.

[0115] Step 4: The server determines whether the type of the changed grid has changed. If there is no change, no processing is performed. If the type of the changed grid has changed, its connection relationship is updated. Among them, the situation where the type changes may include changing from a dense grid to a sparse grid and changing from a sparse grid to a dense grid. Step 5: The server updates the connection relationship of the changed grid.

[0116] Step 6: Repeat the above steps until all the changed grids are processed.

[0117] The execution steps of the obstacle reduction clustering device are basically the same as the specific process of the obstacle increase clustering device, except for Step 1. As Figure 12 shown, Step 1 of the obstacle reduction clustering device is used to obtain the changed grids of obstacle reduction.

[0118] S205: Update the target clusters in the target area according to the updated connection relationship.

[0119] In this embodiment, the server can obtain the updated connection relationship. The server can obtain the updated target clusters according to the updated connection relationship. The server can display the updated target clusters and the cluster centers of the target clusters through the visualization display unit, so that the user can perform subsequent analysis and decision-making according to the clustering result through the visualization display unit.

[0120] In the clustering method provided by this application, the server can determine the changed grid corresponding to the changed obstacle information. Moreover, the server can also obtain the changed obstacle information in these changed grids. The server can determine the second grid density of the changed grid according to the data point information and the obstacle information in each changed grid. The server can determine the adjacent grids of the changed grid. The server can update the connection relationship between the changed grid and the adjacent grids according to the second grid density of the changed grid, the first grid density of the adjacent grids, and the preset clustering algorithm. The server can obtain the updated connection relationship, and the server can obtain the updated target cluster according to the updated connection relationship. In this application, by updating the connection relationship of the changed grid, the update of the target cluster is realized, and the update efficiency of the target cluster is improved.

[0121] Figure 13 FIG. shows a schematic structural diagram of a clustering device provided by an embodiment of this application, as Figure 13 shown, the clustering device 10 in this embodiment is used to implement the operations corresponding to the server in any of the above method embodiments. The clustering device 10 in this embodiment includes:

[0122] An obtaining module 11, configured to divide a target area into multiple grids, and data points and obstacles in the target area are distributed in the grids.

[0123] A processing module 12, configured to determine the first grid density of each grid according to the grid information, the data point information, and the obstacle information in each grid, where the grid information includes the grid length and the grid width. Perform clustering calculation on all grids in the target area according to the first grid density of each grid and the preset clustering algorithm to obtain at least one target cluster, and each target cluster includes at least one grid.

[0124] In one example, the obtaining module 11 is specifically configured to determine grid information according to the data point information of all data points in the target area and the step mapping table, where the grid information includes the grid length and the grid width, and the step mapping table is used to indicate the mapping relationship between the second data point information and the grid length and the grid width. Divide the target area into multiple grids according to the grid length and the grid width.

[0125] In one example, the processing module 12 is specifically configured to divide all grids in the target area into dense grids and sparse grids according to the first grid density of each grid and the grid density threshold. Traverse all grids in the target area, and establish a connection relationship for two adjacent dense grids. Determine that multiple grids with a connection relationship form a target cluster according to the connection relationship of the grids in the target area.

[0126] In one example, when the first grid includes polygon obstacles, the processing module 12 is specifically configured to determine the available area of the first grid according to the grid information of the first grid and the obstacle information within the first grid; determine the number of data points of the first grid according to the data point information within the first grid; and determine the first grid density of the first grid according to the number of data points of the first grid and the available area of the first grid.

[0127] In one example, when the second grid includes linear obstacles, the processing module 12 is specifically configured to divide the grid into at least one sub-grid according to the obstacle information and the grid information within the second grid, and determine the sub-grid area of each sub-grid; determine the number of sub-grid data points within each sub-grid according to the data point information of the second grid and the sub-grids of the second grid; and determine the sub-grid density of each sub-grid according to the sub-grid area and the number of sub-grid data points of each sub-grid.

[0128] In one example, the processing module 12 is specifically configured to, when the first grid or the sub-grid is a dense grid, determine the adjacent grids of the first grid or the sub-grid according to the grids within the target area and the first grid or the sub-grid. When the adjacent grid is a dense grid, establish a connection relationship between the first grid or the sub-grid and the adjacent grid.

[0129] In one example, the processing module 12 is further configured to obtain the changed grids where the obstacles within the target area have changed, and the obstacle information within the changed grids; determine the second grid density of the changed grids according to the data point information and the obstacle information within each changed grid; determine the adjacent grids of the changed grids; update the connection relationship between the changed grids and the adjacent grids according to the second grid density of the changed grids, the first grid density of the adjacent grids, and a preset clustering algorithm; and update the target clusters within the target area according to the updated connection relationship.

[0130] The clustering device 10 provided in the embodiments of the present application can execute the above method embodiments. For the specific implementation principle and technical effects, reference can be made to the above method embodiments, and details are not described herein again.

[0131] Figure 14 FIG. shows a schematic hardware structure diagram of a server provided in the embodiments of the present application. As Figure 14 shown, the server 20 is used to implement the operations corresponding to the server in any of the above method embodiments. The server 20 in this embodiment may include: a memory 21, a processor 22, and a communication interface 24.

[0132] Among them, the memory 21 is used to store computer programs. The processor 22 is used to execute the computer programs stored in the memory to implement the clustering method in the above embodiments. For specific reference, please refer to the relevant descriptions in the foregoing method embodiments.

[0133] In one example, the memory 21 can be either independent or integrated with the processor 22. When the memory 21 is a device independent of the processor 22, the server 20 can further include a bus 23. The bus 23 is used to connect the memory 21 and the processor 22.

[0134] In one example, the communication interface 24 can be connected to the processor 21 through the bus 23. The communication interface 24 can be used to output the clustering result. Also, the communication interface 24 can be further used to obtain the data point information and obstacle information within the target area.

[0135] The server provided in this embodiment can be used to execute the above clustering method, and its implementation manner and technical effects are similar, which will not be elaborated here in this embodiment.

[0136] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the methods provided by the above various embodiments.

[0137] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device can read the computer program from the computer-readable storage medium, and the execution of the computer program by at least one processor enables the device to implement the methods provided by the above various embodiments.

[0138] The embodiment of this application also provides a chip, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the device installed with the chip executes the methods in the above various possible embodiments.

[0139] In several embodiments provided by this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or module can be in an electrical, mechanical or other form.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A clustering method, characterized in that, The method includes: Dividing a target area into a plurality of grids, where data points and obstacles within the target area are distributed in the grids; Determining a first grid density of each grid according to the grid information of each grid, the data point information within the grid, and the obstacle information within the grid, where the grid information includes a grid length and a grid width; Performing clustering calculation on all grids within the target area according to the first grid density of each grid and a preset clustering algorithm to obtain at least one target cluster, where each target cluster includes at least one grid; When the first grid includes a polygonal obstacle, the determining a first grid density of each grid according to the grid information of each grid, the data point information within the grid, and the obstacle information within the grid includes: Determining an available area of the first grid according to the grid information of the first grid and the obstacle information within the first grid; Determining the number of data points in the first grid according to the data point information within the first grid; Determining the first grid density of the first grid according to the number of data points in the first grid and the available area of the first grid; When the second grid includes a linear obstacle, the determining a first grid density of each grid according to the grid information of each grid, the data point information within the grid, and the obstacle information within the grid includes: Dividing the grid into at least one sub-grid according to the obstacle information within the second grid and the grid information of the second grid, and determining the sub-grid area of each sub-grid; Determining the number of sub-grid data points within each sub-grid according to the data point information of the second grid and the sub-grids of the second grid; Determining the sub-grid density of each sub-grid according to the sub-grid area and the number of sub-grid data points of each sub-grid.

2. The method according to claim 1, wherein The dividing the target area into a plurality of grids includes: Determining grid information according to the data point information of all data points within the target area and a step size mapping table, where the grid information includes a grid length and a grid width, and the step size mapping table is used to indicate the mapping relationship from the data point information to the grid length and the grid width; Dividing the target area into a plurality of grids according to the grid length and the grid width.

3. The method according to claim 1, wherein The performing clustering calculation on all grids within the target area according to the first grid density of each grid and a preset clustering algorithm to obtain at least one target cluster includes: Dividing all the grids within the target area into dense grids and sparse grids according to the first grid density of each grid and a grid density threshold; Traversing all the grids within the target area and establishing a connection relationship for two adjacent dense grids; Determining that a plurality of grids with a connection relationship form a target cluster according to the connection relationship of the grids within the target area.

4. The method according to claim 3, wherein The traversing all the grids within the target area and establishing a connection relationship for two adjacent dense grids includes: When the first grid or the sub-grid is a dense grid, determine the adjacent grids of the first grid or the sub-grid according to the grids in the target area and the first grid or the sub-grid; When the adjacent grid is a dense grid, establish a connection relationship between the first grid or the sub-grid and the adjacent grid.

5. The method according to any one of claims 1 to 3, characterized in that The method further includes: Obtain the changed grids where obstacles in the target area have changed, and the obstacle information in the changed grids; Determine the second grid density of the changed grid according to the data point information and the obstacle information in each changed grid; Determine the adjacent grids of the changed grid; Update the connection relationship between the changed grid and the adjacent grids according to the second grid density of the changed grid, the first grid density of the adjacent grids, and the preset clustering algorithm; Update the target clusters in the target area according to the updated connection relationship.

6. A server, characterized in that, The server includes: a memory, a processor; The memory is used to store a computer program; the processor is used to implement the clustering method according to any one of claims 1-5 based on the computer program stored in the memory.

7. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it is used to implement the clustering method according to any one of claims 1-5.

8. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the clustering method according to any one of claims 1-5.

Citation Information

Cited By

  • Gynecological immunohistochemical double-staining AI detection system based on computer vision and image processing technology

    CN120629151A