Method and device for identifying land use function based on clustering
By using mobile phone signaling data and clustering algorithms, the function of identifying urban land based on travel data is solved, and the problem of low recognition efficiency caused by untimely update of POI data is achieved, real-time and efficient identification of urban land functions is achieved.
Patent Information
- Application Number
- CN202510341683.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
When the existing technology relies on POI data to identify urban land functions, the update is not timely, resulting in the identification results that cannot reflect the dynamic changes in urban land use in real time, reducing the identification efficiency.
The travel data of the target identification area is obtained through mobile phone signaling data, the travel volume proportion sequence is calculated, and the clustering algorithm is used to cluster the grid cells in time and space to identify the land type and functional clustering area.
It realizes timely identification of urban land functions, improves identification efficiency and accuracy, and can better reflect the dynamic changes of urban land.
Smart Images

Figure CN120277283A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly to a method and device for identifying land use functions based on clustering. Background Art
[0002] With the rapid development of cities, accurately identifying urban land use functions is crucial for urban planning, rational allocation of resources, and sustainable development. Traditional methods for identifying urban land use functions face many limitations in the big data era and cannot achieve more efficient identification. Therefore, identifying urban land use functions based on big data has become the focus.
[0003] Most existing methods for identifying urban land use functions based on big data rely on POI (Point of Interest) data. For example, first obtain POI data from relevant map service platforms, then classify the POI data with reference to relevant classification criteria, extract POI data features such as frequency density and type ratio, and finally identify land use functional areas by the ratio of the frequency density of the i-th type of POI to the frequency density of all types of POIs.
[0004] However, since the update of POI data is often not timely enough, the results of land use function identification based on it cannot reflect the dynamic changes of urban land use in real time. For example, some commercial POIs may be idle for a long time due to poor management and other reasons, but they are still marked as commercial use in the POI data. To avoid inaccurate land use information obtained based on POI data, it may be necessary to calibrate the land use information through other channels, and thus spend more time obtaining the identification results of land use functions, reducing the efficiency of identifying land use functions. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a method and device for identifying land use functions based on clustering, and the main purpose is to break through the limitations of the land use function identification approach and provide a method that can obtain land use functions in a timely manner to improve the efficiency of identifying land use functions.
[0006] To achieve the above object, the present invention mainly provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for identifying land use functions based on clustering, and the method includes:
[0008] Obtaining travel data of each grid cell in each time slice of a target identification area based on mobile phone signaling data;
[0009] Calculating a travel volume proportion sequence within each grid cell according to the travel data of each grid cell in each time slice;
[0010] Performing temporal clustering on grid cells by using the travel volume proportion sequence and preset clustering features to obtain multiple land use types and corresponding target grid cells;
[0011] Identifying the target grid cells according to the preset spatial clustering and identification strategy of each land use type to obtain multiple land use functions and the function aggregation areas of each land use function.
[0012] In a second aspect, the present invention provides a land use function identification device based on clustering, and the device includes:
[0013] An acquisition unit, configured to acquire travel data of each grid cell in each time slice of a target identification area based on mobile phone signaling data;
[0014] A calculation unit, configured to calculate the travel volume proportion sequence in each grid cell according to the travel data of each grid cell in each time slice acquired by the acquisition unit;
[0015] A clustering unit, configured to perform temporal clustering on grid cells by using the travel volume proportion sequence obtained by the calculation unit and preset clustering features to obtain multiple land use types and corresponding target grid cells;
[0016] An identification unit, configured to identify the target grid cells according to the preset spatial clustering and identification strategy of each land use type obtained by the clustering unit to obtain multiple land use functions and the function aggregation areas of each land use function.
[0017] In a third aspect, the present invention further provides a computing device, and the computing device includes: at least one processor, and a memory, wherein the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor can execute a land use function identification method based on clustering in the first aspect above.
[0018] In a fourth aspect, the present invention further provides a readable storage medium, and the readable storage medium is used for storing a computer program, wherein when the computer program runs, it controls the device where the storage medium is located to execute a land use function identification method based on clustering in the first aspect above.
[0019] With the above technical solution, a method and device for identifying land use functions based on clustering provided by the present invention obtain the travel data of each grid cell in each time slice by using mobile phone signaling data. Since mobile phone signaling data has the characteristic of real-time acquisition, the obtained travel data is more reliable. Since the population flow in each grid cell within the target recognition area may vary greatly, simply comparing the travel changes in different grid cells based on the number of trips reaching each grid cell is not feasible. The present invention divides a period of time into multiple time slices and calculates the proportion of travel volume in each time slice according to the travel data, avoiding the influence of the difference in population flow in different grid cells. Then, a sequence of travel volume proportions is statistically formed, which can better reflect the travel frequency of each grid cell compared with other grid cells within a preset time. Generally, areas with the same land use function have more similar travel situations. The present invention clusters the grid cells according to the sequence of travel volume proportions of each grid cell and preset clustering features to divide the grid cells with similar travel situations into the same land use type, thereby realizing the identification of land use functions in the overall target recognition area. To identify the concentrated contiguous areas with the same urban land use function formed by the aggregation of grid cells of the same land use type within the target recognition area, since the urban function aggregation areas corresponding to different land use types may have different scale characteristics and spatial distribution characteristics, formulating a preset spatial clustering and recognition strategy for different land use types can more accurately determine the target urban function aggregation area. This solution determines the travel data of each user within the target recognition area based on real-time updated mobile phone signaling data, making the efficiency of time clustering all grid cells using travel data higher, and thus realizing a more efficient identification of different land use functions in the target recognition area. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0021] Figure 1 Schematically shows a flowchart of a method for identifying land use functions based on clustering proposed in an embodiment of the present invention;
[0022] Figure 2 Schematically shows a flowchart of another method for identifying land use functions based on clustering proposed in an embodiment of the present invention;
[0023] Figure 3 Schematically shows a structural diagram of a device for identifying land use functions based on clustering proposed in an embodiment of the present invention;
[0024] Figure 4Schematically shown is a schematic structural diagram of another land use function recognition device based on clustering proposed in an embodiment of the present invention. Detailed implementation manners
[0025] Hereinafter, the exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0026] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those skilled in the art to which the present invention belongs.
[0027] In order to identify urban land use functions using big data, currently, POI data within the research area is usually obtained from relevant map service platforms, including information such as the names, categories, and geographical locations of points of interest. However, usually, the data update frequency of map service platforms is relatively low. The names or categories of some points of interest have actually changed, but the map service platforms have failed to update in a timely manner, resulting in inaccurate POI data obtained. Consequently, the land use functions identified using the POI data are also inaccurate, and more time is required to calibrate the data, making it impossible to identify the land use functions of the target identification area in a timely manner. Therefore, on the basis of wanting to develop a technical method for timely identifying land use functions, the inventor of this case wants to use a type of data that can more realistically reflect residents' travel situations to identify land use functions. Then, the inventor thinks of using users' mobile phone signaling data to reflect users' real-time travel situations, such as the number of trips and travel times. According to the travel situations, it can be reflected whether the user activities in each area are frequent at different times, and further reflect which land use function the area is closer to, so as to achieve more efficient identification of the land use functions of each area within the target identification area.
[0028] For this reason, the inventor of this case proposes a land use function recognition method based on clustering. This method uses mobile phone signaling data to obtain travel data within the target identification area. According to the divided grid cells and time slices, the travel data is converted into a travel volume proportion sequence, and time clustering is performed on grid cells with similar features to form multiple land use types. Then, based on a preset spatial clustering recognition strategy, each grid cell is recognized, and it can more efficiently identify all the land use functions and function aggregation areas within the target identification area. An embodiment of a land use function recognition method based on clustering of the present invention, as Figure 1 shown, at least includes 101 - 104.
[0029] 101. Obtain the travel data of each grid cell in each time slice of the target recognition area based on mobile phone signaling data.
[0030] The target recognition area in this solution is a specific area where land use function recognition needs to be carried out. The target recognition area can be the entire urban area, a specific urban area, an administrative region, a street, etc. in the city. The specific scope depends on the actual application scenario and research purpose. Before obtaining the travel data, it is necessary to first obtain the mobile phone signaling data and perform raster processing on the target recognition area. To obtain the mobile phone signaling data, it can be obtained by cooperating with the local mobile operator to obtain the mobile phone signaling data within a certain time span in the target recognition area. The mobile phone signaling data includes, but is not limited to, the unique identifier of the user, the timestamp when the signaling is generated, and the corresponding location information, etc. The collection frequency of the mobile phone signaling data varies according to the actual usage of the user, and it may be collected once per second or once per minute. To perform raster processing on the target recognition area, according to the geographical scope and analysis accuracy requirements of the target recognition area, the area can be divided into uniformly sized grid cells using Geographic Information System (GIS) technology. The size of the grid cell can be 100 meters × 100 meters or 200 meters × 200 meters, which is not limited here. Each grid cell is assigned a unique grid identifier, and its position on the map is accurately determined to facilitate subsequent recognition of land use functions based on the grid cell.
[0031] To obtain the travel data of each grid cell in each time slice based on the mobile phone signaling data, specifically, the shortest travel distance and the shortest residence duration can be preset first. According to the mobile phone signaling data, each person's continuous trajectory points are segmented into at least one independent traffic travel trajectory, and the arrival location and arrival time of each traffic travel trajectory are identified. Among them, the arrival location is a spatial location. Determine which grid cell's spatial location range the traffic travel trajectory conforms to according to the arrival location of each traffic travel trajectory, and divide the travel data corresponding to the traffic travel trajectory that conforms to the grid cell's spatial location into the travel data of the grid cell. Then, use the arrival time corresponding in the travel data of each grid cell to divide each travel data of each grid cell into the time slice that matches its arrival time. The travel data can include the unique identifier of the traffic travel, the travel arrival timestamp, and the travel arrival spatial location. The period to be recognized in this embodiment is 24 hours of a certain day.
[0032] 102. Calculate the sequence of travel volume ratios within each grid cell according to the travel data of each grid cell in each time slice.
[0033] Before calculating the proportion sequence of travel volume, first, according to the travel data within the period to be recognized, count the total travel volume of each grid cell within the period to be recognized. This total travel volume is the total number of trips reaching each grid cell within the period to be recognized. Then, according to the preset slice duration, divide the period to be recognized into multiple time slices. The preset slice duration can be 1 hour or 2 hours, and is not limited here. According to the travel data, count the travel volume of each grid cell within multiple time slices. This travel volume is the number of trips reaching the grid cell within a specific time slice. In some embodiments, it is also possible to first divide the period to be recognized, and then collect the travel data corresponding to each time slice according to the time slices respectively. First, determine the travel volume of each grid cell within each time slice according to the travel data corresponding to each time slice, and then count the total travel volume of each grid cell within the period to be recognized.
[0034] When calculating the proportion sequence of travel volume in each grid cell, based on a specific grid cell, first calculate the ratio of the travel volume in each of its time slices to the total travel volume, obtain the proportion ratio corresponding to each time slice, and arrange the multiple proportion ratios corresponding to the multiple time slices in the order of the time slices to form the proportion sequence of travel volume corresponding to this grid cell. And so on, calculate and obtain the proportion sequence of travel volume of each grid cell in turn. The grid cell can be denoted as c, and the total number of grid cells divided in the target recognition area is N, that is, grid cell c i (i = 1…N); Divide the period to be recognized at equal intervals with the preset slice duration t, that is, 24 hours, and obtain n time slices, that is, t j (j = 1…n), n = 24 / t; Calculate and obtain that the total travel volume reaching grid cell c j in time slice t i is V ij (i = 1…N, j = 1…n); For each grid cell c i , determine the time sequence of travel volume V i1 , V i2 …, V ij , …V in , and the total travel volume Calculate and obtain that the proportion sequence of travel volume is v i1 , v i2 …, v ij , …v in , where v ij = V ij / V i .
[0035] 103. Use the proportion sequence of travel volume and the preset clustering features to perform time clustering on the grid cells to obtain multiple land use types and the corresponding target grid cells.
[0036] In this step, the preset clustering features are feature indicators that can reflect the travel characteristics of each grid cell. When identifying land use types, the clustering features of grid cells can be geographical location, population density, traffic flow, travel salience, etc. In order to perform clustering based on the proportion of travel volume, clustering features that can reflect travel characteristics need to be selected. In this embodiment, the proportion of travel volume in the morning peak salience, the proportion of travel volume in the evening peak salience, etc. can be selected as the preset clustering features to analyze whether each grid cell has the characteristics of the morning peak or the evening peak, and then judge the corresponding land use type.
[0037] Before performing time clustering using the preset clustering features, the target recognition area can be divided into multiple expected clustering results according to the preset clustering features. For example, when the preset clustering features include the proportion of travel volume in the morning peak salience, the proportion of travel volume in the evening peak salience, and the peak ratio of the morning and evening peaks of the travel volume, the expected clustering results can be divided into morning peak travel type, evening peak travel type, double peak travel type, and no peak travel type. Then, the clustering algorithm is used to identify the target grid cells and land use types corresponding to each expected clustering result. For example, the proportion of evening peak travel in residential land is prominent, so the grid cells of the evening peak travel type are residential land; the proportion of morning peak travel in employment land is prominent, so the grid cells of the morning peak travel type are employment land; the proportion of morning and evening peak travel in mixed residential and employment land is prominent, so the grid cells of the double peak travel type are mixed residential and employment land; there is no obvious travel peak in public service land, so the grid cells of the no peak travel type are public service land.
[0038] When clustering grid cells, the hierarchical clustering algorithm or the K-means clustering algorithm can be used, as long as the clustering of grid cells can be achieved. No further limitation is imposed on the clustering algorithm here.
[0039] 104. According to the preset spatial clustering recognition strategy of each land use type, identify the target grid cells to obtain multiple land use functions and the function aggregation areas of each land use function.
[0040] In this embodiment, in order to identify the urban functional agglomeration areas formed by the aggregation of grid cells of different land use types in the target recognition area, according to the preset spatial clustering recognition strategy for each land use type, further spatial clustering processing is performed on all grid cells to identify all grid cells in the target recognition area as multiple land use functions, and the grid cells with the same land use function are divided into multiple clusters and unclustered independent grid cells. Each cluster represents a functional agglomeration area. Since each grid cell in this solution has a specific spatial position, the functional agglomeration area composed of grid cells also has its specific spatial position. Among them, when performing spatial clustering on all grid cells, it can be based on the land use type as a unit to sequentially identify multiple target grid cells corresponding to each land use type; it can also be to bind the relationship between all grid cells in the target recognition area and the land use type, and sequentially identify all grid cells according to the spatial position. Before identification, based on the corresponding relationship between the current grid cell and the land use type, the corresponding preset spatial clustering recognition strategy is retrieved.
[0041] The spatial clustering in this step can determine the preset spatial clustering recognition strategy according to different clustering algorithms. For example, when identifying according to the DBSCAB algorithm, the scale of the urban functional agglomeration area needs to be considered to set the minimum number of points included in DBCAN and the neighborhood radius. In order to more intuitively reflect the spatial position of each functional agglomeration area, the spatial clustering result in this solution can be displayed in the form of an image. If the target grid cells are identified in units of land use types, multiple functional agglomeration area images can be obtained. Each functional agglomeration area image corresponds to a land use function, and each functional agglomeration area image includes multiple functional agglomeration areas; if all grid cells are sequentially identified according to the spatial position, a functional agglomeration area image can be generated. This functional agglomeration area image includes multiple land use functions and multiple functional agglomeration areas, and the functional agglomeration areas with the same land use function can be reflected by the same color or line.
[0042] Based on the above Figure 1 implementation method, it can be seen that through the rich data source of mobile phone signaling data, combined with the method of dividing the target recognition area into grid cells, it is possible to more comprehensively and timely capture the travel information at different locations in the area. By dividing a day into multiple time slices and calculating the sequence of the proportion of travel volume of each grid cell in different time slices, the change law of the travel activity of each grid cell at different time periods in a day can be clearly shown. Using the sequence of the proportion of travel volume to perform time clustering analysis on each grid cell can effectively classify the grid cells based on travel characteristics, so as to identify different land use types. Through the preset spatial clustering recognition strategy, spatial clustering is performed on the grid cells of each land use type to more efficiently identify the urban land use functions, and the corresponding relationship between each land use function and the geographical location is more clearly displayed in the form of functional agglomeration areas.
[0043] In some embodiments of the above embodiments, according to the Figure 1 embodiment of the present invention shown above, the embodiment of the present invention will provide a more detailed description of the steps for analyzing land use types and identifying land use functions, as Figure 2 shown, including at least 201-204.
[0044] 201. Obtain the travel data of each grid cell in each time slice of the target recognition area based on mobile phone signaling data.
[0045] In this embodiment, first obtain the mobile phone signaling data corresponding to the period to be recognized, and then divide the target recognition area into multiple grid cells according to a preset side length. Each grid cell can reflect a specific spatial position within the target recognition area. The land use function density characteristics of the target recognition area can be analyzed using Geographic Information System (GIS) data, and the target recognition area can be divided into multiple grid cells with relatively uniform sizes. In this embodiment, the target recognition area is divided into 7556 grid cells, and the size of each grid cell is 500 meters × 500 meters. Select 0:00-24:00 of the previous day as the period to be recognized, and divide the 24 hours of a day into multiple time slices according to a preset slice duration. The preset slice duration in this embodiment is 1 hour, and the day is divided into 24 time slices. Finally, taking each grid cell in each time slice as the acquisition unit, at least one user travel data of each user is extracted based on the preset travel distance and preset residence duration from the mobile phone signaling data, and a unique identifier is assigned. The arrival position and arrival time of at least one user travel data are identified. Using the unique identifier of the user travel data and the arrival position and arrival time, each user travel data is associated with the corresponding grid cell and time slice; based on the target recognition area, all the user travel data associated with the same time slice and the same grid cell are combined to form the travel data of each grid cell in each time slice of the target recognition area. Among them, the preset travel distance is the shortest travel distance that can be regarded as each traffic trip. If the travel distance is less than this preset travel distance, it is not regarded as a traffic trip, that is, the corresponding travel data is not acquired; the preset residence duration can indicate whether the current traffic trip has ended. If the residence duration of the current traffic trip is less than this preset residence duration, the current traffic trip has not ended. If the residence duration of the current traffic trip is greater than or equal to this preset residence duration, the current traffic trip has ended, that is, a set of user travel data can be extracted based on the current traffic trip.
[0046] 202. Calculate the sequence of travel volume ratios of each grid cell.
[0047] Count the number of arriving trips in each grid cell within each time slice according to the unique identifier corresponding to each group of trip data, which is the first trip volume in each grid cell within each time slice. For each grid cell, accumulate the trip volumes within all time slices to obtain the total daily trip volume of each grid cell. Calculate the ratio of the corresponding first trip volume in each time slice within each grid cell to the total daily trip volume of the grid cell to obtain the proportion of the trip volume in each time slice within the grid cell. Sort the proportions of the trip volumes in each time slice in chronological order to form a sequence of trip volume proportions.
[0048] 203. According to the eigenvalue of multiple preset clustering features corresponding to each grid cell, use the K-means clustering algorithm and the expected clustering result to perform time clustering on all grid cells.
[0049] In this embodiment, the land use function characteristics of each grid cell are highlighted according to the morning peak, evening peak, and the peak ratio of morning and evening peaks. Therefore, the preset clustering features in this step include: the significance of the morning peak proportion of trip volume, the significance of the evening peak proportion of trip volume, and the peak ratio of morning and evening peak proportions of trip volume. Due to the differences in residents' commuting times, in some embodiments, the preset morning peak period is selected as 6:00 - 10:00, the preset evening peak period is selected as 16:00 - 20:00, and the preset off-peak period is selected as 11:00 - 15:00 based on historical trip data. Also considering that most land uses have different degrees of mixed functions, when determining the land use type, it is determined not only by the significant degree of the proportion of trip volume during the morning and evening peak periods, but also by the peak ratio of the morning and evening peaks. Specifically, first calculate the eigenvalue of multiple preset clustering features corresponding to each grid cell according to the sequence of the proportion of trip volume in each grid cell; then, according to the eigenvalue of multiple preset clustering features corresponding to each grid cell, use the K-means clustering algorithm and the expected clustering result to perform time clustering on all grid cells. The expected clustering results include morning peak trip class, evening peak trip class, double-peak trip class, and no-peak class; finally, determine multiple land use types and target grid cells based on the expected clustering results.
[0050] When calculating the eigenvalue of multiple preset clustering features corresponding to each grid cell, first determine the first ratio belonging to the preset morning peak period, the second ratio belonging to the preset evening peak period, and the third ratio belonging to the preset off-peak period in the travel volume ratio sequence. Among them, when determining the first ratio and the second ratio, it can be the travel volume ratio of the highest peak 1 hour selected from the preset morning peak period and the preset evening peak period, or the average travel volume ratio calculated in the preset morning peak period and the preset evening peak period; when determining the third ratio, calculate the average value according to all the travel volume ratios within the preset off-peak period to obtain the third ratio. Then, according to the first ratio, the second ratio, and the third ratio, calculate the eigenvalue of multiple preset clustering features corresponding to each grid cell, specifically including: calculating the first eigenvalue of the morning peak significance of the travel volume ratio of each grid cell according to the first ratio and the third ratio; calculating the second eigenvalue of the evening peak significance of the travel volume ratio of each grid cell according to the second ratio and the third ratio; calculating the third eigenvalue of the peak ratio of the morning and evening peak travel volume ratios of each grid cell according to the first ratio and the second ratio. The calculation formulas are as follows:
[0051] f(v im ,v iu )=v im -v iu ;
[0052] f(v ie ,v iu )=v ie -v iu ;
[0053] f(v im ,v ie )=v im / v ie ;
[0054] Wherein, f(v im ,v iu ) is the first eigenvalue of the morning peak significance of the travel volume ratio of grid cell c i , v im is the first ratio, v iu is the third ratio, f(v ie ,v iu ) is the second eigenvalue of the evening peak significance of the travel volume ratio of grid cell c i , v ie is the second ratio, f(v im ,v ie ) is the third eigenvalue of the peak ratio of the morning and evening peak travel volume ratios of grid cell c i .
[0055] After the time clustering is completed, multiple land use types and target grid cells are determined based on the expected clustering results, which specifically include: determining the grid cells with the expected clustering result of the morning peak travel category as the employment land use type with the morning peak arrival travel characteristics; determining the grid cells with the expected clustering result of the evening peak travel category as the residential land use type with the evening peak arrival travel characteristics; determining the grid cells with the expected clustering result of the bimodal travel category as the mixed employment and residential land use type with the bimodal arrival travel characteristics; determining the grid cells with the expected clustering result of the no-peak category as the public service land use type with the no-peak arrival travel characteristics. When determining the land use type of each grid cell, a type identifier that can reflect its land use type is assigned to it, and multiple target grid cells with the same land use type have the same type identifier.
[0056] 204. Based on the correspondence between the land use type and the target grid cell, according to the preset spatial clustering recognition strategy, the DBSCAN spatial clustering algorithm is used to identify the target grid cells of multiple land use types, and multiple land use functions corresponding to the multiple land use types are obtained, and the target grid cells of each land use function are divided into multiple function aggregation areas.
[0057] In this embodiment, the DBSCAN spatial clustering algorithm is used to identify each target grid cell corresponding to each land use type and reflect its spatial position in the form of an image. Therefore, it is necessary to determine the preset spatial clustering recognition strategy according to the spatial scale characteristics of different land use types. For example, when the target recognition area is Beijing, considering the scale characteristics of the functional areas of Beijing, a megacity, the focus is on identifying the function aggregation areas with a certain scale and volume above 3 km², and the minimum number of points MinPts included in DBSCAN is calculated to be 12, and the neighborhood radius Eps is 0.71. Among them, the minimum number of points included in DBSCAN and the neighborhood radius are the recognition parameters in the preset spatial clustering recognition strategy.
[0058] According to the preset spatial clustering recognition strategy, based on the DBSCAN clustering algorithm, identify the urban functional agglomeration areas corresponding to all target grid cells of each land use type. First, determine multiple land use functions according to the land use type, with each land use type corresponding to one land use function. Then, map the target grid cells corresponding to the land use type onto the land use function to construct the corresponding relationship between the land use function and the target grid cells, identify the target grid cells corresponding to each land use function, divide the grid cells of each land use function into multiple clusters and unclustered independent grid cells, with each cluster representing a functional agglomeration area, so as to achieve the purpose of identifying multiple functional agglomeration areas for each land use function, and display the results in the form of an image. The functional agglomeration area image can reflect the spatial location of the target recognition area. Taking Beijing as the target recognition area as an example, according to the preset spatial clustering recognition strategy, based on the DBSCAN spatial clustering algorithm, the target grid cells corresponding to employment land are clustered into 11 urban employment functional agglomeration areas, and the spatial locations included in the generated functional agglomeration area image respectively correspond to CBD, Financial Street, Zhongguancun, Yizhuang, Wangjing, Shangdi, Fengtai Science and Technology Park, etc.; the target grid cells corresponding to residential land are clustered into 65 urban residential functional agglomeration areas; the target grid cells corresponding to mixed employment and residence land are clustered into 90 urban mixed employment and residence functional agglomeration areas.
[0059] Based on the above Figure 2 implementation method, it can be seen that this solution not only takes into account the specific travel values during the morning and evening rush hours, but also considers the different densities in different land use areas. Although the travel volume is low, it may still present the characteristics of morning rush hour travel. Therefore, cluster the land use types of grid cells more accurately according to the proportion of travel volume. On the basis of determining the land use type and the corresponding target grid cells, the functional areas agglomerated by each land use type in the target area are also identified to identify more practically applicable urban functional agglomeration areas.
[0060] Furthermore, as an implementation of the above Figure 1 , 2 shown method embodiments, the embodiment of the present invention provides a land use function recognition device based on clustering, which is used to recognize land use functions based on mobile phone signaling data. The embodiment of this device corresponds to the foregoing method embodiments. For the convenience of reading, the details of the foregoing method embodiments will not be described one by one in this embodiment, but it should be clear that the device in this embodiment can correspondingly implement all the contents of the foregoing method embodiments. Specifically as Figure 3 shown, this device includes:
[0061] An acquisition unit 31, configured to acquire travel data of each grid cell in each time slice of the target recognition area based on mobile phone signaling data;
[0062] A calculation unit 32, configured to calculate a travel volume proportion sequence within each grid cell according to the travel data of each grid cell within each time slice obtained by the acquisition unit 31;
[0063] A clustering unit 33, configured to perform time clustering on grid cells by using the travel volume proportion sequence obtained by the calculation unit 32 and a preset clustering feature, so as to obtain multiple land use types and corresponding target grid cells;
[0064] An identification unit 34, configured to identify target grid cells according to a preset spatial clustering identification strategy for each land use type obtained by the clustering unit 33, so as to obtain multiple land use functions and a function aggregation area for each land use function.
[0065] Further, as Figure 4 shown, the acquisition unit 31 includes:
[0066] A division module 311, configured to divide a target recognition area into multiple grid cells according to a preset side length, and the grid cells can reflect specific spatial positions within the target recognition area;
[0067] A slicing module 312, configured to slice 24 hours of a day into multiple time slices according to a preset slice duration;
[0068] An extraction module 313, configured to extract at least one set of user travel data of each user based on mobile phone signaling data according to a preset travel distance and a preset residence duration, and assign a unique identifier;
[0069] An identification module 314, configured to identify the arrival location and arrival time of at least one set of user travel data extracted by the extraction module 313, and the arrival location is a spatial position;
[0070] An association module 315, configured to associate each set of user travel data with the corresponding grid cell and time slice by using the unique identifier of the user travel data obtained by the extraction module 313 and the arrival location and arrival time obtained by the identification module 314;
[0071] A combination module 316, configured to combine all the user travel data associated with the same time slice and the same grid cell obtained by the association module 315 based on the target recognition area, so as to form the travel data of each grid cell within each time slice of the target recognition area.
[0072] Further, as Figure 4 shown, the calculation unit 32 includes:
[0073] A calculation module 321, configured to calculate a first travel volume of each grid cell within each time slice and the total travel volume of each grid cell in a day according to the travel data of each grid cell within each time slice;
[0074] The calculation module 321 is further configured to calculate, based on each grid cell, the proportion of the first travel volume in the total travel volume within each time slice.
[0075] The sorting module 322 is configured to sort the travel volume proportions obtained by the calculation module 321 in chronological order to form a travel volume proportion sequence.
[0076] Furthermore, as Figure 4 shown, the clustering unit 33 includes:
[0077] The calculation module 331 calculates the eigenvalue of multiple preset clustering features corresponding to each grid cell according to the travel volume proportion sequence of each grid cell.
[0078] The clustering module 332 is configured to perform time clustering on all grid cells according to the eigenvalues of multiple preset clustering features corresponding to each grid cell obtained by the calculation module 331, using the K-means clustering algorithm and the expected clustering results. The expected clustering results include morning peak travel class, evening peak travel class, double-peak travel class, and non-peak class.
[0079] The determination module 333 is configured to determine multiple land use types and target grid cells based on the expected clustering results obtained by the clustering module 332.
[0080] Furthermore, as Figure 4 shown, the calculation module 331 includes:
[0081] The determination sub-module 3311 determines the first proportion belonging to the preset morning peak period, the second proportion belonging to the preset evening peak period, and the third proportion belonging to the preset flat peak period in the travel volume proportion sequence.
[0082] The first calculation sub-module 3312 is configured to calculate the first eigenvalue of the morning peak significance of the travel volume proportion of each grid cell according to the first proportion and the third proportion determined by the determination sub-module 3311.
[0083] The second calculation sub-module 3313 is configured to calculate the second eigenvalue of the evening peak significance of the travel volume proportion of each grid cell according to the second proportion and the third proportion determined by the determination sub-module 3311.
[0084] The third calculation sub-module 3314 is configured to calculate the third eigenvalue of the peak ratio of the morning and evening peak travel volume proportions of each grid cell according to the first proportion and the second proportion determined by the determination sub-module 3311.
[0085] Furthermore, as Figure 4 shown, the determination module 333 includes:
[0086] The first determination sub-module 3331 is configured to determine grid cells with an expected clustering result of morning peak travel type, which are employment land types with morning peak arrival travel characteristics;
[0087] The second determination sub-module 3332 is configured to determine grid cells with an expected clustering result of evening peak travel type, which are residential land types with evening peak arrival travel characteristics;
[0088] The third determination sub-module 3333 is configured to determine grid cells with an expected clustering result of bimodal travel type, which are mixed employment and residential land types with bimodal arrival travel characteristics;
[0089] The fourth determination sub-module 3334 is configured to determine grid cells with an expected clustering result of no peak type, which are public service land types with no peak arrival travel characteristics.
[0090] Further, as Figure 4 shown, the recognition unit 34 includes:
[0091] A determination module 341, configured to determine a preset spatial clustering recognition strategy for each land use type according to the spatial scale characteristics of the urban functional aggregation area;
[0092] A recognition module 342, configured to, based on the correspondence between the land use type and the target grid cell, according to the preset spatial clustering recognition strategy determined by the determination module 341, use the DBSCAN spatial clustering algorithm to recognize the target grid cells of multiple land use types, obtain multiple land use functions corresponding one by one to the multiple land use types, and divide the target grid cells of each land use function into multiple functional aggregation areas.
[0093] Further, an embodiment of the present invention further provides a computing device, where the computing device includes: at least one processor, and a memory, where the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor can execute the one based on clustering as described above Figure 1 、 2 land use function recognition method.
[0094] Further, an embodiment of the present invention further provides a readable storage medium, where the readable storage medium is used to store a computer program, where when the computer program runs, it controls the device where the storage medium is located to execute the one based on clustering as described above Figure 1 、 2 land use function recognition method.
[0095] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0096] It can be understood that the relevant features in the above methods and devices can be referred to each other. In addition, the "first", "second", etc. in the above embodiments are used to distinguish each embodiment, and do not represent the advantages or disadvantages of each embodiment.
[0097] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0098] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The structure required to construct such systems will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is for the purpose of disclosing the best mode of the present invention.
[0099] In addition, the memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.
[0100] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to generate a computer-implemented process, thereby providing steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0104] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0105] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0106] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0107] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0108] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0109] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for identifying land use functions based on clustering, characterized in that, The method includes: Obtaining travel data of each grid cell in each time slice of the target recognition area based on mobile phone signaling data; Calculating a travel volume proportion sequence within each grid cell according to the travel data of each grid cell in each time slice; Performing time clustering on the grid cells by using the travel volume proportion sequence and preset clustering features to obtain multiple land use types and corresponding target grid cells; Identifying the target grid cells according to the preset spatial clustering recognition strategy for each land use type to obtain multiple land use functions and function aggregation areas for each land use function.
2. The method according to claim 1, characterized in that, Obtaining travel data of each grid cell in each time slice of the target recognition area based on mobile phone signaling data, including: Dividing the target recognition area into multiple grid cells according to a preset side length, and the grid cells can reflect specific spatial positions within the target recognition area; Dividing 24 hours of a day into multiple time slices according to a preset slice duration; Extracting at least one set of user travel data for each user based on the mobile phone signaling data according to a preset travel distance and a preset stay duration, and assigning a unique identifier; Identifying the arrival location and arrival time of the at least one set of user travel data, where the arrival location is a spatial position; Associating each set of user travel data with the corresponding grid cell and time slice by using the unique identifier of the user travel data and the arrival location and arrival time; Based on the target recognition area, combining all user travel data associated with the same time slice and the same grid cell to form travel data of each grid cell in each time slice of the target recognition area.
3. The method according to claim 2, wherein The calculating the travel volume proportion sequence within each grid cell according to multiple time slices and the travel data includes: Calculating a first travel volume of each grid cell in each time slice and the total travel volume of each grid cell in a day according to the travel data of each grid cell in each time slice; Calculating a travel volume proportion of the first travel volume in each time slice to the total travel volume based on each grid cell; Sorting the travel volume proportions in chronological order to form the travel volume proportion sequence.
4. The method according to claim 1, wherein The preset clustering features include the morning peak significance of the travel volume proportion, the evening peak significance of the travel volume proportion, and the peak ratio of the morning and evening peaks of the travel volume proportion. Performing time clustering on the grid cells by using the travel volume proportion sequence and the preset clustering features to obtain multiple land use types and corresponding target grid cells, including: Calculating feature values of multiple preset clustering features corresponding to each grid cell according to the travel volume proportion sequence of each grid cell; Performing the time clustering on all grid cells by using the K-means clustering algorithm and an expected clustering result according to the feature values of multiple preset clustering features corresponding to each grid cell, where the expected clustering result includes a morning peak travel class, an evening peak travel class, a double-peak travel class, and a no-peak class; Determining the multiple land use types and the target grid cells based on the expected clustering result.
5. The method according to claim 4, characterized in that, Calculating the eigenvalues of multiple preset clustering features corresponding to each grid cell according to the travel volume proportion sequence of each grid cell, including: Determining a first proportion belonging to a preset morning peak period, a second proportion belonging to a preset evening peak period, and a third proportion belonging to a preset flat peak period in the travel volume proportion sequence; Calculating a first eigenvalue of the morning peak significance of the travel volume proportion of each grid cell according to the first proportion and the third proportion; Calculating a second eigenvalue of the evening peak significance of the travel volume proportion of each grid cell according to the second proportion and the third proportion; Calculating a third eigenvalue of the peak ratio of the morning and evening peak travel volume proportions of each grid cell according to the first proportion and the second proportion.
6. The method according to claim 4, wherein Determining the multiple land use types and the target grid cells based on the expected clustering result, including: Determining the grid cells with the expected clustering result of the morning peak travel type as the employment land use type with the morning peak arrival travel characteristics; Determining the grid cells with the expected clustering result of the evening peak travel type as the residential land use type with the evening peak arrival travel characteristics; Determining the grid cells with the expected clustering result of the double peak travel type as the mixed employment and residential land use type with the double peak arrival travel characteristics; Determining the grid cells with the expected clustering result of the no peak type as the public service land use type with the no peak arrival travel characteristics.
7. The method according to claim 1, characterized in that, Identifying the target grid cells according to the preset spatial clustering identification strategy of each land use type, and obtaining multiple land use functions and the function aggregation areas of each land use function, including: Determining the preset spatial clustering identification strategy of each land use type according to the spatial scale characteristics of the urban function aggregation area; Based on the correspondence between the land use type and the target grid cell, according to the preset spatial clustering identification strategy, using the DBSCAN spatial clustering algorithm to identify the target grid cells of multiple land use types, obtaining multiple land use functions corresponding one by one to the multiple land use types, and dividing the target grid cells of each land use function into multiple function aggregation areas.
8. A land use function recognition device based on clustering, characterized in that The device includes: An acquisition unit for acquiring the travel data of each grid cell in each time slice of the target recognition area based on mobile phone signaling data; A calculation unit for calculating the travel volume proportion sequence in each grid cell according to the travel data of each grid cell in each time slice acquired by the acquisition unit; A clustering unit for performing time clustering on the grid cells by using the travel volume proportion sequence obtained by the calculation unit and preset clustering features, to obtain multiple land use types and corresponding target grid cells; An identification unit for identifying the target grid cells according to the preset spatial clustering identification strategy of each land use type obtained by the clustering unit, and obtaining multiple land use functions and the function aggregation areas of each land use function.
9. A computing device, characterized in that, The computing device includes: at least one processor, and a memory, wherein the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor can execute a method for identifying land use functions based on clustering as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium is used to store a computer program, wherein when the computer program runs, it controls the device where the storage medium is located to execute a clustering-based land use function recognition method according to any one of claims 1-7.