Multi-scale city function classification method based on landmark constraint
Through the multi-scale urban functional classification method based on landmark constraints, the multi-scale image segmentation and deep learning framework are used to model object spatial relationships, which solves the problem of inconsistent with the actual distribution of the existing urban functional area planning map and achieves higher classification accuracy.
Patent Information
- Application Number
- CN202510073171.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing urban functional area planning map is inconsistent with the actual distribution, making it difficult to accurately manage urban functions. Traditional field measurements are costly and time-consuming, and methods based on geographic big data have problems with insufficient data quality and coverage.
A multi-scale urban functional classification method based on landmark constraints is adopted to obtain segmented areas through multi-scale image segmentation, and a deep learning framework is combined with a deep learning framework to model object spatial relationships and classify functional areas to improve the accuracy of classification.
It realizes a more accurate description of urban functional area boundaries, improves the accuracy of urban functional area classification, and makes up for the shortcomings of existing methods that only focus on the internal characteristics of the analysis unit.
Smart Images

Figure CN120014444A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geographic information technology, and in particular to a multi-scale urban function classification method based on landmark constraints. Background Art
[0002] Urbanization is an important factor in measuring the level of social development of a country. Reasonable urban spatial layout helps to promote and improve the research theory of existing urban functional division, and is of great significance to ubiquitous computing and urban planning. Urban functional areas refer to areas that bear specific social and economic functions formed in the process of urban development, including industrial areas, commercial areas, residential areas, etc. Accurately grasping their spatial distribution is of great significance to the sustainable development of cities. The existing urban functional area planning map is inconsistent with the actual distribution of urban functional areas due to the influence of external factors such as residents' living habits and economic development level, and it is difficult to directly use it as a reference for urban functional management.
[0003] In recent years, with the further demand for refined and multi-scale urban functional zoning results in urban development, more and more researchers have begun to pay attention to the refined functional classification of urban spatial units. The traditional field measurement and survey methods have the disadvantages of high labor costs and long time consumption, and it is difficult to conduct a large-scale survey of the distribution of urban functional areas. In recent years, some urban functional area classification methods based on geographic big data, such as points of interest and street view images, have been proposed. Including using the heavy tail interruption method and density analysis to statistically model the points of interest to achieve the classification of functional areas; the top-down semantic information detection model extracts 4 types of typical urban functional areas from the street view. However, these data all have some shortcomings. For example, most of the points of interest data are uploaded by users. Due to the large number of uploads, it is difficult to realize manual verification and its quality cannot be guaranteed. The street view data is not comprehensive, and it is difficult to obtain street view data for internal roads such as urban communities and parks and suburban areas. Therefore, it is urgent to study a new method for fast and accurate classification of urban functional areas. Summary of the invention
[0004] The present invention provides a multi-scale urban function classification method based on landmark constraints. The multi-scale image segmentation method is used to extract the segmented area as the minimum analysis unit, and the object spatial relationship modeling and refined classification of functional areas are realized under the deep learning framework, so as to more accurately describe the boundaries of urban functional areas and improve the accuracy of urban functional area classification.
[0005] The present invention provides a multi-scale urban function classification method based on landmark constraints, comprising:
[0006] Acquire image data of a set range of the city, and perform multi-scale segmentation on the image data using set scale parameters to obtain multiple segmented areas;
[0007] Calculating the frequency density of each segmented area, and determining the functional form of the corresponding segmented area according to the frequency density; wherein the functional form includes a single functional area, a mixed functional area and a no-data area;
[0008] When the functional form is a single functional area, determining the functional type of the segmented area according to the landmarks;
[0009] When the functional form is not a single functional area, a CNN model using ResNet50 as the backbone network of the feature encoder is used to extract internal features of the segmented area, and the internal features of the object are fused with the geographic attributes;
[0010] A multi-layer Transformer model is used to perform spatial relationship modeling and urban functional area classification on the fused internal features of the object.
[0011] Furthermore, the step of calculating the frequency density of each segmented region and determining the functional form of the corresponding segmented region according to the frequency density includes:
[0012] The land use function type, POI type and POI quantity of the segmented area are obtained, the segmented area is divided into spatial grids of a set scale using a fishing net tool, and the frequency density of the segmented area is calculated, and the calculation formula is:
[0013]
[0014] Among them, F i is the frequency density of the number of POIs of type i land use in a unit grid to the total number of POIs of this type of land use; i is the land use type; n i N is the number of POIs of type i land in a single grid; i is the total number of POIs for type i land; C i is the ratio of the frequency density of the i-th type of land use in the unit grid to the total frequency density of all land use types in this segmented area;
[0015] When C i When the value is greater than or equal to the set threshold, the land in this segmented area is considered to be single-function land. i When all values are less than the set threshold, the land use in this segmented area is considered to be mixed-function land. If there is no data in this segmented area, it is considered to be a no-data area.
[0016] Furthermore, when the functional form is a single functional area, the step of determining the functional type of the segmented area according to the landmarks includes:
[0017] Acquire evaluation image data within the segmented area, and extract the building signs that appear most frequently in the evaluation image data;
[0018] The building sign is used as a landmark of the segmented area, and the type of businesses with the most businesses in the landmark is determined as the type of the segmented area.
[0019] Furthermore, when the functional form is not a single functional area, the step of extracting internal features of the segmented area using a CNN model with ResNet50 as the backbone network of the feature encoder and fusing the internal features of the object with the geographic attributes includes:
[0020] Pre-training of internal feature extractor for objects: For K image samples and the corresponding one-hot encoded label y i Urban functional area classification dataset The CNN model consists of four parts. The first part is the 50 convolutional layers in ResNet50 and the associated pooling, normalization, and activation functions, which convert x i Encoded as a size The second part is a point-to-point convolution layer, which is used to compress high-dimensional features to 64 dimensions to reduce the burden of subsequent operations; the third part is a global average pooling layer GAP to aggregate features; the fourth part is a fully connected layer with a SoftMax function to calculate x i The category probability distribution p i ;
[0021] Object internal feature extraction: Object features are extracted by feature image F * and the prior probability P * The two parts are aggregated, and the convolution layer and point-to-point convolution in the previous step are used to extract the feature image. f encodes it into a size of The feature image is then restored to the same size as the input image through bilinear interpolation. At the same time, using g(·|θ g ) * Perform pixel-by-pixel calculations to obtain the prior probability of each pixel in image X corresponding to the category of the urban functional area To help the subsequent classification; therefore, the mixed feature FP of image X is expressed as:
[0022]
[0023] Among them, the urban functional areas are divided into 10 categories, so cls = 10 + 1, where 1 represents the "unlabeled" pixels in the dataset; based on multi-scale segmentation, the object mask of X is obtained and the corresponding set of objects Where K represents the number of objects after multi-scale segmentation of X, and the value range of the mask M pixel is an integer from 1 to K, which represents the number of the object to which the pixel belongs;
[0024] Fusion of object internal features and geographic attributes: Using the coordinates of the object's annotation points (m i , n i ), area s i and the perimeter c i The location encoding of the alternative ViT model is used to model the spatial relationship between objects in urban functional areas, that is, objects o i The geographical attribute characteristics are represented by g i =[m i ,n i ,s i ,c i ] represents; the geometric center coordinates of the object are determined by the following method: In the example, a rectangular coordinate system is established with the geometric center of the object set as the origin, the east direction as the x-axis, and the south direction as the y-axis. i The center coordinates of the annotation (m i , n i ) represents the spatial location of the object, and the geographic information G of the object set O can be expressed as [g1,g2,…g K ], by concatenating the internal features of the object set with the geographic information features, we can obtain the object set features F with location coding:
[0025]
[0026] Among them, D=64+cls+4.
[0027] Furthermore, in the step of using a multi-layer Transformer model to perform spatial relationship modeling and urban functional zone classification on the internal features of the fused object,
[0028] Each layer of the Transformer encoder in the Transformer model structure contains two normalization modules Norm, a multi-head attention module MHA and a multi-layer perceptron module MLP. The normalization module is used to control the feature value range obtained during the calculation process and eliminate the effect of dimension.
[0029] In the first layer of Transformer, for the object set feature F, the Transformer model first performs layer normalization on it to obtain the normalized object set feature Fnorm, and then calculates the relationship weights between different objects through multiple self-attention modules to mine the spatial relationship between different object features. The calculation formula is:
[0030]
[0031] F′=concatente(F′1,F′2,…,F′ n )W F
[0032] Among them, K, Q, and V are three linear mappings of Fnorm, QK T What is calculated is the dot product of the features between two objects. Since this feature contains both the internal features of the object and the geographical attributes of the object, the result can be regarded as the spatial correlation between two objects. The normalized result multiplied by V on the left embeds this relationship into the original object set features. Concatenate means matrix concatenation.
[0033] The result of multiple Transformer encoders is recorded as Z. At the end of the model, Z will be classified through a linear layer to obtain the final category probability distribution of each object.
[0034] Furthermore, the cross entropy loss function is used to optimize the ViT model:
[0035]
[0036] Among them, y i,c If and only if the object o i If it belongs to category c, it is 1, otherwise it is 0; p i,c Represents object o i The predicted probability of belonging to class c.
[0037] Furthermore, after the multi-layer Transformer model is used to perform spatial relationship modeling and urban functional zone classification on the internal features of the fused objects, the method further includes:
[0038] When the segmented area is a mixed functional area, the functional mixing degree of the segmented area is calculated according to the POI data of the segmented area; the formula is:
[0039]
[0040] Among them, H is the mixed degree of land use functions in the unit grid; n is the total number of all POI types contained in this segmented area; P m The ratio of the total amount of POIs of the mth type in this segmented area to the total amount of all POIs in this grid;
[0041] The functional mixing degree is marked and displayed on the urban functional area classification of the divided area.
[0042] The present invention also provides a multi-scale urban function classification device based on landmark constraints, comprising:
[0043] An acquisition module is used to acquire image data within a set range of a city, and perform multi-scale segmentation on the image data using set scale parameters to obtain multiple segmented areas;
[0044] A calculation module, used for calculating the frequency density of each segmented area, and determining the functional form of the corresponding segmented area according to the frequency density; wherein the functional form includes a single functional area, a mixed functional area and a no-data area;
[0045] A determination module, used for determining the functional type of the segmented area according to the landmarks when the functional form is a single functional area;
[0046] An extraction module, for extracting internal features of the object from the segmented area using a CNN model with ResNet50 as the backbone network of the feature encoder when the functional form is not a single functional area, and fusing the internal features of the object with the geographic attributes;
[0047] The classification module is used to use a multi-layer Transformer model to perform spatial relationship modeling and urban functional area classification on the fused internal features of the object.
[0048] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0049] The present invention also provides a computer-readable storage medium on which a computer program is stored, and the computer program implements the steps of the above method when executed by a processor.
[0050] The beneficial effects of the present invention are:
[0051] The present invention adopts a multi-scale image segmentation method to extract the segmented area as the minimum analysis unit to obtain relatively accurate boundaries of urban functional areas. On this basis, the functional form is determined, including single functional areas, mixed functional areas and no-data areas. An object-oriented Transformer model is proposed for mixed functional areas and no-data areas. The object's geographic information is used as the position code to model the spatial relationship between objects, so as to make up for the shortcomings of the existing urban functional area classification methods that often only focus on the internal characteristics of the analysis unit, and realize the object spatial relationship modeling and refined classification of functional areas under the deep learning framework. The boundaries of urban functional areas can be described more accurately. At the same time, the spatial relationship between analysis units can effectively improve the accuracy of urban functional area classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1The figure is a schematic diagram of a method flow according to an embodiment of the present invention.
[0053] Figure 2 FIG. 1 is a schematic diagram of a device structure according to an embodiment of the present invention.
[0054] Figure 3 The figure is a schematic diagram of the internal structure of a computer device according to an embodiment of the present invention.
[0055] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0056] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0057] like Figure 1 As shown, the present invention provides a multi-scale urban function classification method based on landmark constraints, comprising:
[0058] S1. Obtain image data of a set range of the city, and perform multi-scale segmentation on the image data using the set scale parameters to obtain multiple segmented areas. The impact data comes from BING maps. Based on the urban functional area classification system used in existing studies, 10 types of urban functional areas are designed, including commercial, residential, institutional, industrial, transportation, green space, undeveloped, forest land, agricultural land and water bodies. In order to train the model, polygon data of the corresponding area was collected from OpenStreetMap (OSM), and reclassified mainly based on 10 fields such as facility type (Amenity), building type (Building), and land use type (Landuse). Taking "commercial area" as an example, any of the following conditions is met and it is marked as a commercial area: (1) Landuse includes commercial or retail; (2) Building is commercial; (3) Shop is wholesale. Specific reclassification basis and the final number of samples collected (70% as training samples and 30% as test samples).
[0059] S2. Calculate the frequency density of each segmented area, and determine the functional form of the corresponding segmented area according to the frequency density; wherein the functional form includes a single functional area, a mixed functional area and a no-data area.
[0060] POI data is obtained by calling the API interface provided by the AutoNavi open platform, including 23 major categories such as catering services, scenic spots, hotels, banks, etc. The urban functional areas are divided into the above 10 categories. Combined with the nature of POI subcategories, and referring to the reclassification principles of POI data by previous scholars, the POIs are cleaned, deleted and reclassified. In the reclassification process, the subcategory POIs contained in some major categories do not conform to the functional types of the major categories. For example, movies, bars, etc. in sports and leisure services belong to commercial service facilities, not sports functions. At the same time, shared power banks, ATMs, public toilets, etc. have no significant significance for identifying functional areas. Therefore, it is necessary to delete them during the reclassification process to improve the accuracy of functional area identification.
[0061] S201, obtaining the land use function type, POI type and POI quantity of the segmented area, dividing the segmented area into spatial grids of a set scale using a fishing net tool, and calculating the frequency density of the segmented area, the calculation formula of which is:
[0062]
[0063] Among them, F i is the frequency density of the number of POIs of type i land use in a unit grid to the total number of POIs of this type of land use; i is the land use type; n i N is the number of POIs of type i land in a single grid; i is the total number of POIs for type i land; C i is the ratio of the frequency density of the i-th type of land use in the unit grid to the total frequency density of all land use types in this segmented area;
[0064] S202, when C i When the value is greater than or equal to the set threshold, the land in this segmented area is considered to be single-function land. i When all values are less than the set threshold, the land use in this segmented area is considered to be mixed-function land. If there is no data in this segmented area, it is considered to be a no-data area.
[0065] S3. When the functional form is a single functional area, determining the functional type of the segmented area according to the landmarks;
[0066] S301, obtaining evaluation image data within the segmented area, and extracting the building signs that appear most frequently in the evaluation image data;
[0067] S302: Use the building sign as a landmark of the segmented area, and determine the type of businesses with the most businesses in the landmark as the type of the segmented area.
[0068] S4. When the functional form is not a single functional area, a CNN model using ResNet50 as the backbone network of the feature encoder is used to extract internal features of the segmented area, and the internal features of the object are integrated with the geographic attributes.
[0069] S401, object internal feature extractor pre-training: for K image samples and the corresponding one-hot encoded label y i Urban functional area classification dataset The CNN model consists of four parts. The first part is the 50 convolutional layers in ResNet50 and the associated pooling, normalization, and activation functions, which convert x i Encoded as a size The second part is a point-to-point convolution layer, which is used to compress high-dimensional features to 64 dimensions to reduce the burden of subsequent operations; the third part is a global average pooling layer GAP to aggregate features; the fourth part is a fully connected layer with a SoftMax function to calculate x i The category probability distribution p i ; Let the first and second parts be denoted as f(·|θ f ), the third part is denoted as GAP(·), and the fourth part is denoted as g(·|θ g ), θ f and θ g Represent the parameters of the two models respectively. The forward propagation process of the model can be expressed by the following formula:
[0070] p i =g(GAP(f(x i |θ f ))|θ g )
[0071] The cross entropy loss function is used to optimize the feature extractor end-to-end, and its expression is as follows:
[0072]
[0073] Among them, cls represents the total number of categories, p i,c Represents sample x i is the predicted probability of category c.
[0074] S402, object internal feature extraction: object features are extracted from feature image F * and the prior probability P * The two parts are aggregated, and the convolution layer and point-to-point convolution in the previous step are used to extract the feature image. f encodes it into a size of The feature image is then restored to the same size as the input image through bilinear interpolation. At the same time, using g(·|θ g ) * Perform pixel-by-pixel calculations to obtain the prior probability of each pixel in image X corresponding to the category of the urban functional area To help the subsequent classification; therefore, the mixed feature FP of image X is expressed as:
[0075]
[0076] Among them, the urban functional areas are divided into 10 categories, so cls = 10 + 1, where 1 represents the "unlabeled" pixels in the dataset; based on multi-scale segmentation, the object mask of X is obtained and the corresponding set of objects Among them, K represents the number of objects after multi-scale segmentation of X, and the value range of the mask M pixel is an integer from 1 to K, which represents the number of the object to which the pixel belongs. j , pooling the FP according to its object mask, i.e. its internal features It can be expressed as the average pooling value of all its corresponding pixel features:
[0077]
[0078] Among them, the T function is used to determine whether the pixel (x, y) belongs to the object o j The characteristics of the object set O can be expressed as
[0079] S403, object internal features and geographic attributes fusion: using the object's annotation point coordinates (m i , n i ), area s i and the perimeter c i The location encoding of the alternative ViT model is used to model the spatial relationship between objects in urban functional areas, that is, objects o i The geographical attribute characteristics are represented by g i =[m i ,n i ,s i ,c i ] represents; the geometric center coordinates of the object are determined by the following method: In the example, a rectangular coordinate system is established with the geometric center of the object set as the origin, the east direction as the x-axis, and the south direction as the y-axis. i The center coordinates of the annotation (m i , n i) represents the spatial location of the object, and the geographic information G of the object set O can be expressed as [g1,g2,…g K ], by concatenating the internal features of the object set with the geographic information features, the object set features F with location encoding can be obtained:
[0080]
[0081] Among them, D=64+cls+4.
[0082] S5. Use a multi-layer Transformer model to perform spatial relationship modeling and urban functional area classification on the fused internal features of the object.
[0083] Each layer of the Transformer encoder in the Transformer model structure contains two normalization modules Norm, a multi-head attention module MHA and a multi-layer perceptron module MLP. The normalization module is used to control the feature value range obtained during the calculation process and eliminate the effect of dimension.
[0084] In the first layer of Transformer, for the object set feature F, the Transformer model first performs layer normalization on it to obtain the normalized object set feature Fnorm, and then calculates the relationship weights between different objects through multiple self-attention modules to mine the spatial relationship between different object features. The calculation formula is:
[0085]
[0086] F′=concatenate(F′1,F′2,…,F′ n )W F
[0087] Among them, K, Q, and V are three linear mappings of Fnorm, QK T What is calculated is the dot product of the features between each pair of objects. Since this feature contains both the internal features of the object and the geographical attributes of the object, the result can be regarded as the spatial correlation between the two objects. The normalized result multiplied by V on the left embeds this relationship into the original object set features. Concatenate means matrix concatenation, which can obtain richer spatial relationship information.
[0088] In the actual ViT structure, multiple Transformer encoders are often stacked to improve the feature extraction effect. The result of multiple Transformer encoders is recorded as Z. At the end of the model, Z will be classified through a linear layer to obtain the final category probability distribution of each object. The cross entropy loss function is used to optimize the ViT model:
[0089]
[0090] Among them, y i,c If and only if the object o i If it belongs to category c, it is 1, otherwise it is 0; p i,c Represents object o i The predicted probability of belonging to class c.
[0091] S6. When the segmented area is a mixed functional area, the absolute information entropy in information theory is used to calculate the diversity of POI types in the unit space, so as to calculate the functional mixing degree of this space unit; the functional mixing degree of the segmented area is calculated according to the POI data of the segmented area, and the formula is:
[0092]
[0093] Among them, H is the mixed degree of land use functions in the unit grid; n is the total number of all POI types contained in this segmented area; P m The ratio of the total amount of POIs of the mth type in this segmented area to the total amount of all POIs in this grid.
[0094] S7. Marking and displaying the functional mixing degree on the urban functional area classification of the divided area.
[0095] The present invention adopts a multi-scale image segmentation method to extract the segmented area as the minimum analysis unit to obtain relatively accurate boundaries of urban functional areas. On this basis, the functional form is determined, including single functional areas, mixed functional areas and no-data areas. An object-oriented Transformer model is proposed for mixed functional areas and no-data areas. The object's geographic information is used as the position code to model the spatial relationship between objects, so as to make up for the shortcomings of the existing urban functional area classification methods that often only focus on the internal characteristics of the analysis unit, and realize the object spatial relationship modeling and refined classification of functional areas under the deep learning framework. The boundaries of urban functional areas can be described more accurately. At the same time, the spatial relationship between analysis units can effectively improve the accuracy of urban functional area classification.
[0096] like Figure 2 As shown, the present invention also provides a multi-scale urban function classification device based on landmark constraints, comprising:
[0097] The acquisition module 1 is used to acquire image data within a set range of the city, and perform multi-scale segmentation on the image data using set scale parameters to obtain multiple segmented areas;
[0098] Calculation module 2, used to calculate the frequency density of each segmented area, and determine the functional form of the corresponding segmented area according to the frequency density; wherein the functional form includes a single functional area, a mixed functional area and a no-data area;
[0099] A determination module 3, for determining the functional type of the segmented area according to the landmarks when the functional form is a single functional area;
[0100] Extraction module 4, used for extracting internal features of the segmented area by using ResNet50 as the CNN model of the backbone network of the feature encoder when the functional form is not a single functional area, and fusing the internal features of the object with the geographical attributes;
[0101] The classification module 5 is used to use a multi-layer Transformer model to perform spatial relationship modeling and urban functional area classification on the fused internal features of the object.
[0102] In one embodiment, the calculation module 2 includes:
[0103] The acquisition unit is used to obtain the land use function type, POI type and POI quantity of the segmented area, divide the segmented area into spatial grids of a set scale using a fishing net tool, and calculate the frequency density of the segmented area, and the calculation formula is:
[0104]
[0105] Among them, F i is the frequency density of the number of POIs of type i land use in a unit grid to the total number of POIs of this type of land use; i is the land use type; n i N is the number of POIs of type i land in a single grid; i is the total number of POIs for type i land; C i is the ratio of the frequency density of the i-th type of land use in the unit grid to the total frequency density of all land use types in this segmented area;
[0106] Determination unit, used when C i When the value is greater than or equal to the set threshold, the land in this segmented area is considered to be single-function land. i When all values are less than the set threshold, the land use in this segmented area is considered to be mixed-function land. If there is no data in this segmented area, it is considered to be a no-data area.
[0107] In one embodiment, the determination module 3 includes:
[0108] An image data acquisition unit, used to acquire evaluation image data within the segmented area, and extract the building signs that appear most frequently in the evaluation image data;
[0109] The determining unit is used to use the building sign as a landmark of the segmented area, and determine the type of the most businesses in the landmark as the type of the segmented area.
[0110] In one embodiment, the extraction module 4 comprises:
[0111] Training unit, used for pre-training of internal feature extractor of the object: for K image samples and the corresponding one-hot encoded label y i Urban functional area classification dataset The CNN model consists of four parts. The first part is the 50 convolutional layers in ResNet50 and the associated pooling, normalization, and activation functions, which convert x i Encoded as a size The second part is a point-to-point convolution layer, which is used to compress high-dimensional features to 64 dimensions to reduce the burden of subsequent operations; the third part is a global average pooling layer GAP to aggregate features; the fourth part is a fully connected layer with a SoftMax function to calculate x i The category probability distribution p i ;
[0112] Feature extraction unit, used for object internal feature extraction: object features are obtained by feature image F * and the prior probability P * The two parts are aggregated, and the convolution layer and point-to-point convolution in the previous step are used to extract the feature image. f encodes it into a size of The feature image is then restored to the same size as the input image through bilinear interpolation. At the same time, using g(·|θ g ) * Perform pixel-by-pixel calculations to obtain the prior probability of each pixel in image X corresponding to the category of the urban functional area To help the subsequent classification; therefore, the mixed feature FP of image X is expressed as:
[0113]
[0114] Among them, the urban functional areas are divided into 10 categories, so cls = 10 + 1, where 1 represents the "unlabeled" pixels in the dataset; based on multi-scale segmentation, the object mask of X is obtained and the corresponding set of objects Where K represents the number of objects after multi-scale segmentation of X, and the value range of the mask M pixel is an integer from 1 to K, which represents the number of the object to which the pixel belongs;
[0115] Fusion unit, used for the fusion of internal features and geographic attributes of objects: using the coordinates of the object's annotation points (m i , n i ), area s i and the perimeter c i The location encoding of the alternative ViT model is used to model the spatial relationship between objects in urban functional areas, that is, objects o i The geographical attribute characteristics are represented by g i =[m i ,n i ,s i ,c i ] represents; the geometric center coordinates of the object are determined by the following method: In the example, a rectangular coordinate system is established with the geometric center of the object set as the origin, the east direction as the x-axis, and the south direction as the y-axis. i The center coordinates of the annotation (m i , n i ) represents the spatial location of the object, and the geographic information G of the object set O can be expressed as [g1,g2,…g K ], by concatenating the internal features of the object set with the geographic information features, the object set features F with location encoding can be obtained:
[0116]
[0117] Among them, D=64+cls+4.
[0118] In one embodiment, in the classification module 5,
[0119] Each layer of the Transformer encoder in the Transformer model structure contains two normalization modules Norm, a multi-head attention module MHA and a multi-layer perceptron module MLP. The normalization module is used to control the feature value range obtained during the calculation process and eliminate the effect of dimension.
[0120] In the first layer of Transformer, for the object set feature F, the Transformer model first performs layer normalization on it to obtain the normalized object set feature Fnorm, and then calculates the relationship weights between different objects through multiple self-attention modules to mine the spatial relationship between different object features. The calculation formula is:
[0121]
[0122] F′=concatenate(F′1,F′2,…,F′ n )W F
[0123] Among them, K, Q, and V are three linear mappings of Fnorm, QK T What is calculated is the dot product of the features between two objects. Since this feature contains both the internal features of the object and the geographical attributes of the object, the result can be regarded as the spatial correlation between two objects. The normalized result multiplied by V on the left embeds this relationship into the original object set features. Concatenate means matrix concatenation.
[0124] The result of multiple Transformer encoders is recorded as Z. At the end of the model, Z will be classified through a linear layer to obtain the final category probability distribution of each object.
[0125] In one embodiment, in the classification module 5, the cross entropy loss function is used to optimize the ViT model:
[0126]
[0127] Among them, y i,c If and only if the object o i If it belongs to category c, it is 1, otherwise it is 0; p i,c Represents object o i The predicted probability of belonging to class c.
[0128] In one embodiment, it further includes:
[0129] The calculation module is used to calculate the functional mixing degree of the segmented area according to the POI data of the segmented area when the segmented area is a mixed functional area; the formula is:
[0130]
[0131] Among them, H is the mixed degree of land use functions in the unit grid; n is the total number of all POI types contained in this segmented area; P m The ratio of the total amount of POIs of the mth type in this segmented area to the total amount of all POIs in this grid;
[0132] A display module is used to mark and display the functional mixing degree on the urban functional area classification of the divided area.
[0133] The above modules and units are used to execute the corresponding steps in the above multi-scale urban function classification method based on landmark constraints. The specific implementation method thereof is described in the above method embodiment and will not be repeated here.
[0134] like Figure 3 As shown, the present invention also provides a computer device, which can be a server, and its internal structure can be as shown in Figure 3As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all data required for the process of the multi-scale urban function classification method based on landmark constraints. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the multi-scale urban function classification method based on landmark constraints is implemented.
[0135] Those skilled in the art will understand that Figure 3 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.
[0136] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, any of the above-mentioned multi-scale urban function classification methods based on landmark constraints is implemented.
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0138] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0139] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multi-scale urban function classification method based on landmark constraints, characterized in that: include: Acquire image data of a set range of the city, and perform multi-scale segmentation on the image data using set scale parameters to obtain multiple segmented areas; Calculating the frequency density of each segmented area, and determining the functional form of the corresponding segmented area according to the frequency density; wherein the functional form includes a single functional area, a mixed functional area and a no-data area; When the functional form is a single functional area, determining the functional type of the segmented area according to the landmarks; When the functional form is not a single functional area, a CNN model using ResNet50 as the backbone network of the feature encoder is used to extract internal features of the segmented area, and the internal features of the object are fused with the geographic attributes; A multi-layer Transformer model is used to perform spatial relationship modeling and urban functional area classification on the fused internal features of the object.
2. The multi-scale urban function classification method based on landmark constraints according to claim 1 is characterized in that: The step of calculating the frequency density of each segmented region and determining the functional form of the corresponding segmented region according to the frequency density comprises: The land use function type, POI type and POI quantity of the segmented area are obtained, the segmented area is divided into spatial grids of a set scale using a fishing net tool, and the frequency density of the segmented area is calculated, and the calculation formula is: Among them, F i is the frequency density of the number of POIs of type i land use in a unit grid to the total number of POIs of this type of land use; i is the land use type; n i N is the number of POIs of type i land in a single grid; i is the total number of POIs for type i land; C i is the ratio of the frequency density of the i-th type of land use in the unit grid to the total frequency density of all land use types in this segmented area; When C i When the value is greater than or equal to the set threshold, the land in this segmented area is considered to be single-function land. i When all values are less than the set threshold, the land use in this segmented area is considered to be mixed-function land. If there is no data in this segmented area, it is considered to be a no-data area.
3. The multi-scale urban function classification method based on landmark constraints according to claim 1 is characterized in that: When the functional form is a single functional area, the step of determining the functional type of the segmented area according to the landmarks comprises: Acquire evaluation image data within the segmented area, and extract the building signs that appear most frequently in the evaluation image data; The building sign is used as a landmark of the segmented area, and the type of businesses with the most businesses within the landmark is determined as the type of the segmented area.
4. The multi-scale urban function classification method based on landmark constraints according to claim 1 is characterized in that: When the functional form is not a single functional area, the step of extracting internal features of the object from the segmented area using a CNN model with ResNet50 as the backbone network of the feature encoder and fusing the internal features of the object with the geographic attributes includes: Pre-training of internal feature extractor for objects: For K image samples and the corresponding one-hot encoded label y i Urban functional area classification dataset The CNN model consists of four parts. The first part is the 50 convolutional layers in ResNet50 and the associated pooling, normalization, and activation functions, which convert x i Encoded as a size The second part is a point-to-point convolution layer, which is used to compress high-dimensional features to 64 dimensions to reduce the subsequent computing burden; the third part is a global average pooling layer GAP to aggregate features; the fourth part is a fully connected layer with a SoftMax function to calculate x i The category probability distribution p i ; Object internal feature extraction: Object features are extracted by feature image F * and the prior probability P * The two parts are aggregated, and the convolution layer and point-to-point convolution in the previous step are used to extract the feature image. f encodes it into a size of The feature image is then restored to the same size as the input image through bilinear interpolation. At the same time, using g(·|θ g ) * Perform pixel-by-pixel calculations to obtain the prior probability of each pixel in image X corresponding to the category of the urban functional area To help the subsequent classification; therefore, the mixed feature FP of image X is expressed as: Among them, the urban functional areas are divided into 10 categories, so cls = 10 + 1, where 1 represents the "unlabeled" pixels in the dataset; based on multi-scale segmentation, the object mask of X is obtained and the corresponding set of objects Where K represents the number of objects after multi-scale segmentation of X, and the value range of the mask M pixel is an integer from 1 to K, which represents the number of the object to which the pixel belongs; Fusion of object internal features and geographic attributes: Using the coordinates of the object's annotation points (m i , n i ), area s i and the perimeter c i The location encoding of the alternative ViT model is used to model the spatial relationship between objects in urban functional areas, that is, objects o i The geographical attribute characteristics are represented by g i =[m i ,n i ,s i ,c i ] represents; the geometric center coordinates of the object are determined by the following method: In the example, the geometric center of the object set is taken as the origin, the east direction is the x-axis, and the south direction is the y-axis to establish a plane rectangular coordinate system. The center coordinates (m i , n i ) represents the spatial location of the object, and the geographic information G of the object set O can be expressed as [g1,g2,…g K ], by concatenating the internal features of the object set with the geographic information features, we can obtain the object set features F with location coding: Among them, D=64+cls+4.
5. The multi-scale urban function classification method based on landmark constraints according to claim 4 is characterized in that: In the step of using a multi-layer Transformer model to perform spatial relationship modeling and urban functional zone classification on the internal features of the fused object, Each layer of the Transformer encoder in the Transformer model structure contains two normalization modules Norm, a multi-head attention module MHA and a multi-layer perceptron module MLP. The normalization module is used to control the feature value range obtained during the calculation process and eliminate the effect of dimension. In the first layer of Transformer, for the object set feature F, the Transformer model first performs layer normalization on it to obtain the normalized object set feature Fnorm, and then calculates the relationship weights between different objects through multiple self-attention modules to mine the spatial relationship between different object features. The calculation formula is: F′=concatenate(F′1,F′2,…,F′ n )W F Among them, K, Q, and V are three linear mappings of Fnorm, QK T What is calculated is the dot product of the features between two objects. Since this feature contains both the internal features of the object and the geographical attributes of the object, the result can be regarded as the spatial correlation between two objects. The normalized result multiplied by V on the left embeds this relationship into the original object set features. Concatenate means matrix concatenation. The result of multiple Transformer encoders is recorded as Z. At the end of the model, Z will be classified through a linear layer to obtain the final category probability distribution of each object.
6. The multi-scale urban function classification method based on landmark constraints according to claim 5 is characterized in that: The cross entropy loss function is used to optimize the ViT model: Among them, y i,c If and only if the object o i If it belongs to category c, it is 1, otherwise it is 0; p i,c Represents object o i The predicted probability of belonging to class c.
7. The multi-scale urban function classification method based on landmark constraints according to claim 2 is characterized in that: After the multi-layer Transformer model is used to perform spatial relationship modeling and urban functional zone classification on the internal features of the fused objects, the method further includes: When the segmented area is a mixed functional area, the functional mixing degree of the segmented area is calculated according to the POI data of the segmented area; the formula is: Among them, H is the mixed degree of land use functions in the unit grid; n is the total number of all POI types contained in this segmented area; P m The ratio of the total amount of POIs of the mth type in this segmented area to the total amount of all POIs in this grid; The functional mixing degree is marked and displayed on the urban functional area classification of the divided area.
8. A multi-scale urban function classification device based on landmark constraints, characterized in that: include: An acquisition module is used to acquire image data within a set range of a city, and perform multi-scale segmentation on the image data using set scale parameters to obtain multiple segmented areas; A calculation module, used for calculating the frequency density of each segmented area, and determining the functional form of the corresponding segmented area according to the frequency density; wherein the functional form includes a single functional area, a mixed functional area and a no-data area; A determination module, used for determining the functional type of the segmented area according to the landmarks when the functional form is a single functional area; An extraction module, for extracting internal features of the object from the segmented area using a CNN model with ResNet50 as the backbone network of the feature encoder when the functional form is not a single functional area, and fusing the internal features of the object with the geographic attributes; The classification module is used to use a multi-layer Transformer model to perform spatial relationship modeling and urban functional area classification on the fused internal features of the object.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.