Urban land utilization identification method based on BERT model text classification algorithm

By using a text classification algorithm based on the BERT model and a POI data cleaning strategy, the problems of accuracy and cross-city adaptability in urban land use identification in existing technologies are solved, and high-precision, adaptive urban land use identification is achieved.

CN122020361APending Publication Date: 2026-05-12TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2025-12-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies for identifying urban land use using POI data suffer from problems such as inappropriate data selection principles, inaccurate identification unit division, lack of context sensitivity in model algorithms, inaccurate translation of identification results, and inability to adapt to mixed land use, resulting in low identification accuracy and a lack of cross-city adaptability.

Method used

A text classification algorithm based on the BERT model was adopted. Urban land use identification units were divided into 50×50 meter grids. Combined with POI data cleaning strategy and data augmentation technology, an urban land use identification model was established. The context capture capability and multi-layer semantic understanding of the BERT model were used to achieve high-precision identification.

Benefits of technology

It significantly improves the accuracy and precision of urban land use identification, can adapt to urban spatial structure, support land use compatibility and mixed policies, has adaptive optimization capabilities, and achieves efficient and accurate urban land use identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020361A_ABST
    Figure CN122020361A_ABST
Patent Text Reader

Abstract

The invention relates to an urban land utilization recognition method based on a BERT model text classification algorithm, and the method is characterized in that the method comprises the following steps: 1, determining an urban land utilization recognition type, determining an urban land utilization recognition region range, obtaining POI data in the region range, and carrying out the data screening; step 2, establishing an urban land utilization identification unit in an urban land utilization identification model based on a BERT model text classification algorithm, and associating the screened POI data with the urban land utilization identification unit; step 3, carrying out geographic text mining on the POI data; and step 4, performing urban land utilization identification of the urban land utilization identification model based on the BERT model text classification algorithm to obtain a high-precision urban land utilization identification result. According to the method, high-precision urban land utilization identification can be timely and accurately carried out by utilizing the available POI data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of urban land use technology, and relates to an urban land use identification method, particularly an urban land use identification method based on the BERT model text classification algorithm. Background Technology

[0002] Urban land use is the result of the combined effects of various urban elements and human activities, playing a crucial role in urban planning, design, activity guidance, and management. However, in practice, urban land use planning is often based on previous versions, relying on top-down, experience-based adjustments for optimization, lacking a bottom-up feedback mechanism from grassroots practice and social needs. This limits the scientific rigor and accuracy of the planning. Furthermore, because urban land use planning requires dynamic adjustments, related data updates are often delayed, and information frequently fails to reflect the current situation in a timely manner. Therefore, accurately and promptly obtaining high-precision urban land use information to support government decision-making and management has become one of the core tasks of current urban spatial planning.

[0003] With the rapid development of information technology, diversified social perception data effectively compensates for the shortcomings of traditional remote sensing technology and high-resolution satellite imagery in land cover identification, which cannot obtain socio-economic information caused by human activities. In particular, POI (Point of Interest) data has significant advantages such as high accuracy, wide coverage, rapid updates, large data volume, and ease of acquisition. POI data describes geographic entities through rich semantic information such as coordinates, addresses, names, and categories, providing strong support for the generation of urban social activity maps and finding widespread application in urban land use identification.

[0004] Current technologies for urban land use identification based on POI data typically involve several key steps: First, identification units are divided by defining grids with sides of 200-1000 meters or traffic analysis zones based on road networks. This division method directly affects the accuracy of urban land use identification. Second, natural language processing models such as Word2Vec, Place2Vec, and GeoSemantic2Vec are used to extract geographic text information from the POI data. Finally, by comparing with the latest urban land use planning maps and manually labeling land use categories, training samples are generated to establish the relationship between geographic text information in the POI data and urban land use. Then, machine learning algorithms such as Support Vector Machine, Random Forest, and XGBoost (Extreme Gradient Boosting) are used to finally output the identification results of urban land use.

[0005] The shortcomings and deficiencies of existing technologies in using POI data to identify urban land use are specifically manifested in the following aspects: (1) Data screening principle problem: In the process of using POI data to identify urban land use, the existing technology often adopts a unified identification path for the eight types of urban construction land as specified in the "Classification of Urban Land Use and Standard for Planning and Construction Land" (GB50137-2011), namely residential land, public management and public service land, commercial service facilities land, industrial land, logistics and warehousing land, road and transportation facilities land, public facilities land, and green space and square land. This ignores the network or area distribution characteristics of spatial entities such as road and transportation facilities land and green space and square land. It faces the problem of limited POI data and large plot area, which affects the accuracy of establishing the relationship between POI data and urban land use type, resulting in a large deviation in the identification of urban land use for large plots.

[0006] (2) Problem of identification unit division: The current technology has low accuracy of identification units and generally only relies on the POI data within the identification unit to determine the urban land use category, while ignoring the influence of the POI data around the identification unit. This results in the identification units being independent of each other, ignoring the spatial relationship and interrelationship between urban lands, thus limiting the accuracy and comprehensiveness of urban land use identification.

[0007] (3) Model algorithm defects: Existing natural language processing models generally lack context sensitivity and deep semantic understanding capabilities, and cannot accurately capture the changes of words in different contexts. At the same time, these models perform poorly in capturing long-distance semantic relationships, especially when faced with polysemous words and complex contexts, where the recognition effect is poor.

[0008] (4) Problem of translation of identification results: Due to the spatial distribution and quantitative characteristics of POI data, coupled with the accuracy of identification units, the generated urban land use identification results often show the characteristics of mottled distribution of various land uses, resulting in a lack of agglomeration of urban land use areas of the same category, which cannot effectively reflect the overall structure and land use patterns of urban space.

[0009] (5) Inability to adapt to mixed land use and compatibility issues: In the existing technology, the scope of the identification unit is significantly different from the actual plot scale, making it difficult to match the plot size in the real environment. In addition, the existing technology lacks sufficiently refined urban land use identification capabilities and cannot effectively handle mixed land use forms such as commercial and residential, industrial and residential, as well as complex situations where a single building contains multiple functions.

[0010] (6) Non-replicability of identification results: Existing technologies usually only identify land use in specific cities, and the identification accuracy is low. They fail to delve into the underlying logic between POI data and urban land use. Therefore, their application is often limited to a single city, lacking cross-city adaptability and replicability, resulting in low identification efficiency and limiting their promotion and application in a wider range of urban environments.

[0011] These limitations severely restrict the widespread application and development of existing technologies in rapid and high-precision urban land use identification, and urgently require improvement and optimization.

[0012] Therefore, this invention proposes an urban land use identification method based on the BERT model text classification algorithm.

[0013] A search revealed no publicly available literature of the same or similar prior art as this invention. Summary of the Invention

[0014] The purpose of this invention is to overcome the shortcomings of existing technologies and propose an urban land use identification method based on the BERT model text classification algorithm. This method can utilize available POI data to perform timely and accurate high-precision urban land use identification, alleviating the technical problem of difficulty in obtaining urban land use data in urban development and construction and in-depth urban research.

[0015] The present invention solves its practical problem by adopting the following technical solution: A method for urban land use identification based on the BERT model text classification algorithm includes the following steps: Step 1: Determine the urban land use identification category, define the urban land use identification area, and obtain POI data within the area before filtering the data. Step 2: Establish urban land use identification units in the urban land use identification model based on the BERT model text classification algorithm, and associate the POI data filtered in Step 1 with the urban land use identification units; Step 3: Based on the association results between urban land use identification units and POI data obtained in Step 2, perform geographic text mining on the POI data; Step 4: Based on the geographic text mining results of the POI data obtained in Step 3, perform urban land use identification using an urban land use identification model based on the BERT model text classification algorithm to obtain high-precision urban land use identification results.

[0016] Furthermore, the specific method of step 1 is as follows: First, the urban land use identification category of the urban land use identification model is constructed using six categories of urban construction land: residential land, public management and public service land, commercial service land, industrial land, logistics and warehousing land, and public facilities land, as the text classification algorithm of the BERT model. Then, the scope of the urban land use identification area is determined, and POI data within this area is obtained, including serial number, administrative division, address, name, latitude and longitude coordinates, and the inherent attributes of the following 23 categories: catering services, road ancillary facilities, address and place name information, scenic spots, public facilities, companies and enterprises, shopping services, transportation facilities services, financial and insurance services, science, education and culture services, motorcycle services, automobile services, automobile repair, automobile sales, commercial and residential, living services, events and activities, indoor facilities, sports and leisure services, access facilities, medical and health care services, government agencies and social organizations, and accommodation services.

[0017] Next, POI data located in green spaces, water bodies, and roads within the identified area are removed. Based on the correlation between POI data categories and six types of urban construction land, POI data in six major categories—road ancillary facilities, address and place name information, transportation facility services, events and activities, indoor facilities, and access facilities—are also removed. Finally, the POI data, including 17 major categories, 229 medium categories, and 753 minor categories, were filtered and retained.

[0018] Furthermore, the specific steps of step 2 include: First, urban land use identification units are established in the urban land use identification model based on the BERT model text classification algorithm using a subset of the traffic analysis area. The study area is divided into 50×50 meter grids, with the grid center as the sampling location and a buffer with a radius of 50 meters set. The buffers of adjacent sampling locations overlap to ensure that the model can obtain information around each sampling location.

[0019] Then, for each POI data within the sampling location buffer, the distance between the POI data and the corresponding sampling location is calculated, and the POI data is sorted in order from near to far to generate a POI data list including fields such as sampling location number, POI number, POI address, POI name, POI latitude and longitude coordinates, POI category, and distance between POI and sampling location; each grid corresponds to a POI data list, which serves as the feature information of that grid, representing the association result between urban land use identification units and POI data.

[0020] Furthermore, the specific steps of step 3 include: (1) Based on the association results of urban land use identification units and POI data obtained in step 2, the Chinese names in the POI data list are concatenated to remove duplicates and separated by spaces, serving as the representative attribute of each sampling location buffer; each grid number corresponds to an input data consisting of concatenated "statements".

[0021] (2) Input the input data obtained in step (1) into the BERT model encoder Transformer Encoder to obtain the high-dimensional embedding vectors of these information as the geographic text mining results of POI data; The specific steps of step 3, step (2) include: First, the Tokenizer in the BERT model segments the input data into words and converts each word into its corresponding numeric code. , which serves as the input to the BETR model encoder, Transformer Encoder; Then, the BETR model encoder (Transformer Encoder) maps the input to a high-dimensional embedding vector. And use it as the result of geographic text mining of POI data, that is:

[0022] Furthermore, the specific steps of step 4 include: (1) First, a portion of POI data is randomly selected and land use categories are manually labeled by comparing it with the latest urban land use planning map to form a labeled dataset for model training and validation; the remaining unlabeled POI data only contain basic information and do not involve urban land use categories, and will be used as input data for model prediction.

[0023] (2) Using data augmentation techniques, the words in the “sentence” corresponding to each grid number are copied multiple times and randomly shuffled to increase the robustness of the model; (3) Input the training set and validation set data into the model training program, use the backpropagation mechanism of the neural network to iteratively optimize the model, and generate a trained urban land use identification model based on the BERT model text classification algorithm.

[0024] Moreover, the specific method of step (3) of step 4 is as follows: The high-dimensional embedding vector obtained in step 3 is input into the multilayer perceptron-based classifier Classifier to obtain a score for each urban land use category. The SoftMax function is then used to convert the score into a probability value for that category. ,Right now:

[0025] Finally, the category with the highest probability value is returned as the final prediction result:

[0026] In summary, the complete mathematical expression of the trained urban land use identification model based on the BERT model text classification algorithm is as follows:

[0027] (4) Finally, the test dataset is input into the trained final model to obtain high-precision urban land use identification results with the identification area divided into 50×50 meter grids, and its performance is evaluated and the effectiveness of the model identification is verified.

[0028] Advantages and beneficial effects of the present invention: 1. This invention proposes an urban land use identification method based on the BERT model text classification algorithm. It uses a 50×50 meter grid as the urban land use identification unit, significantly improving the accuracy of urban land use identification compared to existing technologies. By setting a buffer zone with a radius of 50 meters around each sampling location, with overlapping buffer zones for adjacent sampling locations, the model can acquire POI data around the sampling location. The interaction between POI data within the buffer zone jointly determines the urban land use category of the grid where the sampling location is located. This approach effectively establishes spatial connections between urban land areas, significantly improving the scientific rigor and accuracy of the urban land use identification method.

[0029] 2. This invention proposes a POI data cleaning strategy, which solves the problem of scarce POI data in areas with large-scale, areal or network distribution, such as green spaces and plazas, roads and transportation facilities, and water bodies, by removing POI data located in green spaces, water bodies, and roads. This avoids the negative impact of insufficient POI data in these areas on the accuracy of urban land use identification.

[0030] 3. This invention leverages the powerful contextual capture capabilities of the BERT model to comprehensively understand the meaning and context of words, and significantly improves the processing performance of polysemous words through a multi-layered attention mechanism. This enables the model to more accurately capture complex semantic relationships, demonstrating higher accuracy in polysemous word disambiguation and contextual understanding. Unlike traditional methods that rely on manual feature extraction and shallow models, this invention significantly reduces human intervention, improving the model's automation and generalization capabilities. Through training and optimization based on backpropagation, the model's classification accuracy and adaptability are significantly improved, achieving more accurate and efficient identification of urban land use.

[0031] 4. This invention proposes a data augmentation correction method based on the training dataset. By repeatedly copying the words within the "sentence" corresponding to each grid number and randomly changing their order, not only is the size of the training dataset effectively expanded, but the generalization ability of the model is also enhanced. This solves the performance degradation problem caused by the randomness of the sampling position leading to changes in the order of the grid "sentences" in the existing technology.

[0032] 5. The identification results of this invention can effectively map the distribution characteristics of urban spatial structure in land use. Furthermore, the 50×50 meter grid is generally smaller than the actual size of urban plots in terms of spatial scale. Therefore, the identification results can not only be used to analyze the composition of land use functions, but are also suitable for evaluating urban plots implementing land use compatibility and mixed-use policies in the context of increasingly diversified and complex land use patterns. Compared with traditional land use planning maps, the identification results of this invention can more accurately reflect the interrelationship between urban elements and human activities, providing more realistic image support for refined urban renewal.

[0033] 6. The model of this invention has significant adaptive optimization capabilities, enabling continuous updates and iterations during practical applications. A dynamic model optimization strategy is employed, transforming the test set from each testing phase into the training set required for the next round of model training. This cyclical update mechanism not only improves the model's adaptability to new data and scenarios but also continuously optimizes model performance during iteration. Attached Figure Description

[0034] Figure 1 This refers to the urban land use identification area of ​​this invention.

[0035] Figure 2 This invention relates to the POI data distribution in the Hankou area of ​​Wuhan City.

[0036] Figure 3 This is a schematic diagram of the 50×50 meter grid division of the Hankou area of ​​Wuhan City according to the present invention.

[0037] Figure 4 This is a schematic diagram showing the distance between the sampling location and the POI data in this invention.

[0038] Figure 5 This is a schematic diagram of the text classification model based on BERT embedding vectors of the present invention.

[0039] Figure 6 This is a schematic diagram illustrating the generation of the labeled dataset according to the present invention.

[0040] Figure 7 This is a map showing the urban land use identification results in Hankou area of ​​Wuhan City, as presented in this invention. Detailed Implementation

[0041] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings: A method for urban land use identification based on the BERT model text classification algorithm includes the following steps: Step 1: Determine the urban land use identification category, define the urban land use identification area, and obtain POI data within the area before filtering the data. The specific method for step 1 is as follows: First, based on the "Classification of Urban Land Use and Standards for Planning and Construction Land" (GB50137-2011), and excluding the interference of road and transportation facility land use and green space and square land use, the urban land use identification model is constructed using six categories of urban construction land: residential land use, public management and public service land use, commercial service land use, industrial land use, logistics and warehousing land use, and public facility land use as the urban land use identification categories for the BERT model text classification algorithm.

[0042] Then, the scope of the urban land use identification area is determined, and POI data within this area is obtained, including serial number, administrative division, address, name, latitude and longitude coordinates, and the inherent attributes of the following 23 categories: catering services, road ancillary facilities, address and place name information, scenic spots, public facilities, companies and enterprises, shopping services, transportation facilities services, financial and insurance services, science, education and culture services, motorcycle services, automobile services, automobile repair, automobile sales, commercial and residential, living services, events and activities, indoor facilities, sports and leisure services, access facilities, medical and health care services, government agencies and social organizations, and accommodation services.

[0043] Next, POI data located in green spaces, water bodies, and roads within the identified area are removed. Based on the correlation between POI data categories and six types of urban construction land, POI data in six major categories—road ancillary facilities, address and place name information, transportation facility services, events and activities, indoor facilities, and access facilities—are also removed. Finally, the POI data, which includes 17 major categories, 229 medium categories, and 753 minor categories, was filtered and retained. For example, the POI data with serial number 14 contains the following information: administrative division (Hanyang District, Wuhan City, Hubei Province), address (Building A, 5th-18th Floor, Pengjialing No. 1 Miling City Plaza), name (Vienna International Hotel (Wuhan Mengjiapu Metro Station Branch)), latitude and longitude coordinates (114.162929427774, 30.5664412435903), major category (accommodation services), medium category (hotels), and minor category (hotels).

[0044] Step 2: Establish urban land use identification units in the urban land use identification model based on the BERT model text classification algorithm, and associate the POI data filtered in Step 1 with the urban land use identification units; The specific steps of step 2 include: First, urban land use identification units are established in the urban land use identification model based on the BERT model text classification algorithm using a subset of the traffic analysis area. The study area is divided into 50×50 meter grids, with the grid center as the sampling location and a buffer with a radius of 50 meters set. The buffers of adjacent sampling locations overlap to ensure that the model can obtain information around each sampling location.

[0045] Then, for each POI data within the sampling location buffer, the distance between the POI data and the corresponding sampling location is calculated, and the POI data is sorted in order from near to far to generate a POI data list including fields such as sampling location number, POI number, POI address, POI name, POI latitude and longitude coordinates, POI category, and distance between POI and sampling location; each grid corresponds to a POI data list, which serves as the feature information of that grid, representing the association result between urban land use identification units and POI data.

[0046] Step 3: Based on the association results between urban land use identification units and POI data obtained in Step 2, perform geographic text mining on the POI data; The specific steps of step 3 include: (1) Based on the association results of urban land use identification units and POI data obtained in step 2, the Chinese names in the POI data list are concatenated to remove duplicates and separated by spaces, serving as the representative attribute of each sampling location buffer; each grid number corresponds to an input data consisting of concatenated "statements".

[0047] (2) Input the input data obtained in step (1) into the BERT model encoder Transformer Encoder to obtain the high-dimensional embedding vectors of these information as the geographic text mining results of POI data; The specific steps of step 3, step (2) include: First, the Tokenizer in the BERT model segments the input data into words and converts each word into its corresponding numeric code. , which serves as the input to the BETR model encoder, Transformer Encoder; Then, the BETR model encoder (Transformer Encoder) maps the input to a high-dimensional embedding vector. And use it as the result of geographic text mining of POI data, that is:

[0048] In this embodiment, the BERT model can effectively capture the contextual information of words, fully understand word meaning and context, and process different contexts of polysemous words through a multi-layer attention mechanism.

[0049] Step 4: Based on the geographic text mining results of the POI data obtained in Step 3, perform urban land use identification using an urban land use identification model based on the BERT model text classification algorithm to obtain high-precision urban land use identification results.

[0050] The specific steps of step 4 include: (1) First, a portion of POI data is randomly selected and land use categories are manually labeled by comparing it with the latest urban land use planning map to form a labeled dataset for model training and validation; the remaining unlabeled POI data only contain basic information and do not involve urban land use categories, and will be used as input data for model prediction.

[0051] (2) Using data augmentation techniques, the words in the “sentence” corresponding to each grid number are copied multiple times and randomly shuffled to increase the robustness of the model; (3) Input the training set and validation set data into the model training program, use the backpropagation mechanism of the neural network to iteratively optimize the model, and generate a trained urban land use identification model based on the BERT model text classification algorithm.

[0052] The specific method for step (3) of step 4 is as follows: The embedding vector generated in this model consists of 768 numbers. Next, the high-dimensional embedding vector obtained in step 3 is input into a multilayer perceptron-based classifier (Classifier) ​​to obtain a score for each urban land use category. The SoftMax function is then used to convert the score into a probability value for that category. ,Right now:

[0053] Finally, the category with the highest probability value is returned as the final prediction result:

[0054] In summary, the complete mathematical expression of the trained urban land use identification model based on the BERT model text classification algorithm is as follows:

[0055] (4) Finally, the test dataset is input into the trained final model to obtain high-precision urban land use identification results with the identification area divided into 50×50 meter grids, and its performance is evaluated and the effectiveness of the model identification is verified.

[0056] In this embodiment, the selection of suitable land parcels for both land use compatibility and mixed land use can also be achieved.

[0057] First, urban road network-based land parcels are overlaid on the urban land use identification results. Then, by analyzing the proportion of various urban land use areas to the total area of ​​each parcel in the real environment, and based on the relevant provisions of the "Classification of Urban and Rural Land Use and Standards for Planning and Construction Land" GB50137- (2018 Draft for Comments) regarding land use compatibility and mixed land use, the parcels in the city suitable for implementing land use compatibility and mixed land use policies are determined.

[0058] In this embodiment, cross-city adaptive optimization of the model can also be achieved. By transforming the results of land use identification in different cities into a training set, the training set can be continuously optimized, the model can be continuously improved, gradually adapt to the needs of different cities, and ultimately achieve the reproducibility of the identification results.

[0059] The invention will be further illustrated below with specific examples: This invention uses the Hankou area of ​​Wuhan City as an example to identify urban land use. Before using the model, the computer should meet the following requirements: Operating System: Windows 10 or later; or macOS 13.0 or later; or Linux Software platform: ArcGIS Hardware requirements: The CPU and memory must be sufficient for the normal operation of the operating system and ArcGIS; the graphics card must be from the NVIDIA series, with a minimum of 4GB of video memory for model inference and a minimum of 12GB of video memory for model training.

[0060] Python environment: Python 3.10 and above; PyTorch 1.10 and above; CUDA 12.4; tqdm 4.66.2; sklearn 1.2.2; Pandas 2.2.2; Numpy 1.26.4; boto3 1.34.154; The specific steps for performing inference using a pre-trained model are as follows: Step 1: Data Acquisition. First, the Hankou area of ​​Wuhan City is selected as the area for urban land use identification, such as... Figure 1As shown, this includes Jiang'an District, Jianghan District, and Hankou District. Then, POI data for the Hankou area of ​​Wuhan City was obtained. The sample data includes 17 major categories, covering 229 subcategories and 753 subcategories, such as catering services, scenic spots, public facilities, companies and enterprises, shopping services, financial and insurance services, science, education and culture services, motorcycle services, automobile services, automobile repair, automobile sales, commercial and residential properties, living services, sports and leisure services, medical and health services, government agencies and social organizations, and accommodation services. The data includes text information such as serial number, address, name, latitude and longitude, major category, intermediate category, and minor category. The processed POI data was then loaded into the ArcGIS platform, as shown... Figure 2 As shown.

[0061] Step 2: Data Cleaning and Loading. Since green spaces and plazas, roads and transportation facilities, and water areas are distributed in a planar or grid pattern, and these areas are relatively large with limited POI data, the accuracy of urban land use identification would be affected. Therefore, in the ArcGIS platform, POI data located in green spaces, water bodies, and roads were removed from the original POI data, ultimately yielding 126,505 valid POI data.

[0062] Step 3: Data Preprocessing. The Hankou area of ​​Wuhan City was divided into 61,233 grids using a 50×50 meter grid, as shown below. Figure 3 As shown, the center point of the grid is used as the sampling location. A buffer with a radius of 50 meters is set at each sampling location, and the buffers of adjacent sampling locations overlap to ensure that the model can acquire surrounding POI data, such as... Figure 4 As shown. Based on the distance from the sampling location, the POI data within each buffer are sorted from closest to furthest. The preprocessed data is exported as a .csv file, containing the following: TARGET_FID (grid number), type_ (category, including major, intermediate, and minor categories, separated by semicolons), name (POI name), and land use (land use category). The land use column is empty; the inference result corresponds to this column. The csv file is saved in the poi_data folder; see the example hankou_POI_predict.csv for the format.

[0063] Step 4: Model Inference. This model consists of two main parts—a BERT-based pre-trained Chinese encoder (Transformer Encoder) and a multilayer perceptron-based classifier (Classifier). The inference process is as follows: Figure 5As shown, by combining keywords containing rich information from POI data and inputting them into the encoder of the BERT model, a high-dimensional representation of this useful information, namely the embedding vector, can be obtained. After passing through a classifier composed of fully connected neural networks, the embedding vector is mapped into meaningful urban land use categories.

[0064] First, the tokenizer in the model segments the input text into words and converts each word into its corresponding numerical code. As input to the Transformer Encoder, the word segmentation corresponds one-to-one with the numeric encoding, and the correspondence is stored in the . / bert_pretrain / vocab.txt file, which is the vocabulary included with the BERT Chinese pre-trained model. Then, the BERT model's Transformer Encoder maps this input to embedding vectors. ,Right now:

[0065] The embedding vectors generated in this model consist of 768 numbers. These embedding vectors are then fed into the classifier to obtain a score for each land use category. The SoftMax function is then used to convert the scores into probability values ​​for that category. ,Right now:

[0066] Finally, the category with the highest probability value is returned as the final prediction result:

[0067] In summary, the complete mathematical expression of the model is as follows:

[0068] To do this, open a command line / terminal in the root directory of the code file, activate the Python environment, and type the following command: python predict.py --model “bert.ckpt” --csv_name “hankou_POI_predict.csv” --save_name “pred.csv” The `--model` parameter represents the model filename, located in the `. / poi_data / saved_dict / ` folder; `--csv_name` is the name of the data file saved in step 3, located in the `. / poi_data / ` folder; and `--save_name` is the name of the file saved after inference, located in the `. / poi_data / ` folder.

[0069] Step 5: Screening suitable land use. Overlay the urban road network onto the urban land use identification results. By analyzing the proportions of various urban land use category grids contained in each plot in the real environment, and based on the relevant provisions on land use compatibility and mixed land use in the "Classification of Urban and Rural Land Use and Planning Construction Land Standard" GB50137- (2018 Draft for Comments), determine the urban plots suitable for implementing land use compatibility and mixed land use policies.

[0070] This model also supports fine-tuning training with user-defined data, enhancing its robustness by training the model with user-defined labeled data. The specific steps are as follows: Steps 1-3 are the same as steps 1-3 above when using the pre-trained model. The CSV file generated in step 3 needs to be renamed, and the "Land Use" column needs to provide manually labeled data. Providing as much labeled data as possible will greatly improve the model's accuracy and robustness. Taking Hankou area of ​​Wuhan as an example, a portion of POI data is randomly sampled and manually labeled according to the "Wuhan City Planning Map" to form a labeled dataset; the remaining POI data without land use type information is used as an unlabeled dataset for model prediction. The labeled dataset contains 17,011 records distributed across 6,296 grids; the unlabeled dataset contains 375,847 records distributed across 33,724 grids, as shown below. Figure 6 As shown. Labeled data will be used for fine-tuning the model during training, while unlabeled data will be used for inference with the fine-tuned model. Save both types of data as separate CSV files in the . / poi_data / folder.

[0071] Step 4: Data Segmentation and Data Augmentation. The preprocessed labeled POI data is segmented to generate training, validation, and test datasets in an 8:1:1 ratio. The labeled training dataset is used for model training, the validation dataset is used for performance evaluation during model training, and the model with the highest accuracy in the validation dataset is saved as the final model. The test dataset is used to test the model's performance on unknown datasets. The labeled dataset is divided into training, validation, and test datasets in an 8:1:1 ratio. The training dataset contains 5,036 grids, while the validation and test datasets each contain 630 grids.

[0072] Considering the BERT model's sensitivity to word order and the inherent randomness of urban land use grid division, which could lead to shifts in center point positions and thus affect the "sentence" order, data augmentation techniques were employed on the training set to mitigate its impact on model performance. First, the 5,036 merged sentences were copied 1,000 times, generating 5,036,000 data points. Second, the word order of the "sentence" portion in each data point was randomly changed. Third, duplicate sentences were removed. Finally, the augmented training, validation, and test sets were saved to three separate text files, separated by tabs. This method not only effectively increased the size of the training dataset but also improved the model's generalization ability.

[0073] This process requires executing the `data_preprocess.py` file. Before execution, open this file and modify the `csv_path` variable in line 10 to the filename of the completed file. Also, modify the `aug_times` variable in line 13 to the number of copies needed (default is 1000). Then, open a command line / terminal in the root directory of the code file, activate the Python environment, and type the following command: python data_preprocess.py After execution, the code will generate three txt files in the . / poi_data / data / directory: train.txt; dev.txt; test.txt.

[0074] Step 5: Model Training. Model initialization was completed by loading the BERT Chinese pre-trained model, released by Google and trained based on Chinese Wikipedia. The BERT-base version was selected, containing 12 heads and 768 hidden layers, with approximately 110M parameters. The experiment was conducted on a single NVIDIA L40 graphics card using batch training with a batch size of 1024. During training, the model's VRAM usage was approximately 41GB, lower than the 48GB total VRAM of the L40 graphics card. Users need to select an appropriate batch size based on the VRAM size of their training graphics card. This parameter is modified in the `self.batch_sizebian` variable on line 23 of the `. / models / bert.py` file, with a default value of 1024. The optimizer used was the BERT Adam function, with a learning rate set to 5e-5 and the cross-entropy loss function used for performance evaluation of the multi-class classification model. The final classification output consisted of five categories: residential land, commercial service facilities land, public management and public service land, industrial land, and public facilities land. The model training epochs are set to 10, and the early stop threshold is set to 5000 packets. This means training will stop early if the model performance does not improve after 5,000,000 consecutive training data points. The final model stops training after 4 epochs, saving the current optimal parameters. All parameters are adjusted in the `. / models / bert.py` file. After adjusting the parameters, open a command line / terminal in the root directory of the code file, activate the Python environment, and type the following command: python run.py The trained model will be saved in the file . / poi_data / saved_dict / bert.ckpt.

[0075] Step 6: Model Performance Testing. The model is evaluated using the F1 score and confusion matrix to analyze its accuracy and classification ability. After model training, validation and test results will be automatically output. The F1 score shows that the model achieved an overall accuracy of 77.94%, indicating strong overall predictive ability for the test data. The "Residential Land" and "Commercial Service Facilities Land" categories performed best, achieving F1 scores of 83.33% and 80.11% respectively, demonstrating reliable predictive ability for the main categories, especially in terms of recall, where it captures the vast majority of samples.

[0076] After obtaining the trained model and achieving satisfactory test results, perform inference on the unlabeled data following the steps of "using the pre-trained model for inference," such as... Figure 7 As shown, the results of urban land use identification in the Hankou area of ​​Wuhan City were obtained.

[0077] It should be emphasized that the embodiments described in this invention are illustrative rather than limiting. Therefore, this invention includes, but is not limited to, the embodiments described in the specific implementation. Any other implementations derived by those skilled in the art based on the technical solutions of this invention are also within the scope of protection of this invention.

Claims

1. A method for urban land use identification based on the BERT model text classification algorithm, characterized in that: Includes the following steps: Step 1: Determine the urban land use identification category, define the urban land use identification area, and obtain POI data within the area before filtering the data. Step 2: Establish urban land use identification units in the urban land use identification model based on the BERT model text classification algorithm, and associate the POI data filtered in Step 1 with the urban land use identification units; Step 3: Based on the association results between urban land use identification units and POI data obtained in Step 2, perform geographic text mining on the POI data; Step 4: Based on the geographic text mining results of the POI data obtained in Step 3, perform urban land use identification using an urban land use identification model based on the BERT model text classification algorithm to obtain high-precision urban land use identification results.

2. The urban land use identification method based on the BERT model text classification algorithm according to claim 1, characterized in that: The specific method for step 1 is as follows: First, the urban land use identification category of the urban land use identification model is constructed using six categories of urban construction land: residential land, public management and public service land, commercial service land, industrial land, logistics and warehousing land, and public facilities land, as the text classification algorithm of the BERT model. Then, the scope of the urban land use identification area is determined, and POI data within this area is obtained, including serial number, administrative division, address, name, latitude and longitude coordinates, and the following 23 categories of inherent attributes: catering services, road ancillary facilities, address and place name information, scenic spots, public facilities, companies and enterprises, shopping services, transportation facilities services, financial and insurance services, science, education and culture services, motorcycle services, automobile services, automobile repair, automobile sales, commercial and residential, life services, events and activities, indoor facilities, sports and leisure services, access facilities, medical and health care services, government agencies and social organizations, and accommodation services. Next, POI data located in green spaces, water bodies, and roads within the identified area are removed. Based on the correlation between POI data categories and six types of urban construction land, POI data in six major categories—road ancillary facilities, address and place name information, transportation facility services, events and activities, indoor facilities, and access facilities—are also removed. Finally, the POI data, including 17 major categories, 229 medium categories, and 753 minor categories, were filtered and retained.

3. The urban land use identification method based on the BERT model text classification algorithm according to claim 1, characterized in that: The specific steps of step 2 include: First, urban land use identification units are established in the urban land use identification model based on the BERT model text classification algorithm using a subset of the traffic analysis area. The study area is divided into 50×50 meter grids, with the grid center as the sampling location and a buffer with a radius of 50 meters set. The buffers of adjacent sampling locations overlap to ensure that the model can obtain information around each sampling location. Then, for each POI data within the sampling location buffer, the distance between the POI data and the corresponding sampling location is calculated, and the POI data is sorted in order from near to far to generate a POI data list including fields such as sampling location number, POI number, POI address, POI name, POI latitude and longitude coordinates, POI category, and distance between POI and sampling location; each grid corresponds to a POI data list, which serves as the feature information of that grid, representing the association result between urban land use identification units and POI data.

4. The urban land use identification method based on the BERT model text classification algorithm according to claim 1, characterized in that: The specific steps of step 3 include: (1) Based on the association results of urban land use identification units and POI data obtained in step 2, the Chinese names in the POI data list are concatenated to remove duplicates and separated by spaces, serving as the representative attribute of each sampling location buffer; each grid number corresponds to an input data consisting of concatenated "statements"; (2) Input the input data obtained in step (1) into the BERT model encoder Transformer Encoder to obtain the high-dimensional embedding vectors of these information as the geographic text mining results of POI data; The specific steps of step 3, step (2) include: First, the Tokenizer in the BERT model segments the input data into words and converts each word into its corresponding numeric code. , which serves as the input to the BETR model encoder, Transformer Encoder; Then, the BETR model encoder (Transformer Encoder) maps the input to a high-dimensional embedding vector. And use it as the result of geographic text mining of POI data, that is: 。 5. The urban land use identification method based on the BERT model text classification algorithm according to claim 1, characterized in that: The specific steps of step 4 include: (1) First, a portion of POI data is randomly selected and the land use categories are manually labeled by comparing it with the latest urban land use planning map to form a labeled dataset for model training and validation; the remaining unlabeled POI data only contains basic information and does not involve urban land use categories, and will be used as input data for model prediction. (2) Using data augmentation techniques, the words in the "sentence" corresponding to each grid number are copied multiple times and randomly shuffled to increase the robustness of the model; (3) Input the training set and validation set data into the model training program, use the backpropagation mechanism of the neural network to iteratively optimize the model, and generate a trained urban land use identification model based on the BERT model text classification algorithm.

6. The urban land use identification method based on the BERT model text classification algorithm according to claim 5, characterized in that: The specific method for step (3) of step 4 is as follows: The high-dimensional embedding vector obtained in step 3 is input into the multilayer perceptron-based classifier Classifier to obtain a score for each urban land use category. The SoftMax function is then used to convert the score into a probability value for that category. ,Right now: ; Finally, the category with the highest probability value is returned as the final prediction result: ; In summary, the complete mathematical expression of the trained urban land use identification model based on the BERT model text classification algorithm is as follows: ; (4) Finally, the test dataset is input into the trained final model to obtain high-precision urban land use identification results with the identification area divided into 50×50 meter grids, and its performance is evaluated and the effectiveness of the model identification is verified.