Land use classification method and apparatus
By combining multi-source data fusion of point cloud data, image data, and point of interest data, and using machine learning algorithms for land use classification, this technology solves the problem that existing land use classification cannot reflect the impact of socio-economic activities, and achieves a more comprehensive reflection and dynamic monitoring of land use status.
Patent Information
- Application Number
- CN202510399554.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing land use classification methods are mainly based on natural attributes, which makes it difficult to fully reflect the specific role of land in socio-economic activities. As a result, the classification results cannot accurately reflect the actual situation and use needs of the land.
Natural attribute features are extracted by combining point cloud data and image data, and social attribute features are extracted by combining point of interest data. Multi-source data fusion is performed through machine learning algorithms to classify multi-contour regions and finally obtain the target category classification result.
This enables land classification results to more comprehensively reflect the actual situation and utilization needs of land, accurately count the area, distribution and quality of various types of land, and keep abreast of dynamic changes in land resources.
Smart Images

Figure CN120495718B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology for land classification, and in particular to a method and apparatus for classifying land use categories. Background Technology
[0002] Human production activities and population growth influence global land use changes, and these dynamic changes, in turn, affect and guide human production and living practices by altering climate and the ecological environment. Land use classification is a crucial indicator of urban morphology and forms the basis for extracting two-dimensional and three-dimensional urban morphological characteristics. Therefore, conducting land use classification research is fundamental to studying urban two-dimensional and three-dimensional morphological characteristics. However, current land use classification methods are typically based on natural attribute characteristics, and the results fail to reflect the specific role of land in socio-economic activities, thus failing to comprehensively characterize the actual situation and utilization needs of the land. Summary of the Invention
[0003] This invention provides a land use classification method and apparatus to at least solve the technical problem in related technologies where land classification based on natural attributes cannot comprehensively represent the actual situation and use needs of the land. The technical solution of this invention is as follows:
[0004] According to a first aspect of the present invention, a land use category classification method is provided. The method includes: extracting natural attribute feature information of a target area from point cloud data and image data; and extracting social attribute feature information of the target area from point of interest data of the target area; dividing the target area into multiple first contour areas according to the natural attribute feature information; and dividing the target area into multiple second contour areas according to the social attribute feature information; each first contour area corresponds to a first land cover category, the first land cover category representing a natural attribute land cover category; the correlation degree of natural attribute feature information between pixels in the first contour areas is greater than or equal to a first correlation degree threshold; each second contour area corresponds to a second land cover category, the second land cover category representing a social attribute land cover category; and performing secondary classification on the multiple first contour areas of the target area based on the multiple second contour areas to obtain a target category classification result of the target area; the target category classification result includes the first land cover category and / or the second land cover category.
[0005] According to a second aspect of the present invention, a land use category classification apparatus is provided. The apparatus includes: an extraction unit, configured to extract natural attribute feature information of a target area from point cloud data and image data, and to extract social attribute feature information of the target area from point of interest data of the target area; a first classification unit, configured to divide the target area into multiple first contour areas according to the natural attribute feature information; and to divide the target area into multiple second contour areas according to the social attribute feature information; each first contour area corresponds to a first land cover category, the first land cover category representing a natural attribute land cover category; the correlation degree of natural attribute feature information between pixels in the first contour areas is greater than or equal to a first correlation degree threshold; each second contour area corresponds to a second land cover category, the second land cover category representing a social attribute land cover category; and a second classification unit, configured to perform secondary classification on the multiple first contour areas of the target area based on the multiple second contour areas to obtain a target category classification result of the target area; the target category classification result includes the first land cover category and / or the second land cover category.
[0006] According to a third aspect of the present invention, a land use category classification system is provided, configured to perform a land use category classification method as described in the first aspect and any possible implementation thereof.
[0007] According to a fourth aspect of the present invention, a computer device is provided, comprising: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement a land use category classification method as described in the first aspect and any possible implementation thereof.
[0008] According to a fifth aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, such that when the instructions in the computer-readable storage medium are executed by a processor of a computer device, the computer device is able to perform a land use category classification method as described in the first aspect and any possible implementation thereof.
[0009] According to a sixth aspect of the embodiments of this application, a computer program product is provided, the computer program product including computer instructions, which, when executed on a computer device, cause the computer device to perform the land use category classification method of the first aspect and any possible implementation thereof described above.
[0010] The technical solution provided by the embodiments of the present invention brings at least the following beneficial effects: Natural attribute feature information of the target area is extracted from point cloud data and image data to classify the target area according to its natural attribute features, dividing it into multiple first contour areas; simultaneously, social attribute feature information of the target area is extracted from the point of interest data of the target area to classify it according to its social attribute features, dividing it into multiple second contour areas. Furthermore, the multiple first contour areas are further classified based on the multiple second contour areas, so that the final target classification result can accurately count the area, distribution, and quality of various types of land, promptly grasp the dynamic changes of land resources, and reflect the specific role of land in socio-economic activities, thereby making the land classification result more comprehensively represent the actual situation and utilization needs of land.
[0011] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application, and do not constitute an undue limitation of this application.
[0013] Figure 1 This is a flowchart illustrating a land use category classification method according to an exemplary embodiment;
[0014] Figure 2 This is a graphical illustration of the result of data fusion with different feature dimensions according to an exemplary embodiment;
[0015] Figure 3 This is a schematic block diagram illustrating a land use classification device according to an exemplary embodiment;
[0016] Figure 4 This is a schematic diagram of a computer device according to an exemplary embodiment. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0018] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0019] Before providing a detailed introduction to the land use category classification method provided in the embodiments of this application, let's first briefly introduce the application scenarios involved in the embodiments of this application.
[0020] Human production activities and population growth influence global land use changes, and these dynamic changes, in turn, affect and guide human production and living practices by altering climate and the ecological environment. Land use classification is a crucial indicator of urban morphology and forms the basis for extracting two-dimensional and three-dimensional urban morphological characteristics. Therefore, conducting land use classification research is fundamental to studying urban two-dimensional and three-dimensional morphological characteristics. With the acceleration of urbanization in my country, the disorderly expansion of construction land, the encroachment on green space resources, and environmental pollution have brought significant challenges to urban planning and development. Therefore, conducting land use classification research to accurately grasp the utilization status of land resources provides fundamental data and decision-making basis for the scientific planning of urban development, the rational utilization of land resources, and the improvement of the living environment, possessing significant theoretical and practical value. However, current land use classification methods are usually based on natural attribute characteristics, and the results fail to reflect the specific role of land in socio-economic activities, and cannot comprehensively characterize the actual situation and utilization needs of land.
[0021] Furthermore, classifying land solely based on its natural attributes fails to capture its specific role in socio-economic activities. For instance, a piece of land designated for construction could be a commercial area, an industrial area, or a residential area, each with significant differences in economic output and job creation. However, categorizing it solely by its natural attributes would simply classify them all as construction land, failing to distinguish their distinct economic functions and values.
[0022] Based on this, the following research was conducted: Satellite remote sensing technology can acquire Earth's surface cover information in a wide range, quickly and accurately. It has advantages such as high efficiency, accuracy, comprehensiveness, real-time performance and periodicity. Remote sensing sensors can sense and record the spectral reflectance of different land cover characteristics from visible light to infrared bands and from medium to extremely high spatial resolution. This is conducive to better image interpretation and processing and is an important technical means to understand the current status of land use.
[0023] Currently, high spatial resolution imagery, multi / hyperspectral imagery, and lidar have been widely used in land use classification research. However, with the continuous development of cities, land cover types are becoming increasingly diverse and complex, and single remote sensing data sources are no longer sufficient to extract complex urban land use and land cover categories. The fusion of multi-source remote sensing data can effectively achieve complementary advantages. Multi / hyperspectral remote sensing imagery can acquire more or continuous spectral information about land covers, but most imagery has low spatial resolution, making it difficult to meet the requirements of fine land use classification. Laser scanners can acquire high-precision three-dimensional spatial information of the land surface without being affected by lighting conditions, but lidar data lacks sufficient spectral information, making it difficult to identify land covers with similar height characteristics. Airborne aerial imagery can acquire high-resolution images in the visible light band and can achieve synchronous acquisition and acquisition with LiDAR point cloud data.
[0024] Therefore, the fusion of airborne LiDAR and imagery data not only achieves complementary advantages but also eliminates the problem of asynchronous data acquisition time, making high-resolution urban land use classification research possible. However, the land cover features extracted from both remote sensing imagery and airborne LiDAR point cloud data only describe the natural attributes of land cover, lacking a description of its socio-economic attributes. Land use classification, on the other hand, is the result of the combined effects of the natural world and human activities.
[0025] Therefore, in refined land use classification studies considering the fusion of multi-source data, it is also necessary to introduce socioeconomic data that can characterize human activities. With the rapid development of information science and technology, Points of Interest (POIs) based on big data represent the rapid development of Internet technology and the ever-updating of online electronic maps. They describe the spatial and attribute information of geographic entities, greatly enhancing the ability to acquire entity locations and linking human daily life trajectories with the geographic information world. They have strong computational and expressive capabilities and can better reflect human activity information in cities, providing important reference value for geospatial data mining and location-based regional population analysis.
[0026] Based on this, we will further conduct land use classification research by integrating high-resolution remote sensing data and POI data. This research will not only focus on the differences in the natural attributes of land use types but also consider the varying impacts of human socio-economic activities on different land cover types. We will fully explore the direct and indirect features of airborne LiDAR point cloud data, the spectral and spatial features of aerial imagery, and extract features from POI data. We will design classification schemes for single and multi-source data fusion, compare the impact of different input features on classification accuracy, and analyze the advantages and disadvantages of different data fusion classification schemes. We will conduct classification experiments using machine learning algorithms, including the widely used K-Nearest Neighbor (KNN) algorithm, the well-performing Random Forest (RF) algorithm, and the relatively new and widely used eXtreme Gradient Boosting (XBGoost) algorithm in various competitions. Analysis of the results from these three classifier algorithms will serve two purposes: firstly, to verify the stability of the classification scheme; and secondly, to select the highest classification accuracy from all the results of the three classifiers for land use classification output. Finally, the classification scheme with the highest accuracy is selected to generate the land use classification map output.
[0027] Based on the aforementioned research findings, this application provides a land use classification method. It extracts natural attribute features from point cloud data and image data to classify the target area according to these features, dividing it into multiple first contour regions. Simultaneously, it extracts social attribute features from point-of-interest data of the target area to classify it according to these features, dividing it into multiple second contour regions. Furthermore, it reclassifies the multiple first contour regions based on the multiple second contour regions. This ensures that the final classification result accurately reflects the area, distribution, and quality of various land types, allows for timely monitoring of dynamic changes in land resources, and demonstrates the specific role of land in socio-economic activities. Consequently, the land classification result more comprehensively represents the actual situation and utilization needs of the land.
[0028] The land use classification method provided in this application can be applied to a land use classification system or a computer device used for land use classification. For ease of understanding, the land use classification method provided in this application will be described in detail below with reference to the accompanying drawings.
[0029] Figure 1 This is a flowchart illustrating a land use category classification method according to an exemplary embodiment, such as... Figure 1 As shown, the land use classification method includes the following steps.
[0030] S11, extract the natural attribute feature information of the target area from point cloud data and image data, and extract the social attribute feature information of the target area from the point of interest data of the target area.
[0031] In some implementations, natural attribute feature information includes indirect feature information, direct feature information, spectral feature information, and image spatial feature information. Direct feature information includes elevation feature information and intensity feature information; indirect feature information includes spatial geometric feature information of the point cloud and texture feature information used to describe the correlation between the gray levels of two points within a preset distance and in a preset direction; spatial geometric feature information includes the undulation of ground features on the ground surface, the flatness of the ground feature surface, and the dispersion of the point cloud; image spatial feature information includes: morphological building index, morphological shadow index, and urban complexity.
[0032] Step S11 is implemented in detail through the following steps.
[0033] Firstly, as a data acquisition method, point cloud data is acquired from the lidar scanner of the airborne LiDAR measurement system, and image data is acquired from the airborne digital camera.
[0034] In order to ensure the accuracy of image data, the resolution of the airborne digital camera is set to be higher than the resolution threshold.
[0035] Secondly, indirect and direct feature information of each pixel at each location in the target area is extracted from the point cloud data, and spectral and spatial feature information of each pixel at each location is extracted from the image data. It is understood that the point cloud data and image data mentioned above are image data of the land area to be classified.
[0036] In some implementations, point cloud data refers to airborne LiDAR point cloud data, and image data can also be called airborne image data. Point cloud data includes ground points and non-ground points, and is discrete point cloud data.
[0037] In some embodiments, the airborne LiDAR measurement system can simultaneously acquire LiDAR point cloud data and aerial imagery data. The airborne laser scanning system integrates hardware such as a RIEGL VUX-1LR LiDAR scanner, a PHASE ONE IXU1000-R high-resolution digital camera, and a POS system consisting of an IMU and differential GPS. The entire system is mounted on a flight platform or aircraft (such as a manned aircraft or a drone) for data acquisition. During data acquisition, the LiDAR scanner acquires the X, Y, and Z three-dimensional coordinates and intensity information of the target ground features in real time; the digital camera acquires high-resolution imagery data of the target ground features in the visible light band; and the POS system uses a combination of IMU inertial navigation information and differential GPS positioning information for navigation, acquiring the attitude and position information of the carrier's motion (i.e., POS trajectory information), which is used for subsequent point cloud fusion and calculation, as well as image correction.
[0038] The point cloud data used above has undergone preliminary preprocessing, including the acquisition of raw data, calculation of POS trajectory, and fusion of POS and laser point cloud data to generate point cloud data in geographic coordinate system, which will not be elaborated here.
[0039] In the embodiments of this application, the orthorectification of airborne image data is achieved by using model key points extracted from point cloud ground points based on point cloud data as ground control points. Therefore, the geographic coordinates of various feature information extracted from point cloud data and image data are consistent, so as to unify the coordinate system of feature information in each dimension.
[0040] S12, according to natural attribute information, the target area is divided into multiple first contour areas; and according to social attribute information, the target area is divided into multiple second contour areas.
[0041] Each first contour region corresponds to a first land feature category, and the first land feature category represents the natural attribute land feature category.
[0042] The correlation of natural attribute features between pixels in the first contour region is greater than or equal to the first correlation threshold; each second contour region corresponds to a second land cover category, and the second land cover category represents the social attribute land cover category.
[0043] S13, based on multiple second contour regions, perform secondary classification on multiple first contour regions of the target region to obtain the target category classification result of the target region.
[0044] The target category classification results include a first land cover category and / or a second land cover category.
[0045] Through the above implementation methods, natural attribute feature information of the target area is extracted from point cloud data and image data to classify the target area according to its natural attribute features, dividing it into multiple first contour regions. Simultaneously, social attribute feature information of the target area is extracted from the point of interest data of the target area to classify it according to its social attribute features, dividing it into multiple second contour regions. Furthermore, the multiple first contour regions are further classified based on the multiple second contour regions, so that the final target classification result can accurately count the area, distribution, and quality of various types of land, promptly grasp the dynamic changes of land resources, and reflect the specific role of land in socio-economic activities. This results in a more comprehensive representation of the actual situation and utilization needs of land.
[0046] As one implementation method, in step S12 above, the target area is divided into multiple second contour areas according to social attribute feature information, and the specific process is as follows.
[0047] First, determine the kernel density of the interest point data at each location in the target area.
[0048] Point of Interest (POI) data consists of a series of point data with attributes such as name, address, coordinates, and category. It possesses strong computational and expressive capabilities and is widely used in urban spatial analysis and visualization. Kernel density analysis is used to calculate the unit density of point and line feature measurements within a specified neighborhood. It can intuitively reflect the distribution of discrete measurements within a continuous area, ultimately producing a smooth surface with large median values and small periphery values; the raster value is the unit density. Kernel density analysis effectively expresses the spatial distribution of POI point data and is a method used for POI data representation. The specific representation is shown in formula (1-1).
[0049]
[0050] In the above formula (1-1), i = 1, 2, ..., n represent the points in the input data. If these points are within the radius of the (x, y) position, only the points in the sum are included; D(x, y) is the density prediction value of the new (x, y) point; r is the search radius; pop i The weight value of the point (can be ignored, and can be assigned a value of 1); d i Let be the distance between point i and (x, y).
[0051] Specifically, this can be understood as spatially transforming POI data that expresses human activities, social and economic phenomena, and using kernel density analysis to convert discrete point data into raster data that can be spatially analyzed and expressed.
[0052] In some implementations, the POI data feature extraction process for the target area is as follows.
[0053] First, the POI data needs to be reclassified. The POI data types in the target area mainly include 20 categories: automotive services, car sales, car repair, motorcycle services, catering services, shopping services, lifestyle services, sports and leisure services, healthcare services, accommodation services, scenic spots, commercial and residential properties, government agencies and social organizations, science, education and culture services, transportation facilities services, financial and insurance services, companies and enterprises, road ancillary facilities, place name and address information, and public facilities. This data is mainly distributed in built-up areas. During the reclassification process, these 20 categories of data are re-divided according to their social function attributes into residential areas, commercial areas, industrial areas, public services, and open spaces.
[0054] Then, kernel density analysis is performed on the POI data containing functional area category attributes, and the search radius is set to a preset length (e.g., 500m).
[0055] Finally, the POI data, including POI data features after classification, is rasterized to ensure that the generated POI data has the same resolution as the airborne image data, resulting in a 1m resolution raster image, also known as a POI distribution heatmap. Based on the block area, the mean (mn), standard deviation (std), and sum (sum) of the kernel density of residential areas, commercial areas, industrial areas, public services, and open spaces within each block area are extracted to obtain 15 POI kernel density features.
[0056] Secondly, referring to the preset kernel density range corresponding to different second land cover categories, the second target land cover category corresponding to the target kernel density range to which the kernel density of the point of interest data at each location in the target area belongs is divided.
[0057] The second target feature category mentioned above is any feature category within the second feature category.
[0058] Third, candidate regions consisting of consecutive locations belonging to the same second target feature category within the target area are identified as the second contour regions corresponding to the same second feature category.
[0059] Optionally, in step S22 above, the target area is divided into multiple first contour areas according to the natural attribute feature information. Specifically, this includes: determining the natural attribute feature information of each location in the target area in multiple dimensions; dividing multiple locations in the target area whose similarity of the natural attribute feature information in multiple dimensions is higher than the corresponding similarity threshold in each dimension and whose locations are consecutive into the same first contour area; and determining the first target land cover category corresponding to the same first contour area based on the natural attribute feature information of the same first contour area; the first target land cover category is any land cover category in the first land cover category.
[0060] Optionally, to improve classification accuracy, the division of the target region into multiple first contour regions according to natural attribute information in step S22 above can be implemented in the following way: One pixel represents one location.
[0061] First, multiple different land use classification models are used to classify the land cover categories to which each pixel belongs based on the similarity between the indirect feature information, direct feature information, spectral feature information, visible light vegetation index, and image spatial feature information of each pixel in the image corresponding to the target area, which is higher than the corresponding similarity threshold. This results in multiple first classification results. The land use classification model represents the relationship between the feature information of point cloud data and image data and the land cover category to which each pixel belongs.
[0062] Secondly, based on the frequency of the land cover category to which each pixel belongs in the multiple first classification results, the land cover category to which each pixel belongs in the multiple first classification results is evaluated and analyzed, and the land cover category to which each pixel belongs with the highest evaluation result is determined as the first target land cover category of each pixel, thus obtaining the second classification result including each first target land cover category to which each pixel belongs.
[0063] Third, according to the preset segmentation scale, the image including each pixel is segmented, and the elevation feature information in the first multi-dimensional feature information and the spectral feature information in the second multi-dimensional feature information are used as the basis for determining the object contour features. Based on the preset merging conditions, the segmented regions in the image are merged to obtain an image including multiple contour regions corresponding to multiple target objects; multiple pixels included in a contour region belong to one target object.
[0064] The above-mentioned preset merging conditions include one or more of the following: minimizing object heterogeneity or minimizing image smoothness heterogeneity, minimizing image compactness heterogeneity, minimizing image shape heterogeneity, and minimizing the average heterogeneity of all objects; the above-mentioned preset segmentation scale is determined according to the object type of the segmented objects.
[0065] Fourth, based on the target object to which the pixels in the contour region belong, the first target feature category to which the pixels in the second classification result that are the same as the pixels in the contour region belong is corrected, resulting in a third classification result that includes the second target feature category of each pixel. The third classification result includes multiple first contour regions.
[0066] This classification method determines land cover categories based on the feature similarity between different dimensional feature information from multiple data sources, including the first multi-dimensional feature information of each pixel in point cloud data and the second multi-dimensional feature information of each pixel in image data. This ensures that each first classification result is based on data from multiple sources with different dimensions, thereby guaranteeing the accuracy of each first classification result. Furthermore, each first classification result is evaluated, and the land cover category classification result with the best evaluation result is used as the second classification result. This reduces the impact of instability in the application of individual classification algorithms on the first classification result, thus ensuring that the second classification result is more accurate. Further, to avoid the "salt and pepper" phenomenon in classification results caused by considering too many data dimensions (i.e., the distribution of two-dimensional and three-dimensional target land cover is relatively fragmented), an object-oriented classification method is adopted based on two dimensions of information: elevation feature information and spectral feature information. This further classifies the land cover categories in the image to determine the target objects to which multiple pixels in the contour region belong. Based on the target object classification result, the "salt and pepper" phenomenon that may occur in the second classification result is corrected, ultimately resulting in a more accurate third classification result.
[0067] Optionally, in order to improve the representation dimension of the classification results, step S23 above performs secondary classification on multiple first contour regions of the target region based on multiple second contour regions to obtain the target category classification result of the target region, specifically including the following steps.
[0068] First, it is determined that the first contour region and the second contour region of the target area have overlapping areas.
[0069] Secondly, if the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region are the same, and the granularity of the first target feature category is higher than that of the second target feature category, then the repeated region in the first contour region shall be classified according to the second target feature category of the repeated region.
[0070] The larger the granularity of a land feature category, the lower the precision of its land division, and the higher the granularity priority of its corresponding land feature category.
[0071] Third, if the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region are the same, and the granularity of the first target feature category is lower than that of the second target feature category, then the first target feature category of the repeated region in the first contour region is retained.
[0072] Fourth, if the category attributes of the first target feature category in the first contour area and the second target feature category in the second contour area are different, the repeated areas in the first contour area shall be classified according to the second target feature category of the repeated areas.
[0073] The first land feature category mentioned above includes one or more of the following: built-up land, bare land, cultivated land, grassland, roads, woodland, and water bodies; the second land feature category includes one or more of the following: residential land, commercial land, industrial land, public service and management land, educational and research land, green space, and plaza land. The first target land feature category is any one of the species in the first land feature category mentioned above. The second target land feature category is any one of the species in the second land feature category mentioned above.
[0074] Building land belongs to the same category as residential land, commercial land, and industrial land. The granularity of building land is higher than that of residential land, commercial land, and industrial land. Therefore, if the first target land feature category of the overlapping area is building land, and the second target land feature category of the overlapping area is residential land, commercial land, or industrial land, the overlapping area will be reclassified into the second target land feature category.
[0075] Grassland, woodland, and water bodies share the same category attributes as green space and plaza land. Grassland, woodland, and water bodies have a lower granularity than green space and plaza land. Therefore, if the first target feature category of a duplicate area is grassland, woodland, or water body, and the second target feature category of the duplicate area is green space or plaza land, the duplicate area will be reclassified into the first target feature category.
[0076] Bare land and public service and management land belong to different category attributes. If the first target feature category of the overlapping area is bare land, and the second target feature category of the overlapping area is public service or management land, the overlapping area should be reclassified into the second target feature category.
[0077] As another land classification method, to ensure rapid classification, land is classified according to a trained land classification model. Specifically, the natural and social attribute characteristics of each location in the target area are input into multiple different land use classification models to obtain classification results for different land cover categories. The classification results of different land cover categories are evaluated according to the frequency of the land cover category to which the same location belongs in the different land cover category classification results, and the land cover category classification result with the highest frequency in the evaluation results is determined as the target category classification result.
[0078] The aforementioned different land use classification models include the K-nearest neighbor classifier algorithm model, the random forest classifier algorithm model, and the extreme gradient boosting classifier algorithm model.
[0079] As a classification model, the above-mentioned different land use classification models are different machine learning models. The specific process of obtaining a land use classification model includes the following steps.
[0080] First, we construct several different initial models to represent the relationship between the various feature information of point cloud data, image data, and POI data and the land cover category to which each pixel belongs; these initial models include the K-nearest neighbor classifier algorithm model, the random forest classifier algorithm model, and the extreme gradient boosting classifier algorithm model.
[0081] Secondly, historical indirect and direct feature information of each historical pixel is extracted from multiple sets of historical point cloud data. Historical spectral and spatial feature information of each historical pixel is extracted from multiple sets of historical image data corresponding to multiple sets of historical point cloud data. Social attribute feature information is obtained from POI data, and multiple sets of historical images corresponding to multiple sets of historical point cloud data are obtained. The land cover category to which each historical pixel belongs in the historical images has been marked.
[0082] Third, according to the preset ratio rules, multiple sets of historical point cloud data, multiple sets of historical image data, historical POI data and multiple sets of historical images are selected to form training sample datasets and validation sample datasets; the preset ratio rules refer to selecting the number of samples for each land cover category based on the proportion of each land cover category in the historical image area.
[0083] Fourth, multiple initial models are trained separately based on the training sample dataset, and the initial models after training are validated using the validation sample dataset, so that the initial models with the best validation results can be used as multiple different land use classification models.
[0084] As a specific implementation method, the selection of the training sample dataset mainly includes the following steps: First, we randomly create sample points covering the entire study area; these samples do not contain any feature attribute information. Second, we combine high-resolution Google Earth imagery and aerial true-color imagery from similar imaging times to label the sample points by category. Then, we use random sampling again to divide the sample points into an 80% training dataset and a 20% test dataset (with no overlap between the training and test datasets). Finally, based on the synthetic images generated in the data fusion scheme, we extract the feature information (band information) contained in the synthetic images of each classification scheme and assign it to the corresponding training and test sample points as their attribute values. Based on the above steps, we obtain data fusion results with different feature dimensions, as detailed below. Figure 2 The chart contains sample datasets for fourteen classification schemes, where each sample point contains corresponding input feature information.
[0085] Optionally, the target region can be divided into multiple first contour regions according to natural attribute feature information, specifically implemented in the following manner.
[0086] Based on the data characteristics of point cloud data and image data of the target area, a first target model is determined from multiple first preset classification models; the natural attribute feature information of the target area is input into the first target model, and the target area is divided into multiple first contour areas.
[0087] Optionally, the target region can be divided into multiple second contour regions according to social attribute feature information. Specifically, this can be implemented as follows: based on the data features of the points of interest in the target region, a second target model is determined from multiple second preset classification models; the social attribute feature information of the target region is input into the second target model, and the target region is divided into multiple second contour regions.
[0088] The aforementioned multiple first-preset classification models and multiple second-preset classification models can be K-Nearest Neighbors (KNN), Random Forest (RF), and Extreme Gradient Boosting (XGBoost) algorithms, respectively. The methods for obtaining and training these multiple first-preset classification models are the same as those for training the land use classification model, and will not be repeated here.
[0089] The aforementioned data characteristics include the amount of data indicating the size of the target area, the number of feature dimensions indicating the level of feature information in the data, the degree of randomness characterizing whether the data distribution is regular, and the computing resources required for data processing.
[0090] The randomness of data distribution can be understood as the degree of fit to obtain multiple data distribution patterns. The higher the fit, the more regular the data distribution, and the lower the corresponding randomness.
[0091] To ensure that the model selection is more accurate and relevant to the current data, a correspondence between data features and preset classification models was established.
[0092] For example, if the amount of data in the target region is less than or equal to the first data size, the K-nearest neighbor classifier algorithm is used. If the amount of data in the target region is greater than the first data size, a random forest classifier algorithm and an extreme gradient boosting classifier algorithm are used.
[0093] If the number of feature dimensions in the target region is greater than or equal to the first dimension, a random forest classifier algorithm and an extreme gradient boosting classifier algorithm are used. If the number of feature dimensions in the target region is less than the first dimension, a K-nearest neighbor classifier algorithm is used.
[0094] If the randomness of the data in the target region is greater than or equal to the first randomness, the K-nearest neighbor classifier algorithm is used. If the randomness of the data in the target region is less than the first randomness, the random forest classifier algorithm and the extreme gradient boosting classifier algorithm are used.
[0095] If the computational resources required for data processing in the target region are greater than the first preset resource amount, the extreme gradient boosting classifier algorithm model is adopted; if the computational resources required for data processing are greater than the second preset resource amount but less than or equal to the first preset resource amount, the random forest classifier algorithm model is adopted; if the computational resources required for data processing are less than or equal to the second preset resource amount, the K-nearest neighbor classifier algorithm model is adopted.
[0096] This application selects three classifier algorithms from machine learning—KNN, RF, and XGBoost—for land use classification. Further explanation of the classification model involved in this application is provided below.
[0097] On the one hand, by analyzing the classification accuracy of the three classifier algorithms, we can verify the consistency of the impact of feature addition on classification accuracy; on the other hand, by comparing the classification results of the three classifier algorithms, we select the classification scheme with the highest classification accuracy and output the classification results of land use classification in the study area.
[0098] The first method of using KNN as a classifier is explained below.
[0099] The KNN algorithm is a powerful non-parametric learning algorithm widely used in pattern recognition. Due to its simplicity, ease of understanding, good predictive performance, insensitivity to outliers, and ability to determine the uniformity of sample distribution based on algorithm accuracy, KNN is often used in conjunction with other machine learning algorithms in remote sensing image classification research. KNN is also an instance-based lazy algorithm model; it does not require pre-training a classifier from a large number of samples. Instead, it stores all available instances and measures the similarity between samples based on distance calculation. The implementation principle of KNN is as follows: For the sample set prepared for training, each sample contains features and a target variable labeled as a classification value. For new input sample data without a classification value, the features in this sample data are compared with the features of every sample in the training sample set. The K most similar (closest) data are found, and the classification value of the most frequent category among these K most similar data is used as the classification value of the new input prediction data.
[0100] Specifically, KNN classification mainly includes the following steps.
[0101] Step 1: Sample data preparation.
[0102] Training sample set T = {(x1, y1), (x2, y2), ..., (x i y i), ..., (x n y n )}, where n is the number of training samples, and the test sample set {(x1, y1), (x2, y2), ..., (x j y j ), ..., (x m y m )},x i ,x j ∈F d Represents the eigenvector, y i y j ∈C={c1,c2,…,c l} represents the classification label, where the categories in the test sample set are used for post-classification cross-validation. Before training, we first perform comparable quantization on all features of all samples, that is, quantize non-numerical features into numerical values; secondly, we normalize all features. Since there are multiple input feature parameters, each parameter has its own domain and value range, which will have a certain impact on the subsequent distance calculation. For example, parameters with larger values will have a greater impact than parameters with smaller values. Formula 2-1 is used to normalize all feature values of the samples.
[0103]
[0104] Step 2: Calculate the distance between each sample in the test set and the training set. KNN distance calculation typically uses Euclidean distance or Manhattan distance; this paper uses the Euclidean distance algorithm (calculation formula 2-2) to calculate the distance between sample data in the multidimensional feature space. After calculating the distance from all test samples to the training samples, all sample data are sorted in ascending order of distance.
[0105]
[0106] Step 3: Determining the K value. The choice of the K value is crucial in the classification process. Whether the K value is large or small, it will lead to over-normalization of the pattern or highlighting local differences.
[0107] Step 4: Category Determination. Using the optimal K value as the nearest neighbor number, count the frequency of each category among the K points, and select the category with the highest frequency as the category of the unknown point (test point).
[0108] The second approach, using RF as a classifier, is explained below.
[0109] The RF classifier algorithm is a non-parametric pattern recognition method. It is an ensemble classifier based on decision trees and bagging. The classification prediction result is determined by voting based on the classification results of multiple decision trees. RF can handle input samples with high-dimensional features without the need for dimensionality reduction. It does not require prior assumptions about the data distribution. Even in the case of a limited number of samples, it can efficiently operate on large datasets and obtain good classification results. At the same time, it can evaluate the importance of input variables. Due to these advantages, random forests perform well in remote sensing data classification, such as multispectral data, hyperspectral data, LiDAR data, multi-source remote sensing data, etc.
[0110] Suppose there are N samples and M feature variables. RF randomly and with replacement extracts 2N / 3 independent sample data from the original training samples through the bootstrap sampling method to construct decision trees to generate a random forest; then randomly selects m features (m < M) with replacement as the basis for splitting the branches of this tree, and determines the split of each node based on the Gini criterion. The node is split by the variable that provides the best split.
[0111] The remaining sample data is called out-of-bag (OOB) data, which is used to evaluate the error rate of the random forest and calculate the importance of each feature. Through multiple iterations, OOB gradually removes relatively poor features and selects the best forest. After the OOB predicts the results of all samples and compares them with the true values, the out-of-bag error rate (OOB estimate of error rate) of this forest can be obtained. The classification of new sample data is determined by the majority vote among the classification results of all constructed decision trees.
[0112] The following specifically explains the best split of each node. RF generates classification decision trees based on the Classification and Regression Tree (CART) algorithm. The CART algorithm mainly constructs a binary tree recursively. In terms of feature selection, it adopts the Gini Index (GI) minimization criterion. GI is an impurity splitting method, which represents the probability that a randomly selected sample in the sample set is misclassified. Suppose a given sample dataset T has a total of K classes, and the probability that a sample belongs to the kth class is p k , then the GI of this probability distribution is:
[0113]
[0114] The formula satisfies the condition
[0115] The smaller the GI, the less uncertainty there is in the sample category and the higher the purity. When GI(T) = 0, it means that all samples at that node belong to the same category, indicating that the uncertainty of the sample is 0, and the splitting of the tree will stop.
[0116] Take the sample data set T of a certain child node i It is further divided into j parts T i = {1, 2, ..., j}, where N i For child node T i Let N be the number of samples in set T, then the feature m at the child node... i The GI (Gross Integrity) of each split node will be calculated. The split GINI coefficient for that node is:
[0117]
[0118] For a sample set T, calculate the minimum GINI coefficient for each feature:
[0119]
[0120] The basic idea of GINI minimum splitting is: for each feature, iterate through all possible splitting methods, and if the minimum GINI can be found... split If the purity of the sample is at its maximum at this point, this is used as the splitting criterion for that node. The subtree that divides this feature is the optimal branch. Then, the node continues to split according to the specification with the lowest purity, until the branch stopping rule is met and growth stops. By setting a purity threshold for leaf nodes, the leaf node stops growing when it is greater than or equal to the threshold.
[0121] Two important parameters in Random Forest (RF) are the number of decision trees (ntree) and the number of features randomly selected by each tree at each node (mtry). Generally, the value of ntree is considered more important because a larger number of decision trees increases model complexity and reduces efficiency. For most RF applications, a value of 0 to 1000 for ntree is a good range, while the value of mtry is defaulted to the square root of the number of input features.
[0122] The third approach, using XGBoost as a classifier, is explained below.
[0123] XGBoost is a type of boosting algorithm, belonging to the category of gradient boosting tree models. It's a gradient boosting algorithm based on decision trees. XGBoost undergoes multiple rounds of data iteration, generating a weak classifier in each iteration. The training of each classifier is based on the classification residuals obtained from the previous iteration. Weak classifiers are required to be simple, low-variance, and high-bias (such as the CART classifier, and linear classifiers are also supported). The training process continuously improves classification accuracy by reducing bias. The final classifier is generated by weighted summation of the weak classifiers obtained from each training round using an additive model. XGBoost offers advantages such as regularization to reduce overfitting, parallel processing, customizable optimization objectives and evaluation criteria, and the ability to handle sparse and missing values, making it valuable for remote sensing data processing. XGBoost demonstrates high prediction accuracy and processing efficiency in remote sensing data processing and classification applications. Some researchers have already used XGBoost for land use / cover classification, but there is still room for exploration in classification research within the context of multi-source remote sensing data fusion.
[0124] The implementation process of XGBoost is mainly as follows:
[0125] Given a dataset with n samples and m features, the dataset D can be defined as {(x...} i y i )}(|D|=n,x i ∈R m y i ∈R),x i Let y represent the i-th sample. i Let represent the label of the i-th sample. The algorithm uses K trees to calculate the predicted value of each tree for the sample, and then sums the predicted values of all trees to obtain the predicted value of the sample. The ensemble algorithm function for the trees is defined as follows:
[0126]
[0127] Where F is the space of the decision tree, F = {f(x) = ω} q(x)}(q:R m →T,ω∈R T ), where q represents the tree model, i.e., taking a sample as input, mapping the sample to leaf nodes according to the model, and outputting a predicted score; ω q f(x) represents the set of scores for all leaf nodes of tree q; T is the number of leaf nodes in tree q. The regularization objective function defined for the learning model f(x) is as follows:
[0128]
[0129] In the formula, y represents the model's predicted value;i This represents the category label of the i-th sample; For sample x i Training error; Ω(f k Let represent the regularization term of the k-th tree; T represent the number of leaf nodes in each tree; ω represent the set of scores for each leaf node; γ and λ represent the regularization coefficients, which need to be adjusted in practical applications. The goal is to obtain L(φ) and f(x) and their corresponding models. The model is learned through additive training, meaning that the original model is kept unchanged each time, and a new f(x) is added to the model in each iteration to minimize the objective function. Let... For the predicted value of the i-th instance at the t-th iteration, f needs to be added. t Minimize the following objective as follows:
[0130]
[0131] Indicates sample x i The final prediction is the sum of the predictions from the t-th tree and the predictions from the first t-1 trees. Next, we perform the first derivative g. i and the second reciprocal h i The solution is as follows:
[0132]
[0133] Substituting into (Equation 2-12), we get:
[0134]
[0135] The constant term has no effect on the change in the minimum value of the objective function. We remove the constant term to simplify the calculation, and the objective function simplifies to the following formula:
[0136]
[0137] Next, define I j ={i|q(x i Let ) = j} be an instance of leaf node j. Based on the above inference, we have the following formula:
[0138]
[0139] In the formula, ω j For each leaf node j of a tree, ω represents the score of each tree. j The final predicted score is obtained by summing the results, and the optimal ω is obtained by minimizing the objective function. j Value, and for ω in the above formula j Calculate the partial derivatives and determine the optimal weight values:
[0140]
[0141] The optimal value corresponding to tree q is:
[0142]
[0143] The above formula (2-15) can be used to evaluate the quality q of the tree, which is similar to the impurity evaluation score of a decision tree. In general, it is impossible to compute all possible tree structures. XGBoost uses a greedy algorithm, starting with a single leaf node and iteratively adding branches to the tree. Assume that set I is partitioned into left leaf nodes I... L and right leaf node I R The loss function after dividing the nodes into left and right leaf nodes is as follows:
[0144]
[0145] In practical applications, this formula is typically used to evaluate the quality of the split tree structure. The XGBoost algorithm sorts the feature samples, divides the features in ascending order, compares the magnitude of the objective function after each division, and finds the feature with the largest decrease in objective function, using it as the optimal split point.
[0146] XGBoost is often used in conjunction with cross-validation and grid search, two crucial components of machine learning. The combination of cross-validation and grid search is a commonly used method for model optimization and parameter evaluation. In machine learning, using the same dataset for both model training and estimation can lead to inaccurate error estimates. Cross-validation addresses this issue by estimating the generalization error more closely to the true model's performance, especially when sufficient sample data is available. However, in practical applications, the amount of sample data is often insufficient, necessitating data reuse. Cross-validation randomly divides the sample data into training, validation, and test sets. The training set is used to train the model, the validation set for model estimation and optimization, and the test set for evaluating the final learning method. In practical applications, we often use k-fold cross-validation. Its basic principle is to evenly divide the original data into k groups, with each subset serving as validation data in turn. The remaining k-1 subsets are used as the training set, resulting in k models. Finally, we use the average classification accuracy of these k models on the validation set as the performance metric for the classifier under this k-fold cross-validation. k-fold cross-validation only needs to be repeated k times, which greatly reduces the computational complexity of the model. In practice, k=10 is generally a good empirical value. The grid search algorithm uses cross-validation to find the optimal model parameters. It is an exhaustive search method, finding the optimal result in an array. That is, among all candidate parameter choices, it iterates through all possibilities, trying each one, and the parameter that performs best is the final result, enabling automatic parameter tuning.
[0147] To achieve the above functions, the land use classification device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art will readily recognize that, based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0148] This application embodiment also provides a method such as Figure 3 The land use classification device shown includes: an extraction unit 301, a first classification unit 302, and a second classification unit 303.
[0149] Extraction unit 301 is used to extract natural attribute feature information of the target area from point cloud data and image data, and to extract social attribute feature information of the target area from point of interest data of the target area.
[0150] The first classification unit 302 is used to divide the target area into multiple first contour areas according to natural attribute feature information; and to divide the target area into multiple second contour areas according to social attribute feature information; each first contour area corresponds to a first land feature category, and the first land feature category represents a natural attribute land feature category; the correlation degree of natural attribute feature information between pixels in the first contour area is greater than or equal to a first correlation degree threshold; each second contour area corresponds to a second land feature category, and the second land feature category represents a social attribute land feature category.
[0151] The second classification unit 303 is used to perform secondary classification on multiple first contour regions of the target region based on multiple second contour regions to obtain the target category classification result of the target region; the target category classification result includes the first land feature category and / or the second land feature category.
[0152] In one implementation, the first classification unit 302 is specifically used to: determine the kernel density of point of interest data at each location in the target area; classify the second target land cover category corresponding to the target kernel density range to which the kernel density of point of interest data at each location in the target area belongs, with reference to the preset kernel density range corresponding to different second land cover categories; the second target land cover category is any land cover category in the second land cover category; and determine the candidate area composed of each location in the target area that belongs to the same second target land cover category and is located consecutively as the second contour area corresponding to the same second land cover category.
[0153] As another implementation, the first classification unit 302 is specifically used to: determine the natural attribute feature information of each location in the target area in multiple dimensions; divide multiple locations in the target area whose similarity of the natural attribute feature information in multiple dimensions is higher than the similarity threshold of each dimension and whose locations are consecutive into the same first contour area; and determine the first target land cover category corresponding to the same first contour area based on the natural attribute feature information of the same first contour area; the first target land cover category is any land cover category in the first land cover category.
[0154] In another implementation, the second classification unit 303 is specifically used to determine that there are overlapping areas between the first contour area and the second contour area of the target area; if the category attributes between the first target feature category of the first contour area and the second target feature category of the second contour area corresponding to the overlapping area are the same, and the granularity priority of the first target feature category is higher than that of the second target feature category, then the overlapping area in the first contour area is classified according to the second target feature category of the overlapping area; if the category attributes between the first target feature category of the first contour area and the second target feature category of the second contour area corresponding to the overlapping area are the same, and the granularity priority of the first target feature category is lower than that of the second target feature category, then the first target feature category of the overlapping area in the first contour area is retained; if the category attributes between the first target feature category of the first contour area and the second target feature category of the second contour area corresponding to the overlapping area are different, then the overlapping area in the first contour area is classified according to the second target feature category of the overlapping area.
[0155] As another implementation, the second classification unit 302 is also used to input the natural attribute characteristic information and social attribute characteristic information of each location in the target area into multiple different land use classification models to obtain different land feature category classification results; according to the frequency of the land feature category to which the same location belongs in the different land feature category classification results, the different land feature category classification results are evaluated, so as to determine the land feature category classification result with the highest frequency in the evaluation results as the target category classification result; wherein, the first land feature category includes one or more of the following: building land, bare land, cultivated land, grassland, roads, forest land and water bodies; the second land feature category includes one or more of the following: residential land, commercial land, industrial land, public service and management land, education and scientific research land, green space and square land.
[0156] As another implementation, the first classification unit 302 is specifically used to: determine a first target model from multiple first preset classification models based on the data characteristics of point cloud data and image data of the target area; input the natural attribute feature information of the target area into the first target model, and divide the target area into multiple first contour areas.
[0157] As another implementation, the first classification unit 302 is specifically used to: determine a second target model from multiple second preset classification models based on the data characteristics of the points of interest in the target area; input the social attribute feature information of the target area into the second target model, and divide the target area into multiple second contour areas.
[0158] As another implementation, data characteristics include the amount of data indicating the size of the data, the number of feature dimensions indicating the level of feature information in the data, the randomness characterizing whether the data distribution is regular, and the computing resources required for data processing.
[0159] As another implementation, the natural attribute feature information includes indirect feature information, direct feature information, spectral feature information, and image spatial feature information; the extraction unit 301 is specifically used to: acquire point cloud data of the target area from the lidar scanner of the airborne LiDAR measurement system, and image data of the target area from the airborne digital camera; the resolution of the airborne digital camera is higher than the resolution threshold; extract indirect feature information and direct feature information of each pixel corresponding to each location in the target area from the point cloud data, and extract spectral feature information and image spatial feature information of each pixel corresponding to each location from the image data; wherein, the direct feature information includes elevation feature information and intensity feature information; the indirect feature information includes spatial geometric feature information of the point cloud and texture feature information used to describe the correlation between the gray levels of two points within a preset distance and in a preset direction; the spatial geometric feature information includes the degree of undulation of the ground features on the ground surface, the flatness of the ground feature surface, and the dispersion of the point cloud; the image spatial feature information includes: morphological building index, morphological shadow index, and urban complexity.
[0160] Regarding the apparatus in the above embodiments, the specific manner in which each unit module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0161] Figure 4 This is a schematic diagram of a computer device provided in this application. Figure 4 The computer device 60 may include at least one processor 601 and a memory 603 for storing processor-executable instructions. The processor 601 is configured to execute the instructions in the memory 603 to implement the land use classification method described in the following embodiments.
[0162] In addition, the computer device 60 may also include a communication bus 602, at least one communication interface 604, an input device 606, and an output device 605.
[0163] The processor 601 may be a processor (central processing unit, CPU), a microprocessor unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application.
[0164] The communication bus 602 may include a path for transmitting information between the aforementioned components.
[0165] Communication interface 604 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0166] Input device 606 is used to receive input signals and output device 605 is used to output signals.
[0167] The memory 603 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory may exist independently and be connected to the processing unit via a bus. The memory may also be integrated with the processing unit.
[0168] The memory 603 stores instructions for executing the scheme of this application, and the processor 601 controls the execution. The processor 601 executes the instructions stored in the memory 603 to realize the functions of the method of this application.
[0169] In a specific implementation, as one embodiment, the processor 601 may include one or more CPUs, for example... Figure 4 CPU0 and CPU1 in the CPU.
[0170] In a specific implementation, as one example, the computer device 60 may include multiple processors, such as... Figure 4 Processors 601 and 607 are described herein. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0171] The computer device, such as Figure 4 The diagram includes a processor 601 and a memory 603 for storing executable instructions of the processor 601; wherein the processor 601 is configured to execute executable instructions to implement the land use category classification method as described in any of the possible embodiments above. And it can achieve the same technical effect, so to avoid repetition, it will not be described again here.
[0172] This application also provides a computer-readable storage medium, which, when executed by a processor of a control device or control apparatus, enables the control device or control apparatus to perform a land use category classification method as described in any of the possible embodiments above. And it achieves the same technical effect; to avoid repetition, it will not be described again here.
[0173] This application also provides a computer program product, including a computer program or instructions, which are executed by a processor using a land use category classification method as described in any of the possible implementations above. This achieves the same technical effect, and to avoid repetition, will not be repeated here.
[0174] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0175] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A land use classification method, characterized in that, The method includes: Natural attribute feature information of the target area is extracted from point cloud data and image data, and social attribute feature information of the target area is extracted from point of interest data of the target area; According to the natural attribute feature information, the target area is divided into multiple first contour areas; and according to the social attribute feature information, the target area is divided into multiple second contour areas; each first contour area corresponds to a first land cover category, the first land cover category representing a natural attribute land cover category; the correlation degree of natural attribute feature information between pixels in the first contour area is greater than or equal to a first correlation degree threshold; each second contour area corresponds to a second land cover category, the second land cover category representing a social attribute land cover category; Based on the plurality of second contour regions, the plurality of first contour regions of the target region are classified in a secondary manner to obtain the target category classification result of the target region; the target category classification result includes the first land feature category and / or the second land feature category. The target region is divided into multiple second contour regions, including: Determine the kernel density of the interest point data at each location in the target region; Referring to the preset kernel density range corresponding to different second land cover categories, the kernel density of the point of interest data at each location in the target area is divided into second target land cover categories corresponding to the target kernel density range; the second target land cover category is any land cover category among the second land cover categories; The candidate region consisting of consecutive locations belonging to the same second target feature category within the target region is determined as the second contour region corresponding to the same second feature category. The step of performing secondary classification on the multiple first contour regions of the target region based on the multiple second contour regions to obtain the target category classification result of the target region includes: It is determined that the first contour region and the second contour region of the target region have overlapping areas; If the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region are the same, and the granularity priority of the first target feature category is higher than that of the second target feature category, then the repeated region in the first contour region is classified according to the second target feature category of the repeated region. If the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region corresponding to the repeated region are the same, and the granularity priority of the first target feature category is lower than that of the second target feature category, then the first target feature category of the repeated region in the first contour region is retained. If the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region are different, the repeated region in the first contour region shall be classified according to the second target feature category of the repeated region. The method further includes: The natural attribute feature information and the social attribute feature information of each location in the target area are respectively input into the multiple different land use classification models to obtain classification results of different land cover categories; The classification results of different land cover categories are evaluated based on the frequency of the land cover category at the same location in the different land cover category classification results, so that the land cover category classification result with the highest frequency in the evaluation results is determined as the target category classification result; The first land feature category includes one or more of the following: building land, bare land, cultivated land, grassland, roads, forest land, and water bodies; the second land feature category includes one or more of the following: residential land, commercial land, industrial land, public service and management land, education and scientific research land, green space, and plaza land.
2. The method according to claim 1, characterized in that, The step of dividing the target region into multiple first contour regions according to the natural attribute feature information includes: Determine the natural attribute feature information of each location in the target region in multiple dimensions; Multiple locations in the target region whose similarity to the natural attribute feature information of multiple dimensions is higher than the corresponding similarity threshold of each dimension and whose positions are consecutive are divided into the same first contour region. Based on the natural attribute feature information of the same first contour region, the first target land cover category corresponding to the same first contour region is determined. The first target land cover category is any land cover category in the first land cover category.
3. The method according to claim 1, characterized in that, The step of dividing the target region into multiple first contour regions according to the natural attribute feature information includes: Based on the data features of the point cloud data and the image data of the target area, a first target model is determined from multiple first preset classification models; The natural attribute feature information of the target region is input into the first target model, and the target region is divided into the plurality of first contour regions.
4. The method according to claim 1, characterized in that, The step of dividing the target region into multiple second contour regions according to the social attribute feature information includes: Based on the data features of the points of interest in the target region, a second target model is determined from multiple second preset classification models; The social attribute feature information of the target region is input into the second target model, and the target region is divided into the plurality of second contour regions.
5. The method according to claim 3 or 4, characterized in that, The data characteristics include the amount of data indicating the size of the data, the number of feature dimensions indicating the level of feature information in the data, the randomness characterizing whether the data distribution is regular, and the computing resources required for data processing.
6. The method according to any one of claims 1 to 4, characterized in that, The natural attribute feature information includes indirect feature information, direct feature information, spectral feature information, and image spatial feature information; The extraction of natural attribute feature information of the target area from point cloud data and image data includes: Point cloud data of the target area is acquired from the lidar scanner of an airborne LiDAR measurement system, and image data of the target area is acquired from an airborne digital camera; the resolution of the airborne digital camera is higher than a resolution threshold. The indirect and direct feature information of each pixel at each location in the target region is extracted from the point cloud data, and the spectral and spatial feature information of each pixel at each location is extracted from the image data. The direct feature information includes elevation feature information and intensity feature information; the indirect feature information includes spatial geometric feature information of point cloud and texture feature information used to describe the correlation between gray levels of two points within a preset distance and in a preset direction; the spatial geometric feature information includes the degree of undulation of ground features on the ground surface, the flatness of ground feature surfaces, and the dispersion of point cloud; the image spatial feature information includes: morphological building index, morphological shadow index, and urban complexity.
7. A land use classification device, characterized in that, The device includes: The extraction unit is used to extract natural attribute feature information of the target area from point cloud data and image data, and to extract social attribute feature information of the target area from point of interest data of the target area; A first classification unit is configured to divide the target region into multiple first contour regions according to the natural attribute feature information; and to divide the target region into multiple second contour regions according to the social attribute feature information; each first contour region corresponds to a first land cover category, the first land cover category representing a natural attribute land cover category; the correlation degree of natural attribute feature information between pixels in the first contour region is greater than or equal to a first correlation degree threshold; each second contour region corresponds to a second land cover category, the second land cover category representing a social attribute land cover category; The second classification unit is used to perform secondary classification on the plurality of first contour regions of the target region based on the plurality of second contour regions, to obtain the target category classification result of the target region; the target category classification result includes the first land feature category and / or the second land feature category. The first classification unit is specifically used for: Determine the kernel density of the interest point data at each location in the target region; Referring to the preset kernel density range corresponding to different second land cover categories, the kernel density of the point of interest data at each location in the target area is divided into second target land cover categories corresponding to the target kernel density range; the second target land cover category is any land cover category among the second land cover categories; The candidate region consisting of consecutive locations belonging to the same second target feature category within the target region is determined as the second contour region corresponding to the same second feature category. The second classification unit is specifically used for: It is determined that the first contour region and the second contour region of the target region have overlapping areas; If the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region are the same, and the granularity priority of the first target feature category is higher than that of the second target feature category, then the repeated region in the first contour region is classified according to the second target feature category of the repeated region. If the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region corresponding to the repeated region are the same, and the granularity priority of the first target feature category is lower than that of the second target feature category, then the first target feature category of the repeated region in the first contour region is retained. If the category attributes of the first target feature category in the first contour region and the second target feature category in the second contour region are different, the repeated region in the first contour region shall be classified according to the second target feature category of the repeated region. The device is also used for: The natural attribute feature information and the social attribute feature information of each location in the target area are respectively input into the multiple different land use classification models to obtain classification results of different land cover categories; The classification results of different land cover categories are evaluated based on the frequency of the land cover category to which the same location belongs in the classification results of different land cover categories, so that the land cover category classification result with the highest frequency in the evaluation results is determined as the target category classification result; The first land feature category includes one or more of the following: building land, bare land, cultivated land, grassland, roads, forest land, and water bodies; the second land feature category includes one or more of the following: residential land, commercial land, industrial land, public service and management land, education and scientific research land, green space, and plaza land.
Citation Information
Patent Citations
Land utilization classification method and system based on remote sensing data
CN117541940A
Land utilization category classification method and device and storage medium
CN119006876A