Attribute partition identification method of existing urban parks based on pre-trained visual large model
Through the method based on pre-trained visual big model, using drone surveying and computer vision technology, the existing urban park attribute zoning identification methods are solved, and efficient and accurate attribute zoning identification is achieved.
Patent Information
- Application Number
- CN202411171190.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-08-26
AI Technical Summary
The existing urban park attribute zoning identification method is inefficient and has low accuracy. It relies on manual field surveys and drawings, and cannot fully utilize the advantages of drone surveying and computer vision.
Using a pre-trained visual big model method, by creating a city park attribute partition identification data set containing simulated and real data, the city park plane image is automatically roughly extracted, and by modifying the structure of the visual big model and calculating feature-related vectors, detailed attribute partition identification is performed.
It greatly improves the accuracy and efficiency of urban park attribute zoning identification, reduces manual intervention, and can quickly and accurately identify different attribute zoning of urban parks.
Smart Images

Figure CN119360232B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of surveying and mapping geographic information services, and specifically relates to a method for identifying attribute zoning of existing urban parks based on a pre-trained visual large model. Background Art
[0002] In recent years, urbanization has accelerated, and urban parks, as the "green lungs" of cities, have achieved rapid development in recent years. At present, most of the existing urban park attribute zoning identification methods rely on manual field surveys and mapping operations. A small number of methods using digital technology cannot avoid manual participation in the identification process, resulting in low efficiency and low accuracy in the identification of existing urban park attribute zoning. The application of current advanced UAV surveying and mapping technology and artificial intelligence technology has brought new means to the field of existing urban park attribute zoning identification. Therefore, how to use the efficient imaging capabilities of drones and the powerful processing capabilities of computer vision to provide a new solution framework and implementation details for the problem of existing urban park attribute zoning identification is a key technical problem that needs to be solved urgently. Summary of the invention
[0003] Based on the above shortcomings, the present invention provides an existing urban park attribute zoning identification method based on a pre-trained visual large model, which solves the problem that the current existing urban park attribute zoning identification method is inefficient and has low accuracy.
[0004] The technical solution adopted by the present invention is as follows: a method for identifying the attribute partitions of existing urban parks based on a pre-trained visual large model, which is implemented according to the following steps:
[0005] Step S1: Create an existing urban park attribute partition identification data set, including a simulation data subset and a real data subset. The plane image of the existing urban park in the simulation data subset is converted from a two-dimensional or three-dimensional rendering generated by a landscape or architectural design software. The two-dimensional rendering is directly saved as a vector diagram, and the three-dimensional rendering needs to save the top view as a vector diagram. The plane image of the existing urban park in the real data subset uses an aerial drone to collect the plane image of the existing urban park, and the image stitching technology is used to synthesize the aerial images of the existing urban park into a plane image. The plane images of the existing urban parks in the two subsets need to be resized to 2024 pixels × 2024 pixels and pixel-level semantic segmentation masks are created; the existing urban park attribute partitions include 5 categories, namely road area, leisure area, vegetation area, water area and building area, and different numerical values are used to represent these 5 types of partitions and manual mask annotation is performed;
[0006] Step S2: Select the Everything mode of the visual large model SAM pre-trained on the large-scale image segmentation dataset SA-1B to automatically perform attribute partitioning and rough extraction on the existing urban park plane image. The generated rough extraction result is saved in the form of color RGB image, and different colors are extracted and counted. There are n kinds of colors, and after sorting according to the R value of the color, the pixels in different color areas are assigned I k , then I k =255k / n, k represents the color number;
[0007] Step S3: modify the embedded part of the visual large model SAM image, delete the original structure and add the existing urban park attribute partition segmentation head, the segmentation head is composed of multiple types of operation blocks connected in series, from front to back, respectively, deconvolution block 1-deconvolution block 2-convolution block 1-interpolation block 1-interpolation block 2-convolution block 2-convolution block 3;
[0008] Step S4: Calculate the color feature correlation vector K of the rough extraction result of the existing urban park attribute partition obtained in step S2 a and the multi-attribute partition feature correlation vector K of the feature map output by interpolation block 2 in the segmentation header in step S3 b , the calculation method is as follows: Assume the rough extraction result is Get the color separation matrix H a =f a (Ψ), where f a is the minimum similarity function between different colors; the feature map output by the interpolation block 2 is Get the multi-attribute partition separation matrix Η b =f b (Ψ), where f b (·) is the minimum similarity function between different attribute partitions; finally calculate K a ,K b = Θ(H a Ψ,H b Z), where Θ(·) is the maximum similarity function, and then the two related vectors are concatenated and input into the convolution block 2 in step S3 to form a fine-extracted visual model;
[0009] Step S5: Use the existing urban park attribute partition recognition data set to train the fine extraction visual large model using the multi-level segmentation loss. During training, the image embedding and previous model parameters of the fine extraction visual large model adopt the model parameters of the visual large model SAM pre-trained on SA-1B. The remaining model parameters are initialized using He. The multi-level segmentation loss is divided into three items, namely pixel-level segmentation loss, edge-level segmentation loss, and region-level segmentation loss. The three types of losses are weighted to form the final multi-level segmentation loss. The weights of the three types of losses are all 1.0. The optimizer used during training is SGD, and the initial learning rate is set to 0.001;
[0010] Step S6: Use the trained fine extraction visual large model to process the existing urban park plane image to be identified, and generate preliminary results of the existing urban park attribute partition identification. The preliminary results are 5 2024 pixel × 2024 pixel matrices, each matrix corresponds to a category of an attribute partition, and the value of each pixel in the matrix is the probability of belonging to the category. For each pixel, the category with the highest probability in the five matrices is taken as the final category of the pixel, which is then converted into the final identification result of the existing urban park attribute partition.
[0011] Furthermore, in step S1, the aerial photography altitude needs to be set at 50 to 100 meters, and the number of routes and images to be taken is selected according to the scale of the urban park, but the number of routes shall not exceed 5, and the number of images for each route shall not exceed 10.
[0012] Furthermore, in step S3, the output size of the deconvolution block 1 and the deconvolution block 2 is twice the input size, the convolution kernel of the convolution block 1 is 1×1, and its purpose is to adjust the number of channels. The output size of the interpolation block 1 and the interpolation block 2 is twice the input size, and the polynomial interpolation method is used. The convolution kernel of the convolution block 2 and the convolution block 3 is 3×3.
[0013] Furthermore, in step S5, the pixel-level segmentation loss, edge-level segmentation loss and region-level segmentation loss use mean square error loss function MSE, image similarity measurement SSIM, loss function and DICE function to calculate the difference between the predicted partition attribute category value and the labeled partition attribute category value.
[0014] Another object of the present invention is to provide a system for identifying attribute zoning of existing urban parks based on a pre-trained visual large model, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program, when executed by the processor, implements a method for identifying attribute zoning of existing urban parks based on a pre-trained visual large model as described above.
[0015] Another object of the present invention is to provide a computer-readable storage medium, on which is stored an implementation program for information transmission, and when the program is executed by a processor, the steps of a method for identifying attribute zoning of existing urban parks based on a pre-trained visual large model as described above are implemented.
[0016] The beneficial effects of the present invention are as follows: the present invention greatly improves the accuracy and efficiency of identifying the attribute zoning of existing urban parks. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a framework diagram of the existing urban park attribute zoning recognition model based on the pre-trained visual large model proposed by the present invention;
[0018] Figure 2 This is a test result diagram of the existing urban park attribute zoning identification model proposed in the present invention. Specific implementation plan
[0019] The specific implementation scheme of the present invention is illustrated by using a data set for identifying attribute zoning of existing urban parks created according to step S1.
[0020] Embodiment 1:
[0021] like Figure 1 As shown, a method for identifying the attribute partitions of existing urban parks based on a pre-trained visual large model is implemented according to the following steps:
[0022] Step S1: Create an existing urban park attribute partition identification dataset, including a simulation data subset and a real data subset. The plane images of existing urban parks in the simulation data subset are converted from two-dimensional renderings generated by SketchUP 2020 software. The two-dimensional renderings can be directly saved as PNG vector graphics, with a total of 500 images. The plane images of existing urban parks in the real data subset are collected by using an aerial drone. The aerial photography height needs to be set at 75 meters. The number of routes and the number of images taken are selected according to the scale of the urban park. The number of routes is 3, and the number of images for each route is 6. The image stitching technology is used to synthesize the aerial images of the existing urban park into plane images, with a total of 80 images. The plane images of existing urban parks in both subsets need to be resized to 2024 pixels × 2024 pixels and pixel-level semantic segmentation masks are created. The attribute partitions of existing urban parks include 5 categories, namely road area, leisure area, vegetation area, water area, and building area. Different numerical values are used to represent these 5 types of partitions and manual mask annotation is performed;
[0023] The existing urban park attribute zoning identification dataset created in this step includes a simulated data subset and a real data subset. The simulated data subset is converted from a two-dimensional or three-dimensional rendering generated by landscape or architectural design software, ensuring that training and testing can be performed even when actual data is lacking. The real data subset collects plane images of existing urban parks through aerial drones and synthesizes them using image stitching technology to ensure the authenticity and integrity of the data. The plane images of existing urban parks in both subsets are resized to 2024 pixels × 2024 pixels and pixel-level semantic segmentation masks are created. At the same time, the attributes of existing urban parks are clearly divided into five categories: road areas, leisure areas, vegetation areas, water areas, and building areas, and artificial mask annotation is performed to make the data more accurate, providing a reliable basis for subsequent identification.
[0024] Step S2: Select the Everything mode of the visual large model SAM pre-trained on the large-scale image segmentation dataset SA-1B to automatically perform attribute partitioning and rough extraction on the existing urban park plane image. The generated rough extraction result is saved in the form of color RGB image, and different colors are extracted and counted. There are n kinds of colors, and after sorting according to the R value of the color, the pixels in different color areas are assigned I k , then I k =255k / n, k represents the color number;
[0025] In this step, the Everything mode of the visual large model SAM pre-trained on the large-scale image segmentation dataset SA-1B is selected to automatically partition and roughly extract the attributes of the existing urban park plane images, generate color RGB image results, improve the recognition efficiency, and reduce the need for manual intervention.
[0026] Step S3: Modify the embedded part of the visual large model SAM image, delete the original structure and add the existing urban park attribute partition segmentation head, the segmentation head is composed of multiple types of operation blocks connected in series, from front to back, respectively, deconvolution block 1-deconvolution block 2-convolution block 1-interpolation block 1-interpolation block 2-convolution block 2-convolution block 3, wherein the output size of the deconvolution block 1 and the deconvolution block 2 is twice the input size, and the output size of the deconvolution block 1 is 256 pixels × 256 pixels , the output size of the deconvolution block 2 is 512 pixels × 512 pixels, the convolution kernel of the convolution block 1 is 1 × 1, the purpose of which is to adjust the number of channels, the adjusted number of channels is 5, the output size of the interpolation block 1 and the interpolation block 2 is twice the input size, using polynomial interpolation, the output size of the interpolation block 1 is 1024 pixels × 1024 pixels, the output size of the interpolation block 2 is 2048 pixels × 2048 pixels, and the convolution kernel of the convolution block 2 and the convolution block 3 is 3 × 3;
[0027] This step modifies the part after the SAM image of the visual large model is embedded, and adds an existing urban park attribute partition segmentation head, which is composed of various types of operation blocks and can better adapt to the task of identifying the existing urban park attribute partitions and improve the accuracy of identification.
[0028] Step S4: Calculate the color feature correlation vector K of the rough extraction result of the existing urban park attribute partition obtained in step S2 a and the multi-attribute partition feature correlation vector K of the feature map output by interpolation block 2 in the segmentation header in step S3 b , the calculation method is as follows: Assume the rough extraction result is Get the color separation matrix H a =f a (Ψ), where f a is the minimum similarity function between different colors; the feature map output by the interpolation block 2 is Get the multi-attribute partition separation matrix Η b =f b (Ψ), where f b (·) is the minimum similarity function between different attribute partitions; finally calculate K a ,K b = Θ(H a Ψ,H b Z), where Θ(·) is the maximum similarity function, and then the two related vectors are concatenated and input into the convolution block 2 in step S3 to form a fine-extracted visual model;
[0029] This step calculates the color feature correlation vector K a and multi-attribute partition feature correlation vector K b, which can fully consider the color information in the rough extraction results of the existing urban park attribute partitions and the multi-attribute partition information of the feature map output by the interpolation block 2 in the segmentation head. Among them, the calculation of the separation matrix between colors and the separation matrix between multi-attribute partitions uses the minimum similarity function, which can effectively highlight the differences between different colors and different attribute partitions, thereby improving the accuracy of subsequent identification. Among them, the use of the maximum similarity function, after the two related vectors are spliced, is input into the convolution block 2, so that the model can comprehensively consider the characteristics of color and attribute partitions, and further enhance the model's ability to identify the attribute partitions of existing urban parks. The calculation method of this step can adapt to different types of existing urban park plane images. Whether it is a simulated data subset or a real data subset, it can be accurately identified by analyzing the color and attribute partition characteristics. For different existing urban park attribute partitions, such as road areas, leisure areas, vegetation areas, water areas and building areas, this method can effectively distinguish them according to their characteristic differences, thereby improving the adaptability and generalization ability of the model. The process of calculating the related vectors and separation matrices in this step is actually a further extraction and optimization of the model features. In this way, the fine-extraction visual large model can be more focused on the identification task of existing urban park attribute partitions, reduce the interference of irrelevant information, and improve the performance and efficiency of the model.
[0030] Step S5: adopting multi-level segmentation loss to train the fine extraction visual large model using the existing urban park attribute partition identification data set, and dividing the data set into training set, validation set and test set (all test sets are real images) in a ratio of 8:1:1. During training, the fine extraction visual large model image embedding and previous model parameters adopt the model parameters of the visual large model SAM pre-trained on SA-1B, and the remaining model parameters adopt the He initialization method. The multi-level segmentation loss is divided into three items, namely pixel-level segmentation loss, edge-level segmentation loss and region-level segmentation loss. The three types of losses are weighted to form the final multi-level segmentation loss. The weights of the three types of losses are all 1.0. The optimizer used during training is SGD, and the initial learning rate is set to 0.001; the pixel-level segmentation loss, edge-level segmentation loss and region-level segmentation loss use mean square error loss function MSE, image similarity measurement SSIM, loss function and DICE function to calculate the difference between the predicted partition attribute category value and the labeled partition attribute category value;
[0031] This step uses multi-level segmentation losses to train the fine-extraction visual large model, including pixel-level segmentation loss, edge-level segmentation loss and region-level segmentation loss. The three types of losses are weighted to form the final multi-level segmentation loss, so that the model is optimized at different levels, further improving the accuracy of recognition.
[0032] Step S6: Use the trained fine extraction visual large model to process the existing urban park plane image to be identified, and generate the preliminary results of the existing urban park attribute partition identification, such as Figure 2 As shown in the figure, the preliminary result is five matrices of 2024 pixels × 2024 pixels. Each matrix corresponds to a category of an attribute partition. The value of each pixel in the matrix is the probability of belonging to that category. For each pixel, the category with the highest probability in the five matrices is taken as the final category of the pixel, which is then converted into the final recognition result of the existing urban park attribute partition.
[0033] In this step, the trained fine-extraction visual large model is used to process the existing urban park plane images to be identified, and the preliminary results are generated as 5 matrices. Each matrix corresponds to an attribute partition category. The category with the highest probability is taken as the final category of the pixel point and converted into the final identification result. This method is flexible and accurate and can adapt to different existing urban park plane images.
[0034] Example 2
[0035] In order to test the effectiveness and accuracy of the trained existing urban park attribute zoning identification model proposed by the present invention, the present invention adopts Deeplab v3p, one of the most advanced methods in the field of semantic segmentation, and a method of building a fine extraction model strategy that fuses coarse extraction results without using step S4 as a comparison method, and uses the test set in the existing urban park attribute zoning identification data set created according to step S1 for evaluation. Table 1 is a comparison of the existing urban garden functional zoning identification results of the two methods on the test set, using the most common DICE coefficient in the semantic segmentation field as the evaluation index. It can be seen that the method of the present invention is significantly superior to the Deeplab v3p model and the method without using step S4 in the identification of 5 types of functional zoning, which illustrates the accuracy of the present invention, especially the necessity of step S4. At the same time, the present invention is also compared with manual investigation and mapping. For a city park of 1.2 km × 1.8 km, the technical route of the present invention only takes about 2 hours (1.5 hours for drone mapping, 25 minutes for data sorting, and 3 minutes for identification), while the manual investigation and mapping method takes about 12 hours, involving 2 investigators and 1 mapper, of which the investigation takes about 9 hours and the mapping takes about 3 hours, which illustrates the high efficiency of the present invention.
[0036] Table 1 Dice scores of the two methods for identifying the functional zoning of existing urban parks on the test set (×100%)
[0037]
[0038] This example verifies the accuracy and effectiveness of the algorithm proposed in the present invention.
Claims
1. A method for identifying the attribute partitions of existing urban parks based on a pre-trained visual large model, characterized in that: Here are the steps: Step S1: Create an existing urban park attribute partition identification data set, including a simulation data subset and a real data subset. The plane image of the existing urban park in the simulation data subset is converted from a two-dimensional or three-dimensional rendering generated by a landscape or architectural design software. The two-dimensional rendering is directly saved as a vector diagram, and the three-dimensional rendering needs to save the top view as a vector diagram. The plane image of the existing urban park in the real data subset uses an aerial drone to collect the plane image of the existing urban park, and the image stitching technology is used to synthesize the aerial images of the existing urban park into a plane image. The plane images of the existing urban parks in the two subsets need to be resized to 2024 pixels × 2024 pixels and pixel-level semantic segmentation masks are created; the existing urban park attribute partitions include 5 categories, namely road area, leisure area, vegetation area, water area and building area, and different numerical values are used to represent these 5 types of partitions and manual mask annotation is performed; Step S2: Select the Everything mode of the visual large model SAM pre-trained on the large-scale image segmentation dataset SA-1B to automatically perform attribute partitioning and rough extraction on the existing urban park plane image. The generated rough extraction result is saved in the form of color RGB image, and different colors are extracted and counted. There are n kinds of colors, and after sorting according to the R value of the color, the pixels in different color areas are assigned I k , then I k =255k / n, k represents the color number; Step S3: modify the embedded part of the visual large model SAM image, delete the original structure and add the existing urban park attribute partition segmentation head, the segmentation head is composed of multiple types of operation blocks connected in series, from front to back, respectively, deconvolution block 1-deconvolution block 2-convolution block 1-interpolation block 1-interpolation block 2-convolution block 2-convolution block 3; Step S4: Calculate the color feature correlation vector K of the rough extraction result of the existing urban park attribute partition obtained in step S2 a and the multi-attribute partition feature correlation vector K of the feature map output by interpolation block 2 in the segmentation header in step S3 b , the calculation method is as follows: Assume the rough extraction result is Get the color separation matrix H a =f a (Ψ), where f a is the minimum similarity function between different colors; the feature map output by the interpolation block 2 is Get the multi-attribute partition separation matrix Η b =f b (Ψ), where f b (·) is the minimum similarity function between different attribute partitions; finally calculate K a ,K b =θ(H a Ψ,H b Z), where Θ(·) is the maximum similarity function, and then the two related vectors are concatenated and input into the convolution block 2 in step S3 to form a fine-extracted visual model; Step S5: Use the existing urban park attribute partition recognition data set to train the fine extraction visual large model using the multi-level segmentation loss. During training, the image embedding and previous model parameters of the fine extraction visual large model adopt the model parameters of the visual large model SAM pre-trained on SA-1B. The remaining model parameters are initialized using He. The multi-level segmentation loss is divided into three items, namely pixel-level segmentation loss, edge-level segmentation loss, and region-level segmentation loss. The three types of losses are weighted to form the final multi-level segmentation loss. The weights of the three types of losses are all 1.
0. The optimizer used during training is SGD, and the initial learning rate is set to 0.001; Step S6: Use the trained fine extraction visual large model to process the existing urban park plane image to be identified, and generate preliminary results of the existing urban park attribute partition identification. The preliminary results are 5 2024 pixel × 2024 pixel matrices, each matrix corresponds to a category of an attribute partition, and the value of each pixel in the matrix is the probability of belonging to the category. For each pixel, the category with the highest probability in the five matrices is taken as the final category of the pixel, which is then converted into the final identification result of the existing urban park attribute partition.
2. According to claim 1, a method for identifying attribute partitions of existing urban parks based on a pre-trained visual large model is characterized in that: In step S1, the aerial photography altitude needs to be set at 50 to 100 meters, and the number of routes and images to be taken is selected according to the scale of the urban park, but the number of routes shall not exceed 5, and the number of images for each route shall not exceed 10.
3. According to claim 1, a method for identifying attribute partitions of existing urban parks based on a pre-trained visual large model is characterized in that: In step S3, the output size of the deconvolution block 1 and the deconvolution block 2 is twice the input size, the convolution kernel of the convolution block 1 is 1×1, and its purpose is to adjust the number of channels. The output size of the interpolation block 1 and the interpolation block 2 is twice the input size, and the polynomial interpolation method is used. The convolution kernel of the convolution block 2 and the convolution block 3 is 3×3.
4. The method for identifying the attribute partitions of existing urban parks based on a pre-trained visual large model according to claim 1 is characterized in that: In step S5, the pixel-level segmentation loss, edge-level segmentation loss and region-level segmentation loss use the mean square error loss function MSE, image similarity measurement SSIM, loss function and DICE function to calculate the difference between the predicted partition attribute category value and the labeled partition attribute category value.
5. A system for identifying the attribute partitions of existing urban parks based on a pre-trained visual large model, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements a method for identifying attribute zoning of existing urban parks based on a pre-trained visual large model as described in any one of claims 1 to 4.
6. A computer-readable storage medium, on which a program for implementing information transmission is stored, and when the program is executed by a processor, the steps of a method for identifying attribute zoning of existing urban parks based on a pre-trained visual large model as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Method for displaying attribute information and related device
CN108737244A
Geographic aerial image segmentation data augmentation method based on visual general large model
CN118447413A