Geographical aerial image segmentation data augmentation method based on visual general large model

CN118447413BActive Publication Date: 2026-10-09YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410420556.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2026-10-09
Estimated Expiration
2044-04-09

AI Technical Summary

Technical Problem

数据量不足:目前的地理航拍影像数据集数量较少,且规模较小

Benefits of technology

第一,本发明使用了一种全新的利用零样本分割大模型对数据进行增广的方式,与传统的基于图像处理的数据增广方式相比,不用从种类繁多的图像操作中选择最合适的一种,极大的降低了人工操作的复杂度;而与新兴的利用深度学习模型和强化学习算法进行数据增广的方法相比,本发明大大减少了计算的复杂度,降低了计算资源的消耗。本发明解决了航拍地理图像上数据量不足、数据质量低、数据分布不均匀的问题,提高了地理分割的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447413B_ABST
    Figure CN118447413B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image segmentation and data augmentation, and discloses a geographic aerial image segmentation data augmentation method based on a visual general large model, which acquires an original geographic aerial image data set; for each image in the original geographic aerial image data set, data augmentation is performed using a zero-shot segmentation large model and original labeling; an aug suffix is added to the generated augmented labeling, an original image data is copied and an aug suffix is added, and then the two data are merged to form an augmented geographic aerial image data set; a deep learning model is used to train and segment the augmented geographic aerial image data set to obtain a final geographic aerial image segmentation model; and the geographic aerial image segmentation model is used to automatically segment new geographic aerial data. The present application reduces the complexity of manual operation, reduces the complexity of calculation, reduces the consumption of computing resources, and improves the accuracy of geographic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image segmentation and data augmentation technology, and in particular relates to a method for segmenting and augmenting geographic aerial images based on a general visual model. Background Technology

[0002] Aerial geographic imagery refers to geographic images captured by drones from the air. Compared to remote sensing satellite images, drone aerial images have advantages such as high resolution, fast response time, and fewer shooting restrictions. Aerial geographic imagery can reflect information such as crop growth status, soil quality, and disaster conditions, which is of great significance for agricultural production monitoring and management. To extract useful information from aerial geographic imagery, image segmentation is necessary, that is, distinguishing different objects in the image (such as farmland, greenhouses, water bodies, and houses) and labeling their categories and boundaries. The traditional method is to manually label different objects after acquiring the data, and then count the categories and quantities of each object. In recent years, with the continuous development of artificial intelligence, some have attempted to apply deep learning methods to geographic imagery, but most of these have focused on the identification of a specific type of crop or scene. Existing deep learning-based aerial geographic imagery segmentation methods face some serious challenges.

[0003] Based on the above analysis, the problems and shortcomings of the existing technology are as follows: Insufficient data: Current geographic aerial image datasets are limited in number and size. This makes it difficult for deep learning models to fully learn the features and patterns of various object categories in the images, easily leading to overfitting or underfitting problems.

[0004] Low data quality: Errors and inconsistencies in the manual annotation process result in low data quality, which affects the performance of deep learning models.

[0005] Uneven data distribution: Due to the different frequencies and proportions of different objects (such as farmland, greenhouses, water bodies, houses, etc.) in geographic aerial images, the data distribution is uneven, making it difficult for deep learning models to process each category in a balanced way, and easily leading to bias towards certain categories or neglect of certain categories.

[0006] Existing technology analysis 1) Deep Learning-Based Aerial Image Segmentation: Currently, deep learning, especially convolutional neural networks (CNNs), has been widely researched and applied in the field of aerial image segmentation. These methods typically rely on large amounts of labeled data for training to achieve high-precision land cover classification and segmentation. These models are able to learn mappings from raw pixels to complex land cover categories, thereby effectively identifying and segmenting different land covers in new aerial imagery.

[0007] Problems with existing technology 1) Dependence on large amounts of labeled data: Traditional deep learning methods typically require a large number of labeled samples for training. In the field of geographic aerial imagery, obtaining large amounts of high-quality, accurately labeled data is extremely difficult and costly. Therefore, this dependence limits the universality and application scope of the model.

[0008] 2) Limitations in generalization ability: Traditional geographic aerial image segmentation models often perform well on specific datasets, but their segmentation accuracy drops significantly when faced with novel or unseen land cover types. This is because the model's generalization ability is limited; it can only identify and segment the categories contained in the training data.

[0009] 3) High computational resource requirements: Deep learning models, especially large CNN models, have high computational resource requirements. This includes not only resource consumption during the training phase, but also resource requirements during model deployment and application. For applications that require real-time or near real-time processing of large amounts of aerial imagery, this becomes a limiting factor.

[0010] 4) Limitations of data augmentation: Traditional data augmentation techniques, such as rotation, cropping, and color adjustment, can increase the robustness of the model to some extent, but they cannot extend the model's ability to identify unknown categories. This limits the effectiveness of the model when facing diverse and dynamically changing real-world environments. Summary of the Invention

[0011] To address the problems existing in the prior art, this invention provides a method for augmenting geographic aerial image segmentation data based on a visual universal large model.

[0012] This invention is implemented as follows: a method for augmenting geographic aerial image segmentation data based on a universal visual model, comprising the following steps: Step 1: Obtain the original geographic aerial image dataset, which contains several geographic aerial images with category labels and boundary annotations; Step 2: For each image in the original geographic aerial imagery dataset, perform data augmentation using a zero-shot segmentation model and the original annotations; Step 3: Add the suffix "aug" to the generated augmented labels, copy the original image data and add the suffix "aug" to it, and then merge the two datasets to form an augmented geographic aerial image dataset. Step 4: Use a deep learning model to train and segment the augmented geographic aerial image dataset to obtain the final geographic aerial image segmentation model. Step 5: Use the geographic aerial image segmentation model to automatically segment the new geographic aerial data.

[0013] Furthermore, the dataset in step one is a general form of image segmentation dataset, which is divided into two parts: image data and labeled data.

[0014] Furthermore, the image data is an RGB three-channel color image, the annotation file is a single-channel grayscale image, and the pixel values ​​start from 1 and represent different categories in sequence, with 0 representing the background.

[0015] Furthermore, the model in step two can segment untrained categories, taking an unknown image as input and segmenting it into various categories, or segmenting a specified category based on a small amount of clue information.

[0016] Furthermore, the data augmentation method in step two includes the following steps: Step 1: Split the different categories based on the differences in the original labeled pixel values, ensuring that each sub-image contains only background pixels (0) and pixels of the other category; Step 2: Extract zero-shot model segmentation hints based on the location information of the category pixel values ​​in each sub-image; Step 3: Input the image and the extracted prompt information into the zero-shot segmentation model. The model will segment the specified region of the image based on the prompt information. Step 4: Merge the segmentation results and the original annotations according to certain judgment rules; Step 5: Merge the fusion results of each category to form the complete augmented annotation for the image.

[0017] Another objective of this invention is to provide a geographic aerial image segmentation data augmentation system based on a visual universal large model, comprising: The dataset acquisition module is used to acquire the original geographic aerial image dataset, which contains several geographic aerial images with category labels and boundary annotations; The data augmentation module is used to augment each image in the original geographic aerial imagery dataset using a zero-shot segmentation large model and the original annotations. The data merging module is used to add the .aug suffix to the generated augmented labels, copy the original image data and add the .aug suffix, and then merge the two data to form an augmented geographic aerial image dataset. The segmentation model building module is used to train and segment the augmented geographic aerial image dataset using a deep learning model to obtain the final geographic aerial image segmentation model. The automatic segmentation module is used to automatically segment new geographic aerial data using a geographic aerial image segmentation model.

[0018] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows: First, this invention employs a novel method for data augmentation using a large-scale zero-shot segmentation model. Compared to traditional image processing-based data augmentation methods, it eliminates the need to select the most suitable method from a wide variety of image operations, significantly reducing the complexity of manual operations. Furthermore, compared to emerging methods utilizing deep learning models and reinforcement learning algorithms for data augmentation, this invention greatly reduces computational complexity and resource consumption. This invention addresses the problems of insufficient data volume, low data quality, and uneven data distribution in aerial geographic images, thereby improving the accuracy of geographic segmentation.

[0019] Secondly, this invention proposes an intelligent segmentation method for geographic aerial imagery based on a general visual model for data augmentation. It utilizes a zero-shot visual segmentation model to augment the original geographic imagery, expanding the scale and diversity of the dataset. Then, a deep learning model is used to train and segment the augmented imagery, thereby improving the accuracy and robustness of geographic aerial imagery segmentation and enabling automatic recognition and segmentation of geographic aerial imagery. This invention can provide more precise information support for agricultural production.

[0020] Third, the expected benefits and commercial value of the technical solution of this invention after transformation are as follows: Currently, the statistics on the quantity and category of geographic information are all based on manual annotation. This invention can automatically complete the segmentation of geographic information, greatly reducing the cost of statistical analysis of the quantity and category of geographic information, and has huge commercial benefits.

[0021] The technical solution of this invention fills a technical gap in the industry both domestically and internationally: This invention proposes a data augmentation method based on a general visual large model, filling the gap in the application of large models in data augmentation.

[0022] The technical solution of this invention solves a technical problem that people have long desired to solve but have never been able to: this invention solves the problem of high data statistics cost and slow statistics speed in the field of geographic information.

[0023] Fourth, the geographic aerial image segmentation data augmentation method based on a general visual model provided by this invention has achieved the following significant technical advancements compared to existing technologies: 1) Enhanced handling of unknown categories: By introducing zero-shot learning, this method can effectively handle categories not present in the training dataset. This is particularly important for segmentation of geographic aerial imagery, as these images often contain diverse features, many of which are not present in the training set. This advancement significantly improves the model's usability and adaptability.

[0024] 2) Reduced reliance on large amounts of labeled data: Traditional deep learning methods typically require a large amount of labeled data for geographic aerial image segmentation. Your method alleviates this requirement through data augmentation, enabling the training of effective models even with limited labeled data. This is particularly valuable when resources are limited.

[0025] 3) Improved generalization ability: In the processing of aerial geographic imagery, generalization ability is particularly crucial when dealing with images taken in different regions and at different times. Your method improves the model's adaptability to diverse inputs through augmentation and zero-shot learning, enabling the model to maintain high performance under different environments and conditions.

[0026] 4) Optimized data processing workflow: Your approach utilizes intelligent data augmentation to more efficiently leverage existing data, thereby optimizing the entire data processing and model training process. This increased efficiency is particularly important for processing large volumes of geographic aerial imagery.

[0027] 5) Improved automation and accuracy: By combining deep learning models and zero-shot learning, your method can automatically and accurately segment new geographic aerial data, improving the automation and accuracy of segmentation tasks.

[0028] The method provided by this invention offers a more flexible, accurate, and efficient solution for geographic aerial image segmentation, especially demonstrating significant technological advancements in handling unknown categories and reducing reliance on large amounts of labeled data. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of the geographic aerial image segmentation data augmentation method based on a visual universal large model provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the data augmentation process provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the specific implementation data provided in the embodiments of the present invention; Figure 4 This is a partial image of a geographic aerial photograph provided in an embodiment of the present invention; Figure 5 This is a diagram showing the intelligent segmentation result of a local image of aerial photography provided in an embodiment of the present invention; Figure 6 This is a structural diagram of a geographic aerial image segmentation data augmentation system based on a general visual model provided in an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0032] Regarding the geographic aerial image segmentation data augmentation method based on a general visual model provided by this invention, two specific embodiments and implementation schemes are given: Example 1: Urban Land Use Classification 1) Data preparation: Collect aerial images of urban areas, including images of different land use types such as residential areas, commercial areas, parks, and roads. These images should be labeled with category labels and boundary markings.

[0033] 2) Zero-shot segmentation model application: The zero-shot learning model is applied to process the original aerial imagery. By analyzing existing category labels and image features, the model can identify new categories in the imagery (such as unlabeled small parks or new building types).

[0034] 3) Data Augmentation: Based on the model's output, each image in the original dataset is augmented. The original image is copied and the .aug suffix is ​​added, resulting in an augmented image that includes the newly identified category.

[0035] 4) Model Training and Segmentation: Deep learning models, such as convolutional neural networks, are used to train the augmented dataset. Through this training, the model can more accurately identify and segment different types of urban land use.

[0036] 5) Application: The trained model is used to automatically segment new urban aerial images to achieve more accurate land use classification.

[0037] Example 2: Agricultural Crop Identification and Analysis 1) Data preparation: Collect aerial images of farmland containing different crop types (such as wheat, corn, and rice). The images should have clear crop category labels and boundary markings.

[0038] 2) Zero-shot segmentation model application: The original aerial imagery is processed using a zero-shot learning model. The model can identify new or uncommon crop types, and even crop areas damaged by disease.

[0039] 3) Data augmentation: Based on the model's output, each original aerial image is augmented to create an augmented image set containing the new recognition categories.

[0040] 4) Model training and segmentation: Deep learning models are used to train the augmented dataset, thereby improving the accuracy and robustness of the model in agricultural crop identification and health status analysis.

[0041] 5) Application: The trained model is used to analyze new aerial images of farmland to identify crop species, monitor crop health status, and predict yield.

[0042] The two embodiments provided in this invention demonstrate the broad applicability and effectiveness of your method in different application areas, particularly its advantages in handling unknown categories and reducing reliance on large amounts of labeled data.

[0043] To address the problems existing in the prior art, this invention provides a method for augmenting geographic aerial image segmentation data based on a visual universal large model.

[0044] like Figure 1 As shown in the embodiment of the present invention, the geographic aerial image segmentation data augmentation method based on a visual general large model first uses a drone to collect a certain number of geographic aerial images and performs geographic image segmentation.

[0045] Step 1: Obtain the original aerial imagery dataset; In the preliminary work, a geographic information dataset for this area was obtained, containing 12,000 images, divided into 5 categories: farmland, greenhouses, ponds, seedlings, and houses. The corresponding pixels are labeled from 1 to 5, with pixel 0 representing the background. The specific statistical data of the dataset are shown in Table 1 below: Table 1. Statistical data of the dataset

[0046] Step 2: For each image in the original geographic aerial imagery dataset, perform data augmentation using a zero-shot segmentation model and the original annotations. Figure 2 ); The general-purpose visual model uses the SAM model, which is a zero-shot segmentation model proposed and open-sourced by Meta AI. It can segment the corresponding target in the image based on simple augmentation information, including points, boxes, coarse masks or text cues. It has the advantages of fast segmentation speed, good segmentation effect and simple use.

[0047] Specific implementation data augmentation, such as Figure 3 As shown, perform the following operations on each image in the dataset: Step 1: Based on the differences in pixel values ​​of the original annotations, split the different categories to ensure that each sub-image contains only background pixels (0) and pixels of the other category. The original annotations were split into two categories: farmland and greenhouse, resulting in two sub-images: a farmland sub-image and a greenhouse sub-image.

[0048] Step 2: Extract the bounding boxes of rectangles from each independent connected component in each subgraph as cue information. In the farmland subgraph, where the farmland pixel value is 1, identify the independent connected components where each pixel is 1, and determine the four boundaries (top, bottom, left, and right). Perform the same operation for other subgraph categories. Then, extract the segmentation cue information based on the bounding boxes. Each subgraph corresponds to a list of bounding boxes, and each list stores several bounding boxes. The number of bounding boxes is the same as the number of independent connected components in the subgraph. Each bounding box is stored in the format [(x1, y1), (x1, y2), (x2, y2), (x2, y1)], representing the coordinates of the top-left, bottom-left, bottom-right, and top-right corners, respectively.

[0049] Step 3: Input the image and the extracted prompt information into the zero-shot segmentation model. The SAM model can accept various forms of prompt information, but the best results are achieved when the region to be segmented is delineated in the form of a prompt box. Therefore, the prompt box segmentation mode is adopted.

[0050] Step 4, the SAM segmentation results are as follows: Figure 3 The results are shown in the "SAM segmentation results".

[0051] Step 5: Merge the fusion results for each category to form the complete augmentation annotation mask for the image. aug .

[0052] Step 3: Add the suffix "aug" to the generated augmented labels to distinguish them from the original labels. Copy the original image data and add the "aug" suffix. Then merge the two datasets to form an augmented geographic information dataset, in which the number of images and labels is doubled. Then divide the dataset into training and validation sets in an 8:2 ratio.

[0053] Step 4: Train the VIT segmentation model using the pre-defined training and dataset sets to obtain a geographic information image segmentation model. Use this model to segment unprocessed areas of the image; the segmentation results are shown in the attached figure. Figure 4 and 5 As shown.

[0054] For the segmentation network model, this invention selects the VIT model, which stands for Vision Transformer. This model was the first to introduce the self-attention mechanism into the field of image segmentation and has powerful image analysis and processing capabilities.

[0055] Step 3: Add the suffix "aug" to the generated augmented labels. Copy the original image data and add the "aug" suffix to it. Then merge the two datasets to form an augmented geographic aerial image dataset. Step four: Use the divided training set and dataset to train the VIT segmentation model to obtain a geographic information image segmentation model.

[0056] The segmentation network model chosen is the VIT model, which stands for Vision Transformer. This model was the first to introduce the self-attention mechanism into the field of image segmentation and has powerful image analysis and processing capabilities.

[0057] Step 5: Use this model to segment other unprocessed regions of the image. The segmentation result is as follows: Figure 4 and 5 As shown.

[0058] like Figure 6 As shown, this embodiment of the invention provides a geographic aerial image segmentation data augmentation system based on a visual universal large model, which includes a method for augmenting geographic aerial image segmentation data based on a visual universal large model. The dataset acquisition module is used to acquire the original geographic aerial image dataset, which contains several geographic aerial images with category labels and boundary annotations; The data augmentation module is used to augment each image in the original geographic aerial imagery dataset using a zero-shot segmentation large model and the original annotations. The data merging module is used to add the .aug suffix to the generated augmented labels, copy the original image data and add the .aug suffix, and then merge the two data to form an augmented geographic aerial image dataset. The segmentation model building module is used to train and segment the augmented geographic aerial image dataset using a deep learning model to obtain the final geographic aerial image segmentation model. The automatic segmentation module is used to automatically segment new geographic aerial data using a geographic aerial image segmentation model.

[0059] This invention can be applied to agriculture. Through automatic segmentation, it can complete the statistical analysis of crop types and quantities in a designated area at extremely low cost, providing scientific basis and data support for agricultural policy making, food market monitoring, agricultural resource protection, and agricultural technological innovation.

[0060] This invention can be used in rural housing statistics. By automatically segmenting houses in the same area at different times, illegal buildings can be quickly identified.

[0061] This invention was tested on multiple datasets. The first dataset was an agricultural aerial photography dataset of a certain area, the second was the open-source dataset ADE20K, and the third was the COCO2017 dataset. On each dataset, this invention also tried using a variety of different deep neural networks. The experimental results show that the data augmentation algorithm of this invention improves the image segmentation effect to varying degrees. The results are shown in Table 2.

[0062] Table 2

[0063] Simultaneously, this invention also utilized the model trained on the first dataset for agricultural geography identification, with the identification results as follows: Figure 5 As shown.

[0064] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for augmenting geographic aerial image segmentation data based on a general visual model, characterized in that, Includes the following steps: Step 1: Obtain the original geographic aerial image dataset, which contains several geographic aerial images with category labels and boundary annotations; Step 2: For each image in the original geographic aerial imagery dataset, perform data augmentation using a zero-shot segmentation model and the original annotations; Step 3: Add the suffix "aug" to the generated augmented labels, copy the original image data and add the suffix "aug" to it, and then merge the two datasets to form an augmented geographic aerial image dataset. Step 4: Use a deep learning model to train and segment the augmented geographic aerial image dataset to obtain the final geographic aerial image segmentation model. Step 5: Use the geographic aerial imagery segmentation model to automatically segment the new geographic aerial data; The data augmentation method in step two includes the following steps: Step 1: Split the different categories based on the differences in the original labeled pixel values, ensuring that each sub-image contains only background pixels (0) and pixels of the other category; Step 2: Extract zero-shot model segmentation hints based on the location information of the category pixel values ​​in each sub-image; Step 3: Input the image and the extracted prompt information into the zero-shot segmentation model. The model will segment the specified region of the image based on the prompt information. Step 4: Merge the segmentation results and the original annotations according to certain judgment rules; Step 5: Merge the fusion results of each category to form the complete augmented annotation for the image; The specific judgment rules are expressed by the following formula: In the formula, This indicates the category label of the original annotation. This represents the category result identified by the large model in zero-shot segmentation; ⊕ is the XOR operation. It is a smoothing function used to remove small, fragmented, isolated regions from the XOR result, thus smoothing the entire expression's result. This indicates a significant difference between the zero-shot model segmentation results and the original annotations; The function is a ratio function, where S represents the area. This represents the original image, i.e. It is the ratio of the area of ​​the absolute difference between the original annotation and the zero-shot model recognition results to the original area of ​​the image; α is a specified threshold, when... When it is less than this threshold, use and The sum is used to represent the augmented label of the category, when When it exceeds this threshold, use To represent the augmentation label of this category, i.e. , where the subscript i represents the category number.

2. The method for augmenting geographic aerial image segmentation data based on a general visual model as described in claim 1, characterized in that, The dataset in step one is a general form of image segmentation dataset, which is divided into two parts: image data and labeled data.

3. The method for augmenting geographic aerial image segmentation data based on a general visual model as described in claim 2, characterized in that, The image data is an RGB three-channel color image, and the annotation file is a single-channel grayscale image. The pixel values ​​start from 1 and represent different categories in sequence, with 0 representing the background.

4. The method for augmenting geographic aerial image segmentation data based on a general visual model as described in claim 1, characterized in that, In step two, the model segments untrained categories. It takes an unknown image as input and segments the image into various categories, or segments a specified category based on a small amount of information.

5. The geographic aerial image segmentation data augmentation system based on a visual universal large model, as described in any one of claims 1 to 4, is characterized in that... The system includes: The dataset acquisition module is used to acquire the original geographic aerial image dataset, which contains several geographic aerial images with category labels and boundary annotations; The data augmentation module is used to augment each image in the original geographic aerial imagery dataset using a zero-shot segmentation large model and the original annotations. The data merging module is used to add the .aug suffix to the generated augmented labels, copy the original image data and add the .aug suffix, and then merge the two data to form an augmented geographic aerial image dataset. The segmentation model building module is used to train and segment the augmented geographic aerial image dataset using a deep learning model to obtain the final geographic aerial image segmentation model. The automatic segmentation module is used to automatically segment new geographic aerial data using a geographic aerial image segmentation model.

Citation Information

Patent Citations

  • Multispectral unmanned aerial vehicle remote sensing image crop segmentation method

    CN116503590A

  • Image classification data augmentation method and system

    CN117593581A