Photograph land type identification method based on artificial intelligence
The field photo recognition is carried out through the deep convolutional neural network model, which solves the efficiency and accuracy of manual investigations in field verification, and realizes efficient and automated photo recognition, which improves the management level of natural resource monitoring.
Patent Information
- Application Number
- CN202510424924.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-09-02
AI Technical Summary
The data of the field verification and evidence-producing photos in natural resource monitoring is large in volume and a wide variety of data, manual field investigation work is large, and the operating standards of different personnel are inconsistent, which affects the quality of work results. How to achieve efficient and automated photo recognition through artificial intelligence technology to reduce labor and time costs.
The photo classification interpretation model based on deep convolutional neural network is adopted, and the training set training optimization is used to identify field photos, including data preprocessing, model selection and optimization, and the cross-entropy loss function and evaluation indicators are used to optimize the model performance.
It realizes automatic identification and probability estimation of field photos, improves work efficiency and accuracy, reduces labor costs, and the photo classification accuracy reaches more than 85%, supporting rapid response and management of natural resource monitoring.
Smart Images

Figure CN120580573A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geographic information system data technology, and in particular relates to a photo land type recognition method based on artificial intelligence. Background Art
[0002] Natural resource surveys and monitoring are major national foundational tasks. The results of natural resource surveys and monitoring serve as crucial foundational data for natural resource management, ecological civilization development, and economic and social development. Fieldwork plays a vital role in natural resource monitoring and supervision, directly impacting data accuracy and the effectiveness of resource management. First, field surveys can collect the most direct and up-to-date on-site information. These raw data are the basis for formulating natural resource management policies and plans. Second, field inspections can uncover activities such as illegal land occupation and illegal mining, providing clues for law enforcement agencies and promoting the effective implementation of laws and regulations. Fieldwork is not only a bridge between theoretical research and practical application, but also an important means of ensuring the sustainable use of natural resources.
[0003] In the natural resource supervision and management of the Guangxi Zhuang Autonomous Region, field verification and evidence photos are of great significance for discovering and proving illegal land use. These photos, taken on the spot, can visually record land use conditions, helping relevant departments identify and confirm whether illegal land use exists. In current natural resource monitoring and supervision, firstly, field verification usually requires a large amount of manual field investigation, especially when involving large areas, remote areas, or multiple reviews, which is a heavy workload and time-sensitive. Secondly, the amount of data collected in the field is huge and diverse, with field evidence photos ranging from millions to tens of millions of photos each year. How to efficiently organize, analyze, and store this data is a considerable challenge. Furthermore, the different levels of professional knowledge and technical skills of different personnel, as well as inconsistent operating standards, will affect the quality of the final work results.
[0004] Therefore, it is very important to comprehensively use modern deep learning technology to perform intelligent and automatic photo recognition, carry out efficient operations, reduce the manpower and time costs of natural resource surveys, and significantly improve the overall level of natural resource management. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention proposes an artificial intelligence-based photo land classification identification method, which can use artificial intelligence recognition technology to quickly identify the land classification of evidence photos based on massive field photo samples with local vegetation geographical characteristics, effectively reducing the manpower required for land classification of field evidence photos. By introducing artificial intelligence technology, the dependence of such work on manpower is reduced, and personnel are freed from such repetitive and mechanical work, reducing labor costs and time costs, thereby improving the efficiency of investigation and monitoring work.
[0006] The present invention provides a method for identifying land types in photos based on artificial intelligence, comprising:
[0007] Obtaining data to be identified;
[0008] The data to be identified is input into a photo classification and interpretation model to obtain a classification result, wherein the photo classification and interpretation model is constructed by a deep convolutional neural network model and is obtained by post-training evaluation optimization of a training set, and the training set is photo data.
[0009] Optionally, obtaining the training set includes:
[0010] Get photo samples;
[0011] The photo samples are labeled according to the land category to which the photos belong, to obtain the training set.
[0012] Optionally, obtaining the photo samples includes: collecting photos at different seasons of the year, different sun exposure angles within a day, different weather conditions, and different shooting angles.
[0013] Optionally, before inputting the data to be identified into the photo classification and interpretation model, the method further includes:
[0014] Preprocessing the data to be identified to obtain processed data;
[0015] Select a pre-trained deep convolutional neural network model based on the processed data.
[0016] Optionally, preprocessing the data to be identified includes: performing random cropping, rotation and flipping, color dithering, grayscale processing, and text information removal on the data to be identified.
[0017] Optionally, the selection of a pre-trained deep convolutional neural network model is based on: dataset size, number of categories, image resolution, sample variance, real-time requirements, and computational resource limitations.
[0018] Optionally, optimizing the deep convolutional neural network model includes:
[0019] Train the existing mature deep convolutional neural network model and use the cross entropy loss function to perform targeted optimization on the selected deep convolutional neural network model.
[0020] Optionally, targeted optimization is performed on the selected deep convolutional neural network model, including at least: feature extraction optimization, output adjustment, data enhancement optimization, loss function optimization, and post-processing optimization.
[0021] Optionally, the cross entropy loss function is:
[0022]
[0023] Where: M is the number of categories; N is the number of samples; y ic is a sign function, which takes 1 if the true category of sample i is equal to c, otherwise it takes 0; p ic is the predicted probability that the observed sample i belongs to category c.
[0024] Optionally, the trained photo classification and interpretation model is evaluated using evaluation indicators, wherein the evaluation indicators include: accuracy, precision, recall rate and F1 score.
[0025] Compared with the prior art, the present invention has the following advantages and technical effects:
[0026] The present invention can automatically and high-quality perform land classification and probability estimation on photos with different pixels during field work. First, in terms of improving work efficiency, the traditional method of relying on the human eye to identify land classification in pictures can only process one picture at a time, while the present invention relies on artificial intelligence to perform batch land classification on field photos, which naturally greatly improves efficiency; in terms of working time, relying on the operator's work to identify the land classification of field photos takes 7-10 hours a day, while the present invention relies on computer intelligent operation to work 24 hours a day; in terms of the accuracy of photo land classification, the trained model has an overall classification accuracy of not less than 85% and a recall rate of not less than 85% for cultivated land, gardens, woodlands, piled earth, roads and other six types of land. Artificial intelligence photo land classification technology greatly improves the speed and accuracy of field photo processing, reduces labor costs, can extract valuable information from massive image data, and quickly and automatically identify photo land classifications, which is of great significance for improving the level of natural resource monitoring and supervision. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0028] Figure 1 This is a flow chart of a method for identifying land types in photos based on artificial intelligence according to an embodiment of the present invention;
[0029] Figure 2 Schematic diagram of the processing results of the embodiment of the present invention. DETAILED DESCRIPTION
[0030] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0031] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0032] This embodiment proposes a method for identifying land types in photos based on artificial intelligence. Figure 1 As shown, the specific steps include:
[0033] Obtaining data to be identified;
[0034] The data to be identified is input into the photo classification and interpretation model to obtain the classification results. The photo classification and interpretation model is constructed through a deep convolutional neural network model and is obtained through post-training evaluation optimization of the training set, where the training set is photo data.
[0035] Specifically, current natural resource monitoring requires the use of scientific methods and technologies for data collection and analysis, the establishment of a long-term, continuous monitoring system covering different temporal and spatial scales, and the realization of instant data transmission and processing whenever possible to ensure a scientific, systematic, accurate, and real-time response to changes. In natural resource monitoring and supervision, field photographs can serve as first-hand information on on-site conditions, providing direct visual evidence of changes in resource conditions. Trends in natural resource changes can also be tracked by comparing photographs taken at different time points. Therefore, during annual change surveys, it is crucial to use artificial intelligence technology to quickly and accurately identify field photographs of suspected change sites.
[0036] Furthermore, obtaining a training set includes:
[0037] Get photo samples;
[0038] According to the land category to which the photo belongs, the photo samples are labeled to obtain the training set.
[0039] Specifically, when preparing intelligent recognition data, regarding the quality of field photos, first, high-resolution photos should be selected. Sufficient resolution allows for clear detail, which is particularly important for identifying specific plant species or subtle land changes. High-resolution photos also facilitate zooming in on local areas, enabling more precise feature extraction and analysis. Second, photos should be stable and clear to clearly capture the target. Third, the target ground features should account for at least 50% of the evidence photos to avoid background interference and model misdetection or omission. Regarding the richness of field photos, the photo data can primarily be from regional land change surveys, including photos from different seasons of the year, different sun angles throughout the day, different weather conditions, and different shooting angles. After collecting a sufficient number of photo samples, manually identify the land type of the photos and indicate the land type to which they belong. Photos with the same land type identified manually are placed in different folders. Sample photos should be in common formats such as JPEG, JPG, PNG, and BMP, and uploaded in a zip file. After collecting and organizing the photo samples, they are divided into training, validation, and test sets. The directory structure and file format for importing the photo classification sample library are as follows:
[0040] ├──<Category 1>#The directory name is the photo category, and the directory contains photos belonging to this category. The photo formats supported are jpg, jpeg, png, and bmp.
[0041] │├──00000000_00000000.jpg
[0042] │├──00000000_00000000.jpg
[0043] │├──......
[0044] ├──<Category 2>#The directory name is the photo category, and the directory contains photos belonging to this category. The photo formats supported are jpg, jpeg, png, and bmp.
[0045] │├──00000000_00000000.png
[0046] │├──00000000_00000000.png
[0047] │├──......
[0048] ├──<Category N>#The directory name is the photo category, and the directory contains photos belonging to this category. The photo formats supported are jpg, jpeg, png, and bmp.
[0049] │├──00000000_00000000.png
[0050] │├──00000000_00000000.png
[0051] │├──......
[0052] Furthermore, obtaining the photo samples includes collecting photos at different seasons of the year, different sun exposure angles within a day, different weather conditions, and different shooting angles.
[0053] Furthermore, before inputting the data to be identified into the photo classification and interpretation model, the following steps are also included:
[0054] Preprocess the data to be identified and obtain the processed data;
[0055] Select a pre-trained deep convolutional neural network model based on the processed data.
[0056] Specifically, we train an intelligent photo classification and interpretation model. Depending on the complexity of the task, we select a mature deep convolutional neural network model, such as ResNet (Residual Network), EfficientNet, or Vision Transformer. These models have strong image feature extraction capabilities. We then perform preprocessing on the sample photo data, including random cropping, rotation and flipping, color dithering, grayscale conversion, and text removal. Use the training set and validation set to fine-tune the selected deep convolutional neural network model, set appropriate hyperparameters, and stop training when the model reaches the target accuracy; the initial learning rate can be set between [10-5, 10-1]; the minimum number of GPUs is 1, which can be set according to the actual number of servers; the single-card batch size represents the number of samples used by a single GPU in the last iteration, with a value range of [1, 4] and a default value of 2; the model save interval can be set as needed, and the recommended number of saves is 10-20; by adjusting the convolution kernel size of the model and adding the attention mechanism, the perception ability of key areas is improved; use the cross-entropy loss function and an appropriate optimizer (such as Adam: Adaptive Moment Estimation adaptive moment estimation algorithm, SGD: Stochastic Gradient Descent random gradient descent, etc.) for training, and use the optimizer to make the cross-entropy loss function smaller when possible, so that the model can more accurately judge the land type of field photos.
[0057] Furthermore, preprocessing the data to be recognized includes: performing random cropping, rotation and flipping, color dithering, grayscale processing, and text information removal on the data to be recognized.
[0058] Furthermore, the selection of pre-trained deep convolutional neural network models is based on: dataset size, number of categories, image resolution, sample differences, real-time requirements, and computing resource limitations.
[0059] Furthermore, the optimization of the deep convolutional neural network model includes:
[0060] Train the existing mature deep convolutional neural network model and use the cross entropy loss function to perform targeted optimization on the selected deep convolutional neural network model.
[0061] Furthermore, the selected deep convolutional neural network model is optimized in a targeted manner, including at least: feature extraction optimization, output adjustment, data enhancement optimization, loss function optimization and post-processing optimization.
[0062] Furthermore, the cross entropy loss function is:
[0063]
[0064] Where: M is the number of categories; N is the number of samples; y ic is a sign function, which takes 1 if the true category of sample i is equal to c, otherwise it takes 0; p ic is the predicted probability that the observed sample i belongs to category c.
[0065] Furthermore, the trained photo classification and interpretation model is evaluated using evaluation indicators, where the evaluation indicators include accuracy, precision, recall rate and F1 score.
[0066] Specifically, to evaluate the intelligent photo classification and interpretation model, we first evaluate the model on a test machine. After obtaining the results, we analyze the existing problems, adjust the structure and parameters of the sample or model, and rely on indicators such as accuracy, precision, recall rate, and F1 score to measure the accuracy of the model.
[0067] More specifically, to use the intelligent photo classification and interpretation model for field photo recognition, the model should first be deployed in a computer environment, enabling the interface to batch upload field photos and receive photos to be classified and interpreted. The model intelligently identifies the land type of field photos and outputs information such as field photo information, predicted land type, and predicted probability. The JSON structure is as follows:
[0068]
[0069] More specifically, for model iteration and optimization, you can use indicators such as accuracy, precision, recall, and F1 score to monitor and collect problems in the actual application of the model; and iteratively train the model based on the solutions to the collected problems. Taking into account the actual use of the model, the scenarios in which the model is used include gardens, woodlands, and cultivated land. Especially in Guangxi, the features are similar and easily confused. It is expected that the DaViT (Dual Attention Vision Transformer) visual analysis model network structure will be used. It combines the advantages of convolution and Transformer, and improves the performance of visual tasks by combining global attention and local attention. It can better capture local and global information in the image, more effectively extract spatial information and contextual information, and has good generalization performance when the amount of data is small. It is a relatively excellent model that can adapt to various image classification scenarios.
[0070] More specifically, the DaViT (Dual Attention Vision Transformer) model adopts a pyramidal hierarchical architecture, achieving multi-scale feature learning through four progressive processing stages. Its core consists of convolution-guided feature embedding and a two-stream attention mechanism: the model first performs efficient initial feature extraction through 7×7 large kernel convolution (stride = 4), then alternates between 2×2 strided convolutions (stride = 2) in four stages to achieve patch embedding with resolution reduction and channel dimension increase, and stacks a dual attention module consisting of windowed spatial attention and global channel attention.
[0071] The processing result diagram of this embodiment is as follows Figure 2 As shown, this embodiment provides an artificial intelligence-based photo land classification identification method, which is based on a deep learning remote sensing interpretation algorithm, uses field photo samples with regional characteristics and a pre-trained deep convolutional neural network model as materials, and learns color, texture, edge and other information of photos at multiple spatial window scales. It can perform multi-directional, automated and high-quality land classification identification on field photos and judge the probability that the photo belongs to the land class, meeting the need for rapid response to change detection in natural resource monitoring and supervision, complementing field work, and jointly providing a solid foundation for the effective management and protection of natural resources.
[0072] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for identifying land types in photos based on artificial intelligence, characterized in that: include: Obtaining data to be identified; The data to be identified is input into a photo classification and interpretation model to obtain a classification result, wherein the photo classification and interpretation model is constructed by a deep convolutional neural network model and is obtained by post-training evaluation optimization of a training set, and the training set is photo data.
2. The method for identifying land types based on photos based on artificial intelligence according to claim 1, characterized in that: Obtaining the training set includes: Get photo samples; The photo samples are labeled according to the land category to which the photos belong, to obtain the training set.
3. The method for identifying land types based on photos according to claim 2, characterized in that: Obtaining photo samples includes: collecting photos at different seasons of the year, different angles of sunlight within a day, different weather conditions, and different shooting angles.
4. The method for identifying land types based on photos based on artificial intelligence according to claim 1, characterized in that: Before inputting the data to be identified into the photo classification and interpretation model, the following steps are also included: Preprocessing the data to be identified to obtain processed data; Select a pre-trained deep convolutional neural network model based on the processed data.
5. The method for identifying land types based on photos using artificial intelligence according to claim 4, characterized in that: The pre-processing of the data to be identified includes: performing random cropping, rotation and flipping, color dithering, grayscale processing and text information removal on the data to be identified.
6. The method for identifying land types based on photos according to claim 4, characterized in that: The selection of a pre-trained deep convolutional neural network model is based on: dataset size, number of categories, image resolution, sample variance, real-time requirements, and computational resource limitations.
7. The method for identifying land types based on photos based on artificial intelligence according to claim 1, characterized in that: Optimizing the deep convolutional neural network model includes: Train the existing mature deep convolutional neural network model and use the cross entropy loss function to perform targeted optimization on the selected deep convolutional neural network model.
8. The method for identifying land types based on photos using artificial intelligence according to claim 7, characterized in that: Targeted optimization of the selected deep convolutional neural network model, including at least: feature extraction optimization, output adjustment, data enhancement optimization, loss function optimization and post-processing optimization.
9. The method for identifying land types based on photos according to claim 8, characterized in that: The cross entropy loss function is: Where: M is the number of categories; N is the number of samples; y ic is a sign function, which takes 1 if the true category of sample i is equal to c, otherwise it takes 0; p ic is the predicted probability that the observed sample i belongs to category c.
10. The method for identifying land types based on photos based on artificial intelligence according to claim 1, characterized in that: The trained photo classification and interpretation model is evaluated using evaluation indicators, wherein the evaluation indicators include: accuracy, precision, recall rate and F1 score.