Crop sample construction method and system based on pseudo-label theory and multi-source remote sensing data collaboration
By combining pseudo-label theory and multi-source remote sensing data, multi-source remote sensing image data preprocessing, phenological feature screening and confidence screening are used to solve the timeliness and accuracy problems of traditional crop sample construction methods, and efficient and accurate crop sample generation is achieved.
Patent Information
- Application Number
- CN202510638503.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional crop sample construction method has problems such as low timeliness, low accuracy and high investment costs, and is difficult to adapt to multiple scenarios, which limits the application of deep learning models in crop sample construction.
The crop sample construction method based on pseudo-label theory and multi-source remote sensing data is adopted to determine the distribution space range, confidence screening and progressive learning strategy expansion samples through multi-source remote sensing image data preprocessing, phenological feature screening, and threshold segmentation methods to generate high-quality crop samples.
It improves the efficiency and accuracy of crop sample construction, reduces manpower investment, reduces errors caused by human factors, and achieves high-precision crop sample generation, which is suitable for large-scale crop monitoring and yield estimation.
Smart Images

Figure CN120182834A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of agricultural remote sensing data processing and analysis, and relates to a method and system for constructing crop samples, and particularly to a method and system for constructing crop samples based on the collaboration of pseudo-label theory and multi-source remote sensing data. Background Art
[0002] The growth of the population and the improvement of living standards have led to a rapid increase in food demand. Therefore, timely and accurately obtaining crop samples for spatial distribution mapping is the premise and foundation of crop yield estimation. Traditional methods for obtaining crop samples include reporting layer by layer and sampling surveys, which have the disadvantages of low timeliness, low accuracy, high input costs, and are easily affected by subjective human factors, restricting the accuracy and efficiency of crop monitoring work.
[0003] Remote sensing technology, with its advantages of strong timeliness, large amount of information, wide coverage, and low cost, provides an effective method and means for obtaining crop spatial distribution data. In early crop monitoring, traditional classification methods showed relatively higher classification accuracy compared to some simple data processing methods, but they also had many defects and deficiencies. For example, the process of making training samples was not only cumbersome but also labor-intensive. In addition, most of these traditional learning algorithms belong to shallow structures, with limited model expression ability and learning ability, making it difficult to fully and effectively capture and express the highly heterogeneous and non-linear characteristics of crops, resulting in the inability to construct a high-quality sample set that meets actual needs. When changes occur in the geographical environment, climate conditions, or crop planting patterns in the study area, it is necessary to readjust and optimize the algorithm, and even possibly reconstruct the entire model, increasing the complexity and cost of the research work and weakening its adaptability and promotion value in multiple scenarios.
[0004] Deep learning models have powerful self-learning and fault-tolerant capabilities. By constructing complex neural network structures, they can deeply explore and efficiently and accurately extract the internal relationships between different band information in remote sensing images and the spectral, texture, shape and other characteristic patterns of crops at different growth stages. Furthermore, they can construct highly abstract and representative feature representations, effectively overcoming the limitations of traditional learning methods in feature extraction. However, deep learning models also face a series of severe challenges in the actual application process. For example, the training process of deep learning models highly depends on large-scale labeled image sets, and the workload of constructing such training sample sets through visual interpretation means is extremely heavy and time-consuming. When dealing with large-scale remote sensing image data processing tasks, the labeling results are extremely prone to inconsistencies and error fluctuations, resulting in poor quality and representativeness of the constructed crop samples. This deep dependence on a large number of high-quality labeled samples and the numerous obstacles in the sample construction process greatly limit the wide popularization and effective implementation of deep learning models in key application fields such as large-scale crop mapping and long-term crop monitoring, making the huge potential of deep learning technology in the application scenario of crop sample construction still in a state of waiting to be deeply explored and fully released.
[0005] Therefore, the pseudo-label construction technology has emerged. This technology can effectively overcome the limitations of traditional methods and the problem of shortage of training samples in the research area, and to a certain extent, replace the training samples that rely on a large number of manual annotations in traditional methods, significantly reducing the manpower input, reducing the errors caused by human factors, greatly improving the efficiency of sample generation, shortening the data acquisition cycle, effectively improving the timeliness, and is an effective solution to promote the construction of the national crop dataset. It is of great significance for improving agricultural production efficiency, ensuring the stable growth of grain output, and providing solid technical support and data guarantee for the implementation of the national food security strategy.
[0006] The combination of remote sensing technology and pseudo-label construction technology provides a new idea for crop sample construction. However, there are still many challenges in the actual application. First of all, remote sensing image data has the characteristics of multi-source, multi-scale, multi-temporal, etc. Its data volume is large and complex. How to efficiently integrate and utilize these multi-source data is a key issue. Secondly, the generation of pseudo-labels depends on the prediction results of the initial model, and the performance of the initial model directly affects the quality of the pseudo-labels. In crop sample construction, due to the long growth cycle and complex spectral characteristics of crops, it is often difficult to guarantee the prediction accuracy of the initial model, resulting in large errors in the generated pseudo-labels. In addition, in the process of pseudo-label construction, how to effectively screen and correct pseudo-labels to improve the accuracy and representativeness of samples still needs in-depth study. Finally, when applying the combination of remote sensing technology and pseudo-label construction technology to large-scale crop mapping, how to balance the computational efficiency and model accuracy, and how to deal with the heterogeneity problems between different regions are all difficulties that need to be overcome in the actual application.
[0007] In view of the above problems, the present invention proposes a method and system for constructing samples by collaborating the pseudo-label theory with multi-source remote sensing data. First, the present invention constructs an initial training sample set by using multi-source remote sensing data and the threshold segmentation method, and screens through confidence. Then, it extracts features such as the spectrum, texture, and shape of crops through a machine learning model to expand the samples and generate an initial pseudo-label sample set. Subsequently, an adaptive pseudo-label screening mechanism is introduced, and combined with the crop growth model and expert knowledge, the pseudo-labels are corrected and optimized to improve the accuracy and representativeness of the samples. The present invention not only overcomes the dependence of traditional methods on a large number of manually labeled samples, but also improves the efficiency and accuracy of crop sample construction, providing reliable technical support for large-scale crop monitoring and yield estimation. Summary of the Invention
[0008] To solve the above problems, the present invention provides a method and system for constructing crop samples by collaborating the pseudo-label theory with multi-source remote sensing data.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A method for constructing crop samples by collaborating the pseudo-label theory with multi-source remote sensing data, comprising the following steps:
[0011] S1. Obtain multi-source remote sensing image data and perform preprocessing;
[0012] S2. Analyze the preprocessed multi-source remote sensing image data, screen the image data that can be used for constructing target crop samples based on phenological characteristics, and use the threshold segmentation method to determine the possible distribution spatial range of the target crops from the image data that can be used for constructing target crop samples;
[0013] S3. Extract and integrate the features of the target crops according to the image data of the possible distribution spatial range of the target crops to obtain rough samples of the target crops;
[0014] S4. Screen the rough samples of the target crops according to the confidence theory to obtain a pseudo-label sample set;
[0015] S5. Based on the initial pseudo-label sample set, adopt a progressive learning strategy to expand the sample set to obtain representative samples;
[0016] S6. Integrate the representative samples to obtain the final target crop sample set.
[0017] Further, in step S1, the multi-source remote sensing image data includes optical remote sensing image data and microwave remote sensing image data; the preprocessing includes geometric correction and radiometric correction for the optical remote sensing image data, and radiometric calibration and noise filtering processing for the microwave remote sensing image data.
[0018] Further, in step S2, the method for screening based on phenological characteristics is as follows:
[0019] Arrange the normalized difference vegetation index (NDVI) at each pixel position in the image data in chronological order to form time series data; perform curve fitting on the time series data, and determine the vegetation growth stage according to the curve change process; screen out the image data that can be used for constructing the target crop samples based on the vegetation growth stage.
[0020] Further, in step S2, the specific process of using the threshold segmentation method to determine the possible distribution spatial range of the target crop from the image data that can be used for constructing the target crop samples includes:
[0021] Use the Otsu method to determine the optimal threshold, and based on the optimal threshold, process the image data that can be used for the target crop
[0022] Further, in step S3, the features of the target crop are extracted from optical remote sensing images, including spectral features, spatial features, and texture features.
[0023] Further, in step S4, the specific steps of using the pseudo-labeling theory to preliminarily label the rough samples of the target crop to obtain a pseudo-label sample set are as follows:
[0024] Calculate the normal distribution value of the rough samples of the target crop, and screen out the samples with a confidence level within the confidence interval according to the standard normal distribution table as the pseudo-label sample set.
[0025] Further, the progressive learning strategy is adopted to optimize the sample set, which specifically includes:
[0026] Based on the random forest algorithm, expand the samples according to the features of the target crop, and then screen the expanded samples based on the confidence theory, and perform multiple iterative optimizations to obtain representative samples.
[0027] A crop sample construction system based on the collaboration of pseudo-labeling theory and multi-source remote sensing data includes:
[0028] A data acquisition and processing module, which is used to obtain multi-source remote sensing image data and perform preprocessing;
[0029] A distribution space determination module, which is used to analyze the preprocessed multi-source remote sensing image data, screen out the image data that can be used for constructing the target crop samples based on phenological characteristics, and use the threshold segmentation method to determine the possible distribution spatial range of the target crop from the image data that can be used for constructing the target crop samples;
[0030] A feature extraction and integration module, configured to extract and integrate features of the target crop based on the image data of the possible distribution space range of the target crop, so as to obtain a rough sample of the target crop;
[0031] A sample screening module, configured to screen the rough sample of the target crop according to the confidence theory to obtain a pseudo-label sample set;
[0032] A sample augmentation module, configured to optimize the sample set based on the initial pseudo-label sample set by adopting a progressive learning strategy to obtain representative samples;
[0033] A sample integration module, configured to integrate the representative samples to obtain a final target crop sample set.
[0034] A computer device, the computer device includes:
[0035] One or more processors;
[0036] A memory, configured to store one or more programs;
[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for constructing crop samples by collaborating with the pseudo-label theory and multi-source remote sensing data.
[0038] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned method for constructing crop samples by collaborating with the pseudo-label theory and multi-source remote sensing data.
[0039] The beneficial effects of the present invention are as follows: The method for constructing crop samples by collaborating with the pseudo-label theory and multi-source data proposed by the present invention combines multi-source remote sensing images and screens high-quality target crop samples based on confidence. The proposed target crop sample generation framework is simple, effective, and easy to implement. In addition, in the target crop sample generation framework, not only the possible distribution range of the target crop is determined from multi-source remote sensing image data by means of threshold segmentation method and crop phenological characteristics. At the same time, considering the particularity of sample generation, using the pseudo-label theory, confidence is used to perform collaborative processing on optical remote sensing image data, radar remote sensing image data and other auxiliary data, fully mining the correlation between multi-source data, and generating high-quality crop samples. In short, the method proposed by the present invention can improve the quality of target crop samples, overcome the problem of poor sample generation quality caused by factors such as noise in image data, limited resolution, and complex crop growth environment when generating crop samples relying on a single remote sensing image traditionally, and achieve high-precision crop sample generation. Therefore, the method proposed by the present invention has significant practical application value. Description of the Drawings
[0040] Figure 1 It is the flowchart of the crop sample construction method in the embodiment of the present invention.
[0041] Figure 2 It is the visualization schematic diagram of winter wheat sample generation in the embodiment of the present invention.
[0042] Figure 3 It is the verification schematic diagram of optimized sample extraction of winter wheat in the embodiment of the present invention. Specific embodiments
[0043] The technical solution of the present invention will be clearly and detailedly described below in conjunction with the accompanying drawings and specific examples.
[0044] A crop sample construction method based on the collaboration of pseudo-label theory and multi-source remote sensing data includes the following steps:
[0045] S1. Obtain multi-source remote sensing image data, including optical remote sensing image data and microwave remote sensing image data (SAR), and perform geometric correction (ensuring accurate spatial coordinates of the image), radiometric correction (converting the DN value of the image into surface reflectance), and atmospheric correction (eliminating the influence of the atmosphere on the image) on the optical remote sensing image data, and perform radiometric calibration, noise filtering processing, etc. on the microwave remote sensing image data to improve the data quality and lay a foundation for subsequent analysis and processing.
[0046] S2. Analyze the preprocessed multi-source remote sensing image data, and screen the image data that can be used for constructing target crop samples based on phenological characteristics. Specifically: Arrange the normalized difference vegetation index (NDVI) of each pixel position (corresponding to a small area in the farmland) in the image data in chronological order to form time series data; perform curve fitting on the time series data, and determine the vegetation growth stage from slow growth in the initial stage, rapid growth in the middle stage to stable or declining stage in the later stage according to the curve change process; screen out the image data that can be used for constructing target crop samples according to the vegetation growth stage.
[0047] The fitting curve formula is:
[0048] ,
[0049] In the formula, y is the NDVI value, K is the maximum value (saturation value) of NDVI, A and r are fitting parameters, and t is time.
[0050] Then, the threshold segmentation method is used to determine the possible distribution space range of the target crop from the image data available for constructing the target crop samples. Specifically, the Otsu method is used to determine the optimal threshold, and the image data available for constructing the target crop samples is binarized according to the optimal threshold. Those within the threshold are retained, and those outside the threshold are excluded, so as to determine the possible distribution space range of the target crop. The threshold calculation formula is as follows:
[0051]
[0052] Wherein, represents the maximum between-class variance, represents the global mean of the image, represents the probability that a pixel is classified into a certain class, represents the cumulative mean of the gray level K.
[0053] S3. Extract and integrate the target crop features from the image data of the possible distribution space range of the target crop to obtain the rough samples of the target crop; the target crop features are extracted through optical remote sensing images, including spectral features, spatial features, and texture features. The joint preparation and application of multiple features can fully characterize the crop features.
[0054] The calculation methods of spectral features are as follows:
[0055] Normalized Difference Vegetation Index (NDVI): ; Enhanced Vegetation Index (EVI): ; Normalized Difference Water Index (NDWI): ; Normalized Difference Built-up Index (NDBI): .
[0056] Wherein, R is the red band in the optical image, B is the blue band in the optical image, G is the green band in the optical image, NIR is the near-infrared band in the optical image, and SWIR is the short-wave infrared band.
[0057] The calculation method of spatial features is:
[0058]
[0059] In the formula, NNI is the nearest neighbor index, which is used to measure the aggregation degree of the distribution of ground objects, is the average value of the actual nearest neighbor distance, is the expected nearest neighbor distance under random distribution.
[0060] If the NNI is less than 1, it indicates that the ground objects are clustered; if the NNI is equal to 1, it is randomly distributed; if the NNI is greater than 1, it is evenly distributed.
[0061] The calculation method of texture features is as follows:
[0062]
[0063] In the formula, C is the obtained texture feature, L is the number of gray levels, is the element at the position in the gray-level co-occurrence matrix, and the contrast reflects the clarity and roughness of the texture in the image.
[0064] S4. Screen the rough samples of the target crop according to the confidence theory to obtain a pseudo-label sample set; specifically: calculate the normal distribution value of the rough samples of the target crop:
[0065] ,
[0066] In the formula, Z is the normal distribution value, is the overall mean, is the sample mean, is the overall standard deviation, and n is the sample size.
[0067] Screen the samples with confidence levels within the confidence interval according to the standard normal distribution table as the pseudo-label sample set. In this embodiment, a confidence level greater than 95% is selected as the confidence interval. To make the confidence level greater than 95%, the corresponding two-sided Z value (found through the standard normal distribution table, which can be found in existing statistical textbooks) should be below 1.96, that is, when |Z| > 1.96, the sample is not within the interval with a confidence level greater than 95%, and when |Z| ≤ 1.96, the sample is within the interval with a confidence level greater than 95%.
[0068] S5. Based on the initial pseudo-label sample set, adopt an incremental learning strategy to optimize the sample set to obtain representative samples. The main purposes of this step are as follows: ① After screening the samples combined with the threshold segmentation method and the confidence level, the number of samples still cannot meet the requirements; ② To expand the sample data and ensure the quality of the samples while expanding the number of samples, we need to continuously purify the quality of the expanded samples. The specific steps are as follows: Based on the random forest algorithm, expand the samples according to the spectral features, spatial features, and texture features of the target crop. At the same time, to ensure the quality of the samples, it is necessary to screen the expanded samples again based on the confidence theory and perform multiple iterative optimizations to improve the accuracy and representativeness of the samples and obtain representative samples.
[0069] Random Forest (RF) is an integrated classifier based on multiple decision trees. The basic principle of its algorithm is as follows: ① Use the bootstrap aggregation method to randomly draw N samples from the original training set with replacement to form a new training sample, and the samples not drawn (about 37%) form the out-of-bag (OOB) data; ② Under the principle of the minimum Gini coefficient, randomly select a subset of each node variable after internal splitting of N decision trees to construct multiple Cart decision trees, and form a random forest with the generated decision trees; the definition formula of the Gini coefficient is as follows:
[0070]
[0071] In the formula, T is the original data set; C i is a certain category that a randomly selected sample is identified as; is the probability that the selected sample is of category C i category.
[0072] The generated random forest classifier classifies the sample data for accuracy evaluation. Whenever a sample belongs to OOB data, its vote count will be statistically recorded, and the classification category will be determined in the form of a majority vote. Since the OOB data does not participate in the construction of the decision tree, the OOB error ( ) can be used to evaluate the performance of the model and the importance of its quantitative variables. The definition formula is:
[0073]
[0074] Among them, N represents the total number of samples, which is the number of all samples used to train the random forest model; , is an indicator function. When the true category ( ) of the i-th sample is not equal to the category ( ) predicted for this sample based on the out-of-bag data, the value of this function is 1, indicating a classification error; conversely, when the true category and the predicted category are the same, the function value is 0, indicating a correct classification.
[0075] Step 6: Use the generated representative samples to form a large-scale and multi-region high-quality winter wheat sample, and form a winter wheat sample set;
[0076] To better explain the purpose and advantages of the technical solution of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and taking the extraction of winter wheat samples as an example.
[0077] The technical solution of the present invention uses Sentinel-1 and Sentinel-2 data as data sources and is implemented through ENVI5.3 software and Python language programming. The following combines Figure 1 to detail the steps of generating winter wheat samples.
[0078] Step 1: Preprocess Sentinel-1 and Sentinel-2 data, including radiometric calibration, atmospheric correction, orthorectification, noise filtering, etc. The specific operations are prior art;
[0079] Step 2: Arrange the normalized difference vegetation index (NDVI) values at each pixel position in chronological order to form time series data. Curve fitting is performed on these time series data. According to the curve change process, determine the stages of vegetation growth from slow growth in the initial stage, rapid growth in the middle stage to stabilization or decline in the later stage.
[0080]
[0081] In the formula, y is the normalized difference vegetation index value, K is the maximum value (saturation value) of the normalized difference vegetation index, A and r are fitting parameters, and t is time.
[0082] Screen the image data that can be used for the construction of winter wheat samples according to the vegetation growth stage, and then calculate the possible distribution range of winter wheat through the threshold segmentation method. The specific steps are as follows: Determine the optimal threshold by the maximum between-class variance method, perform binary segmentation on the image data that can be used for the construction of winter wheat samples according to the optimal threshold, keep the part within the threshold and exclude the part outside the threshold, so as to determine the possible distribution spatial range of winter wheat. The threshold calculation formula is:
[0083]
[0084] In the formula, represents the maximum between-class variance, represents the global mean of the image, represents the probability that a pixel is classified into a certain class, represents the cumulative mean of the gray level K.
[0085] Step 3: Extract and integrate winter wheat features from the image data of the possible distribution spatial range of winter wheat to obtain a rough winter wheat sample; The features include spectral features, spatial features and texture features. The methods for obtaining spectral features are as follows: Normalized Difference Vegetation Index (NDVI): ; Enhanced Vegetation Index (EVI): ; Normalized Difference Water Index (NDWI): ; Normalized Difference Built-up Index (NDBI): .
[0086] Wherein, R is the red light band in the optical image, B is the blue light band in the optical image, G is the green light band in the optical image, and NIR is the near-infrared band in the optical image.
[0087] The spatial feature calculation method is as follows:
[0088]
[0089] Wherein, the nearest neighbor index (NNI) is used to measure the aggregation degree of the distribution of ground objects. is the average value of the actual nearest neighbor distance. is the expected nearest neighbor distance under random distribution.
[0090] If NNI is less than 1, it indicates that the ground objects are aggregated; if NNI is equal to 1, it is a random distribution; if NNI is greater than 1, it is a uniform distribution.
[0091] The texture feature calculation method is as follows:
[0092]
[0093] Wherein, C is the contrast, which reflects the clarity and roughness of the texture in the image. Among them, L is the number of gray levels. is in the gray-level co-occurrence matrix The element at the position.
[0094] Step 4: Screen the rough winter wheat samples according to the confidence theory to obtain a pseudo-label sample set; specifically: calculate the normal distribution value of the rough winter wheat samples:
[0095] Wherein, the overall mean is , the sample mean is , the overall standard deviation is , and the sample size is n.
[0096] Select the samples with the confidence level within the confidence interval according to the standard normal distribution table as the pseudo-label sample set. In this embodiment, the confidence level greater than 95% is selected as the confidence interval. To make the confidence level greater than 95%, the corresponding two-sided Z value (found through the standard normal distribution table, which can be found in existing statistical textbooks) should be below 1.96. That is, when |Z| > 1.96, the sample is not within the confidence interval with a confidence level greater than 95%, and when |Z| ≤ 1.96, the sample is within the confidence interval with a confidence level greater than 95%.
[0097] After screening the samples through the combination of the threshold segmentation method and the confidence level, the number of samples still cannot meet the requirements; to expand the sample data and ensure the quality of the samples while expanding the number of samples, we need to continuously purify the quality of the expanded samples.
[0098] Step 5: Based on the random forest algorithm, sample augmentation is performed according to the spectral characteristics, spatial characteristics, and texture characteristics of winter wheat. At the same time, to ensure the quality of the samples, it is necessary to screen the augmented samples again based on the confidence theory. Repeat the process of Step 4, perform multiple iterations of optimization, improve the accuracy and representativeness of the samples, and obtain representative samples.
[0099] Step 6: Integrate the representative samples to obtain the final target crop sample set.
[0100] Furthermore, to remove potential misclassified land classes in the target crop sample set, non-target crop land classes can be removed by setting a threshold. Taking the other non-winter wheat land classes mixed in winter wheat as an example, based on Sentinel-1 data, the VH pixel values of winter wheat and non-winter wheat plots are extracted using artificial sample points. The optimal threshold for separating the two is determined to be pixel value -16, that is, VH < -16 is winter wheat, and vice versa. Use the raster feature tool in Arcgis10.6 software to convert the above raster extraction results into vectors, then generate a 10 m buffer inward according to the vector boundaries of each patch, and use the erase tool to remove the 10 m buffer part in the vector results. Finally, calculate the area of each patch, remove the patches with too small area, and retain patches larger than 50000 m 2 patches for subsequent construction of training samples.
[0101] To illustrate the effect of the present invention, the sample set generated by the method of the present invention is used to train the winter wheat extraction model, and winter wheat is extracted from the target area, and compared with other existing methods to verify the effectiveness of the method of the present invention. The specific description is as follows:
[0102] As Figure 2 shown, the sample set obtained by the method of the present invention performs well in distinguishing land classes such as winter wheat from construction land, bare land, water bodies, and evergreen vegetation (a - c), and the extraction effect is better than the published research results. Among them, this method has higher extraction accuracy than NDVIs, mainly manifested in that the winter wheat fields identified by the indicators of this method are more complete, and can more accurately extract the ridges and scattered winter wheat fields, while the extraction results of NDVIs are more fragmented.
[0103] As Figure 3As shown, in the areas shown in (a) and (b), winter wheat is intensively planted and covers a large area. In the image, there are only two non-winter wheat land types, namely construction land and a small area of other crops. It can be found from the figure that as the number of training times increases, the boundaries of winter wheat plots in the extraction results of the random forest model gradually become clear and refined, the ridges and roads are further distinguished, and other crops (light green part) that were initially misclassified as winter wheat are gradually excluded. In the areas shown in (c) and (d), the planting area of winter wheat is small and scattered, and in addition to construction land, it also includes various land types such as wetlands, water bodies (ponds), and other crops. In the first-round training results, the random forest model only extracted a small part of winter wheat. As the number of training times increases, the missed winter wheat pixels are correctly classified, the integrity of winter wheat plots is continuously improved, and the extraction results gradually tend to be stable.
[0104] According to the above analysis results, in the cycle process of "generating pseudo-label samples - random forest model training - winter wheat extraction", based on the constructed pseudo-labels, the ability of the random forest model to distinguish between winter wheat and non-winter wheat is gradually enhanced, and the extraction accuracy steadily increases. Whether in the intensive winter wheat planting area or the scattered winter wheat planting area, good classification effects can be achieved, and relatively accurate classification labels can be obtained. The results show that the method for constructing winter wheat pseudo-labels proposed in this study has certain scientificity and feasibility, and can provide effective and reliable sample data for subsequent training of deep learning models.
[0105] It should be noted and understood that various modifications and improvements can be made to the present invention described in detail above without departing from the spirit and scope of the invention as claimed. Therefore, the scope of the claimed technical solution is not limited by any specific exemplary teachings given.
[0106] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0107] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0108] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0110] The above is only the preferred embodiment of the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data, characterized in that: The following steps are involved: S1. Obtain multi-source remote sensing image data and perform preprocessing; S2. Analyze the preprocessed multi-source remote sensing image data, screen the image data that can be used to construct the target crop sample based on the phenological characteristics, and use the threshold segmentation method to determine the possible distribution space range of the target crop from the image data that can be used to construct the target crop sample; S3. Extract and integrate target crop features based on the image data of the possible distribution space range of the target crop to obtain a rough sample of the target crop; S4. Screen the target crop rough samples according to the confidence theory to obtain a pseudo-label sample set; S5. Based on the initial pseudo-label sample set, a progressive learning strategy is used to expand the sample set to obtain representative samples; S6. Integrate the representative samples to obtain the final target crop sample set.
2. The method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data according to claim 1 is characterized in that: In step S1, the multi-source remote sensing image data includes optical remote sensing image data and microwave remote sensing image data; the preprocessing includes geometric correction and radiation correction for the optical remote sensing image data, and radiation calibration and noise filtering for the microwave remote sensing image data.
3. The method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data collaboration according to claim 1, characterized in that: In step S2, the method of screening based on phenological characteristics is: The normalized vegetation index of each pixel position in the image data is arranged in chronological order to form time series data; curve fitting is performed on the time series data to determine the vegetation growth stage according to the curve change process; and image data that can be used to construct target crop samples are screened out according to the vegetation growth stage.
4. The method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data collaboration according to claim 1, characterized in that: In step S2, the use of the threshold segmentation method to determine the possible distribution space range of the target crop from the image data that can be used to construct the target crop sample specifically includes: The maximum inter-class variance method is used to determine the optimal threshold, and the image data that can be used to construct the target crop sample is binarized according to the optimal threshold to determine the possible distribution space range of the target crop.
5. The method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data collaboration according to claim 1, characterized in that: In step S3, the target crop features are extracted through optical remote sensing images, including spectral features, spatial features and texture features.
6. The method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data collaboration according to claim 1, characterized in that: In step S4, the target crop rough samples are preliminarily labeled according to the confidence theory to obtain a pseudo-label sample set. The specific steps include: Calculate the normal distribution value of the target crop rough sample, and select the samples with confidence levels within the confidence interval according to the standard normal distribution table as the pseudo-label sample set.
7. The method for constructing crop samples based on pseudo-label theory and multi-source remote sensing data collaboration according to claim 1, characterized in that: In step S5, the sample set is expanded by using a progressive learning strategy, specifically including: Based on the random forest algorithm, the samples are expanded according to the characteristics of the target crops, and the expanded samples are screened again based on the confidence theory, and multiple iterations of optimization are performed to obtain representative samples.
8. A crop sample construction system based on pseudo-label theory and multi-source remote sensing data collaboration, characterized in that: include: Data acquisition and processing module, used to acquire multi-source remote sensing image data and perform preprocessing; The distribution space determination module is used to analyze the pre-processed multi-source remote sensing image data, screen the image data that can be used to construct the target crop sample based on the phenological characteristics, and use the threshold segmentation method to determine the possible distribution space range of the target crop from the image data that can be used to construct the target crop sample; A feature extraction and integration module is used to extract and integrate target crop features based on image data of the possible distribution space range of the target crop to obtain a rough sample of the target crop; The sample screening module is used to screen the rough samples of target crops according to the confidence theory to obtain a pseudo-label sample set; The sample expansion module is used to optimize the sample set based on the initial pseudo-label sample set using a progressive learning strategy to obtain representative samples; The sample integration module is used to integrate representative samples to obtain the final target crop sample set.
9. A computer device, characterized in that: The computer device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the crop sample construction method based on pseudo-label theory and multi-source remote sensing data collaboration as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for constructing crop samples based on pseudo-label theory and collaboration with multi-source remote sensing data as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Interferometry method based on video synthetic aperture radar
CN113093184A
Multi-temporal active and passive remote sensing random forest crop identification method and system
CN115035413A
Cited By
Small sample crop extraction method using pseudo tag
CN120510521A