Crop planting area extraction method, system and device coupled with self-supervised learning and segmentation model

By combining self-supervised learning and segmentation models, and using unlabeled remote sensing images and a small number of labeled samples to train a crop planting area recognition model, the problems of environmental factors and insufficient sample number were solved, and high-precision crop planting area recognition was achieved.

CN119625516BActive Publication Date: 2025-10-03SURVEYING & MAPPING INST LANDS & RESOURCE DEPT OF GUANGDONG PROVINCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411453440.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-10-03
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

The crop planting area recognition model in the existing technology is limited by complex environmental factors and a small number of training samples, resulting in insufficient recognition accuracy and an inability to accurately identify the crop planting area.

Method used

Combining self-supervised learning and segmentation models, training is performed through the feature extraction module of the self-supervised learning model, and parameters are shared with the segmentation model. The model is trained using unlabeled remote sensing images and a small number of labeled samples, and the parameters of the segmentation model are optimized to achieve high-precision identification of crop planting areas.

Benefits of technology

The model training cycle is shortened, the feature extraction and recognition capabilities of the segmentation model are improved, and efficient and accurate remote sensing identification of crop planting areas is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625516B_ABST
    Figure CN119625516B_ABST
Patent Text Reader

Abstract

The present application relates to a method, system and device for extracting crop planting areas by coupling self-supervised learning and segmentation models, comprising: combining a self-supervised learning model with a segmentation model, performing feature extraction training on a first feature extraction module of the self-supervised learning model using unlabeled remote sensing image samples, and making the first feature extraction module share model parameters with the feature extraction module of the segmentation model, then training the segmentation module of the segmentation model using a training set, and optimizing parameters using a validation set and model prediction results, using the trained segmentation model to perform full-image recognition and prediction on remote sensing images, and obtaining crop planting area recognition and extraction results, thereby achieving the use of unlabeled samples and a small number of labeled samples to train a segmentation model that can identify and segment crop planting areas with high precision, thereby solving the problem of insufficient recognition accuracy of crop recognition and segmentation models in the prior art due to environmental factors and a small number of training samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition and processing technology, and in particular to a method, system, and device for extracting crop planting areas by coupling self-supervised learning and segmentation models. Background Art

[0002] Currently, this technology is crucial for crop yield prediction, production management optimization, and policy formulation. For example, Camellia oleifera, a subtropical evergreen shrub and important oilseed plant, is rich in unsaturated fatty acids and possesses significant nutritional value. It is widely cultivated in many countries around the world, and rapid and accurate identification of its distribution is crucial.

[0003] In the prior art, crop planting area identification usually adopts model recognition, combined with satellite remote sensing technology to achieve crop planting area identification. The model recognition accuracy is highly dependent on training samples. However, due to the complexity of crop growing areas and growing environments, the training samples for model training are relatively small, which can easily lead to limited accuracy when processing large-scale data, affecting the objectivity of the results, and thus resulting in insufficient model accuracy and inability to accurately extract useful information from images. Specifically, taking Camellia oleifera as an example, Camellia oleifera growing areas are usually hilly and mountainous areas with dense vegetation and complex terrain, which increases the difficulty of extracting Camellia oleifera information from remote sensing images. In addition, the complex forest structure in the south and the similarity of spectral characteristics between tree species are also prone to classification errors. Although it is still possible to capture information about Camellia oleifera growing areas, in actual operation, since Camellia oleifera is mostly distributed in mountainous areas with complex terrain, the signals of drones collecting samples are often interfered with, and sample collection faces many challenges, making the number of high-quality labeled samples very limited. It can be seen that the prior art has the problem of insufficient recognition accuracy of crop identification and segmentation models due to the complexity of environmental factors and the small number of training samples, and is unable to accurately identify crop planting areas. Summary of the Invention

[0004] The present application provides a method, system and equipment for extracting crop planting areas by coupling self-supervised learning and segmentation models. By combining the self-supervised learning model and the segmentation model, after the self-supervised learning model completes feature extraction training, the parameters of the self-supervised learning model are shared with the feature extraction module of the segmentation model, thereby shortening the model training cycle. Only unlabeled samples and a small number of labeled samples are required to complete high-precision model training, and a high-precision segmentation model that can accurately identify and segment crop planting areas is obtained, ultimately realizing efficient and accurate remote sensing identification of oil tea planting areas, and solving the problem of insufficient recognition accuracy of crop identification and segmentation models caused by environmental factors and a small number of training samples in existing technologies.

[0005] In a first aspect, the present application provides a method for extracting crop planting areas by coupling self-supervised learning and a segmentation model, comprising:

[0006] Obtain unlabeled remote sensing image samples and labeled crop samples from the sample area, and construct a combined self-supervised learning model and segmentation model;

[0007] According to a preset comparison learning strategy, a first feature extraction module of the self-supervised learning model is trained with the unlabeled remote sensing image samples for feature extraction, wherein the first feature extraction module is used to capture local and global pattern features from the input image in scenes of different complexities;

[0008] Sharing the model parameters of the first feature extraction module to the feature extraction module of the segmentation model, and migrating the training weights of the first feature extraction module to the segmentation module of the segmentation model;

[0009] Performing vectorization processing on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set;

[0010] Performing segmentation training on the segmentation model according to the training set to obtain a model prediction result output by the segmentation model;

[0011] Optimizing the parameters of the segmentation model according to the validation set and the model prediction result until the model accuracy of the segmentation model reaches a preset requirement, thereby obtaining a trained segmentation model;

[0012] The trained segmentation model is used to perform full-image recognition prediction including crop planting area extraction based on the remote sensing image of the target area, and vector conversion is performed based on the prediction results of the full-image recognition prediction to obtain the planting area recognition and extraction results of the crops to be identified in the target area.

[0013] Optionally, obtaining unlabeled remote sensing image samples and labeled crop samples from the sample area includes:

[0014] Acquire remote sensing image data from the sample area;

[0015] Performing correction and enhancement preprocessing on the remote sensing image data to obtain unlabeled remote sensing image samples that meet quality requirements;

[0016] Collecting crop sample images from the sample area based on a preset drone collection instruction, wherein the drone collection instruction is used to control the drone to collect sample images from the sample area;

[0017] Acquire annotation information corresponding to the crop sample image, and generate a crop annotation sample based on the crop sample image and the annotation information.

[0018] Optionally, performing feature extraction training on the first feature extraction module of the self-supervised learning model using the unlabeled remote sensing image samples according to a preset comparison learning strategy includes:

[0019] Obtaining a preset comparison learning strategy and image cropping information, wherein the image cropping information includes input information of the self-supervised learning model and key image information;

[0020] Based on the input information and the image key information, preprocessing the unlabeled remote sensing image sample to obtain an input image sample, wherein the input image sample includes local image details and global scene information;

[0021] According to the comparison learning strategy, the first feature extraction module of the self-supervised learning model extracts and learns the structural features and semantic information in the image from the input image sample.

[0022] Optionally, extracting and learning structural features and semantic information in the image from the input image sample by the first feature extraction module of the self-supervised learning model includes:

[0023] Inputting the input image sample into the self-supervised learning model;

[0024] Performing random enhancement on each sample image in the input sample to obtain sample pair information corresponding to each sample image, wherein the sample pair information includes a positive sample pair and a negative sample pair;

[0025] Using cosine similarity to measure the sample pair information, and obtain the similarity of each sample pair in the sample pair information;

[0026] Perform comparison model training on the first feature extraction module according to the sample pair information and the similarity.

[0027] Optionally, the measuring the sample pair information using cosine similarity to obtain the similarity of each sample pair in the sample pair information includes:

[0028] The first feature extraction module is based on Measuring the similarity of different images in the sample pair information;

[0029] Where A·B represents the dot product obtained by multiplying the corresponding components of vector A and vector B and summing them. It reflects the relative direction relationship between the vectors and serves as the degree of similarity between the two vectors. If the two vectors have the same direction, the dot product value is large; if the two vectors have opposite directions, the dot product value is small. A||||B represents the norm product of vector A and vector B, where ||A|| is the Euclidean norm of vector A and ||B|| is the Euclidean norm of vector B.

[0030] Optionally, performing comparison model training on the first feature extraction module according to the sample pair information and the similarity includes:

[0031] According to the formula guiding the first feature extraction model to distinguish positive sample pairs from negative sample pairs during the comparison learning process;

[0032] Among them, the eigenvector z i and the eigenvector z j are all feature representations of positive sample pairs, sim(z i ,z j ) is z i and z j The cosine similarity between them is used to measure the similarity of positive sample pairs, τ is the temperature parameter used to adjust the similarity softening degree in contrastive learning, N is the number of sample pairs in the batch, exp(sim(z i ,z j ) / τ) represents the similarity score of the positive sample pair, It represents the sum of the similarity scores between the positive sample pair and all negative sample pairs.

[0033] Optionally, performing vectorization processing based on the crop labeled samples to obtain training image pair data includes:

[0034] Performing vectorization based on the crop labeled samples to obtain vector data;

[0035] Affine transformation is performed on the vector data to obtain mapping data of the vector data on an image plane, and a target area associated with the crop sample is cropped from the mapping data to obtain training image pair data.

[0036] Optionally, the performing parameter optimization on the segmentation model according to the validation set and the model prediction result includes:

[0037] according to performing a quantitative evaluation on the segmentation model;

[0038] Among them, C is the number of categories, Prediction i is the prediction result of the i-th category, Ground Truth i is the true label corresponding to the prediction result of the i-th category.

[0039] In a second aspect, the present application provides a crop planting area extraction system, comprising:

[0040] The sample extraction and model building module is used to obtain unlabeled remote sensing image samples and labeled crop samples from the sample area, and to build a combined self-supervised learning model and segmentation model;

[0041] a feature extraction training module, configured to perform feature extraction training on a first feature extraction module of the self-supervised learning model using the unlabeled remote sensing image samples according to a preset comparison learning strategy, wherein the first feature extraction module is configured to capture local and global pattern features from input images in scenes of varying complexity;

[0042] a sharing module, configured to share the model parameters of the first feature extraction module with the feature extraction module of the segmentation model, and to migrate the training weights of the first feature extraction module to the segmentation module of the segmentation model;

[0043] A vectorization processing module, configured to perform vectorization processing based on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set;

[0044] A segmentation training module, configured to perform segmentation training on the segmentation model based on the training set to obtain a model prediction result output by the segmentation model;

[0045] A parameter optimization module is used to optimize the parameters of the segmentation model according to the validation set and the model prediction result until the model accuracy of the segmentation model meets the preset requirements, thereby obtaining a trained segmentation model;

[0046] The full-image recognition prediction module is used to perform full-image recognition prediction including crop planting area extraction based on the remote sensing image of the target area through the trained segmentation model, and to perform vector conversion based on the prediction results of the full-image recognition prediction to obtain the planting area recognition and extraction results of the crops to be identified in the target area.

[0047] In a third aspect, the present application provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0048] Memory for storing computer programs;

[0049] The processor is used to implement the steps of the crop planting area extraction method coupled with self-supervised learning and segmentation model as described in any embodiment of the first aspect when executing the program stored in the memory.

[0050] In summary, the embodiment of the present application constructs a combined self-supervised learning model and a segmentation model, and uses the unlabeled remote sensing image samples obtained from the sample area to perform model training on the first feature extraction module of the self-supervised learning model, so that the first feature extraction module can have the ability to capture global and local pattern features from the input image in scenes of different complexities. The parameters of the first feature extraction module are then shared with the feature extraction module of the segmentation model, and the model training weights and training weights are then transferred to the segmentation module of the segmentation model. The crop labeled samples obtained from the sample area are vector processed and divided into a training set and a validation set. The segmentation module is trained through the training set, and the prediction results output by the segmentation model are verified through the validation set to optimize the parameters of the segmentation model. Therefore, this application uses unlabeled remote sensing images to perform self-supervised learning on the self-supervised learning module, fully explores and extracts multi-scale and multi-type features in the image, and completes the model training of the feature extraction module and the segmentation module of the segmentation model, shortens the model training cycle, improves the feature extraction ability and segmentation recognition ability of the segmentation model, and obtains a high-precision segmentation model that can accurately identify and segment crop planting areas. The trained segmentation model is used to perform full-image recognition prediction and crop planting area extraction on the remote sensing image of the target area, and vector conversion is performed based on the prediction results of the full-image recognition prediction to obtain the planting area identification and extraction results of the crops to be identified in the target area, ultimately realizing efficient and accurate remote sensing identification of crop planting areas, and solving the problem of insufficient recognition accuracy of crop recognition and segmentation models caused by environmental factors and small training samples in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0053] Figure 1 A flowchart of a method for extracting crop planting areas by coupling self-supervised learning and segmentation models provided in an embodiment of the present application;

[0054] Figure 2 This is a flowchart of the steps of a method for extracting crop planting areas by coupling self-supervised learning and segmentation models, provided by an optional embodiment of the present application;

[0055] Figure 3This is a flow chart of a model training method for coupling self-supervised learning and segmentation models provided as an optional example of this application;

[0056] Figure 4 This is a schematic diagram of the combination of the SimCLR model and the Segformer model provided as an optional example of this application;

[0057] Figure 5 This is a spatial distribution diagram of camellia oleifera extraction results provided as an optional example of this application;

[0058] Figure 6 A structural block diagram of a crop planting area extraction system coupled with self-supervised learning and segmentation models provided in an embodiment of the present application;

[0059] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0061] To facilitate understanding of the embodiments of the present application, further explanation will be given below in conjunction with the drawings and specific embodiments. The embodiments do not constitute a limitation on the embodiments of the present application.

[0062] Figure 1 A flow chart of a method for extracting crop planting areas by coupling self-supervised learning and segmentation models provided in an embodiment of the present application. Figure 1 As shown, the crop planting area extraction method coupled with self-supervised learning and segmentation model provided in the embodiment of the present application may specifically include the following steps:

[0063] Step 110 , obtaining unlabeled remote sensing image samples and labeled crop samples from the sample area, and constructing a combined self-supervised learning model and segmentation model.

[0064] In this embodiment, the sample area can be understood as an area where target crops are planted. Unlabeled remote sensing image samples refer to remote sensing images obtained by photographing the sample area, for example, and used as unlabeled remote sensing image samples for subsequent training of the self-supervised learning model. Labeled crop samples include, but are not limited to, remote sensing images of the crops and their corresponding labels.

[0065] For example, taking oil tea as an example, when training a segmentation model to identify oil tea planting areas, the oil tea planting area can be used as a sample area. By photographing the sample area, etc., the remote sensing image corresponding to the sample area can be obtained. Preferably, high-resolution remote sensing images with good quality can be selected; drones are used to collect corresponding remote sensing images containing oil tea planting areas from the sample area as oil tea samples, and the oil tea samples are labeled to form vector results, thereby obtaining high-quality remote sensing images containing oil tea and corresponding labeled crop labeled samples. Only a small number of labeled high-quality samples are needed to effectively improve the model recognition ability.

[0066] In a specific implementation, this embodiment constructs a segmentation model and a self-supervised learning model, and combines or couples the segmentation model and the self-supervised learning model. Among them, the segmentation model can select the Segformer model (Segformer semantic segmentation model), and the self-supervised learning model can select SimCLR self-supervised learning. In the related art, the existing Segformer semantic segmentation model has become an efficient semantic segmentation tool due to its powerful multi-scale feature extraction capability and excellent segmentation performance. However, when training, the Segformer semantic segmentation model relies on a large number of labeled samples to fully exert its performance, and crops such as oil tea are limited by restrictions such as the growing area, and it is impossible to obtain a large number of labeled samples. To this end, this embodiment combines the SimCLR model and the Segformer model to simultaneously train the two models, reduce model parameters, and effectively reduce computing resource requirements. In addition, the use of SimCLR self-supervised learning can effectively reduce the demand for labeled samples, improve the accuracy of the model, and solve the problem of insufficient recognition accuracy of crop recognition segmentation models caused by the limited number of training samples in the existing technology.

[0067] Step 120 : According to a preset comparison learning strategy, feature extraction training is performed on the first feature extraction module of the self-supervised learning model using the unlabeled remote sensing image samples.

[0068] The first feature extraction module is used to capture local and global pattern features from the input image in scenarios of different complexities.

[0069] In this embodiment, the unlabeled remote sensing image is input into the self-supervised learning model. The self-supervised learning model gradually extracts and learns the structural features and semantic information in the image from the unlabeled remote sensing image based on the comparison learning strategy, so that it can effectively capture local and global pattern features in complex scenes, and significantly improve the model's ability in feature extraction and semantic expression.

[0070] Step 130: Share the model parameters of the first feature extraction module with the feature extraction module of the segmentation model, and migrate the training weights of the first feature extraction module to the segmentation module of the segmentation model.

[0071] In this embodiment, since SimCLR is combined with the Segformer model, the parameters of the feature extraction module of SilmCLR can be shared with the feature extraction module (Encoder module) of the Segformer model, thereby completing the training of the feature extraction module of the Segformer model. This embodiment uses SimCLR self-supervised learning to perform self-supervised learning on unlabeled remote sensing images, fully mining and extracting multi-scale and multi-type features in the images. By sharing parameters with the Encoder module of the Segformer model, the Segformer model can also have the ability to fully mine and extract various features of the image, effectively retaining the global and local features learned by SimCLR. This embodiment uses SimCLR self-supervised learning to effectively reduce the demand for labeled samples, and there is no need to simultaneously learn and train the feature extraction capabilities of the two models, reducing parameters and reducing computing resource requirements.

[0072] In a specific implementation, this embodiment adopts a transfer learning strategy during the training process of the Segformer model, and uses the encoder weights obtained by SimCLR pre-training to initialize the weights of the encoder module of Segformer. The weights of the feature extraction part of the model (i.e., the Encoder part of Segformer) are transferred to the Encoder of Segformer used for image segmentation, and the feature representation learned by SimCLR is applied to the segmentation task, thereby improving the model's ability to understand specific targets, and subsequently the model training can be performed for Segformer.

[0073] Step 140 : performing vectorization processing based on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set.

[0074] In a specific implementation, this embodiment can first vectorize the labeled crop samples to obtain vector data, then map the vector data to obtain training image pair data. The training image pair data can then be partitioned according to a partitioning ratio to obtain a training set and a validation set. Preferably, the partitioning ratio can be 7:3, meaning the ratio of the training set to the validation set is 7:3. Of course, in actual implementation, the partitioning ratio can be set based on actual requirements, and this embodiment does not impose any restrictions on this.

[0075] Step 150: Perform segmentation training on the segmentation model based on the training set to obtain a model prediction result output by the segmentation model.

[0076] Step 160 , optimizing the parameters of the segmentation model according to the validation set and the model prediction result, until the model accuracy of the segmentation model reaches a preset requirement, thereby obtaining a trained segmentation model.

[0077] A unified description of steps 150 to 160 is provided:

[0078] In the specific implementation, the Segformer model combines the advantages of Transformer and convolutional neural network (CNN), can efficiently extract multi-scale features, and adopts a simple and efficient decoder to improve the performance of image segmentation, and can effectively deal with crops with complex growth environments. In this embodiment, the segmentation model can be trained using a training set. The segmentation model predicts the segmentation of the planting area based on the input training set to obtain the corresponding model prediction results. The accuracy of the model prediction results is then verified using a validation set to obtain a trained segmentation model. Taking oil tea as an example, this application uses labeled oil tea area sample data to supervise the training of the Segformer model. By continuously optimizing the parameters of the Segformer model, the segmentation accuracy of the model in the oil tea area is improved, and finally high-precision automatic segmentation of the oil tea area is achieved, achieving efficient and accurate segmentation effects, and solving the problem of insufficient recognition accuracy of crop recognition segmentation models caused by the existing technology being limited by environmental factors and a small number of training samples.

[0079] Step 170, using the trained segmentation model to perform full-image recognition prediction including crop planting area extraction based on the remote sensing image of the target area, and performing vector conversion based on the prediction results of the full-image recognition prediction to obtain the planting area recognition and extraction results of the crops to be identified in the target area.

[0080] In actual implementation, after completing the model training of the Segformer semantic segmentation model in this embodiment, the Segformer semantic segmentation model with the highest segmentation accuracy can be selected and applied to the actual crop planting area segmentation and recognition. By obtaining the remote sensing image of the target area and inputting it into the trained Segformer semantic segmentation model, full-image recognition prediction is performed. The full-image recognition prediction includes but is not limited to: feature extraction, crop recognition, and crop planting area extraction, etc., to obtain the prediction result, and then convert the prediction result into a vector to form the final planting area recognition and extraction result of the crop to be identified in the target area.

[0081] It can be seen that the embodiment of the present application constructs a combined self-supervised learning model and a segmentation model, and uses the unlabeled remote sensing image samples obtained from the sample area to train the first feature extraction module of the self-supervised learning model, so that the first feature extraction module can have the ability to capture global and local pattern features from the input image in scenes of different complexities, and then share the parameters of the first feature extraction module with the feature extraction module of the segmentation model, and then transfer the model training weights and training weights to the segmentation module of the segmentation model, use the crop labeled samples obtained from the sample area to perform vector processing and divide them into training sets and validation sets, use the training set to train the segmentation module, and use the validation set to verify the prediction results output by the segmentation model, and optimize the parameters of the segmentation model. Thus, the present application uses unlabeled remote sensing images to perform self-supervised learning on the self-supervised learning module, fully mines and extracts multi-scale and multi-type features in the image, and completes the model training of the feature extraction module and the segmentation module of the segmentation model, shortens the model training cycle, improves the feature extraction ability and segmentation recognition ability of the segmentation model, and obtains a high-precision segmentation model that can accurately identify and segment crop planting areas. The embodiment of the present application utilizes a trained segmentation model to perform full-image recognition prediction on the remote sensing image of the target area, and performs vector conversion based on the prediction results of the full-image recognition prediction to obtain the identification and extraction results of the planting area of ​​the crops to be identified in the target area, thereby ultimately achieving efficient and accurate remote sensing identification of crop planting areas, and solving the problem of insufficient recognition accuracy of crop recognition segmentation models caused by existing technologies being limited by environmental factors and a small number of training samples.

[0082] Reference Figure 2 , shows a schematic flow chart of the steps of a method for extracting crop planting areas by coupling self-supervised learning and segmentation models, provided in an optional embodiment of the present application. The method may specifically include the following steps:

[0083] Step 210 , obtaining unlabeled remote sensing image samples and labeled crop samples from the sample area, and constructing a combined self-supervised learning model and segmentation model.

[0084] In an optional embodiment, the embodiment of the present application obtains unlabeled remote sensing image samples and labeled crop samples from the sample area, which may specifically include: obtaining remote sensing image data from the sample area; performing correction and enhancement preprocessing on the remote sensing image data to obtain unlabeled remote sensing image samples that meet quality requirements; based on preset drone acquisition instructions, collecting crop sample images from the sample area, and the drone acquisition instructions are used to control the drone to collect sample images of the sample area; obtaining labeling information corresponding to the crop sample images, and generating labeled crop samples based on the crop sample images and the labeling information.

[0085] Among related technologies, sample annotation is currently mainly done through visual interpretation and drone photography. The visual interpretation method is highly dependent on the interpreter's professional experience and subjective judgment, which can easily lead to limited accuracy when processing large-scale data, affecting the objectivity of the results. Taking Camellia oleifera as an example, its growing area and growing environment are complex. Camellia oleifera is mostly distributed in mountainous areas with complex terrain. Drone signals are often interfered with, and sample collection faces many challenges, making the number of high-quality labeled samples very limited. Although drone photography can intuitively capture information about the growing area of ​​Camellia oleifera, in actual operation, since Camellia oleifera is mostly distributed in mountainous areas with complex terrain, drone signals are often interfered with, and sample collection faces many challenges, making the number of high-quality labeled samples very limited.

[0086] To address the problem of low model accuracy caused by insufficient sample size, this embodiment captures high-resolution remote sensing images from the sample area. Preferably, this embodiment can use high-resolution satellite images and select high-quality samples from them. Taking the oil-tea plantation area of ​​Long X County as an example, for unlabeled remote sensing image samples, this example selects remote sensing images with a resolution better than 1m in Long X County in autumn (September-November) and performs full-process preprocessing, including but not limited to: radiation correction, geometric correction, and image enhancement. The remote sensing images are then georeferenced and mosaicked to obtain a final single, complete, and color-free remote sensing image file. For labeled crop samples, this embodiment uses oil-tea samples collected by drones and vectorized based on remote sensing images better than one meter produced by this unit. Multiple internal workers, through human-computer interaction, outline the boundaries on the image based on photos and shooting point information to form vector results. In order to improve recognition capabilities, in addition to the 855 outlined oil-tea samples, 41 forestland samples were added.

[0087] Step 220: Obtain a preset comparison learning strategy and image cropping information.

[0088] The image cropping information includes the input information of the self-supervised learning model and the key image information.

[0089] Step 230 : Preprocess the unlabeled remote sensing image sample based on the input information and the image key information to obtain an input image sample.

[0090] The input image sample contains local image details and global scene information.

[0091] A unified description of steps 220 to 230 is provided:

[0092] In this embodiment, the input information of the self-supervised learning model can be understood as the model's input requirements, which are used to determine the image size of the input to the self-supervised learning model. Preferably, the image size of the input to the self-supervised learning model can be a 512*512 slice size. Image key information is primarily used to select remote sensing images that meet quality requirements from unlabeled remote sensing images, ensuring that the selected unlabeled remote sensing images contain local image details while also taking into account global scene information.

[0093] In a specific implementation, the embodiment of the present application can pre-develop a special slicing tool, and use the slicing tool to crop the unlabeled remote sensing image according to the slice size represented by the input information, and crop the image to a size of 512×512 pixels. This size can ensure the local details of the image while taking into account the global scene information, thereby providing a suitable input scale for the SimCLR model, and obtaining input image samples as training data for self-supervised learning of SimCLR.

[0094] Step 240 : According to the comparison learning strategy, the first feature extraction module of the self-supervised learning model extracts and learns the structural features and semantic information in the image from the input image sample.

[0095] The first feature extraction module is used to capture local and global pattern features from the input image in scenarios of different complexities.

[0096] In the specific implementation, refer to Figure 3 In this example, the cropped 512×512 pixel image blocks, also known as the cropped input image samples, are fed into the SimCLR framework for unsupervised training. SimCLR uses a comparative learning strategy to gradually extract and learn structural features and semantic information from the image. This strategy effectively captures both local and global patterns in complex scenes, significantly improving the model's capabilities in feature extraction and semantic representation.

[0097] In an optional embodiment, the above-mentioned extraction and learning of structural features and semantic information in the image from the input image sample by the first feature extraction module of the self-supervised learning model may specifically include: inputting the input image sample into the self-supervised learning model; performing random enhancement based on each sample image in the input sample to obtain sample pair information corresponding to each sample image, the sample pair information including positive sample pairs and negative sample pairs; using cosine similarity to measure the sample pair information to obtain the similarity of each sample pair of the sample pair information; and performing comparison model training on the first feature extraction module based on the sample pair information and the similarity.

[0098] In a specific implementation, the use of SimCLR self-supervised learning can reduce the labeling cost. Compared with other self-supervised learning, its framework is concise, efficient and easy to implement. As long as there is enough input data, more discriminative features can be learned. This embodiment uses SimCLR self-supervised learning to construct a feature space with strong representation capabilities for the target area. When training the SimCLR model in this embodiment, SimCLR can use cosine similarity to measure the similarity of different image representations. Specifically, in the SimCLR framework, each input image generates two different views (called positive sample pairs) through two random data augmentations. The task of the model is to learn to make the feature representations of these two views as close as possible. At the same time, different views of other images are regarded as negative sample pairs, and the model needs to learn to make the feature representations of these negative samples as far apart as possible. This embodiment uses InfoNCE Loss as the core mechanism to guide the model to distinguish between positive and negative samples in the contrastive learning process, which plays a key role in strengthening the feature expression of similar images and suppressing the interference of dissimilar images.

[0099] In an optional embodiment, the embodiment of the present application uses cosine similarity to measure the sample pair information to obtain the similarity of each sample pair of the sample pair information, which may specifically include: the feature extraction module according to This measure measures the similarity between different images in a sample pair. A·B represents the dot product obtained by multiplying the corresponding components of vector A and vector B and summing them. This reflects the relative direction of the two vectors and indicates the degree of similarity between them. If the two vectors have the same direction, the dot product value is large; if the two vectors have opposite directions, the dot product value is small. ||A||||B|| represents the norm product of vector A and vector B, where ||A|| is the Euclidean norm of vector A and ||B|| is the Euclidean norm of vector B. The Euclidean norm can also be understood as the length or modulus of a vector.

[0100] In an optional embodiment, the embodiment of the present application performs comparison model training on the first feature extraction module based on the sample pair information and the similarity, which may specifically include: according to the formula Guide the first feature extraction model to distinguish positive sample pairs from negative sample pairs during the comparison learning process; wherein the feature vector z i and the eigenvector z j are all feature representations of positive sample pairs, sim(z i , z j ) is z i and z j The cosine similarity between them is used to measure the similarity of positive sample pairs, τ is the temperature parameter used to adjust the similarity softening degree in contrastive learning, N is the number of sample pairs in the batch, exp(sim(zi , z j ) / τ) represents the similarity score of the positive sample pair, It represents the sum of the similarity scores between the positive sample pair and all negative sample pairs.

[0101] Step 250: Share the model parameters of the first feature extraction module with the feature extraction module of the segmentation model, and migrate the training weights of the first feature extraction module to the segmentation module of the segmentation model.

[0102] For example, refer to Figure 4 , combining the SimCLR model and the Segformer model. After completing the feature extraction training of the SimCLR model, the Encoder module of Segformer can obtain the parameters of the SimCLR feature extraction module, ensuring that the global and local features learned by SimCLR are retained. Therefore, in the model feature extraction part of this embodiment, only the SimCLR model needs to be trained, so that the Encoder module of Segformer can effectively capture local and global pattern features in complex scenarios, significantly improving the model's capabilities in feature extraction and semantic expression, and ensuring that the subsequent segmentation model still maintains high-quality segmentation effects under limited computing resources.

[0103] Step 260: Vectorize the crop labeled samples to obtain vector data.

[0104] Step 270 , performing affine transformation on the vector data to obtain mapping data of the vector data on the image plane, and cropping a target area associated with the crop sample from the mapping data to obtain training image pair data.

[0105] The training image pair data includes a training set and a validation set.

[0106] A unified description of steps 260 to 270 is provided:

[0107] In actual implementation, refer to Figure 3 ,This embodiment first vectorizes the crop labeled samples, then uses the affine transformation technology to ,map the vector data to the image plane, and finally divides the training image pair data into ,a training set and a validation set according to a certain ratio.

[0108] Step 280: Perform segmentation training on the segmentation model based on the training set to obtain a model prediction result output by the segmentation model.

[0109] Step 290 , optimizing the parameters of the segmentation model according to the validation set and the model prediction result, until the model accuracy of the segmentation model reaches a preset requirement, thereby obtaining a trained segmentation model.

[0110] In step 300, a full-image recognition prediction including crop planting area extraction is performed based on the remote sensing image of the target area through the trained segmentation model, and vector conversion is performed based on the prediction result of the full-image recognition prediction to obtain the planting area recognition extraction result of the crops to be identified in the target area.

[0111] A unified description of steps 280 to 300 is provided:

[0112] Reference Figure 3 As shown, during the training phase of the segmentation model, this embodiment uses the validation set to verify the accuracy of the model prediction results. By presetting an accuracy requirement, it is determined whether the model accuracy meets the accuracy requirement. If the model accuracy does not meet the accuracy requirement, the segmentation model can be retrained using the training set to optimize and adjust the model parameters of the segmentation model until the optimal model that meets the accuracy requirements is obtained.

[0113] Among them, the model recognition accuracy is shown in Table 1 below:

[0114]

[0115] Table 1 Recognition accuracy of the model

[0116] Furthermore, as an example, after completing the Segformer segmentation model training, the Segformer model with the best accuracy can be used to perform full-map recognition prediction on the remote sensing image containing the oil-tea distribution in Long X County, Guangxi Province. The obtained oil-tea prediction results are converted into vector data to obtain the final oil-tea classification results, where the oil-tea extraction results are as follows: Figure 5 shown.

[0117] Therefore, this embodiment addresses the problems of insufficient samples and low precision in the automatic identification of oil-tea planting areas under current high-resolution images. By using SimCLR self-supervised learning and the Segformer semantic segmentation model, a more efficient and accurate recognition method is proposed. It can efficiently extract multi-scale features and adopts a simple and efficient decoder to improve the performance of image segmentation.

[0118] In an optional embodiment, the embodiment of the present application optimizes the parameters of the segmentation model according to the validation set and the model prediction result, which may specifically include: Quantitative evaluation of the segmentation model; where C is the number of categories, Prediction i is the prediction result of the i-th category, GroundTruth i is the true label corresponding to the prediction result of the i-th category.

[0119] In summary, the embodiment of the present application constructs a combined self-supervised learning model and a segmentation model, uses unlabeled remote sensing image samples obtained from the sample area for preprocessing, obtains input image samples, and then extracts features from the first feature extraction module of the self-supervised learning model in combination with the input image samples according to the comparison learning strategy, and learns the structural features and semantic information in the image, so that the self-supervised learning model can effectively capture local and global pattern features in complex scenes, significantly improving the model's ability in feature extraction and semantic expression. Subsequently, the first feature extraction module of the self-supervised learning model is parameter-shared with the feature extraction module of the segmentation model to achieve training and learning of the feature extraction module of the segmentation model, and adopts a transfer learning strategy to transfer the training weights to the segmentation module of the segmentation model, and then uses the crop labeled samples collected from the sample area for vectorization and affine transformation, etc., to obtain a training set and a validation set, uses the training set to train the segmentation model's ability to identify and segment crop planting areas from the input data, uses the validation set to verify the accuracy of the segmentation model, and applies the segmentation model with the best accuracy to the full-map prediction of the remote sensing image of the target area to achieve the segmentation of the crop planting area. Therefore, this embodiment proposes an efficient method for extracting oil-tea camellia planting areas under the condition of limited labeled samples by comprehensively applying SimCLR and Segformer models. First, the SimCLR model is used to perform self-supervised learning on unlabeled remote sensing images, fully mining and extracting multi-scale and multi-type features in the images, and constructing a feature space with strong characterization capabilities for the target area. Subsequently, in the training process of the Segformer model, a transfer learning strategy is adopted, and the encoder weights obtained from SimCLR pre-training are used for weight initialization of the encoder module of Segformer, reducing the dependence on a large amount of labeled data. In the case of only a small number of labeled samples, the Segformer model can still achieve efficient and accurate extraction of oil-tea camellia planting areas, solving the problem of insufficient recognition accuracy of crop recognition and segmentation models caused by the existing technology being limited by environmental factors and a small number of training samples.

[0120] It should be noted that, for the purpose of simple description, the method embodiments are expressed as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited to the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously.

[0121] like Figure 6 As shown, the embodiment of the present application further provides a crop planting area extraction system 600, comprising:

[0122] The sample extraction and model construction module 610 is used to obtain unlabeled remote sensing image samples and labeled crop samples from the sample area, and to construct a combined self-supervised learning model and segmentation model;

[0123] a feature extraction training module 620 for performing feature extraction training on a first feature extraction module of the self-supervised learning model using the unlabeled remote sensing image samples according to a preset comparison learning strategy, wherein the first feature extraction module is configured to capture local and global pattern features from input images in scenes of varying complexity;

[0124] a sharing module 630, configured to share the model parameters of the first feature extraction module with the feature extraction module of the segmentation model, and to migrate the training weights of the first feature extraction module to the segmentation module of the segmentation model;

[0125] A vectorization processing module 640 is configured to perform vectorization processing based on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set;

[0126] A segmentation training module 650 is configured to perform segmentation training on the segmentation model based on the training set to obtain a model prediction result output by the segmentation model;

[0127] a parameter optimization module 660 for optimizing the parameters of the segmentation model according to the validation set and the model prediction result, until the model accuracy of the segmentation model reaches a preset requirement, thereby obtaining a trained segmentation model;

[0128] The full-image recognition prediction module 670 is used to perform full-image recognition prediction including crop planting area extraction based on the remote sensing image of the target area through a trained segmentation model, and to perform vector conversion based on the prediction results of the full-image recognition prediction to obtain the planting area recognition and extraction results of the crops to be identified in the target area.

[0129] Optionally, the sample extraction and model building module 610 includes:

[0130] Remote sensing image data acquisition submodule, used to obtain remote sensing image data from the sample area;

[0131] A first preprocessing submodule is used to perform correction and enhancement preprocessing on the remote sensing image data to obtain unlabeled remote sensing image samples that meet quality requirements;

[0132] a collection submodule, configured to collect crop sample images from the sample area based on a preset drone collection instruction, wherein the drone collection instruction is configured to control the drone to collect sample images from the sample area;

[0133] The labeled sample generation submodule is configured to obtain the labeling information corresponding to the crop sample image, and to generate a labeled crop sample based on the crop sample image and the labeling information.

[0134] Optionally, the feature extraction training module 620 includes:

[0135] An image processing submodule, configured to obtain a preset comparison learning strategy and image cropping information, wherein the image cropping information includes input information of the self-supervised learning model and key image information;

[0136] A third preprocessing submodule is configured to preprocess the unlabeled remote sensing image sample based on the input information and the image key information to obtain an input image sample, where the input image sample contains local image details and global scene information;

[0137] The first learning submodule is used to extract and learn structural features and semantic information in the image from the input image sample through the first feature extraction module of the self-supervised learning model according to the comparison learning strategy.

[0138] Optionally, the first learning submodule includes:

[0139] An input unit, configured to input the input image sample into the self-supervised learning model;

[0140] A random enhancement unit, configured to perform random enhancement based on each sample image in the input sample to obtain sample pair information corresponding to each sample image, wherein the sample pair information includes a positive sample pair and a negative sample pair;

[0141] a measurement processing unit, configured to perform measurement processing on the sample pair information using cosine similarity to obtain similarities between each sample pair in the sample pair information;

[0142] A comparison training unit is used to perform comparison model training on the first feature extraction module according to the sample pair information and the similarity.

[0143] Optionally, the measurement unit is specifically configured to: the first feature extraction module is based on Measures the similarity of different images in the sample pair information; where A·B represents the dot product obtained by multiplying the corresponding components of vector A and vector B and summing them, which reflects the relative direction relationship between the vectors. As the degree of similarity between the two vectors, if the two vectors have the same direction, the value corresponding to the dot product is large; if the two vectors have opposite directions, the value corresponding to the dot product is small. ||A||||B|| represents the norm product of vector A and vector B, ||A|| is the Euclidean norm of vector A, and ||B|| is the Euclidean norm of vector B.

[0144] Optionally, the comparison training unit is specifically used to: guiding the first feature extraction model to distinguish positive sample pairs from negative sample pairs during the comparison learning process;

[0145] Among them, the eigenvector z i and the eigenvector z j are all feature representations of positive sample pairs, sim(z i ,z j ) is z i and z j The cosine similarity between them is used to measure the similarity of positive sample pairs, τ is the temperature parameter used to adjust the similarity softening degree in contrastive learning, N is the number of sample pairs in the batch, exp(sim(z i ,z j ) / τ) represents the similarity score of the positive sample pair, It represents the sum of the similarity scores between the positive sample pair and all negative sample pairs.

[0146] Optionally, the vectorization processing module includes:

[0147] A vectorization submodule, configured to perform vectorization based on the crop labeled samples to obtain vector data;

[0148] an affine transformation submodule, configured to perform affine transformation on the vector data to obtain mapping data of the vector data on an image plane;

[0149] The cropping submodule is used to crop the target area associated with the crop sample from the mapping data to obtain training image pair data.

[0150] Optionally, the parameter optimization module is specifically used to: Quantitative evaluation of the segmentation model; where C is the number of categories, Prediction i is the prediction result of the i-th category, Ground Truth i is the true label corresponding to the prediction result of the i-th category.

[0151] It should be noted that the crop planting area extraction system provided in the embodiment of the present application can execute the crop planting area extraction method coupled with self-supervised learning and segmentation model provided in any embodiment of the present application, and has the corresponding functions and beneficial effects of the execution method.

[0152] In a specific implementation, the above-mentioned crop planting area extraction system can be integrated into a device so that the device can use sample data to train a coupled self-supervised learning model and a segmentation model, and after the segmentation model completes the model training, the segmentation model is used to identify the planting area of ​​the crop to be identified from the remote sensing image. As an electronic device, it can use unlabeled samples and a small number of labeled samples to train a segmentation model that can accurately identify and segment the crop planting area. The electronic device can be composed of two or more physical entities, or it can be composed of one physical entity. For example, the electronic device can be a personal computer (PC), a computer, a server, etc., and the embodiments of the present application do not impose specific restrictions on this.

[0153] like Figure 7 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114; the memory 113 is used to store computer programs; the processor 111 is used to execute the program stored in the memory 113, and implement the steps of the crop planting area extraction method coupled with self-supervised learning and segmentation model provided by any of the aforementioned method embodiments. Exemplarily, the steps of the crop planting area extraction method coupled with self-supervised learning and segmentation models may include the following steps: obtaining unlabeled remote sensing image samples and labeled crop samples from the sample area, and constructing a combined self-supervised learning model and segmentation model; according to a preset comparison learning strategy, performing feature extraction training on the first feature extraction module of the self-supervised learning model through the unlabeled remote sensing image samples, the first feature extraction module being used to capture local and global pattern features from the input image in scenes of different complexities; sharing the model parameters of the first feature extraction module to the feature extraction module of the segmentation model, and migrating the training weights of the first feature extraction module to the segmentation model. The invention provides a segmentation module of the type; performing vectorization processing based on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set; performing segmentation training on the segmentation model according to the training set to obtain a model prediction result output by the segmentation model; optimizing the parameters of the segmentation model according to the validation set and the model prediction result until the model accuracy of the segmentation model reaches a preset requirement, thereby obtaining a trained segmentation model; performing full-map recognition prediction including crop planting area extraction based on the remote sensing image of the target area through the trained segmentation model, and performing vector conversion based on the prediction result of the full-map recognition prediction to obtain a planting area recognition and extraction result of the crop to be identified in the target area.

[0154] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the crop planting area extraction method coupled with self-supervised learning and segmentation model provided in any of the aforementioned method embodiments are implemented.

[0155] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0156] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A method for extracting crop planting areas by coupling self-supervised learning and segmentation models, characterized in that: include: Obtain unlabeled remote sensing image samples and labeled crop samples from the sample area, and construct a combined self-supervised learning model and segmentation model; According to a preset comparison learning strategy, a first feature extraction module of the self-supervised learning model is trained with the unlabeled remote sensing image samples for feature extraction, wherein the first feature extraction module is used to capture local and global pattern features from the input image in scenes of different complexities; Sharing the model parameters of the first feature extraction module to the feature extraction module of the segmentation model, and migrating the training weights of the first feature extraction module to the segmentation module of the segmentation model; Performing vectorization processing on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set; Performing segmentation training on the segmentation model according to the training set to obtain a model prediction result output by the segmentation model; Optimizing the parameters of the segmentation model according to the validation set and the model prediction result until the model accuracy of the segmentation model reaches a preset requirement, thereby obtaining a trained segmentation model; The trained segmentation model is used to perform full-image recognition prediction including crop planting area extraction based on the remote sensing image of the target area, and vector conversion is performed based on the prediction results of the full-image recognition prediction to obtain the planting area recognition and extraction results of the crops to be identified in the target area.

2. The method according to claim 1, characterized in that The step of obtaining unlabeled remote sensing image samples and labeled crop samples from the sample area includes: Acquire remote sensing image data from the sample area; Performing correction and enhancement preprocessing on the remote sensing image data to obtain unlabeled remote sensing image samples that meet quality requirements; Collecting crop sample images from the sample area based on a preset drone collection instruction, wherein the drone collection instruction is used to control the drone to collect sample images from the sample area; Acquire annotation information corresponding to the crop sample image, and generate a crop annotation sample based on the crop sample image and the annotation information.

3. The method according to claim 1, characterized in that The method of performing feature extraction training on the first feature extraction module of the self-supervised learning model using the unlabeled remote sensing image samples according to a preset comparison learning strategy includes: Obtaining a preset comparison learning strategy and image cropping information, wherein the image cropping information includes input information of the self-supervised learning model and key image information; Based on the input information and the image key information, preprocessing the unlabeled remote sensing image sample to obtain an input image sample, wherein the input image sample includes local image details and global scene information; According to the comparison learning strategy, the first feature extraction module of the self-supervised learning model extracts and learns the structural features and semantic information in the image from the input image sample.

4. The method according to claim 3, characterized in that The extracting and learning structural features and semantic information in the image from the input image sample by the first feature extraction module of the self-supervised learning model includes: Inputting the input image sample into the self-supervised learning model; Performing random enhancement on each sample image in the input sample to obtain sample pair information corresponding to each sample image, wherein the sample pair information includes a positive sample pair and a negative sample pair; Using cosine similarity to measure the sample pair information, and obtain the similarity of each sample pair in the sample pair information; Perform comparison model training on the first feature extraction module according to the sample pair information and the similarity.

5. The method according to claim 4, characterized in that The measuring process of the sample pair information by using cosine similarity to obtain the similarity of each sample pair in the sample pair information includes: The first feature extraction module is based on Measuring the similarity of different images in the sample pair information; Where A·B represents the dot product obtained by multiplying and summing the corresponding components of vector A and vector B, which reflects the relative direction relationship between the vectors and serves as the degree of similarity between the two vectors. If the two vectors have the same direction, the dot product value is large; if the two vectors have opposite directions, the dot product value is small. ‖A‖‖B‖ represents the norm product of vector A and vector B, where ||A|| is the Euclidean norm of vector A, and ||B|| is the Euclidean norm of vector B.

6. The method according to claim 5, characterized in that The performing comparison model training on the first feature extraction module according to the sample pair information and the similarity includes: According to the formula guiding the first feature extraction module to distinguish positive sample pairs from negative sample pairs during the comparison learning process; Among them, the eigenvector z i and the eigenvector z j are all feature representations of positive sample pairs, sim(z i ,z j ) is z i and z j The cosine similarity between them is used to measure the similarity of positive sample pairs, τ is the temperature parameter used to adjust the similarity softening degree in contrastive learning, N is the number of sample pairs in the batch, exp(sim(z i ,z j ) / τ) represents the similarity score of the positive sample pair, It represents the sum of the similarity scores between the positive sample pair and all negative sample pairs.

7. The method according to claim 1, characterized in that The vectorization processing based on the crop labeled samples to obtain training image pair data includes: Performing vectorization based on the crop labeled samples to obtain vector data; Affine transformation is performed on the vector data to obtain mapping data of the vector data on an image plane, and a target area associated with the crop sample is cropped from the mapping data to obtain training image pair data.

8. The method according to claim 1, characterized in that Optimizing parameters of the segmentation model according to the validation set and the model prediction result includes: according to performing a quantitative evaluation on the segmentation model; Among them, C is the number of categories, Prediction i is the prediction result of the i-th category, Ground Truth i is the true label corresponding to the prediction result of the i-th category.

9. A crop planting area extraction system, characterized in that: include: The sample extraction and model building module is used to obtain unlabeled remote sensing image samples and labeled crop samples from the sample area, and to build a combined self-supervised learning model and segmentation model; a feature extraction training module, configured to perform feature extraction training on a first feature extraction module of the self-supervised learning model using the unlabeled remote sensing image samples according to a preset comparison learning strategy, wherein the first feature extraction module is configured to capture local and global pattern features from input images in scenes of varying complexity; a sharing module, configured to share the model parameters of the first feature extraction module with the feature extraction module of the segmentation model, and to migrate the training weights of the first feature extraction module to the segmentation module of the segmentation model; A vectorization processing module, configured to perform vectorization processing based on the crop labeled samples to obtain training image pair data, wherein the training image pair data includes a training set and a validation set; A segmentation training module, configured to perform segmentation training on the segmentation model based on the training set to obtain a model prediction result output by the segmentation model; A parameter optimization module is used to optimize the parameters of the segmentation model according to the validation set and the model prediction result until the model accuracy of the segmentation model meets the preset requirements, thereby obtaining a trained segmentation model; The full-image recognition prediction module is used to perform full-image recognition prediction including crop planting area extraction based on the remote sensing image of the target area through the trained segmentation model, and to perform vector conversion based on the prediction results of the full-image recognition prediction to obtain the planting area recognition and extraction results of the crops to be identified in the target area.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the steps of the method for extracting crop planting areas coupled with self-supervised learning and segmentation models as described in any one of claims 1 to 8 when executing the program stored in the memory.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on self-supervised contrast learning

    CN113011427A

  • Remote sensing identification method for agricultural planting structure

    WO2022214039A1