Remote sensing mulching film extraction method and system based on directional target detection and prompt segmentation
By improving the directional target detection and hint segmentation method, combining the Mamba-YOLO and SAM models, and using the iterative mask optimization segmentation strategy, the problems of complex background interference and changes in film properties in remote sensing film extraction are solved, and higher-precision remote sensing film extraction is achieved.
Patent Information
- Application Number
- CN202510801975.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
Existing remote sensing ground film extraction methods are difficult to accurately distinguish ground film from other ground objects in complex backgrounds, and changes in the ground film's own characteristics lead to differences in spectral characteristics, resulting in low extraction accuracy.
A method based on directional target detection and hint segmentation is adopted, combined with the improved Mamba-YOLO model and SAM model. The segmentation strategy is optimized through iterative masking. The directional target detection model is used to accurately interpret the ground film area and perform point sampling in the core area. The hint segmentation model is driven to perform multiple predictions to improve segmentation accuracy.
It improves the accuracy and completeness of mulch film extraction, ensures complete coverage and accurate positioning of mulch film areas in complex scenarios, and is suitable for accurate statistics and change monitoring of agricultural remote sensing mulch films.
Smart Images

Figure CN120708059A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of remote sensing image recognition, in particular to a remote sensing ground film extraction method based on directional target detection and prompt segmentation. Background Art
[0002] Film mulching is widely used in agricultural production, particularly in arid and semi-arid regions, where it plays a vital role in increasing crop yields and improving soil quality. However, with the increasing use of film, the problem of "white pollution" caused by residual film has become increasingly prominent. Therefore, accurate and rapid extraction of film cover information is crucial for agricultural production management, environmental protection, and sustainable development. Remote sensing technology, with its wide coverage, rapid information acquisition, and low cost, has become a key tool for film mulching.
[0003] At present, the research on extraction using remote sensing technology at home and abroad mainly focuses on the following aspects: (1) Methods based on spectral features. Early studies mainly used the difference in spectral reflectance between ground film and other ground objects (such as soil and vegetation) in visible light, near-infrared and other bands to construct various spectral indices (such as plastic ground film index PMI, normalized difference ground film index NDMI, etc.) for ground film extraction. However, due to differences in ground film type, aging degree, soil background, etc., a single spectral index is often difficult to obtain stable and reliable extraction results in different regions and at different times. In addition, there are also studies that use the absorption characteristics of ground film in specific bands to identify ground film, such as the short-wave infrared band. (2) Methods based on machine learning. With the development of machine learning technology, some studies have applied classification algorithms such as random forest (RF) and support vector machine (SVM) to ground film extraction. By combining multiple features such as spectral features, texture features, and vegetation index, the accuracy of ground film extraction has been improved. For example, a study constructed different feature combination schemes based on Sentinel-2 remote sensing images and random forest algorithms to identify film-covered farmland in the Loess Plateau. (3) Methods based on deep learning. In recent years, deep learning has made significant progress in the field of image processing and has also been applied to remote sensing ground film extraction. Convolutional neural networks (CNNs) can automatically extract deep features of images and have strong feature expression capabilities. Some studies have used classic semantic segmentation networks such as U-Net and DeepLab, as well as Transformer-based semantic segmentation algorithms, to achieve automatic ground film extraction. At the same time, there are also studies that apply target detection technology to ground film extraction. For example, a study proposed a method for directional target detection in remote sensing images based on feature reconstruction. (4) Methods based on prompt segmentation. Recently, Meta AI Research proposed the Segment Anything Model (SAM), a new image segmentation method based on a visual basic model. SAM introduces the idea of prompt engineering, realizes prompt segmentation based on points, boxes, masks, and even free-form text, and uses ultra-large-scale data sets for pre-training, with strong generalization capabilities. Some studies have applied the SAM model to the field of remote sensing and achieved certain results.
[0004] However, existing methods for extracting ground film still face some challenges: (1) Complex background interference; the types of ground objects in remote sensing images are diverse, and ground film is often mixed with bare soil, greenhouses, roads and other ground objects, with similar spectral characteristics, making them difficult to distinguish; (2) Changes in the characteristics of the ground film itself; ground films of different types, thicknesses, and aging degrees have different spectral characteristics. Summary of the Invention
[0005] The purpose of the present invention is to address the above-mentioned problems and provide a remote sensing mulch film extraction method and system based on directional target detection and prompt segmentation, so as to more accurately identify and more completely extract remote sensing mulch films, thereby more effectively supporting the accurate statistics and change monitoring of agricultural remote sensing mulch films.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] The remote sensing ground film extraction method based on directional target detection and hint segmentation includes the following contents:
[0008] Step S1, constructing a ground film extraction model, including the following: constructing a directional target detection model for ground film detection, and constructing a prompt segmentation model for ground film segmentation;
[0009] Step S2: Obtain a target image containing a ground film area, use a directional target detection model to detect the target image to obtain a directional detection frame of the ground film target, sample several points within the core area calibrated by the directional detection frame as the first prompt input of the prompt segmentation model, use the prompt segmentation model to predict the target and extract the first target mask, use the first target mask as the second prompt input of the prompt segmentation model, use the prompt segmentation model again to predict the target and extract the second target mask to output the segmentation result of the ground film.
[0010] Among them, the specific processing flow of step S2 is as follows: step 201, initialize the directional target detection model and the prompt segmentation model; step 202, set the prediction parameters of the directional target detection model, perform object detection on the acquired target image and save the detection results; step 203, traverse the detection results, if a target is detected, obtain the coordinates of the directional target detection frame, convert the rotation angle of the directional target detection frame from radians to degrees, and then construct the directional target detection frame; step 204, sample several points in the core area marked by the directional target detection frame; step 205, use the several sampling points as the first prompt input of the prompt segmentation model; step 206, use the prompt segmentation model to predict the target in the area marked by the directional target detection frame; step 207, extract and obtain the first target mask; step 208, use the first target mask as the second prompt input of the prompt segmentation model, return to execute the aforementioned step 206-step 207 operations, and extract the second target mask to output as the segmentation result of the ground film; wherein the first target mask is the mask with the smallest area; the second target mask is the mask with the highest confidence.
[0011] In step S1, the directional target detection model is an optimized and improved Mamba-YOLO model, and its construction process is as follows: Step 101, construct an improved Mamba-YOLO model for ground film detection; Step 102, set the training parameters and hyperparameters of the improved Mamba-YOLO model; Step 103, train and validate the improved Mamba-YOLO model using the training and validation sets of directional target detection in ground film remote sensing images; Step 104, test the trained improved Mamba-YOLO model using the test set of directional target detection in ground film remote sensing images to obtain the optimized and improved Mamba-YOLO model. The hint segmentation model is a SAM model, and its construction process is as follows: set the image segmentation parameters of the SAM model to obtain the SAM model for ground film segmentation in remote sensing images.
[0012] The construction process of the improved Mamba-YOLO model is as follows: Step 111, in the backbone network of the Mamba-YOLO model, a large target detection layer P6 is introduced to generate a smaller feature layer; Step 112, in the head network of the Mamba-YOLO model, upsampling is performed to generate a feature layer, which is then combined with the output of the eighth layer of the backbone network to match the number of channels and size; Step 113, the feature map combined in Step 112 is detected and classified in the head network together with the output of other detection layers; Step 114, in the head network of the Mamba-YOLO model, the first and second Conv modules are replaced with receptive field attention convolution RFAConv modules respectively; Among them, the RFAConv module adopts a fusion design of deformable convolution and attention mechanism, which can effectively enhance the multi-scale feature representation capability; It is expressed as: where Δp k is the deformable offset, α is the adaptive attention weight coefficient; Step 115, replace the Mamba-YOLO model detection head from standard bounding box detection to oriented bounding box detection, and expand the output dimension to support direction angle prediction; Step 116, use the Inner-SIoU loss function to replace the original loss function CIoU in the Mamba-YOLO model to obtain the improved Mamba-YOLO model. The calculation formula of the Inner-SIoU loss function is: L Inner-SIoU =L SIoU +(IoU-IoU inner ), where L SIoU is the original SIoU loss, IoU is the intersection-over-union ratio of the predicted box to the original real box, IoU inner It is the intersection-over-union ratio of the predicted box and the scaled true box.
[0013] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:
[0014] The present invention integrates the optimized and improved directional target detection model with the prompt segmentation model, and uses an iterative mask optimization segmentation strategy to improve the integrity of the film contour extraction and the edge positioning accuracy. First, the improved directional target detection model is used to accurately interpret the direction and boundary of the remote sensing film area; then, point sampling is performed within the core area calibrated by the directional target detection frame to capture the spectral heterogeneity and local texture characteristics of the agricultural film area; these sampling points will be input into the prompt segmentation model as initial prompts to drive the model to make a preliminary mask prediction for the target area; from these prediction results, the mask with the smallest area is parsed and extracted as a new prompt for the prompt segmentation model, and then input into the prompt segmentation model again for re-prediction and refinement; the mask with the highest confidence is selected from the optimized prediction results and used as the final film segmentation result. This iterative process aims to use the macro information of the previous segmentation to guide the model to fit the target boundary more accurately in the next round of prediction, so as to improve the segmentation accuracy in complex scenes and ensure that the final mask has complete coverage and accurate positioning of the film area. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of the remote sensing image ground film area extraction method of the present invention.
[0016] Figure 2 It is a schematic diagram of the improved Mamba-YOLO model structure. DETAILED DESCRIPTION
[0017] The specific implementation of the invention is further described below with reference to the accompanying drawings.
[0018] Example 1
[0019] As mentioned above, the solution selected in this application adopts the optimized and improved Mamba-YOLO model and the SAM model, and the directional target detection model can also be a YOLO series model such as YOLO v8, YOLO v9, YOLO v10, YOLO v11 or YOLO v12, and the prompt segmentation model can also be a model improved based on SAM such as TinySAM, PointSAM or SAM2. Features can be combined according to specific application examples. The following embodiment will take the directional target detection model as the optimized and improved Mamba-YOLO model and the prompt segmentation model as the SAM model as an example for specific explanation.
[0020] See also Figure 1 and Figure 2 In this embodiment 1, a remote sensing ground film extraction method based on directional target detection and prompt segmentation includes the following steps:
[0021] Step 1: Collect high-resolution remote sensing images of the area covered by agricultural mulch film and stitch the images to obtain a complete remote sensing image of the mulch film. The details are as follows: (1) Obtain high-resolution remote sensing images of agricultural mulch film based on the UAV RGB visible light sensor; (2) Stitch the high-resolution remote sensing images of agricultural mulch film obtained in (1).
[0022] Step 2: Preprocess the stitched agricultural film remote sensing images (high-resolution agricultural film remote sensing images) to eliminate interference from sensors and environmental factors. Specifically, perform geometric correction and other preprocessing on the agricultural film remote sensing images obtained in Step 1 to obtain preprocessed agricultural film remote sensing images. This corrects or compensates for image information loss during the remote sensing imaging process.
[0023] Step 3: Based on the image enhancement index, the agricultural film remote sensing image is subjected to boundary enhancement processing to obtain the agricultural film remote sensing image after boundary enhancement processing. Specifically, the agricultural film remote sensing image obtained in step 2 is processed by the image enhancement index to obtain the agricultural film remote sensing image after boundary enhancement processing to highlight the boundary features of the agricultural film, including: (1) obtaining the green band and the blue band in the agricultural film remote sensing image obtained in step 2; (2) calculating the image enhancement index based on the green band and the blue band obtained in (1); (3) enhancing and highlighting the boundary feature information in the agricultural film remote sensing image obtained in step 2 based on the image enhancement index obtained in (2).
[0024] Step 4: Use an image data enhancement algorithm to enhance the agricultural film remote sensing image obtained in Step 3 to obtain a data-enhanced agricultural film remote sensing image. Specifically, the agricultural film remote sensing image obtained in Step 3 is enhanced using three image data enhancement methods: random rotation, brightness adjustment, and regional occlusion. This enhances sample diversity and alleviates the small sample size training problem.
[0025] Step 5: Use the image annotation tool X-Anylabeling to annotate the agricultural film remote sensing image obtained in step 4 to obtain the agricultural film remote sensing image annotation file. Specifically: Use the image annotation tool X-Anylabeling to annotate the agricultural film remote sensing image obtained in step 4 with a directional target rectangular box at the pixel level to obtain the directional target detection annotation file corresponding to the agricultural film remote sensing image.
[0026] Step 6: Using a multi-scale sliding window mechanism, the original image is systematically sampled at different scales and intervals, and a dataset with YOLO OBB format annotations is generated simultaneously, achieving data enhancement while maintaining the target spatial distribution characteristics. Specifically, a multi-scale sliding window strategy is used to systematically sample the agricultural film remote sensing image obtained in Step 4. Using a predefined 640x640 pixel window size and a 100 pixel step size, grid cropping is performed at three scales: 0.5×, 1.0×, and 1.5×. Image segmentation and label mapping are performed simultaneously to obtain a structured dataset that meets the YOLO OBB dataset format requirements.
[0027] Step 7: Using a randomization method, randomly divide the images and label files from step 6 into training, validation, and test sets in a ratio of 7:2:1. Then, organize them into training, validation, and test sets according to the organizational structure of the directional target detection dataset to obtain the directional target detection dataset for agricultural remote sensing mulch films. Specifically, a fixed random seed (seed = 42) is used to randomly divide the images and label files from step 6 into training, validation, and test sets in a ratio of 7:2:1 to ensure reproducible division results. Then, the three datasets are adjusted according to the organizational structure of the directional target detection dataset to obtain the directional target detection dataset for agricultural remote sensing mulch films.
[0028] Step 8: Improve the backbone network and head network of the Mamba-YOLO model to build an improved Mamba-YOLO model for mulch film detection, and obtain an improved Mamba-YOLO model suitable for agricultural mulch film remote sensing images. The improved target detection model Mamba-YOLO includes:
[0029] (1) In the backbone network of the Mamba-YOLO model, a large target detection layer P6 is introduced to generate smaller feature layers to better capture the characteristics of agricultural remote sensing mulch films. P6 is usually located in the higher layers of the network. The feature maps at this layer help capture larger and more distant targets in the image. For objects that occupy a large area in the image, by adding the P6 layer, the network can capture the global information of these large targets at a higher level, thereby improving detection accuracy.
[0030] (2) Upsampling is performed in the head network of the Mamba-YOLO model to generate a feature layer with 768 channels, which is then combined with the eighth layer output of the backbone network to match the number of channels and size.
[0031] (3) The feature map combined in (2) is combined with the output of other detection layers and processed in the head network for detection and classification. This allows larger objects to be detected on smaller feature maps, thereby enhancing the representation ability of objects in the image.
[0032] (4) In the head network of the Mamba-YOLO model, the first and second Conv modules are replaced with Receptive Field Attention Convolution (RFAConv) to enhance the detection capability of multi-scale targets. The RFAConv module adopts a fusion design of deformable convolution and attention mechanism, which can effectively enhance the multi-scale feature representation capability. It is expressed as: where Δp k is the deformable offset, and α is the adaptive attention weight coefficient, which effectively enhances the multi-scale feature representation capability.
[0033] (5) The Mamba-YOLO model detection head is replaced from standard bounding box detection (Detect) to oriented bounding box detection (OBB), and the output dimension is expanded to support directional angle prediction to achieve oriented target detection of agricultural remote sensing mulch. Oriented bounding box detection (OBB) inherits from standard bounding box detection (Detect). Oriented bounding box detection (OBB) mainly adds a new convolution block to predict the angle (Angle), and finally adds an angle information during forward propagation.
[0034] (6) The Inner-SIoU loss function is used to replace the original loss function CIoU in the Mamba-YOLO model to solve the problem of target detection degradation of the CIoU loss function in complex scenes, so as to improve the performance of agricultural remote sensing mulch target detection. The calculation formula of the Inner-SIoU loss function is: L Inner-SIoU =L SIoU +(IoU-IoU inner ), where L SIoU is the original SIoU loss (including angle, distance, and shape penalties), IoU is the intersection-over-union ratio between the predicted box and the original real box, IoU inner It is the intersection-over-union ratio of the predicted box and the scaled true box.
[0035] (7) The structure of the improved Mamba-YOLO model described above includes: The improved Mamba-YOLO model includes a backbone network and a head network. The backbone network is responsible for feature extraction and includes: a SimpleStem module (which performs preliminary processing on the input image, including two convolution operations with a stride of 2 and a kernel size of 3), multiple VSSBlock modules (for feature extraction), a VisionClueMerge module (for merging feature maps of different scales, where the sixth VisionClueMerge module generates the P6 layer for large object detection), and an SPPF module (for fusing features of different scales). The head network is responsible for target detection, which includes: upsampling operation (enlarging the size of the feature map), Concat operation (joining the feature map with the output of the corresponding layer in the backbone network, specifically including the 11th layer and the 7th layer, the 14th layer and the 5th layer, and the 17th layer and the 3rd layer), XSSBlock module (feature processing), RFAConv module (replacing the original first and second Conv modules to enhance receptive field attention), Conv module (convolution operation), OBB module (replacing the original Detect module for directional target detection, and using the output of the 19th, 22nd, 25th, and 28th layers as the input of the OBB module). In addition, the model uses the Inner-SIoU loss function instead of the original CIoU loss function.
[0036] Step 9: Set the training parameters and hyperparameters of the improved Mamba-YOLO model from step 8, and train and validate the model using the training set and validation set obtained from step 7. The steps for training the improved object detection model Mamba-YOLO are as follows:
[0037] (1) Set the training parameters in the training script: set epoch to 300; set imgsz to 640; set optimizer to AdamW; and set amp to False.
[0038] (2) Set the corresponding model method to be executed according to the task type (train, val) in the training script.
[0039] (3) Set the training and validation parameters in the default.py script: set patience to 100; set close_mosaic to 10; set pretrained to False; set lr to 0.01; set momentum to 0.937; and set weight_decay to 0.0005.
[0040] (4) Train and verify the improved target detection model Mamba-YOLO on the training set and validation set obtained in step 7.
[0041] Step 10: Use the test dataset obtained in step 7 to test the generalization ability of the model trained in step 9. If the test result does not meet expectations, readjust the dataset, training parameters, and model parameters and retrain to obtain an optimized agricultural remote sensing mulch film extraction model.
[0042] If the test results meet expectations, the next step is carried out; if the test results do not meet expectations, the dataset, training parameters, and model parameters are readjusted and retrained to obtain an optimized agricultural remote sensing film extraction model. Specifically, the following steps are performed: (1) Optimizing the training set, validation set, and test set partitioning algorithm. (2) Optimizing the training parameters of the agricultural remote sensing film extraction model. (3) Optimizing the model parameters of the agricultural remote sensing film extraction model.
[0043] Step 11: Modify the image segmentation parameters in the suggested segmentation model SAM and build a SAM model for ground film in order to better predict the target mask.
[0044] Modify the image segmentation parameters in the SAM model to make it suitable for remote sensing ground film segmentation, including: (1) Change the value of the parameter crop_n_layers in the generate() function from the default value of 0 to 1, increase the number of cropped layers, improve the density of the cropped area, and capture finer-grained targets. (2) Change the parameter crop_overlap_ratio in the generate() function from the default value of 512 / 1500 to 640 / 1500, increase the overlap area between cropped blocks, and reduce the probability of target edges being cut. (3) Change the value of the parameter conf_thres in the generate() function from the default value of 0.88 to 0.90, increase the confidence threshold, and filter the mask predicted by the model.
[0045] Step 12: Integrate the optimized model obtained in step 10 with the SAM model with modified parameters in step 11 for prediction.
[0046] Step 13: After the target is detected in the improved Mamba-YOLO model, several points are sampled within the core area of the oriented target frame calibration. This includes: (1) Initializing the oriented target detection model optimized in step 8 and the SAM model based on the ViT skeleton. (2) Setting the prediction parameters of the oriented target detection model initialized in (1), performing object detection on the acquired image and saving the detection results. (3) Traversing the detection results, if a target is detected, obtaining the coordinates of the oriented target detection frame, converting the rotation angle of the oriented target detection frame from radians to degrees, and then constructing the oriented target detection frame. (4) Sampling 5 points within the core area of the oriented target detection frame calibration constructed in (3).
[0047] Step 14: Use the sampled points obtained in step 13 as the first prompt input to the SAM model, perform preliminary mask prediction on the area marked by the target detection box, and parse and extract the smallest mask from the prediction results. Specifically, use the five sampled points obtained in step 13 as the first prompt input to the SAM model, use the SAM model to perform preliminary target mask prediction on the area marked by the target detection box, and then parse and extract the smallest mask from the prediction results.
[0048] Step 15: The mask with the smallest area obtained in step 14 is used as the second prompt input of the SAM model, and the area marked by the target detection frame is re-predicted and refined. The mask with the highest confidence is parsed and extracted from the optimized prediction result, and used as the final ground film segmentation result. Specifically, the SAM model in step 14 will segment three masks. Here, the mask with the smallest area is preferably taken as the first target mask to ensure that the mask does not overflow the actual ground film range; then the mask with the smallest area (first target mask) obtained in step 14 is used as the second prompt input of the SAM model, and is used again for target prediction by SAM. The area marked by the target detection frame is re-predicted and refined to obtain an optimized prediction result. Then, the mask with the highest confidence (second target mask) is parsed and extracted from the optimized prediction result, and used as the final ground film segmentation result.
[0049] As mentioned above, the problems existing in the existing technology are: (1) Complex background interference; There are various types of ground objects in remote sensing images, and ground films are often mixed with bare soil, greenhouses, roads and other ground objects, with similar spectral characteristics, which are difficult to distinguish; (2) Changes in the characteristics of the ground film itself; Ground films of different types, thicknesses, and aging degrees have different spectral characteristics. This application provides a solution for remote sensing ground film extraction based on directional target detection and prompt segmentation, which improves the accuracy and efficiency of ground film extraction by combining the advantages of directional target detection and prompt segmentation technology. This method integrates the optimized and improved directional target detection model with the prompt segmentation model, and uses an iterative mask optimization segmentation strategy to improve the integrity of ground film contour extraction and edge positioning accuracy. First, an improved directional target detection model is used to accurately interpret the direction and boundary of the remote sensing film area; then, point sampling is performed within the core area calibrated by the directional target detection frame to capture the spectral heterogeneity and local texture characteristics of the agricultural film area; these sampling points are input into the prompt segmentation model as initial prompts to drive the model to make preliminary mask predictions for the target area; from these prediction results, the mask with the smallest area is parsed and extracted as a new prompt for the prompt segmentation model, and then input into the prompt segmentation model again for re-prediction and refinement; the mask with the highest confidence is selected from the optimized prediction results and used as the final film segmentation result. This iterative process aims to use the macro information of the previous segmentation to guide the model to fit the target boundary more accurately in the next round of prediction, so as to improve the segmentation accuracy in complex scenes and ensure that the final mask has complete coverage and accurate positioning of the film area. The specific advantages are as follows:
[0050] 1. This paper proposes a remote sensing ground film directional target detection method based on an optimized and improved Mamba-YOLO model to more accurately adapt to the arbitrary directional distribution of remote sensing ground film.
[0051] 2. The present invention integrates the optimized and improved Mamba-YOLO model with the suggested segmentation SAM model, combining the efficient detection capability of Mamba-YOLO with the ability of SAM to segment objects in a zero-shot manner, so that the remote sensing ground film can be segmented and extracted after it is detected.
[0052] 3. The present invention can use several points within the core area of the directional target detection frame as input cues for the SAM model's segmentation prompts, leveraging its powerful visual generalization capabilities to perform refined segmentation of the film-film area. Furthermore, the present invention's iterative film mask optimization segmentation strategy effectively utilizes the macroscopic information obtained from the first segmentation prompt model to guide the model to more accurately fit the target boundary in the second prediction, ensuring that the final mask fully covers the film-film area.
[0053] 4. The present invention can use this method to identify the ground film in the study area when the ground film is applied, that is, during the crop sowing period, and can quickly and timely obtain the distribution of the ground film.
[0054] Example 2
[0055] Based on the remote sensing ground film extraction method based on directional target detection and prompt segmentation of the above-mentioned embodiment 1, a remote sensing ground film extraction system based on directional target detection and prompt segmentation can be formed. Its application examples and feature combinations can be found in the above-mentioned embodiment 1, which will be briefly described below.
[0056] The remote sensing ground film extraction system based on directional target detection and prompt segmentation of the second embodiment is characterized by including the following contents:
[0057] A construction module is used to construct a ground film extraction model, including the following: constructing a directional target detection model for ground film detection, and constructing a prompt segmentation model for ground film segmentation;
[0058] The extraction module is used to acquire a target image containing a ground film area, use a directional target detection model to detect the target image to obtain a directional detection frame, sample several points in the core area calibrated by the directional detection frame as the first prompt input of the prompt segmentation model, use the prompt segmentation model to predict the target and extract the first mask, use the first target mask as the second prompt input of the prompt segmentation model, and use the prompt segmentation model again to predict the target and extract the second target mask to output the segmentation result of the ground film.
[0059] Among them, the specific processing flow of the extraction module is as follows: step 201, initialize the directional target detection model and the prompt segmentation model; step 202, set the prediction parameters of the directional target detection model, perform object detection on the acquired target image and save the detection results; step 203, traverse the detection results, if a target is detected, obtain the coordinates of the directional target detection frame, convert the rotation angle of the directional target detection frame from radians to degrees, and then construct the directional target detection frame; step 204, sample several points in the core area marked by the directional target detection frame; step 205, use the several sampling points as the first prompt input of the prompt segmentation model; step 206, use the prompt segmentation model to predict the target in the area marked by the directional target detection frame; step 207, extract and obtain the first target mask; step 208, use the first target mask as the second prompt input of the prompt segmentation model, return to execute the aforementioned step 206-step 207 operations, and extract the second target mask to output as the segmentation result of the ground film; wherein the first target mask is the mask with the smallest area; the second target mask is the mask with the highest confidence.
[0060] In the construction module, the directional target detection model is the optimized and improved Mamba-YOLO model, and the hint segmentation model is the SAM model. The construction process for the optimized and improved Mamba-YOLO model is as follows: Step 101: Construct the improved Mamba-YOLO model for ground film detection; Step 102: Set the training parameters and hyperparameters of the improved Mamba-YOLO model; Step 103: Train and validate the improved Mamba-YOLO model using the training and validation sets of directional target detection from ground film remote sensing images; Step 104: Test the trained improved Mamba-YOLO model using the test set of directional target detection from ground film remote sensing images to obtain the optimized and improved Mamba-YOLO model. The construction process for the SAM model is as follows: Set the image segmentation parameters of the SAM model to obtain the SAM model for ground film segmentation from remote sensing images.
[0061] The construction process of the improved Mamba-YOLO model is as follows: Step 111, in the backbone network of the Mamba-YOLO model, introduce a large target detection layer P6 to generate a smaller feature layer; Step 112, upsample the feature layer in the head network of the Mamba-YOLO model, and then combine it with the output of the eighth layer of the backbone network to match the number of channels and size of the two; Step 113, the feature map combined in step 112 is detected and classified in the head network together with the output of other detection layers; Step 114, in the head network of the Mamba-YOLO model, replace the first and second Conv modules with receptive field attention convolution RFAConv modules respectively; Among them, the RFAConv module adopts a fusion design of deformable convolution and attention mechanism, which can effectively enhance the multi-scale feature representation capability. It is expressed as: where Δp k is the deformable offset, α is the adaptive attention weight coefficient; step 115, the Mamba-YOLO model detection head is replaced from the standard bounding box detection to the oriented bounding box detection, and the output dimension is expanded to support the direction angle prediction; step 116, the Inner-SIoU loss function is used to replace the original loss function CIoU in the Mamba-YOLO model to obtain the improved Mamba-YOLO model, where the calculation formula of the Inner-SIoU loss function is: L Inner-SIoU =L SIoU +(IoU-IoU inner ), where L SIoU is the original SIoU loss, IoU is the intersection-over-union ratio of the predicted box to the original real box, IoU inner It is the intersection-over-union ratio of the predicted box and the scaled true box.
[0062] As mentioned above, the improved directional target detection model is first used to accurately interpret the direction and boundaries of the remotely sensed ground film area. Point sampling is then performed within the core area calibrated by the directional target detection frame. The sampled points are input into the hint segmentation model as initial cues, driving the model to make a preliminary mask prediction for the target area. The mask with the smallest area is parsed and extracted as a new cue for the hint segmentation model, which is then input into the hint segmentation model for re-prediction and refinement. The mask with the highest confidence score is selected from the optimized prediction results and used as the final ground film segmentation result. This iterative process aims to leverage the macroscopic information from the previous segmentation to guide the model to more accurately fit the target boundary in the next round of prediction, thereby improving segmentation accuracy in complex scenarios and ensuring that the final mask fully covers and accurately locates the ground film area.
[0063] It should be pointed out that the examples of the above embodiments can be preferably combined with one or more of them according to actual needs, and multiple examples use a set of drawings to illustrate the combined technical features, which will not be explained one by one here.
[0064] The above description is a detailed description and illustration of the preferred embodiments of the present invention, but these descriptions are not intended to limit the scope of protection claimed by the present invention. Any equivalent changes or modifications completed under the technical teachings suggested by the present invention should fall within the scope of patent protection covered by the present invention.
Claims
1. A remote sensing ground film extraction method based on directional target detection and prompt segmentation, characterized in that: Includes the following: Step S1, constructing a ground film extraction model, including the following: constructing a directional target detection model for ground film detection, and constructing a prompt segmentation model for ground film segmentation; Step S2: Obtain a target image containing a ground film area, use a directional target detection model to detect the target image to obtain a directional detection frame of the ground film target, sample several points within the core area calibrated by the directional detection frame as the first prompt input of the prompt segmentation model, use the prompt segmentation model to predict the target and extract the first target mask, use the first target mask as the second prompt input of the prompt segmentation model, use the prompt segmentation model again to predict the target and extract the second target mask to output the segmentation result of the ground film.
2. The remote sensing ground film extraction method based on directional target detection and prompt segmentation according to claim 1 is characterized in that: The specific processing flow of step S2 is as follows: step 201, initializing the directional target detection model and the prompt segmentation model; step 202, setting the prediction parameters of the directional target detection model, performing object detection on the acquired target image and saving the detection results; Step 203: traverse the detection results. If a target is detected, obtain the coordinates of the directional target detection frame, convert the rotation angle of the directional target detection frame from radians to degrees, and then construct the directional target detection frame. Step 204: sample several points within the core area of the directional target detection frame. Step 205: use the several sampled points as the first prompt input of the prompt segmentation model. Step 206: use the prompt segmentation model to predict the target in the area marked by the directional target detection frame. Step 207: extract and obtain a first target mask; Step 208: use the first target mask as a second prompt input for prompting the segmentation model, return to execute the aforementioned steps 206-207, and extract and obtain a second target mask to output as the segmentation result of the ground film; The first target mask is the mask with the smallest area; the second target mask is the mask with the highest confidence.
3. The remote sensing ground film extraction method based on directional target detection and prompt segmentation according to claim 2 is characterized in that: In step S1, the directional target detection model is an optimized and improved Mamba-YOLO model, and the prompt segmentation model is a SAM model.
4. The remote sensing ground film extraction method based on directional target detection and prompt segmentation according to claim 3 is characterized in that: In the step S1, The optimized and improved Mamba-YOLO model construction process is as follows: Step 101, constructing an improved Mamba-YOLO model for ground film detection; Step 102, setting the training parameters and hyperparameters of the improved Mamba-YOLO model; Step 103, using the training set and validation set of directional target detection of ground film remote sensing images to train and validate the improved Mamba-YOLO model; Step 104, using the test set of directional target detection of ground film remote sensing images to test the trained improved Mamba-YOLO model, thereby obtaining the optimized and improved Mamba-YOLO model; The SAM model construction process is as follows: set the image segmentation parameters of the SAM model to obtain the SAM model for remote sensing image ground film segmentation.
5. The remote sensing ground film extraction method based on directional target detection and prompt segmentation according to claim 4 is characterized in that: The construction process of the improved Mamba-YOLO model is as follows: Step 111, in the backbone network of the Mamba-YOLO model, a large target detection layer P6 is introduced to generate a smaller feature layer; Step 112, upsampling is performed in the head network of the Mamba-YOLO model to generate a feature layer, which is then combined with the output of the eighth layer of the backbone network to match the number of channels and size; Step 113, the feature map combined in Step 112 is detected and classified in the head network together with the output of the other detection layers; Step 114, in the head network of the Mamba-YOLO model , respectively replace the first and second Conv modules with receptive field attention convolution RFAConv modules; wherein, the RFAConv module adopts a fusion design of deformable convolution and attention mechanism; step 115, replace the Mamba-YOLO model detection head from standard bounding box detection to oriented bounding box detection, and expand the output dimension to support direction angle prediction; step 116, use the Inner-SIoU loss function to replace the original loss function CIoU in the Mamba-YOLO model to obtain the improved Mamba-YOLO model, wherein the calculation formula of the Inner-SIoU loss function is: L Inner-SIoU =L SIoU +(IoU-IoU inner ), where L SIoU is the original SIoU loss, IoU is the intersection-over-union ratio of the predicted box to the original real box, IoU inner It is the intersection-over-union ratio of the predicted box and the scaled true box.
6. A remote sensing ground film extraction system based on directional target detection and prompt segmentation, characterized in that: Includes the following: A construction module is used to construct a ground film extraction model, including the following: constructing a directional target detection model for ground film detection, and constructing a prompt segmentation model for ground film segmentation; The extraction module is used to acquire a target image containing a ground film area, use a directional target detection model to detect the target image to obtain a directional detection frame, sample several points in the core area calibrated by the directional detection frame as the first prompt input of the prompt segmentation model, use the prompt segmentation model to predict the target and extract the first mask, use the first target mask as the second prompt input of the prompt segmentation model, and use the prompt segmentation model again to predict the target and extract the second target mask to output the segmentation result of the ground film.
7. The remote sensing ground film extraction system based on directional target detection and prompt segmentation according to claim 6 is characterized in that: The specific processing flow of the extraction module is as follows: Step 201, initialize the directional target detection model and the prompt segmentation model; Step 202, set the prediction parameters of the directional target detection model, perform object detection on the acquired target image and save the detection results; Step 203: traverse the detection results. If a target is detected, obtain the coordinates of the directional target detection frame, convert the rotation angle of the directional target detection frame from radians to degrees, and then construct the directional target detection frame. Step 204: sample several points within the core area of the directional target detection frame. Step 205: use the several sampled points as the first prompt input of the prompt segmentation model. Step 206: use the prompt segmentation model to predict the target in the area marked by the directional target detection frame. Step 207: extract and obtain a first target mask; Step 208: use the first target mask as a second prompt input for prompting the segmentation model, return to execute the aforementioned steps 206-207, and extract and obtain a second target mask to output as the segmentation result of the ground film; The first target mask is the mask with the smallest area; the second target mask is the mask with the highest confidence.
8. The remote sensing ground film extraction system based on directional target detection and prompt segmentation according to claim 7 is characterized in that: In the construction module, the directional target detection model is an optimized and improved Mamba-YOLO model, and the prompt segmentation model is a SAM model.
9. The remote sensing ground film extraction system based on directional target detection and prompt segmentation according to claim 8 is characterized in that: In the building blocks, The optimized and improved Mamba-YOLO model construction process is as follows: Step 101, constructing an improved Mamba-YOLO model for ground film detection; Step 102, setting the training parameters and hyperparameters of the improved Mamba-YOLO model; Step 103, using the training set and validation set of directional target detection of ground film remote sensing images to train and validate the improved Mamba-YOLO model; Step 104, using the test set of directional target detection of ground film remote sensing images to test the trained improved Mamba-YOLO model, thereby obtaining the optimized and improved Mamba-YOLO model; The SAM model construction process is as follows: set the image segmentation parameters of the SAM model to obtain the SAM model for remote sensing image ground film segmentation.
10. The remote sensing ground film extraction system based on directional target detection and prompt segmentation according to claim 9 is characterized in that: The construction process of the improved Mamba-YOLO model is as follows: Step 111, in the backbone network of the Mamba-YOLO model, a large target detection layer P6 is introduced to generate a smaller feature layer; Step 112, upsampling is performed in the head network of the Mamba-YOLO model to generate a feature layer, which is then combined with the output of the eighth layer of the backbone network to match the number of channels and size; Step 113, the feature map combined in Step 112 is detected and classified in the head network together with the output of the other detection layers; Step 114, in the head network of the Mamba-YOLO model , respectively replace the first and second Conv modules with receptive field attention convolution RFAConv modules; wherein, the RFAConv module adopts a fusion design of deformable convolution and attention mechanism; step 115, replace the Mamba-YOLO model detection head from standard bounding box detection to oriented bounding box detection, and expand the output dimension to support direction angle prediction; step 116, use the Inner-SIoU loss function to replace the original loss function CIoU in the Mamba-YOLO model to obtain the improved Mamba-YOLO model, wherein the calculation formula of the Inner-SIoU loss function is: L Inner-SIoU =L SIoU +(IoU-IoU inner ), where L SIoU is the original SIoU loss, IoU is the intersection-over-union ratio of the predicted box to the original real box, IoU inner It is the intersection-over-union ratio of the predicted box and the scaled true box.
Citation Information
Cited By
Digitized full-automatic wall painting repairing method fused with feature extractor
CN121544496A