Cross-domain cultivated land extraction method and device, equipment, storage medium and program product
By using pre-trained image segmentation models and semantic information of cultivated land areas, combined with boundary cues and buffer zone visual cue strategies, the high cost and insufficient cross-regional adaptability problems caused by the cultivated land extraction method's reliance on specific area training are solved, and high-precision, automated cultivated land identification and positioning are achieved.
Patent Information
- Application Number
- CN202510632218.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-10-17
AI Technical Summary
Existing farmland extraction methods rely on training with labeled samples from specific regions, resulting in high data acquisition and labeling costs, limited cross-regional adaptability, and poor generalization ability.
By using a pre-trained image segmentation model and semantic information of cultivated land areas, combined with boundary hinting engineering and buffer visual hinting strategies, the model is guided to focus on the edge features of cultivated land through bounding boxes and buffer zones, thereby improving the model's ability to understand contextual information.
It achieves high-precision farmland extraction and identification in different regions and complex backgrounds, enhances the generalization and robustness of the model, and reduces dependence on training data.
Smart Images

Figure CN120808133A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a cross-domain cultivated land extraction method, device, equipment, storage medium and program product. Background Art
[0002] Among related technologies, the use of remote sensing satellite imagery for farmland extraction has been widely applied in fields such as agricultural monitoring, land management, and ecological protection. By processing and analyzing remote sensing image data, this technology can automatically identify and extract farmland information over large areas, significantly improving the efficiency and accuracy of farmland identification and reducing reliance on manual on-site surveys, thereby effectively reducing labor costs and work time.
[0003] However, in related technologies, most extraction methods rely on a large number of high-quality labeled samples for training, resulting in high data acquisition and labeling costs. In addition, the models are usually trained based on data from a specific region, and their cross-regional adaptability is limited. It is difficult to maintain stable recognition performance in different geographical environments or crop distribution backgrounds, and the generalization ability is poor. These problems need to be urgently addressed. Summary of the Invention
[0004] The present application provides a cross-domain cultivated land extraction method, device, equipment, storage medium and program product to solve the technical problems in related technologies that cultivated land extraction methods are generally trained based on specific areas and a large number of labeled samples, resulting in high data acquisition and labeling costs, limited cross-regional adaptability, and poor pan-China capabilities.
[0005] The first aspect of the present application provides a cross-domain cultivated land extraction method, comprising the following steps: obtaining a remote sensing image to be processed, inputting the remote sensing image to be processed into a pre-trained image segmentation model to segment at least one target cultivated land; based on the at least one target cultivated land, determining the cultivated land area semantic information of each target cultivated land; and obtaining a binary map of the cultivated land based on the cultivated land area semantic information of each target cultivated land.
[0006] By using the above technical means, combined with the pre-trained image segmentation model and the semantic information of the cultivated land area, a binary map of the cultivated land is generated, which can be used for zero-sample cultivated land extraction, effectively reducing the dependence on training data, thereby improving the accuracy of cultivated land extraction and recognition accuracy, and enhancing the generalization ability of the model in different regions and different remote sensing images, with good adaptability and practicality.
[0007] Optionally, in one embodiment of the present application, the semantic information of the cultivated land area of each target cultivated land is determined based on the at least one target cultivated land, including: creating a buffer zone centered on the cultivated land area of each target cultivated land according to the resolution of the image and the size of the target cultivated land; and inputting the image area within the buffer zone as a visual cue into a pre-trained contrastive language-image pre-training model to output the semantic information of the cultivated land area.
[0008] Through the above technical means, a buffer zone is created according to the resolution of the image and the size of the target cultivated land, and is input as a visual cue into the pre-trained contrast language-image pre-training model. This can guide the model to focus on a certain cultivated land area and suppress background interference information, which helps the model to more effectively extract semantic features related to cultivated land, enabling the pre-trained model to more fully understand the semantic relationship of the image context, so that it can still have good discrimination ability in complex terrain, different landforms or low-contrast images, further improving the accuracy, stability and generalization ability of cultivated land recognition.
[0009] Optionally, in one embodiment of the present application, the image area in the buffer zone is input as a visual cue into a pre-trained contrastive language-image pre-training model to output the cultivated land area semantic information, including: obtaining a pre-constructed text description related to cultivated land; inputting the image area in the buffer zone and the text description into the image encoder and text encoder of the contrastive language-image pre-training model to extract image features and text features; calculating the similarity between the image features and the text features; and judging whether it is cultivated land based on the similarity to obtain the cultivated land area semantic information of the cultivated land.
[0010] Through the above technical means, whether it is cultivated land is judged based on the similarity between image features and text features, and the semantic information of the cultivated land area is obtained. This can more effectively mine the semantic content related to cultivated land in the image, improve the ability to understand and distinguish cultivated land targets, and thus achieve high-precision cultivated land extraction and identification under complex backgrounds, lighting changes or diverse landform conditions, significantly improving the robustness and generalization ability of the system.
[0011] Optionally, in one embodiment of the present application, the remote sensing image to be processed is obtained, and the remote sensing image to be processed is input into a pre-trained image segmentation model to segment at least one target cultivated land, including: obtaining historical vector cultivated land boundary data; converting the historical vector cultivated land boundary data into a bounding box prompt signal; inputting the bounding box prompt signal into a prompt encoder of the image segmentation model to obtain the segmented cultivated land target area from the mask decoder to determine the at least one target cultivated land.
[0012] Through the above technical means, the bounding box prompt signal is input into the prompt encoder of the image segmentation model, and the segmented cultivated land target area is obtained from the mask decoder. This can effectively guide the model to focus on potential cultivated land areas, enhance the model's perception of target boundaries, and reduce the interference of background noise, thereby improving the segmentation accuracy and boundary precision of cultivated land targets. It helps to enhance the model's positioning sensitivity to target areas, thereby having the ability to locate zero-sample objects. Without specific sample training, it can achieve accurate recognition and positioning of cultivated land targets, further improving the intelligence level of cultivated land extraction and the generalization performance of the model.
[0013] Optionally, in one embodiment of the present application, converting the historical vector cultivated land boundary data into a boundary box prompt signal includes: defining the minimum circumscribed rectangle of the cultivated land area as a basis, and adjusting the size and position of the rectangle until a preset prompt condition is met to generate the boundary box prompt signal.
[0014] Through the above technical means, the size and position of the minimum enclosing rectangle of the cultivated land area are adjusted to generate a bounding box signal, which can more accurately reflect the spatial distribution characteristics of the cultivated land area, provide effective positioning guidance for subsequent models, guide the model to focus on key areas, reduce attention to non-target areas, improve the positioning accuracy of the target area, enhance the model's ability to perceive the edges and shapes of cultivated land, and bring higher-quality image segmentation results, thereby realizing accurate extraction and identification of cultivated land under complex background conditions.
[0015] The second aspect of the present application provides a cross-domain cultivated land extraction device, including: a segmentation module, used to obtain a remote sensing image to be processed, and input the remote sensing image to be processed into a pre-trained image segmentation model to segment at least one target cultivated land; a determination module, used to determine the cultivated land area semantic information of each target cultivated land based on the at least one target cultivated land; an extraction module, used to obtain a binary map of the cultivated land based on the cultivated land area semantic information of each target cultivated land.
[0016] By using the above technical means, combined with the pre-trained image segmentation model and the semantic information of the cultivated land area, a binary map of the cultivated land is generated, which can be used for zero-sample cultivated land extraction, effectively reducing the dependence on training data, thereby improving the accuracy of cultivated land extraction and recognition accuracy, and enhancing the generalization ability of the model in different regions and different remote sensing images, with good adaptability and practicality.
[0017] Optionally, in one embodiment of the present application, the determination module includes: a creation unit: used to create a buffer zone centered on the cultivated land area of each target cultivated land according to the resolution of the image and the size of the target cultivated land; an output unit, used to input the image area in the buffer zone as a visual cue into a pre-trained contrast language-image pre-training model to output the semantic information of the cultivated land area.
[0018] Through the above technical means, a buffer zone is created according to the resolution of the image and the size of the target cultivated land, and is input as a visual cue into the pre-trained contrast language-image pre-training model. This can guide the model to focus on a certain cultivated land area and suppress background interference information, which helps the model to more effectively extract semantic features related to cultivated land, enabling the pre-trained model to more fully understand the semantic relationship of the image context, so that it can still have good discrimination ability in complex terrain, different landforms or low-contrast images, further improving the accuracy, stability and generalization ability of cultivated land recognition.
[0019] Optionally, in one embodiment of the present application, the output unit includes: an acquisition subunit for acquiring a pre-constructed text description related to cultivated land; an extraction subunit for inputting the image area in the buffer and the text description into the image encoder and text encoder of the contrastive language-image pre-training model to extract image features and text features; a calculation subunit for calculating the similarity between the image features and the text features; and a judgment subunit for judging whether it is cultivated land based on the similarity to obtain the cultivated land area semantic information of the cultivated land.
[0020] Through the above technical means, whether it is cultivated land is judged based on the similarity between image features and text features, and the semantic information of the cultivated land area is obtained. This can more effectively mine the semantic content related to cultivated land in the image, improve the ability to understand and distinguish cultivated land targets, and thus achieve high-precision cultivated land extraction and identification under complex backgrounds, lighting changes or diverse landform conditions, significantly improving the robustness and generalization ability of the system.
[0021] Optionally, in one embodiment of the present application, the segmentation module includes: an acquisition unit for acquiring historical vector cultivated land boundary data; a conversion unit for converting the historical vector cultivated land boundary data into a bounding box prompt signal; and a determination unit for inputting the bounding box prompt signal into a prompt encoder of an image segmentation model to obtain the segmented cultivated land target area from a mask decoder and determine the at least one target cultivated land.
[0022] Through the above technical means, the bounding box prompt signal is input into the prompt encoder of the image segmentation model, and the segmented cultivated land target area is obtained from the mask decoder. This can effectively guide the model to focus on potential cultivated land areas, enhance the model's perception of target boundaries, and reduce the interference of background noise, thereby improving the segmentation accuracy and boundary precision of cultivated land targets. It helps to enhance the model's positioning sensitivity to target areas, thereby having the ability to locate zero-sample objects. Without specific sample training, it can achieve accurate recognition and positioning of cultivated land targets, further improving the intelligence level of cultivated land extraction and the generalization performance of the model.
[0023] Optionally, in one embodiment of the present application, the conversion unit includes: a generation subunit, which is used to define the minimum circumscribed rectangle of the cultivated land area as a basis, and to generate the bounding box prompt signal by adjusting the size and position of the rectangle until a preset prompt condition is reached.
[0024] Through the above technical means, the size and position of the minimum enclosing rectangle of the cultivated land area are adjusted to generate a bounding box signal, which can more accurately reflect the spatial distribution characteristics of the cultivated land area, provide effective positioning guidance for subsequent models, guide the model to focus on key areas, reduce attention to non-target areas, improve the positioning accuracy of the target area, enhance the model's ability to perceive the edges and shapes of cultivated land, and bring higher-quality image segmentation results, thereby realizing accurate extraction and identification of cultivated land under complex background conditions.
[0025] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the cross-domain cultivated land extraction method as described in the above embodiment.
[0026] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above cross-domain farmland extraction method.
[0027] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above cross-domain cultivated land extraction method.
[0028] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0030] Figure 1 A flowchart of a cross-domain farmland extraction method provided according to an embodiment of the present application;
[0031] Figure 2 This is a schematic diagram of the model architecture of the cross-domain farmland extraction method according to one embodiment of the present application;
[0032] Figure 3 Schematic diagram of a cross-domain farmland extraction device provided according to an embodiment of the present application;
[0033] Figure 4The figure is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application.
[0034] Reference numerals:
[0035] 10-cross-domain cultivated land extraction device; 100-segmentation module, 200-determination module and 300-extraction module; 401-memory, 402-processor and 403-communication interface. DETAILED DESCRIPTION
[0036] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0037] The following describes the cross-domain farmland extraction method, device, electronic device and storage medium of the embodiment of the present application with reference to the accompanying drawings. In view of the fact that the cross-domain farmland extraction technology mentioned in the above background technology is generally based on a specific area and a large number of labeled samples for training, resulting in high data acquisition and labeling costs, as well as limited cross-regional adaptability and poor pan-China capabilities, the present application provides a cross-domain farmland extraction method, in which, based on a pre-trained image segmentation model and farmland area semantic information, a boundary prompting project and a buffer zone visual prompting strategy are introduced to extract a binary map of the farmland, which can guide the model to focus on the edge features of the farmland, improve the accuracy of boundary recognition, and effectively enhance the model's ability to understand contextual information, thereby improving the spatial accuracy and boundary clarity of farmland recognition, enhancing the model's adaptability in complex backgrounds, bringing higher recognition robustness and generalization performance, and achieving high-precision, automated positioning and recognition of farmland. Thus, the technical problem that the farmland extraction method in the related art is generally based on a specific area and a large number of labeled samples for training, resulting in high data acquisition and labeling costs, as well as limited cross-regional adaptability and poor pan-China capabilities is solved.
[0038] Specifically, Figure 1 A flow chart of a cross-domain cultivated land extraction method provided in an embodiment of the present application.
[0039] like Figure 1 As shown, the cross-domain cultivated land extraction method includes the following steps:
[0040] In step S101 , a remote sensing image to be processed is obtained, and the remote sensing image to be processed is input into a pre-trained image segmentation model to segment at least one target cultivated land.
[0041] Remote sensing imagery refers to image data of the Earth's surface and its environment acquired using remote sensing technology. It is widely used in areas such as land cover classification, target identification, change detection, resource monitoring, and environmental assessment. Remote sensing imagery can include, but is not limited to, visible light imagery, multispectral imagery, and hyperspectral imagery. The imagery comes from various satellite platforms, including the domestically produced Gaofen series of satellites, the Sentinel series of satellites, and the Landsat (Land Satellite) series of satellites.
[0042] In an embodiment of the present application, the pre-trained image segmentation model may be SAM (Segment Anything Model), which is a general image segmentation model with powerful zero-sample segmentation capabilities and can perform high-precision and flexible segmentation of any target in any image. SAM is pre-trained based on a large-scale image dataset and adopts a "prompt-driven" segmentation mechanism. It can automatically segment the target area based on the input point, box or text prompts without the need for retraining for specific tasks. SAM has powerful image segmentation and generalization capabilities and can handle remote sensing image segmentation tasks of various types of land objects.
[0043] Optionally, in one embodiment of the present application, a remote sensing image to be processed is obtained and input into a pre-trained image segmentation model to segment at least one target cultivated land, including: obtaining historical vector cultivated land boundary data; converting the historical vector cultivated land boundary data into a bounding box prompt signal; inputting the bounding box prompt signal into a prompt encoder of the image segmentation model to obtain the segmented cultivated land target area from the mask decoder to determine at least one target cultivated land.
[0044] Among them, boundary data is vector or raster data used to accurately describe the contour and edge position of the target object. The acquisition of boundary data can be achieved by manual visual interpretation, classification algorithm or deep learning methods.
[0045] It should be noted that the bounding box prompt signal generally refers to the use of spatial information of a rectangular box as a prompt to tell the model the area that needs to be focused on, so as to prevent the model from blindly searching the entire image.
[0046] A hint encoder generally refers to a module that encodes hint signals (such as bounding boxes, points, masks, etc.) into high-dimensional feature vectors. It can integrate the hint information obtained a priori into the model, enabling the model to better understand and focus on key information areas.
[0047] The mask decoder is a module that combines the encoded image features and hint features (such as bounding box hints, point hints, etc.) to output a fine mask, which can improve boundary precision and overall extraction accuracy.
[0048] Specifically, the embodiment of the present application extracts potential cultivated land boundary information from remote sensing images and converts it into prompt signals that can be understood by SAM, which can guide SAM to accurately focus on cultivated land areas, achieve zero-sample object positioning of cultivated land targets, and further complete high-precision segmentation processing.
[0049] Optionally, in one embodiment of the present application, historical vector cultivated land boundary data is converted into a boundary box prompt signal, including: based on the minimum circumscribed rectangle defining the cultivated land area, adjusting the size and position of the rectangle until a preset prompt condition is met to generate a boundary box prompt signal.
[0050] The preset prompt condition can include that the proportion of the adjusted rectangle covering the cultivated land area reaches a certain threshold, ensuring that the prompt box fully covers the target area; or that the position offset between the center of the rectangle and the centroid (center point) of the cultivated land is less than a certain distance, ensuring accurate positioning of the bounding box. Preset prompt conditions can be set by those skilled in the art based on actual circumstances and are not specifically limited here.
[0051] It should be noted that the size and position of the rectangles can be optimized based on the resolution of the remote sensing image, the typical size of the cultivated land target, and the segmentation characteristics of the SAM. By properly setting the size and layout of the rectangles, SAM can improve its focus on cultivated land boundary areas and recognition accuracy, further enhancing the accuracy and completeness of the segmentation mask.
[0052] In step S102, based on at least one target farmland, farmland region semantic information of each target farmland is determined.
[0053] Among them, semantic information refers to descriptive data with specific meaning extracted from the target area (such as cultivated land area) based on high-level understanding, which may include but is not limited to category information (such as cultivated land categories such as rice fields, corn fields, wheat fields, fallow land, etc.), attribute information (such as cultivated land area, boundary contours, shape characteristics, spatial distribution position, adjacent land object relationships, etc.) and status information (such as crop growth stage, vegetation coverage, tillage conditions, irrigation status, soil moisture, etc.).
[0054] With the help of semantic information, the embodiments of the present application can better distinguish cultivated land from other similar areas (such as bare land, grassland, water bodies, etc.), improve the accuracy and robustness of cultivated land extraction, and significantly improve the reliability of cultivated land identification and classification, especially in complex backgrounds or with diverse crop types.
[0055] Optionally, in one embodiment of the present application, based on at least one target cultivated land, the semantic information of the cultivated land area of each target cultivated land is determined, including: creating a buffer zone centered on the cultivated land area of each target cultivated land according to the resolution of the image and the size of the target cultivated land; and inputting the image area within the buffer zone as a visual cue into a pre-trained contrast language-image pre-training model to output the semantic information of the cultivated land area.
[0056] A buffer zone is an area around a spatial target that expands or contracts at a set distance. A buffer zone can be used to add a certain width to the original target boundary to account for factors such as position error, extraction tolerance, and contextual information.
[0057] It should be noted that the CLIP (Contrastive Language-Image Pre-Training) model provided in the embodiment of the present application can learn a unified multimodal feature space by comparing images with corresponding text descriptions, so that the input image and text can be naturally aligned in the feature space. The embodiment of the present application can utilize the characteristics of the CLIP model to encode semantic information such as cultivated land boundaries, crop types, and plot status into text prompt signals, and compare them with cultivated land image features to improve the accuracy of cultivated land extraction.
[0058] Furthermore, based on the preliminary farmland mask generated by SAM, the embodiment of the present application creates a buffer zone centered on the farmland area, and inputs the image area within the buffer zone into the CLIP model as a visual cue to enhance the CLIP model's understanding of the contextual information of the farmland scene and suppress its misjudgment of areas outside the farmland.
[0059] The size of the buffer zone (radius r) can be adaptively adjusted according to the resolution of the image and the size of the target cultivated land.
[0060] Specifically, the adaptive adjustment strategy of the buffer radius r may include but is not limited to: determining the minimum radius of the buffer based on the resolution of the remote sensing image; dynamically adjusting the radius of the buffer based on the statistical distribution of the target cultivated land area so that the buffer can cover the main characteristic areas of the cultivated land; verifying the classification accuracy under different buffer radii through experiments, and selecting the radius value with the highest classification accuracy as the final buffer radius.
[0061] Optionally, in one embodiment of the present application, the image area in the buffer zone is input as a visual cue into a pre-trained contrastive language-image pre-training model to output semantic information of the cultivated land area, including: obtaining a pre-constructed text description related to cultivated land; inputting the image area and text description in the buffer zone into the image encoder and text encoder of the contrastive language-image pre-training model to extract image features and text features; calculating the similarity between the image features and the text features; judging whether it is cultivated land based on the similarity to obtain the cultivated land area semantic information of the cultivated land.
[0062] For example, the text description can be refined from multiple perspectives such as crop type, farming status, spatial form, natural conditions, etc., and can be a description such as "a regularly distributed rice farmland" to more accurately guide the model to understand and extract the target cultivated land area.
[0063] Furthermore, the embodiment of the present application inputs the image area and text description of the buffer into the image encoder and text encoder of the CLIP model respectively to extract image features and text features; calculates the similarity between the image features and the text features; and determines whether the area is cultivated land based on the similarity value. When the similarity exceeds a preset threshold, the area is determined to be cultivated land.
[0064] In step S103, a binary map of the cultivated land is obtained according to the semantic information of the cultivated land area of each target cultivated land.
[0065] Among them, a binary image means that each pixel in the image has only two possible values (usually 0 and 1), representing two different categories or states.
[0066] In one embodiment of the present application, a binary image can be used to represent the segmentation results of the target area and the background area, where: 1 (white) can represent the target area (such as corn fields, wheat fields, etc.); 0 (black) can represent the non-target area (such as bare land, water bodies, woodlands, etc.).
[0067] like Figure 2 As shown in FIG, a model architecture of the cross-domain farmland extraction method is described in detail using a specific embodiment. The model architecture mainly includes the following parts:
[0068] a. Design prompt word engineering.
[0069] In the embodiment of the present application, the pre-trained SAM is first loaded. As a pre-trained model, SAM has powerful image segmentation capabilities and can handle various types of ground objects.
[0070] In order to enable SAM to automatically locate the cultivated land area without manual interaction, the embodiment of the present application designs a boundary prompt project to convert the historical vector boundary information into a bounding box prompt signal that SAM can understand.
[0071] b. Region of interest segmentation.
[0072] Furthermore, when SAM segments an image, it will pay more attention to the above-mentioned boundary areas, making it easier to segment the farmland targets.
[0073] Specifically, the embodiment of the present application inputs the high-resolution remote sensing image to be processed and the designed bounding box prompt into the prompt encoder of SAM, utilizes SAM's powerful ability to represent and segment everything, and finally obtains the target cultivated land segmented by SAM from the mask decoder.
[0074] However, since SAM itself performs segmentation based only on visual features, the farmland region lacks actual semantics.
[0075] c. Determine the category.
[0076] Furthermore, in order to provide the semantic information of the cultivated land area obtained above, that is, whether it is cultivated land or non-cultivated land, the embodiment of the present application designs a buffer zone visual prompt.
[0077] Specifically, based on the preliminary farmland mask generated by SAM, a buffer zone is created with the farmland area as the center, and the image area within the buffer zone is input into the CLIP model as a visual cue to more accurately determine whether the area is farmland.
[0078] d. Cultivated land extraction output.
[0079] Furthermore, the embodiment of the present application can use the powerful zero-sample classification capability of the CLIP model to determine whether the segmented area is farmland.
[0080] Specifically, the CLIP model can map image information and text information into the same semantic space. Therefore, even if the model has not seen any training samples of cultivated land, it can determine whether the area is cultivated land by comparing the similarity between the image and the text, and finally output a binary image.
[0081] The embodiment of the present application combines boundary hint engineering with buffer zone visual hint strategy. Boundary hint engineering can guide SAM to accurately locate cultivated land areas, and buffer zone visual hint can enhance CLIP's understanding of the scene. The synergistic effect of the two can significantly improve the overall performance of the algorithm.
[0082] According to the cross-domain cultivated land extraction method proposed in the embodiment of the present application, based on the pre-trained image segmentation model and the semantic information of the cultivated land area, the boundary prompt project and the buffer zone visual prompt strategy are introduced to extract the binary image of the cultivated land. This can guide the model to focus on the edge features of the cultivated land, improve the accuracy of boundary recognition, and effectively enhance the model's ability to understand contextual information, thereby improving the spatial accuracy and boundary clarity of cultivated land recognition, enhancing the model's adaptability to complex backgrounds, bringing higher recognition robustness and generalization performance, and realizing high-precision, automated positioning and recognition of cultivated land.
[0083] Next, the cross-domain farmland extraction device proposed according to the embodiment of the present application will be described with reference to the accompanying drawings.
[0084] Figure 3 It is a block diagram of the cross-domain farmland extraction device of an embodiment of the present application.
[0085] like Figure 3 As shown, the cross-domain farmland extraction device 10 includes: a segmentation module 100, a determination module 200 and an extraction module 300.
[0086] The segmentation module 100 is used to obtain a remote sensing image to be processed and input the remote sensing image to be processed into a pre-trained image segmentation model to segment at least one target cultivated land.
[0087] The determination module 200 is configured to determine the farmland region semantic information of each target farmland based on at least one target farmland.
[0088] The extraction module 300 is used to obtain a binary map of the cultivated land according to the semantic information of the cultivated land area of each target cultivated land.
[0089] Optionally, in one embodiment of the present application, the determination module 200 includes: a creation unit and an output segment unit.
[0090] Among them, the creation unit is used to create a buffer zone centered on the cultivated land area of each target cultivated land according to the resolution of the image and the size of the target cultivated land.
[0091] The output unit is used to input the image area in the buffer as a visual cue into a pre-trained contrastive language-image pre-training model to output the semantic information of the cultivated land area.
[0092] Optionally, in one embodiment of the present application, the output unit includes: an acquisition subunit, an extraction subunit and a calculation subunit.
[0093] The acquisition subunit is used to obtain pre-built text descriptions related to cultivated land.
[0094] The extraction subunit is used to input the image area and text description in the buffer into the image encoder and text encoder of the contrastive language-image pre-training model to extract image features and text features.
[0095] The calculation subunit is used to calculate the similarity between the image features and the text features; the judgment subunit is used to judge whether it is cultivated land based on the similarity, so as to obtain the cultivated land area semantic information of the cultivated land.
[0096] Optionally, in one embodiment of the present application, the segmentation module 100 includes: an acquisition unit, a conversion unit, and a determination unit.
[0097] Among them, the acquisition unit is used to obtain historical vector cultivated land boundary data.
[0098] The conversion unit is used to convert the historical vector cultivated land boundary data into a bounding box prompt signal.
[0099] The determination unit is configured to input the bounding box prompt signal into a prompt encoder of the image segmentation model to obtain the segmented farmland target area from the mask decoder and determine at least one target farmland.
[0100] Optionally, in one embodiment of the present application, the conversion unit includes: a generation subunit
[0101] Among them, the generation subunit is used to define the minimum circumscribed rectangle of the cultivated land area as a basis, and to generate a boundary box prompt signal by adjusting the size and position of the rectangle until the preset prompt condition is met.
[0102] It should be noted that the above explanation of the embodiment of the cross-domain farmland extraction method is also applicable to the cross-domain farmland extraction device of this embodiment, and will not be repeated here.
[0103] According to the cross-domain cultivated land extraction device proposed in the embodiment of the present application, based on the pre-trained image segmentation model and the semantic information of the cultivated land area, the boundary prompt project and the buffer zone visual prompt strategy are introduced to extract the binary image of the cultivated land. This can guide the model to focus on the edge features of the cultivated land, improve the accuracy of boundary recognition, and effectively enhance the model's ability to understand contextual information, thereby improving the spatial accuracy and boundary clarity of cultivated land recognition, enhancing the model's adaptability to complex backgrounds, bringing higher recognition robustness and generalization performance, and realizing high-precision, automated positioning and recognition of cultivated land.
[0104] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0105] Memory 401 , processor 402 , and computer programs stored in the memory 401 and executable on the processor 402 .
[0106] When the processor 402 executes the program, the cross-domain farmland extraction method provided in the above embodiment is implemented.
[0107] Furthermore, the electronic device further includes:
[0108] The communication interface 403 is used for communication between the memory 401 and the processor 402 .
[0109] The memory 401 is used to store computer programs that can be run on the processor 402 .
[0110] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0111] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0112] Optionally, in a specific implementation, if the memory 401 , the processor 402 and the communication interface 403 are integrated on a chip, the memory 401 , the processor 402 and the communication interface 403 can communicate with each other through an internal interface.
[0113] The processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0114] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which implements the above cross-domain cultivated land extraction method when executed by a processor.
[0115] An embodiment of the present application also provides a computer program product, including a computer program, which can run computer instructions. When the computer instructions are executed by a processor, the cross-domain cultivated land extraction method provided in the embodiment of the present application is implemented.
[0116] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0117] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0118] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0119] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0120] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0121] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0122] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0123] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A cross-domain cultivated land extraction method, characterized in that: The following steps are involved: Acquiring a remote sensing image to be processed, and inputting the remote sensing image to be processed into a pre-trained image segmentation model to segment at least one target cultivated land; Based on the at least one target cultivated land, determining cultivated land area semantic information of each target cultivated land; A binary map of the cultivated land is obtained according to the cultivated land area semantic information of each target cultivated land.
2. The method according to claim 1, characterized in that The determining of the cultivated land area semantic information of each target cultivated land based on the at least one target cultivated land includes: Taking the cultivated land area of each target cultivated land as the center, a buffer zone is created according to the resolution of the image and the size of the target cultivated land; The image region in the buffer is input as a visual cue into a pre-trained contrastive language-image pre-training model to output semantic information of the cultivated land region.
3. The method according to claim 2, characterized in that The step of inputting the image region in the buffer as a visual cue into a pre-trained contrastive language-image pre-training model to output semantic information of the cultivated land region includes: Get pre-built text descriptions related to farmland; Inputting the image area and the text description in the buffer into the image encoder and the text encoder of the comparative language-image pre-training model to extract image features and text features; Calculating the similarity between the image feature and the text feature; Whether it is cultivated land is determined according to the similarity, so as to obtain cultivated land area semantic information of the cultivated land.
4. The method according to claim 1, wherein The step of obtaining a remote sensing image to be processed and inputting the remote sensing image to be processed into a pre-trained image segmentation model to segment at least one target cultivated land includes: Obtain historical vector farmland boundary data; Converting the historical vector farmland boundary data into a boundary box prompt signal; The bounding box hint signal is input into a hint encoder of an image segmentation model to obtain a segmented farmland target area from a mask decoder to determine the at least one target farmland.
5. The method according to claim 4, characterized in that The converting the historical vector farmland boundary data into a boundary box prompt signal comprises: Based on the minimum circumscribed rectangle defining the cultivated land area, the size and position of the rectangle are adjusted until a preset prompt condition is met, so as to generate the boundary box prompt signal.
6. A cross-domain farmland extraction device, characterized in that: include: a segmentation module, configured to obtain a remote sensing image to be processed, and input the remote sensing image to be processed into a pre-trained image segmentation model to segment at least one target cultivated land; a determination module, configured to determine the cultivated land area semantic information of each target cultivated land based on the at least one target cultivated land; The extraction module is used to obtain a binary map of the cultivated land according to the semantic information of the cultivated land area of each target cultivated land.
7. The device according to claim 6, characterized in that The determining module includes: Creation unit: used for creating a buffer zone centered on the cultivated land area of each target cultivated land according to the resolution of the image and the size of the target cultivated land; An output unit is used to input the image area in the buffer as a visual cue into a pre-trained contrastive language-image pre-training model to output the semantic information of the cultivated land area.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the cross-domain farmland extraction method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the cross-domain cultivated land extraction method according to any one of claims 1 to 5.
10. A computer program product comprising a computer program, characterized in that The computer program is executed to implement the cross-domain cultivated land extraction method according to any one of claims 1 to 5.
Citation Information
Cited By
Method and system for intelligently identifying ground feature elements in irrigation area
CN121170658A