Sea-land image segmentation method and device, electronic equipment and computer readable medium

By registering and fusing multimodal remote sensing data and processing multimodal interactive reasoning, a set of conditional feature information is generated, which solves the problem of insufficient precision in land and sea image segmentation in existing technologies and achieves accurate segmentation of complex land features.

CN121884343APending Publication Date: 2026-04-17HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2025-12-10
Publication Date
2026-04-17

Smart Images

  • Figure CN121884343A_ABST
    Figure CN121884343A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a sea-land image segmentation method and device, electronic equipment and a computer readable medium. One specific embodiment of the method comprises the following steps: acquiring multi-modal remote sensing data; performing registration fusion processing on the multi-modal remote sensing data to obtain registered multi-modal remote sensing data; inputting the registered multi-modal remote sensing data into a visual encoder to obtain a visual feature map; based on the registration multi-modal remote sensing data, generating optimal prompt text information; performing microscopic feature extraction processing on the optimal prompt text information to obtain global text feature information and a word-level text feature information set; performing multi-modal interactive reasoning processing on the visual feature map, the global text feature information and the word-level text feature information set to generate a conditional feature information set; and inputting the conditional feature information set into a pre-trained sea-land fine segmentation model to obtain a sea-land segmentation image. According to the embodiment, the fine segmentation capability of sea-land image segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to land and sea image segmentation methods, apparatuses, electronic devices, and computer-readable media. Background Technology

[0002] With the increasing demands for marine resource management, environmental monitoring, land surveying, and military applications, higher requirements are being placed on the automated processing of remote sensing images. Land-sea image segmentation is a fundamental technology for accurately dividing ocean and land regions within remote sensing images. Currently, the common method for segmenting land-sea images is a pixel labeling method based on thresholds set according to grayscale values ​​or color features, dividing pixels into ocean or land. For example, the grayscale difference between water and land is determined by statistical histogram distribution, and then the image is segmented using a fixed threshold.

[0003] However, when segmenting land and sea images using the above method, the following technical problems often arise: Pixel labeling methods based on grayscale values ​​or color features can only perform basic land-sea dichotomy, simply dividing the image into ocean and land. When dealing with land areas with complex spatial structures and rich semantic information, such as roads, bridges, ports, and agricultural land, they lack an effective mechanism to distinguish between features with similar morphological characteristics but different functional attributes (such as ports and ordinary embankments, both of which are artificially constructed gray linear structures near water). It is difficult to accurately define the sub-types and spatial boundaries of agricultural land such as farmland, orchards, and greenhouses, resulting in poor fine segmentation capabilities for land-sea image segmentation.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide methods, apparatuses, electronic devices, and computer-readable media for land and sea image segmentation to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a land-sea image segmentation method, which includes: acquiring multimodal remote sensing data; performing registration and fusion processing on the multimodal remote sensing data to obtain registered multimodal remote sensing data; inputting the registered multimodal remote sensing data into a pre-trained visual encoder to obtain a visual feature map; generating optimal prompt text information based on the registered multimodal remote sensing data; performing micro-feature extraction processing on the optimal prompt text information to obtain a set of global text feature information and a set of word-level text feature information; performing multimodal interactive reasoning processing on the visual feature map, the global text feature information, and the set of word-level text feature information to generate a set of conditional feature information; and inputting the set of conditional feature information into a pre-trained fine-grained land-sea segmentation model to obtain a land-sea segmented image.

[0008] Secondly, some embodiments of this disclosure provide a land-sea image segmentation apparatus, comprising: an acquisition unit configured to acquire multimodal remote sensing data; a registration and fusion processing unit configured to perform registration and fusion processing on the multimodal remote sensing data to obtain registered multimodal remote sensing data; a first input unit configured to input the registered multimodal remote sensing data into a pre-trained visual encoder to obtain a visual feature map; a generation unit configured to generate optimal prompt text information based on the registered multimodal remote sensing data; a micro-feature extraction processing unit configured to perform micro-feature extraction processing on the optimal prompt text information to obtain a set of global text feature information and a set of word-level text feature information; a multimodal interactive reasoning processing unit configured to perform multimodal interactive reasoning processing on the visual feature map, the global text feature information, and the set of word-level text feature information to generate a set of conditional feature information; and a second input unit configured to input the set of conditional feature information into a pre-trained land-sea fine segmentation model to obtain a land-sea segmented image.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above-described embodiments of this disclosure have the following beneficial effects: the land-sea image segmentation method of some embodiments of this disclosure improves the fine segmentation capability of land-sea image segmentation. Specifically, the reason for the poor fine segmentation capability of land-sea image segmentation is that the pixel labeling method based on grayscale values ​​or color features to set thresholds can only complete basic land-sea dichotomy, simply dividing the image into ocean and land. When facing land areas with complex spatial structures and rich semantic information, such as roads, bridges, ports, and agricultural land, it lacks an effective mechanism to distinguish between land features with similar morphological features but different functional attributes (such as ports and ordinary embankments, both of which are artificially constructed gray linear structures near water). It is difficult to accurately define the subdivision types and spatial boundaries of agricultural land such as farmland, orchards, and greenhouses, resulting in poor fine segmentation capability of land-sea image segmentation. Based on this, the land-sea image segmentation method of some embodiments of this disclosure first acquires multimodal remote sensing data. Then, the multimodal remote sensing data is registered and fused to obtain registered multimodal remote sensing data. This allows for the registration of the spatial positions of pixels in each multimodal remote sensing data set, reducing positional bias and resulting in registered multimodal remote sensing data. Next, this registered multimodal remote sensing data is input into a pre-trained visual encoder to obtain a visual feature map. This yields a visual feature map that retains rich spatial details and visual information. Then, based on the registered multimodal remote sensing data, optimal prompt text information is generated. This allows for the generation of corresponding optimal prompt text information based on the registered multimodal remote sensing data. Next, micro-feature extraction processing is performed on the optimal prompt text information to obtain global text feature information and word-level text feature information sets. This allows for the extraction of global and specific semantic concepts at a micro-level to obtain global and word-level text feature information sets for multimodal interactive reasoning processing. Finally, multimodal interactive reasoning processing is performed on the visual feature map, the global text feature information, and the word-level text feature information set to generate a conditional feature information set. Therefore, the visual modality of the visual feature map can interact and reason with the linguistic modality of the global text feature information and the aforementioned word-level text feature information set, generating a conditional feature information set containing rich semantic information and enhanced visual information. This lays the foundation for achieving fine-grained land feature classification and boundary delineation. Finally, the aforementioned conditional feature information set is input into a pre-trained fine-grained land-sea segmentation model to obtain a land-sea segmentation image. Thus, a land-sea segmentation image capable of segmenting complex land features can be obtained.Because it first generates optimal prompt text information corresponding to the scene of the registered multimodal remote sensing data based on the registered multimodal remote sensing data, and then generates global text feature information and word-level text feature information sets based on the optimal prompt text information. The global text feature information guides the model to focus on global semantic features, and the word-level text feature information sets guide the model to focus on specific and refined semantic features. In this way, it can guide the model to distinguish visually similar but semantically different land features. Therefore, it can improve the differentiation of land features with similar morphological features but different functional attributes (such as port docks and ordinary embankments), thereby achieving fine segmentation of land and sea images and improving the fine segmentation capability of land and sea image segmentation. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the land-sea image segmentation method according to the present disclosure; Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the land and sea image segmentation apparatus according to this disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0016] Figure 1 A flowchart 100 of some embodiments of the land-sea image segmentation method according to the present disclosure is shown. The land-sea image segmentation method includes the following steps: Step 101: Acquire multimodal remote sensing data.

[0017] In some embodiments, the entity executing the land-sea image segmentation method (e.g., a computing device) can acquire optical satellite imagery data, synthetic aperture radar (SAR) data, and digital elevation model (DEM) data through a data platform (e.g., USGS EarthExplorer, Copernicus Open Access Hub, USGS EarthExplorer platform, etc.). Then, the aforementioned optical satellite imagery data, SAR data, and DEM data are identified as multimodal remote sensing data. The optical satellite imagery data can be a digital image model representing the Earth's surface reflection characteristics of solar electromagnetic waves. The optical satellite imagery data includes an image map, a band DN value information set, a scene classification band information set, and a geographic coordinate system for the optical satellite imagery data. Each pixel in the image map corresponds to a band DN value in the band DN value information set, and each pixel in the image map corresponds to a scene classification band in the scene classification band information set. The aforementioned band DN value information includes: aerosol band DN values, blue band DN values, green band DN values, red band DN values, red-edge 1 band DN values, red-edge 2 band DN values, red-edge 3 band DN values, near-infrared band DN values, narrow near-infrared band DN values, water vapor band DN values, shortwave infrared band DN values, shortwave infrared 1 band DN values, and shortwave infrared 2 outer band DN values. Each scene classification band information in the aforementioned scene classification band information set includes values ​​mapped according to a preset mapping rule. The preset mapping rule can be 0-no data, 1-saturation defect, 2-dark pixel, 3-cloud shadow, 4-vegetation, 5-bare ground, 6-water body, 7-snow / ice, 8-cloud, 9-high cloud, 10-thin cloud. The aforementioned image can be an image generated from image data of the Earth's surface observed and recorded by satellites in Earth orbit using various sensors. The geographic coordinate system of the aforementioned optical satellite image data can be the geographic reference system used for the optical satellite image data. The aforementioned synthetic aperture radar (SAR) data may include radar imagery, individual SAR sub-data sets, and a geographic coordinate system for the SAR data. Each pixel in the radar imagery corresponds to one of the SAR sub-data sets, which includes the backscattering coefficient. The geographic coordinate system for the SAR data can be the geographic reference system used for the SAR data. The radar imagery can be a high-resolution image obtained by processing actively transmitted and received microwave signals from a satellite. The backscattering coefficient can be data representing the intensity of radar wave energy reflected from ground features, typically measured in decibels (dB), including VV and VH polarization. The digital elevation model (DEM) data can be a digital model representing surface elevation (altitude) information.The aforementioned digital elevation model (DEM) data includes various DEM sub-data and a geographic coordinate system for the DEM data. Each pixel in the aforementioned image corresponds to one of the aforementioned DEM sub-data, which includes elevation values. These elevation values ​​can represent ground elevation information. The geographic coordinate system for the aforementioned DEM data can be the geographic coordinate system used by the DEM data. For example, the aforementioned band DN value information (containing 4 out of 13 different band DN values) can be [Band2 (blue) DN value: 1250, Band3 (green) DN value: 980, Band4 (red) DN value: 650, Band8 (near-infrared) DN value: 4200]. The aforementioned synthetic aperture radar (SAR) data can be SAR (Synthetic Aperture Radar) data. The aforementioned DEM data can be ALOSDSM data. In practice, firstly, the aforementioned implementing entity can acquire optical satellite imagery data, synthetic aperture radar data, and digital elevation model data collected by the data platform. Then, the aforementioned optical satellite imagery data, synthetic aperture radar data, and digital elevation model data are identified as multimodal remote sensing data.

[0018] Step 102: Perform registration and fusion processing on the multimodal remote sensing data to obtain registered multimodal remote sensing data.

[0019] In some embodiments, the aforementioned execution entity may perform registration and fusion processing on multimodal remote sensing data to obtain registered multimodal remote sensing data.

[0020] In some optional implementations of certain embodiments, the aforementioned execution entity may perform registration and fusion processing on the aforementioned multimodal remote sensing data through the following steps to obtain registered multimodal remote sensing data: The first step involves radiometric calibration and atmospheric correction of the optical satellite imagery data included in the aforementioned multimodal remote sensing data to obtain corrected optical satellite imagery data. This corrected optical satellite imagery data may include a corrected image map, a reflectance information set, a scene classification band information set, and a geographic coordinate system for the optical satellite imagery data. Each pixel in the corrected image map corresponds to a reflectance information in the reflectance information set. The corrected image map can be an image obtained by radiometric calibration and atmospheric correction of the aforementioned image map. Each reflectance information in the reflectance information set includes: aerosol band surface reflectance, blue band surface reflectance, green band surface reflectance, red band surface reflectance, red-edge 1 band surface reflectance, red-edge 2 band surface reflectance, red-edge 3 band surface reflectance, near-infrared band surface reflectance, narrow near-infrared band surface reflectance, water vapor band surface reflectance, shortwave infrared band surface reflectance, shortwave infrared 1 band surface reflectance, and shortwave infrared 2 outer band surface reflectance. The aforementioned aerosol band surface reflectance represents the surface reflectance of electromagnetic bands sensitive to atmospheric aerosols after atmospheric correction. The aforementioned blue band surface reflectance represents the surface reflectance of the blue visible electromagnetic band after atmospheric correction. The aforementioned green band surface reflectance represents the surface reflectance of the green visible electromagnetic band after atmospheric correction. The aforementioned red band surface reflectance represents the surface reflectance of the red visible electromagnetic band after atmospheric correction. The aforementioned red-edge band 1, red-edge band 2, and red-edge band 3 surface reflectance represent the surface reflectance of electromagnetic bands located in the transition region (red edge) between red visible light and near-infrared light after atmospheric correction. The aforementioned near-infrared band surface reflectance represents the surface reflectance of the near-infrared electromagnetic band after atmospheric correction. The aforementioned narrow near-infrared band surface reflectance represents the surface reflectance of the near-infrared electromagnetic band after atmospheric correction. The aforementioned water vapor band surface reflectance can represent the surface reflectance obtained after atmospheric correction for electromagnetic bands sensitive to water vapor absorption in the atmosphere. The aforementioned shortwave infrared band surface reflectance, shortwave infrared band 1 surface reflectance, and shortwave infrared band 2 surface reflectance can represent the surface reflectance obtained after atmospheric correction for shortwave infrared electromagnetic bands. Each pixel in the corrected image corresponds to a scene classification band in the scene classification band information set. For example, the aforementioned reflectance information (four of the 13 different band reflectance information contained in the reflectance information) could be [blue band surface reflectance: 0.073, green band surface reflectance: 0.063, red band surface reflectance: 0.060, near-infrared band surface reflectance: 0.449].

[0021] The second step involves calibrating and correcting the synthetic aperture radar data included in the multimodal remote sensing data, based on the digital elevation model data. This process yields corrected radar data.

[0022] The third step involves performing a unified coordinate transformation on the aforementioned calibrated optical satellite imagery data, calibrated radar data, and digital elevation model data included in the multimodal remote sensing data to obtain unified optical satellite imagery data, unified radar data, and unified digital elevation model data. Specifically, the unified optical satellite imagery data corresponds to the calibrated optical satellite imagery data, the unified radar data corresponds to the calibrated radar data, and the unified digital elevation model data corresponds to the digital elevation model data included in the multimodal remote sensing data. In practice, the implementing entity can use remote sensing processing software (such as ENVI, ERDAS Imagine, QGIS, ArcGIS Pro, etc.) to perform Reproject Raster, Project Raster, or Warp (Reproject) operations on the aforementioned calibrated optical satellite imagery data, calibrated radar data, and digital elevation model data included in the multimodal remote sensing data to ensure that all data have the same geographic reference system (such as WGS84 UTM) to obtain unified optical satellite imagery data, unified radar data, and unified digital elevation model data. The aforementioned unified optical satellite imagery data can represent data that, after undergoing the same coordinate transformation, shares the same geographic coordinates, spatial resolution, and image range as unified radar data and unified digital elevation model data. Similarly, the aforementioned unified radar data can represent data that, after undergoing the same coordinate transformation, shares the same geographic coordinates, spatial resolution, and image range as unified optical satellite imagery data and unified digital elevation model data. The aforementioned unified digital elevation model data can represent data that, after undergoing the same coordinate transformation, shares the same geographic coordinates, spatial resolution, and image range as unified optical satellite imagery data and unified radar data, including digital elevation model data within the aforementioned multimodal remote sensing data. The aforementioned unified optical satellite imagery data can include a unified image map, a reflectivity information set, a scene classification band information set, a unified geographic reference system, and spatial resolution. The aforementioned unified image map can be an image obtained by performing the same coordinate transformation on the corrected image map. Each pixel in the aforementioned unified image map corresponds to a reflectivity information in the reflectivity information set. Each pixel in the aforementioned unified image map corresponds to a scene classification band information in the scene classification band information set. The aforementioned unified geographic reference system can refer to a geographic reference system (such as WGS84 UTM) jointly used by unified optical satellite imagery data, unified radar data, and unified digital elevation model data. The aforementioned spatial resolution can be the actual ground size represented by each pixel in the unified imagery map. The aforementioned unified radar data can include unified radar imagery maps, various unified radar sub-data, spatial resolution, and a geographic reference system (such as WGS84 UTM).Each pixel in the aforementioned unified radar image corresponds to one of the aforementioned unified radar sub-data sets. The aforementioned unified radar image can be an image obtained by performing the same coordinate transformation on a calibrated radar image. The aforementioned unified radar sub-data sets include: a first backscattering coefficient and a second backscattering coefficient. The aforementioned first backscattering coefficient can be a filtered VV polarization value. The aforementioned second backscattering coefficient can be a filtered VH polarization value. The aforementioned filtered VV polarization value can be a filtered backscattering coefficient of vertical transmission and vertical reception polarization. The aforementioned filtered VH polarization value can be a filtered backscattering coefficient of vertical transmission and horizontal reception polarization. The aforementioned unified digital elevation model (DEM) data can include various DEM sub-data sets and a geographic reference system (such as WGS84 UTM). Each pixel in the aforementioned image corresponds to one of the aforementioned unified DEM sub-data sets, which includes elevation values. The aforementioned elevation values ​​can be data representing ground elevation information.

[0023] The fourth step involves geometric registration of the aforementioned unified optical satellite imagery data, unified radar data, and unified digital elevation model (DEM) data to obtain registered optical satellite imagery data, registered radar data, and registered DEM data. The registered optical satellite imagery data corresponds to the aforementioned unified optical satellite imagery data. The registered radar data corresponds to the aforementioned unified radar data. The registered DEM data corresponds to the aforementioned unified DEM data. The registered optical satellite imagery data can be optical data that, using the aforementioned unified optical satellite imagery data as a reference, achieves spatial accuracy alignment with the unified DEM data through geometric registration. The registered radar data can be radar data that, using the aforementioned unified radar data as input, achieves spatial accuracy alignment with the registered optical satellite imagery data through geometric registration. The registered DEM data can be elevation data that, using the aforementioned unified DEM data as input, is typically used as a geometric reference or control basis in the geometric registration process. The aforementioned registered optical satellite imagery data may include a registered image map, a reflectivity information set, a scene classification band information set, a geographic reference system, and spatial resolution. The aforementioned registered image map may be an image whose spatial positional discrepancies between a unified image map and a unified radar image map have been eliminated through geometric transformation. Each pixel in the aforementioned registered image map corresponds to a reflectivity information in the reflectivity information set. Each pixel in the aforementioned registered image map corresponds to a scene classification band information in the scene classification band information set. The aforementioned registered radar data may include a registered radar image map, various registered radar sub-data sets, and a geographic coordinate system. The aforementioned registered radar image map may be an image whose spatial positional discrepancies between a unified radar image map and a unified image map have been eliminated through geometric transformation. Each pixel in the aforementioned registered radar image map corresponds to one of the aforementioned registered radar sub-data sets. The aforementioned registered radar sub-data sets include: a first backscattering coefficient and a second backscattering coefficient. The aforementioned first backscattering coefficient may be a filtered VV polarization value. The aforementioned second backscattering coefficient may be a filtered VH polarization value. The filtered VV polarization value can be the filtered backscattering coefficient of vertical transmission and vertical reception polarization. The filtered VH polarization value can be the filtered backscattering coefficient of vertical transmission and horizontal reception polarization. The registered digital elevation model data can include various registered digital elevation model sub-data and a geographic coordinate system. Each pixel in the registered radar image corresponds to one of the registered digital elevation model sub-data. The registered digital elevation model sub-data includes: elevation values. The elevation values ​​can be data representing ground elevation information.

[0024] The fifth step involves performing feature extraction and stacking processing on the above-mentioned registered optical satellite image data, registered radar data, and registered digital elevation model data to obtain registered multimodal remote sensing data.

[0025] In some optional implementations of certain embodiments, the aforementioned execution entity may perform calibration and correction processing on the synthetic aperture radar data included in the multimodal remote sensing data based on the digital elevation model data included in the multimodal remote sensing data, thereby obtaining corrected radar data: The first step involves radiometric calibration of the synthetic aperture radar (SAR) data within the aforementioned multimodal remote sensing data, based on the digital elevation model (DEM) data. This calibrated SAR data yields precisely calibrated radar data. This calibrated radar data can be data with accurate geometric location obtained after radiometric and geometric calibration (especially orthorectification). The calibrated SAR data can include a calibrated radar image, individual calibrated radar sub-data, and a geographic coordinate system. The calibrated radar image can be an image where geometric distortions (such as overlay, shadows, and top-bottom displacement) caused by terrain undulations have been precisely corrected using the aforementioned DEM data, ensuring that each pixel in the image is located at its correct surface position. Each pixel in the calibrated radar image corresponds to one of the aforementioned calibrated radar sub-data. The calibrated radar sub-data includes: polarization values ​​before and after calibration, and elevation. The polarization values ​​before and after calibration described above characterize the original measurements without radiometric calibration and the radar scattering coefficients obtained after radiometric calibration of the original measurements during the radiometric calibration process of the synthetic aperture radar data. The original measurements can be dimensionless digital values ​​(DN values), reflecting the relative intensity of the echo signals received by the radar system from a specific polarization channel (e.g., VH, VV). The polarization values ​​before and after calibration can include the VV polarization value before calibration, the VV polarization value after calibration, the VH polarization value before calibration, and the VH polarization value after calibration. The VV polarization value before calibration and the VH polarization value before calibration can represent the radar scattering coefficient before calibration. The VV polarization value after calibration and the VH polarization value after calibration can represent the radar scattering coefficient after calibration, in dB. The elevation values ​​described above can be elevation values ​​from digital elevation model data, in meters.

[0026] The second step involves noise filtering of the calibration radar data to obtain filtered radar data. In practice, the executing entity can use adaptive filtering algorithms (such as Lee filtering, Gamma MAP filtering, or Refined Lee filtering) to suppress noise in the calibration radar data, thus obtaining filtered radar data. This filtered radar data can be the data obtained after noise filtering of the calibration radar data, which suppresses the inherent speckle noise in the radar signal while preserving image details and edge information. The filtered radar data can include a filtered radar image, various filtered radar sub-data, and a geographic coordinate system. The filtered radar image can be an image with improved signal-to-noise ratio and more prominent ground features, obtained by noise filtering of the calibration radar image. Each pixel in the filtered radar image corresponds to one of the filtered radar sub-data. The filtered radar sub-data includes: pre-filter polarization value, post-filter polarization value, and elevation. The pre-filter polarization values ​​can be the radar scattering coefficients after calibration of the calibration radar data before noise filtering. The aforementioned post-filter polarization values ​​can be radar scattering coefficients that suppress speckle noise after noise filtering of the aforementioned calibrated radar data. The aforementioned pre-filter polarization values ​​can include pre-filter VV polarization and pre-filter VH polarization. The aforementioned post-filter polarization values ​​can include post-filter VV polarization and post-filter VH polarization. The aforementioned pre-filter VV polarization and pre-filter VH polarization can represent the radar scattering coefficient before filtering. The aforementioned post-filter VV polarization and post-filter VH polarization can represent the radar scattering coefficients after filtering that suppress speckle noise, with units in decibels (dB).

[0027] The third step involves performing terrain correction processing on the filtered radar data based on the digital elevation model data included in the aforementioned multimodal remote sensing data, resulting in corrected radar data. This corrected radar data can be data obtained by eliminating the influence of terrain undulations after terrain correction processing of the filtered radar data. The corrected radar data can include a corrected radar image, various corrected radar sub-data, and a geographic coordinate system. The corrected radar image can be an image obtained by eliminating terrain distortion and terrain brightness effects from the filtered radar image. Each pixel in the corrected radar image corresponds to one of the filtered radar sub-data in the various corrected radar sub-data. The corrected radar sub-data includes: filtered polarization values ​​(filtered VV polarization, filtered VH polarization) and terrain-corrected equivalent views. The filtered VV polarization and filtered VH polarization represent the filtered radar scattering coefficient, measured in decibels (dB). The terrain-corrected equivalent views represent the terrain correction processing effect.

[0028] In some optional implementations of certain embodiments, the aforementioned execution entity may perform feature extraction and stacking processing on the aforementioned registered optical satellite image data, the aforementioned registered radar data, and the aforementioned registered digital elevation model data through the following steps to obtain registered multimodal remote sensing data: The first step involves extracting the reflectance information set from the registered optical satellite imagery data, including the surface reflectance of each blue band, green band, red band, and near-infrared band, to obtain blue band matrices, green band matrices, red band matrices, and near-infrared band matrices. The registered optical satellite imagery data can be optical data that, based on the unified optical satellite imagery data, has undergone geometric registration processing to achieve spatial accuracy alignment with the unified digital elevation model data. Each blue band surface reflectance represents the surface reflectance obtained after atmospheric correction for the blue visible electromagnetic band. Each green band surface reflectance represents the surface reflectance obtained after atmospheric correction for the green visible electromagnetic band. Each red band surface reflectance represents the surface reflectance obtained after atmospheric correction for the red visible electromagnetic band. Each of the aforementioned near-infrared band surface reflectance values ​​represents the surface reflectance obtained after atmospheric correction of the near-infrared electromagnetic band. The aforementioned blue band matrix can be a matrix formed by the positions of the pixels corresponding to the reflectance information of each blue band surface reflectance in the aforementioned registered image map. The aforementioned green band matrix can be a matrix formed by the positions of the pixels corresponding to the reflectance information of each green band surface reflectance in the aforementioned registered image map. The aforementioned red band matrix can be a matrix formed by the positions of the pixels corresponding to the reflectance information of each red band surface reflectance in the aforementioned registered image map. The aforementioned near-infrared band matrix can be a matrix formed by the positions of the pixels corresponding to the reflectance information of each near-infrared band surface reflectance in the aforementioned registered image map. The aforementioned registered optical satellite image data can include a registered image map, a reflectance information set, a scene classification band information set, a geographic reference system, and spatial resolution. The aforementioned registered image can be an image whose spatial positional discrepancy between a unified image and a unified radar image has been eliminated through geometric transformation. Each pixel in the aforementioned registered image corresponds to a reflectance information in the reflectance information set. Each pixel in the aforementioned registered image also corresponds to a scene classification band information in the scene classification band information set. In practice, firstly, the executing entity can construct a blue band matrix based on the positions of the pixels corresponding to the reflectance information of each blue band surface reflectance in the aforementioned registered image. Then, it can construct a green band matrix based on the positions of the pixels corresponding to the reflectance information of each green band surface reflectance in the aforementioned registered image.Next, a red band matrix is ​​constructed by assigning the pixel information corresponding to the reflectance information of each red band surface reflectance to the positions of the corresponding pixels in the aforementioned registered image. Finally, a near-infrared band matrix is ​​constructed by assigning the pixel information corresponding to the reflectance information of each near-infrared band surface reflectance to the positions of the corresponding pixels in the aforementioned registered image.

[0029] The second step involves extracting the first backscattering coefficients and second backscattering coefficients from each of the registered radar sub-data in the aforementioned registered radar data, resulting in a first backscattering coefficient matrix and a second backscattering coefficient matrix. The aforementioned registered radar data can be radar data that, through geometric registration processing, achieves spatial alignment with the registered optical satellite image data using the aforementioned unified radar data as input. Each of the aforementioned first backscattering coefficients can be a filtered VV polarization value. The filtered VV polarization value can be a filtered backscattering coefficient with vertical transmission and vertical reception polarization. Each of the aforementioned second backscattering coefficients can be a filtered VH polarization value. The filtered VH polarization value can be a filtered backscattering coefficient with vertical transmission and horizontal reception polarization. The aforementioned first backscattering coefficient matrix can be a matrix obtained by arranging the first backscattering coefficients according to the position of the pixel information corresponding to the registered radar sub-data in the aforementioned registered radar image. The aforementioned second backscattering coefficient matrix can be a matrix obtained by arranging the pixel information corresponding to each second backscattering coefficient in the registered radar sub-data according to the position of the pixel information in the registered radar image. The registered radar data can include the registered radar image, each registered radar sub-data, and a geographic coordinate system. The registered radar image can be an image whose spatial positional deviation between unified radar images has been eliminated through geometric transformation. Each pixel information in the registered radar image corresponds to one of the registered radar sub-data. The registered radar sub-data includes: a first backscattering coefficient and a second backscattering coefficient. In practice, firstly, the executing entity can arrange the first backscattering coefficients in the registered radar image according to the position of the pixel information corresponding to the registered radar sub-data, thus obtaining the first backscattering coefficient matrix. Then, the second backscattering coefficients are arranged in the registered radar image according to the position of the pixel information corresponding to the registered radar sub-data, thus obtaining the second backscattering coefficient matrix.

[0030] The third step involves extracting the elevation values ​​from each of the registered digital elevation model (DEM) data sets to obtain an elevation matrix. The registered DEM data can be elevation data that uses unified digital elevation model (DEM) data as input and is typically used as a geometric reference in geometric registration. Each elevation value can represent ground elevation information. The elevation matrix can be a matrix obtained by arranging the pixel information corresponding to each elevation value in the registered DEM sub-data set within the registered radar image. The registered DEM data can include each registered DEM sub-data set and a geographic coordinate system. The registered radar image can be an image where spatial positional discrepancies between the unified radar image and the unified image have been eliminated through geometric transformation. Each pixel information in the registered radar image corresponds to one of the registered DEM sub-data sets. The registered DEM sub-data set includes: elevation values. In practice, the aforementioned executing entity can arrange the various elevation values ​​according to the position of the pixel information corresponding to the registered digital elevation model sub-data in the registered radar image map, thereby obtaining an elevation matrix.

[0031] The fourth step involves generating a quality grade label matrix based on the scene classification band information set included in the aforementioned registered optical satellite image data. The aforementioned registered optical satellite image data can be optical data that, using the aforementioned unified optical satellite image data as a reference, undergoes geometric registration processing to achieve spatial accuracy alignment with the unified digital elevation model data. Each scene classification band in the aforementioned scene classification band information set can be a scene classification layer (SCL band). Each pixel in the aforementioned registered image map corresponds to a scene classification band in the scene classification band information set. The aforementioned quality grade label can be a digital tag representing the quality of the registered optical satellite image data, mapped according to a preset rule for each scene classification band in the scene classification band information set. The aforementioned quality grade label matrix can be a matrix obtained by arranging the mapped quality grade labels according to the position of the pixel information corresponding to the scene classification band information in the aforementioned registered image map. For example, the above preset rule could be to assign a value of 2 (low quality) to scene classification band information of 0, 1, 2, 8, and 9, assign a value of 1 (medium quality) to scene classification band information of 3, 7, and 10, and assign a value of 0 (best quality) to scene classification band information of 4, 5, and 6.

[0032] The fifth step involves arranging the aforementioned blue band matrix, green band matrix, red band matrix, near-infrared band matrix, first backscattering coefficient matrix, second backscattering coefficient matrix, elevation matrix, and quality grade marker matrix in a preset order to obtain registered multimodal remote sensing data. This preset order can be the sequential arrangement of the blue band matrix, green band matrix, red band matrix, near-infrared band matrix, first backscattering coefficient matrix, second backscattering coefficient matrix, elevation matrix, and quality grade marker matrix. The registered multimodal remote sensing data can be a data structure that spatially aligns and standardizes optical satellite imagery data, synthetic aperture radar data, and digital elevation model data, allowing for direct joint analysis. In practice, the execution entity can first create an empty three-dimensional array. Then, the aforementioned blue band matrix, green band matrix, red band matrix, near-infrared band matrix, first backscattering coefficient matrix, second backscattering coefficient matrix, elevation matrix, and quality grade marker matrix are added sequentially to the empty three-dimensional array. Finally, the added 3D array is identified as the registered multimodal remote sensing data.

[0033] Step 103: Input the registered multimodal remote sensing data into the pre-trained visual encoder to obtain a visual feature map.

[0034] In some embodiments, the execution entity can input registered multimodal remote sensing data into a pre-trained visual encoder to obtain a visual feature map. This visual feature map can be a low-resolution two-dimensional feature map.

[0035] In some optional implementations of certain embodiments, the aforementioned execution entity may input the registered multimodal remote sensing data into a pre-trained visual encoder through the following steps to obtain a visual feature map: The first step is to convert the aforementioned registered multimodal remote sensing data into a tensor. This registered multimodal remote sensing data can be a data structure that spatially aligns and standardizes optical satellite imagery data, synthetic aperture radar data, and digital elevation model data, allowing for direct joint analysis. The tensor can be a standardized multidimensional array converted from the registered multimodal remote sensing data. In practice, the executing entity first stacks the matrices included in the registered multimodal remote sensing data along the third dimension (depth / channel dimension) to obtain a three-dimensional array. Then, this three-dimensional array is defined as a tensor. As an example, the dimensions of each matrix are [height (H), width (W)], and after stacking 8 matrices, the resulting three-dimensional array has dimensions [H, W, 8].

[0036] The second step is to generate a 3D feature map based on the tensor described above. In practice, firstly, the execution entity can define and initialize each convolutional layer. Each convolutional layer contains a series of learnable parameters, called a convolutional kernel (or filter). The convolutional kernel can be a 3D weight matrix (e.g., 3x3xC_in, where C_in is the number of channels in the input tensor, i.e., 8 modes). Then, the convolutional kernels are applied to the tensor to perform convolution operations, resulting in convolutional tensors. Next, a non-linear activation function (such as ReLU) is applied to the convolutional tensors to obtain 2D feature maps. Finally, the obtained 2D feature maps are stacked along the channel dimension to obtain a 3D feature map. This 3D feature map can be a high-level, abstract feature representation obtained after the tensor undergoes multiple transformations through a convolutional neural network.

[0037] The third step involves performing image patch embedding processing on the aforementioned 3D feature map to obtain a set of embedding vectors. Each embedding vector in this set can be used to represent the visual features of an image patch in the aforementioned 3D feature map. The embedding vectors may include a vector index, a semantic feature vector, and a similarity vector. The vector index is used to identify the embedding vector. The similarity vector can be other embedding vectors in the set that are similar to the original embedding vector.

[0038] The fourth step involves performing position embedding processing on each embedding vector in the aforementioned embedding vector set to obtain a position embedding vector. In practice, the execution entity first processes the semantic feature vectors included in each embedding vector in the aforementioned embedding vector set using position encoding to obtain a position encoded vector. Then, the obtained position encoded vector is added to the aforementioned embedding vector to obtain the position embedding vector. The position encoded vector can be a vector used in the Transformer architecture to represent the position information of each embedding vector in the embedding vector set. As an example, the semantic features (examples of the first few dimensions) included in the above embedding vector can be [0.85, 0.12, -0.23, 0.45, ...], the above position encoding vector (based on position (1,1)) can be [0.00, 0.84, 0.91, -0.42, 0.14, 0.99, -0.76, 0.65, ...], and the above position embedding vector can be [0.85, 0.96, 0.68, 0.03, 0.81, 0.91, 0.16, 0.50, ...].

[0039] The fifth step is to determine the obtained position embedding vectors as a set of position embedding vectors.

[0040] Step 6: Based on the aforementioned set of positional embedding vectors, generate a set of contextual encoding information. In practice, the execution entity can input the aforementioned set of positional embedding vectors into a Transformer encoder to obtain the set of contextual encoding information. Each piece of contextual encoding information in this set can be generated by the Transformer encoder through deep understanding and information integration of each positional embedding vector in the set, producing a feature vector rich in global contextual information for each positional embedding vector.

[0041] Step 7: Reshape and combine the context encoding information in the aforementioned context encoding information set to obtain a visual feature map. In practice, the executing entity can arrange the context encoding information according to the spatial position of each pixel in the aforementioned image to obtain the visual feature map. The visual feature map can be a deep feature representation with a clear spatial structure that incorporates global context information. Each context encoding information corresponds to one pixel in the aforementioned image.

[0042] Step 104: Based on the above-mentioned registered multimodal remote sensing data, generate the optimal prompt text information.

[0043] In some embodiments, the aforementioned execution entity may generate optimal prompt text information based on the aforementioned registered multimodal remote sensing data.

[0044] In some optional implementations of certain embodiments, the aforementioned execution entity can generate optimal prompt text information based on the aforementioned registered multimodal remote sensing data through the following steps: The first step is to input the registered multimodal remote sensing data into a preset scene classifier to obtain scene labels. The preset scene classifier can be MobileViT. The scene labels can be the overall scene type corresponding to the registered multimodal remote sensing data. For example, when segmenting a land-sea image containing coastal windbreak mangroves, the scene label could be "coastal windbreak mangroves."

[0045] The second step involves retrieving prompts that meet preset conditions from a preset prompt information library based on the aforementioned scene labels. These preset conditions correspond to the scene labels, and the preset prompt information library can be a set of key-value pairs. Each key-value pair in this library can include a key and a value. The key is the scene label. The value corresponds to the instruction for that scene label. This instruction can be information guiding the language model to perform in-depth analysis of the visual feature map. The preset prompt information library can be a collection of information containing various preset prompts. These preset prompts can include information about the segmented region, land cover features, and areas requiring special attention during segmentation. For example, the preset prompt could be {“Mangroves in a River Delta”: “Please accurately segment the mangroves in the river delta region. Note that the mangroves are distributed in strips along the river channel, with a well-developed tidal channel network; the boundary between the mangroves and the turbid river water needs to be carefully distinguished.”}

[0046] The third step is to obtain the category name sequence. In practice, the aforementioned execution entity can obtain the category name sequence corresponding to the registered multimodal remote sensing data provided by the user. This category name sequence can be composed of category names arranged in the order input by the user. These category names can be strings used to identify different object or scene categories.

[0047] Fourth, based on the above category name sequence, populate the preset basic prompt template to obtain the basic prompt template. The above category name sequence includes the names of each category. The preset basic prompt template can be something like "Please separate..." <list> , <list> , <list>…”. In practice, the aforementioned executing entity can fill the placeholders with the various category names included in the category name sequence. <list>This yields a basic prompt template. As an example, the above category name sequence could be ['Mangrove', 'Turbid Seawater', 'Tidal Flat Silt', 'Clear Seawater']. After filling in the blanks, the resulting basic prompt template could be "Please separate mangroves, turbid seawater, tidal flat silt, and clear seawater".

[0048] Fifth, based on the aforementioned prompt information and the preset structured prompt information, the basic prompt template is subjected to discrete category prior knowledge fusion processing to obtain the optimal prompt text information. In practice, the aforementioned executing entity can use language models (such as GPT-3, GPT-4, ChatGPT, etc.) to generate the optimal prompt text information. The aforementioned preset structured prompt information can include information such as the spectral characteristics, morphological features, and spatial relationships of ground features. The aforementioned optimal prompt text information can be structured text containing scene labels, category names, prompt information, and preset structured prompt information. As an example, the aforementioned preset structured prompt information could be "Note: Category Name: 'Mangrove' has uniform texture and high near-infrared reflectivity." The optimal prompt text information mentioned above could be: "The current segmentation object is a coastal windbreak mangrove forest. The categories to be identified include: mangroves, turbid seawater, clear seawater, and tidal flat silt. Please segment all land features. Among them, 'mangroves' have high reflectivity in the near-infrared band, have a rough texture, and are distributed in the intertidal zone; 'turbid seawater' contains suspended sediment, has high reflectivity, and is mostly distributed near the shore of estuaries; 'clear seawater' has low reflectivity, strong absorption in the near-infrared band, and is mostly found in open sea areas; 'tidal flat silt' is periodically exposed in the intertidal zone and has moderate reflectivity."

[0049] Step 105: Perform micro-feature extraction processing on the optimal prompt text information to obtain a set of global text feature information and word-level text feature information.

[0050] In addressing the technical problems mentioned above, and considering the application scenario—after natural disasters such as floods and hailstorms, insurance companies or government agencies need to conduct large-scale assessments of damaged farmland—image segmentation is often required. However, this process is hampered by the following technical issues: current methods, when assessing farmland damage after such disasters, can only extract the general feature of "farmland" from the prompt "segment the flooded fields and intact fields," failing to identify key modifiers like "flooded" and "intact." This leads to the system performing indiscriminate, detailed calculations on the entire image without outputting the required specific subdivisions, resulting in wasted computational resources. To address the specific requirements of this application scenario, we aim to reduce the inefficient consumption of computational resources. Therefore, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity can perform micro-feature extraction processing on the aforementioned optimal prompt text information through the following steps to obtain a set of global text feature information and word-level text feature information: The first step is to perform text cleaning and standardization on the optimal prompt text information to obtain standard prompt text information. This standard prompt text information can include category names but exclude illegal characters. In practice, the executing entity can use regular expressions to remove illegal characters from the optimal prompt text information to obtain the standard prompt text information.

[0051] The second step is to obtain the vocabulary. In practice, the aforementioned execution entity can load the dictionary provided by the word segmentation tool (such as jieba or Stanford Word Segmenter) that corresponds to the optimal prompt text information. The vocabulary can be the dictionary provided by the word segmentation tool (such as jieba or Stanford Word Segmenter) that corresponds to the optimal prompt text information.

[0052] The third step involves performing tokenized segmentation on the standard prompt text information based on the aforementioned vocabulary, resulting in a sequence of sub-word units. In practice, the executing entity can use the Byte Pair Encoding algorithm to process the standard prompt text information and obtain the sub-word unit sequence. This sub-word unit sequence can be a sequence of words obtained by segmenting the standard prompt text information and arranging the sub-word units in the order they appear in the standard prompt text information.

[0053] Fourth, add a start marker and an end marker to the above sub-word unit sequence to obtain a word marker information sequence. The start marker corresponds to a word marker in the above word marker information sequence. The end marker corresponds to a word marker in the above word marker information sequence. The start marker can be the character at the beginning of the marker sequence. The above word marker information sequence can be composed of the start marker, sub-word unit sequence, and end marker in that order. The end marker can be the character at the end of the marker sequence. For example, the start marker could be [CLS], and the end marker could be [EOS].

[0054] Fifth, perform the following steps sequentially on each word tag in the word tag information sequence: The first sub-step involves, in response to determining that the aforementioned word tag information is identical to one of the category names included in the aforementioned standard prompt text information, adding a category tag to the aforementioned word tag information sequence to update the word tag information sequence. In practice, the executing entity can add category tags before and after the aforementioned word tag information in the word tag information sequence. The category tag can be a character that marks a category. For example, the category tag could be... <obj>.

[0055] Step 6: Determine the updated word tag information sequence as the category tag information sequence. Specifically, the word tag information corresponding to each category tag in the above-mentioned category tags corresponds to a category tag information in the above-mentioned category tag information sequence. The word tag information corresponding to the above-mentioned end tag corresponds to a category tag information in the above-mentioned category tag information sequence.

[0056] Step 7: Input the above category label information sequence into a pre-trained text encoder to obtain a label feature vector sequence, where each label feature vector in the above label feature vector sequence corresponds to a category label in the above category label information sequence. The above text encoder can be a Transformer-based encoder. The above label feature vector sequence can be a high-dimensional vector sequence arranged in the order of the category label information sequence.

[0057] Step 8: Sequentially traverse the above-mentioned marker feature vector sequence to store the position information of each marker feature vector corresponding to each category marker, thus obtaining a category marker position information set. Each category marker position information in the above category marker position information set can be the position information of the marker feature vector corresponding to the category marker in the above-mentioned marker feature vector sequence. This position information can be the position of each marker feature vector corresponding to each category marker within the above-mentioned marker feature vector sequence.

[0058] Step nine involves extracting category marker feature vectors from the aforementioned category marker location information set, and using each extracted feature vector as a set of word-level text feature information. In practice, firstly, the executing entity can determine the corresponding feature vectors based on the positions of the feature vectors corresponding to each category marker as represented by the aforementioned category marker location information set. Then, it extracts the corresponding feature vectors from the aforementioned feature vector sequence. Each word-level text feature information in the aforementioned set of word-level text feature information can be a feature vector corresponding to one of the category markers.

[0059] Step 10: Extract the end-marker feature vector from the above-mentioned marker feature vector sequence, and use the extracted marker feature vector as global text feature information. In practice, the execution entity can extract the marker feature vector corresponding to the end mark from the above-mentioned marker feature vector sequence, and then use the extracted marker feature vector as global text feature information. The global text feature information can be the marker feature vector corresponding to the end mark from the above-mentioned marker feature vector sequence.

[0060] The above technical solution, combined with steps 106-107 and related content, serves as an inventive point of this disclosure, solving the technical problem of "ineffective consumption of computing resources." Factors leading to wasted computing resources often include: after natural disasters such as floods and hailstorms, when insurance companies or government agencies need to conduct large-scale assessments of damaged farmland, current methods can only extract the general feature of "farmland" from the prompt "segment the flooded fields and intact fields," failing to identify key modifiers such as "flooded" and "intact." This causes the system to perform indiscriminate fine calculations on the entire image, ultimately failing to output the required specific subdivision results, resulting in ineffective consumption of computing resources. Solving these factors can reduce the ineffective consumption of computing resources. To achieve this, firstly, the optimal prompt text information is cleaned and standardized to obtain standard prompt text information. This removes illegal characters from the optimal prompt text information, resulting in standard prompt text information that does not contain illegal characters. Then, a vocabulary is obtained. From this vocabulary, word-level text features can be extracted, resulting in word-level text features composed of general vocabulary, rather than those composed of non-general vocabulary or those derived from splitting a group of consecutive words with modifiers into individual words. Furthermore, based on the vocabulary, the standard prompt text information can be segmented using tokenization to obtain a sequence of sub-word units. This yields a correct sequence of sub-word units composed of general words. Then, a start tag and an end tag are added to the above sub-word unit sequence to obtain a word tag information sequence. This yields a word tag information sequence with a start tag and an end tag added. Next, the following steps are performed sequentially on each word tag in the word tag information sequence. Then, it is determined that the word tag information is identical to one of the category names included in the standard prompt text information. A category tag is added to the word tag information sequence to update the word tag information sequence. This yields a word tag information sequence with added category tags, allowing the extraction of tag feature vectors based on the category tags. Then, the updated word tag information sequence is determined as the category tag information sequence. This yields the category tag information sequence. Next, the above category label information sequence is input into a pre-trained text encoder to obtain a label feature vector sequence. This yields a label feature vector sequence containing contextual encoding information. Then, the above label feature vector sequence is sequentially traversed to store the positional information of each label feature vector corresponding to each category label, resulting in a category label positional information set. Based on this set, each label feature vector corresponding to each category label can be extracted. Finally, based on this set of category label positional information, the above label feature vector sequence undergoes category label feature vector extraction processing, and the extracted label feature vectors are used as a word-level text feature information set.Therefore, a set of word-level text feature information corresponding to each category marker can be obtained. Finally, the above marker feature vector sequence is processed by extracting end marker feature vectors, and the extracted marker feature vectors are used as global text feature information. Thus, global text feature information corresponding to the end marker can be obtained. Because the optimal prompt text information is extracted based on the vocabulary and end markers, a set of global text feature information and word-level text feature information is obtained. The extracted global text feature information is not randomly extracted but extracted based on the end markers. The resulting set of word-level text feature information is a general combination of words rather than a set of words with modifiers that are split into individual words. Combining steps 106-107, multimodal interactive reasoning processing is performed on the visual feature map, global text feature information, and word-level text feature information set to generate a conditional feature information set. Thus, multimodal interactive reasoning processing can be performed on the visual feature map, the global text feature information extracted based on the end markers, and the word-level text feature information set composed of a general vocabulary to generate an accurate conditional feature information set. The conditional feature information set is input into a pre-trained segmentation decoder to output a land-sea segmentation image. This allows the processing unit to segment land and sea images, avoiding the pitfall of using a set of word-level text features—comprising individual words—to guide segmentation. This prevents the model from performing indiscriminate, fine-grained calculations across the entire image region, ultimately failing to output the desired detailed segmentation. This reduces the inefficient consumption of computational resources.

[0061] Step 106: Perform multimodal interactive reasoning processing on the visual feature map, global text feature information and word-level text feature information set to generate a conditional feature information set.

[0062] In addressing the aforementioned technical problems in the application scenario of coastal topographic map updating and land-sea image segmentation, the following technical issues often arise: In land-sea topographic map updating tasks, the low matching degree between textual and visual information leads to coarse image segmentation results, unclear feature boundaries, and missing details. Directly using multimodal interactive inference methods often results in overly broad global textual features and submerged word-level features, making it difficult for the model to accurately identify key areas and subtle feature categories. This leads to poor image segmentation quality and reliability. To address the following requirements for this application scenario: improving the quality and reliability of image segmentation and enhancing the quality and reliability of topographic maps, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the execution entity may perform multimodal interactive reasoning processing on the visual feature map, the global text feature information, and the word-level text feature information set through the following steps to generate a conditional feature information set: The first step involves performing a linear transformation on at least one word-level text feature from the aforementioned global text feature information and word-level text feature information sets to generate a global query matrix and at least one word-level query matrix. The global query matrix can be a matrix obtained by performing a linear transformation on the global text feature information. Similarly, the word-level query matrix can be a matrix obtained by performing a linear transformation on the word-level text feature information.

[0063] The second step involves performing a linear transformation on the aforementioned visual feature map to generate a key matrix and a value matrix. The key matrix can be a matrix encoding key visual features extracted from the visual feature map, where each key corresponds to a location or feature region in the image, briefly describing "what visual features are present here" (e.g., "red area," "texture edge," "dark blue water"). The value matrix can be a matrix encoding the complete visual features pointed to by the key matrix.

[0064] The third step is to generate a global similarity matrix based on the global query matrix and the key matrix described above. This global similarity matrix can be a matrix composed of the similarities between the global query matrix and the key matrix.

[0065] Fourth, based on the above-mentioned at least one word-level query matrix and the above-mentioned key matrix, generate at least one word-level similarity matrix. One of the above-mentioned word-level similarity matrices can be a matrix composed of the similarities between one of the word-level query matrices and the above-mentioned key matrix.

[0066] The fifth step involves generating global attention weights based on the aforementioned global similarity matrix. In practice, the executing entity can input the global similarity matrix into the Softmax function to obtain the global attention weights. These global attention weights represent the correlation between the entire scene and the task's global intent.

[0067] Step 6: Based on the above at least one word-level similarity matrix, generate at least one word-level attention weight. One of the word-level attention weights can represent the relevance between a region in the visual feature map and each category name in the category names.

[0068] Step 7: Perform a weighted fusion process on the aforementioned global attention weights and value matrix to obtain global conditional feature information. This global conditional feature information can be a visual feature representation enhanced by global textual feature information.

[0069] Step 8: Perform a weighted fusion process on the at least one word-level attention weight and the value matrix to obtain at least one word-level conditional feature information. This word-level conditional feature information can be a visual feature representation enhanced by word-level textual feature information.

[0070] The ninth step is to determine the above-mentioned global conditional feature information and at least one word-level conditional feature information as a set of conditional feature information.

[0071] The above technical solution, combined with step 107 and related content, serves as an inventive point of this disclosure, solving the technical problem of "poor quality and reliability of image segmentation." Factors leading to poor image segmentation quality and reliability often include: in land and sea topographic map update tasks, low matching between textual and visual information results in coarse image segmentation results, unclear feature boundaries, and missing details. Directly using multimodal interactive inference methods often leads to overly broad global textual features and submerged word-level features, making it difficult for the model to accurately identify key regions and subtle feature categories. This results in poor image segmentation quality and reliability. Solving these factors can improve information granularity matching and enhance the quality and reliability of topographic maps. To achieve this, firstly, a linear transformation is performed on at least one word-level textual feature in the global textual feature information and word-level textual feature information sets to generate a global query matrix and at least one word-level query matrix. This allows textual features to be mapped to a semantic space. Then, a linear transformation is performed on the visual feature map to generate a key matrix and a value matrix, allowing visual features to be mapped to the same semantic space for comparison. Furthermore, a global similarity matrix can be generated based on the global query matrix and key matrix. This allows us to determine the correlation strength between textual and visual information; textual and visual information with high correlation strength will have higher similarity scores, while those with low correlation strength will have lower similarity scores. Then, global attention weights are generated based on the global similarity matrix. This allows the model to understand instructions from a holistic scene perspective. Next, at least one word-level attention weight is generated based on at least one word-level similarity matrix. This allows the model to focus on specific categories with finer granularity. Then, the global attention weights and value matrix are weighted and fused to obtain global conditional feature information. This allows us to emphasize visual information based on the importance indicated by the global attention weights, resulting in global conditional feature information that incorporates all visual information related to the global instruction. Finally, at least one word-level attention weight and the aforementioned value matrix are weighted and fused to obtain at least one word-level conditional feature. This allows us to emphasize visual information based on the importance indicated by the word-level attention weight, resulting in word-level conditional feature information highly correlated with the category. Finally, the global conditionalized feature information and at least one word-level conditionalized feature information are determined as the conditionalized feature information set. This yields a set of conditionalized feature information with high fine-grained matching. Furthermore, because a linear transformation is used to align visual and textual information to the same semantic space, similarity calculations are performed on the global textual feature information and the aforementioned word-level textual feature information set. An attention mechanism is employed to reduce the possibility of word-level features (such as "intertidal zone," "reef," and "beach") being overlooked, resulting in a set of conditionalized feature information with high fine-grained matching.In step 107, the conditional feature information set is input into a pre-trained segmentation decoder, outputting a land-sea segmentation image. This allows for the input of a highly matched set of conditional feature information with fine granularity into a pre-trained segmentation decoder, resulting in a land-sea segmentation image with accurate segmentation of subtle land features, thus improving the quality and reliability of image segmentation.

[0072] Step 107: Input the set of conditional feature information into the pre-trained fine-grained land-sea segmentation model to obtain the land-sea segmentation image.

[0073] In addressing the technical challenges of the aforementioned background technologies, and considering the specific application scenario—when a vision-based navigation system (such as an unmanned surface vessel) malfunctions (e.g., accidentally colliding with a beach)—engineers need to conduct an accident review. The land-sea segmentation module is the most crucial perception component in a visual navigation system and a key focus of the investigation. However, land-sea image segmentation often presents the following technical challenges: directly outputting segmented images through an end-to-end "black box" model lacks intermediate state information from the decision-making process (such as category confidence), making it difficult to analyze the cause of the malfunction and preventing engineers from accurately locating the fault through land-sea image segmentation results. This application scenario requires the following characteristics: accurate location of the fault through land-sea image segmentation results. Faced with these technical challenges, we decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity can input the aforementioned conditional feature information set into a pre-trained fine-grained land-sea segmentation model through the following steps to obtain a land-sea segmentation image: The first step involves inputting the aforementioned conditional feature information set into the pre-trained fine-grained land-sea segmentation model to obtain groupings for each category. This fine-grained land-sea segmentation model can be a segmentation decoder. The segmentation decoder can consist of a series of upsampling layers (such as bilinear interpolation or transposed convolution) and convolutional layers. Each category score in the aforementioned groupings corresponds to a category name in the aforementioned category name sequence. In practice, the executing entity first inputs the aforementioned conditional feature information set into the upsampling layer of the segmentation decoder to obtain a feature map. Then, a convolutional layer (typically a 1x1 convolution) is applied to the feature map to obtain groupings for each category. The aforementioned feature map is a high-resolution semantic feature representation that recovers spatial information, obtained through upsampling operations (such as bilinear interpolation or transposed convolution). Each category grouping can be the confidence score of each pixel in the aforementioned feature map corresponding to each category name in the aforementioned category name sequence.

[0074] The second step is to perform the following steps for each of the resulting category groups: The first sub-step involves normalizing the score of each category within the category grouping to obtain initial category probability information. In practice, the aforementioned execution entity can use the Softmax function to process each category score within the category grouping to obtain the initial category probability information. This initial category probability information can be the predicted probability of belonging to each category.

[0075] The second sub-step involves determining the obtained initial category probability information as an initial category probability information group, wherein each initial category probability information in the initial category probability information group corresponds to one category score in the above category scores.

[0076] The third step is to define the determined initial category probability information groups as the initial category probability information group set.

[0077] Fourth, perform the following steps for each initial category probability information group in the above initial category probability information group set: The first sub-step is to determine the initial category probability information with the highest probability in the initial category probability information group as the category probability information.

[0078] The fifth step is to determine the obtained category probability information into category probability information groups, wherein each category probability information in the above category probability information group corresponds to one category score in the above category scores.

[0079] The sixth step is to determine each category name in the category name sequence corresponding to each category probability information in the above category probability information group as the category label information.

[0080] Step 7: Based on the determined category label information, output the land-sea segmentation image. In practice, firstly, the aforementioned execution entity can convert the category label information into corresponding colors using a predefined category color mapping table. Then, output the land-sea segmentation image. This land-sea segmentation image can be an image where different colors correspond to different categories, mapped through the defined category color mapping table. As an example, when dealing with a coastal area arranged in a small spatial location (3x3 grid), the predefined category color mapping table above can be {1: [61 / 255, 145 / 255, 64 / 255, 1.0], (Mangrove - Dark Green); 2: [166 / 255, 124 / 255, 82 / 255, 1.0], (Turbid Seawater - Yellowish Brown); 3: [230 / 255, 201 / 255, 163 / 255, 1.0], (Tidal Flat Silt - Light Beige); 4: [135 / 255, 206 / 255, 235 / 255, 1.0], (Clear Seawater - Sky Blue)}.

[0081] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "accurately locating faulty links through the results of land-sea image segmentation". Factors that prevent accurate location of faulty links through land-sea image segmentation results are often as follows: directly outputting segmented images through an end-to-end "black box" model, lacking intermediate state information (such as category confidence) in the decision-making process, makes fault cause analysis impossible, and engineers cannot accurately locate faulty links through land-sea image segmentation results. If these factors are resolved, the effect of accurately locating faulty links through land-sea image segmentation results can be achieved. To achieve this effect, firstly, the above-described conditional feature information set is input into the pre-trained land-sea fine segmentation model to obtain groupings for each category. This allows the model's decision-making process to be digitized, resulting in groupings for each category. Thus, the faulty link can be identified through the category scores of each category group. Then, the following steps are performed on each category group. This ensures that each category group is processed. Then, the scores of each category within the category group are normalized to obtain initial category probability information. This converts the original scores into a standard probability distribution. Next, the obtained initial category probability information is defined as initial category probability information groups. This yields initial category probability information groups. Then, the defined initial category probability information groups are defined as a set of initial category probability information groups. This yields a set of initial category probability information groups. Then, the following steps are performed on each initial category probability information group in the set of initial category probability information groups. Afterwards, the initial category probability information with the highest probability in the initial category probability information group is defined as the category probability information. This yields the initial category probability information with the highest probability as the category probability information. Then, the obtained category probability information is defined as a category probability information group. This yields a category probability information group composed of the highest-probability category information. Afterwards, each category name in the category name sequence corresponding to each category probability information in the category probability information group is defined as each category label information. This yields each category label information corresponding to each category probability information in the category probability information group. Finally, based on the determined category label information, a land-sea segmentation image is output. Thus, a land-sea segmentation image can be output based on the category label information. Because a fine-grained land-sea segmentation model is used to output groupings for each category, and then a probabilistic decision-making process is applied to these groupings to ultimately output category label information, a quantitative indicator of the classification reliability of each pixel is provided for the ship's decision-making system.This avoids the problem of directly outputting segmented images through an end-to-end "black box" model, which lacks intermediate state information (such as category confidence) in the decision-making process, making it impossible to analyze the cause of the fault. Thus, the fault link can be accurately located through the results of land and sea image segmentation.

[0082] The above-described embodiments of this disclosure have the following beneficial effects: the land-sea image segmentation method of some embodiments of this disclosure improves the fine segmentation capability of land-sea image segmentation. Specifically, the reason for the poor fine segmentation capability of land-sea image segmentation is that the pixel labeling method based on grayscale values ​​or color features to set thresholds can only complete basic land-sea dichotomy, simply dividing the image into ocean and land. When facing land areas with complex spatial structures and rich semantic information, such as roads, bridges, ports, and agricultural land, it lacks an effective mechanism to distinguish between land features with similar morphological features but different functional attributes (such as ports and ordinary embankments, both of which are artificially constructed gray linear structures near water). It is difficult to accurately define the subdivision types and spatial boundaries of agricultural land such as farmland, orchards, and greenhouses, resulting in poor fine segmentation capability of land-sea image segmentation. Based on this, the land-sea image segmentation method of some embodiments of this disclosure first acquires multimodal remote sensing data. Then, the multimodal remote sensing data is registered and fused to obtain registered multimodal remote sensing data. This allows for the registration of the spatial positions of pixels in each multimodal remote sensing data set, reducing positional bias and resulting in registered multimodal remote sensing data. Next, this registered multimodal remote sensing data is input into a pre-trained visual encoder to obtain a visual feature map. This yields a visual feature map that retains rich spatial details and visual information. Then, based on the registered multimodal remote sensing data, optimal prompt text information is generated. This allows for the generation of corresponding optimal prompt text information based on the registered multimodal remote sensing data. Next, micro-feature extraction processing is performed on the optimal prompt text information to obtain global text feature information and word-level text feature information sets. This allows for the extraction of global and specific semantic concepts at a micro-level to obtain global and word-level text feature information sets for multimodal interactive reasoning processing. Finally, multimodal interactive reasoning processing is performed on the visual feature map, the global text feature information, and the word-level text feature information set to generate a conditional feature information set. Therefore, the visual modality of the visual feature map can interact and reason with the linguistic modality of the global text feature information and the aforementioned word-level text feature information set, generating a conditional feature information set containing rich semantic information and enhanced visual information. This lays the foundation for achieving fine-grained land feature classification and boundary delineation. Finally, the aforementioned conditional feature information set is input into a pre-trained fine-grained land-sea segmentation model to obtain a land-sea segmentation image. Thus, a land-sea segmentation image capable of segmenting complex land features can be obtained.Because it first generates optimal prompt text information corresponding to the scene of the registered multimodal remote sensing data based on the registered multimodal remote sensing data, and then generates global text feature information and word-level text feature information sets based on the optimal prompt text information. The global text feature information guides the model to focus on global semantic features, and the word-level text feature information sets guide the model to focus on specific and refined semantic features. In this way, it can guide the model to distinguish visually similar but semantically different land features. Therefore, it can improve the differentiation of land features with similar morphological features but different functional attributes (such as port docks and ordinary embankments), thereby achieving fine segmentation of land and sea images and improving the fine segmentation capability of land and sea image segmentation.

[0083] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a plate information recognition device, which are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0084] like Figure 2 As shown, the land and sea image segmentation apparatus 200 in some embodiments includes: an acquisition unit 201, a registration and fusion processing unit 202, a first input unit 203, a generation unit 204, a micro-feature extraction processing unit 205, a multimodal interactive reasoning processing unit 206, and a second input unit 207. The acquisition unit 201 is configured to acquire multimodal remote sensing data; the registration and fusion processing unit 202 is configured to perform registration and fusion processing on the multimodal remote sensing data to obtain registered multimodal remote sensing data; the first input unit 203 is configured to input the registered multimodal remote sensing data into a pre-trained visual encoder to obtain a visual feature map; the generation unit 204 is configured to generate optimal prompt text information based on the registered multimodal remote sensing data; the micro-feature extraction processing unit 205 is configured to perform micro-feature extraction processing on the optimal prompt text information to obtain a set of global text feature information and a set of word-level text feature information; the multimodal interactive reasoning processing unit 206 is configured to perform multimodal interactive reasoning processing on the visual feature map, the global text feature information, and the set of word-level text feature information to generate a set of conditional feature information; and the second input unit 207 is configured to input the set of conditional feature information into a pre-trained fine-grained land-sea segmentation model to obtain a land-sea segmentation image.

[0085] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0086] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0087] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0088] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0089] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.< / obj> < / list> < / list> < / list> < / list>

Claims

1. A land-sea image segmentation method, comprising: Acquire multimodal remote sensing data; The multimodal remote sensing data is registered and fused to obtain registered multimodal remote sensing data; The registered multimodal remote sensing data is input into a pre-trained visual encoder to obtain a visual feature map; Based on the registered multimodal remote sensing data, the optimal prompt text information is generated; Micro-feature extraction processing is performed on the optimal prompt text information to obtain a set of global text feature information and word-level text feature information; Multimodal interactive reasoning processing is performed on the visual feature map, the global text feature information, and the word-level text feature information set to generate a conditional feature information set; The conditional feature information set is input into a pre-trained fine-grained land-sea segmentation model to obtain a land-sea segmentation image.

2. The method of claim 1, wherein, The multimodal remote sensing data includes optical satellite imagery data, synthetic aperture radar data, and digital elevation model data. The multimodal remote sensing data undergoes registration and fusion processing to obtain registered multimodal remote sensing data, including: Radiometric calibration and atmospheric correction are performed on the optical satellite image data included in the multimodal remote sensing data to obtain corrected optical satellite image data; Based on the digital elevation model data included in the multimodal remote sensing data, the synthetic aperture radar data included in the multimodal remote sensing data is calibrated and corrected to obtain corrected radar data. The corrected optical satellite image data, the corrected radar data, and the digital elevation model data included in the multimodal remote sensing data are subjected to the same coordinate transformation process to obtain unified optical satellite image data, unified radar data, and unified digital elevation model data. Geometric registration processing is performed on the unified optical satellite image data, the unified radar data, and the unified digital elevation model data to obtain registered optical satellite image data, registered radar data, and registered digital elevation model data. Feature extraction and stacking processing are performed on the registered optical satellite image data, the registered radar data, and the registered digital elevation model data to obtain registered multimodal remote sensing data.

3. The method of claim 2, wherein, The process involves calibrating and correcting the synthetic aperture radar (SAR) data included in the multimodal remote sensing data based on the digital elevation model data, resulting in corrected SAR data, including: Based on the digital elevation model data included in the multimodal remote sensing data, radiometric calibration processing is performed on the synthetic aperture radar data included in the multimodal remote sensing data to obtain calibrated radar data. The calibration radar data is subjected to noise filtering to obtain filtered radar data; Based on the digital elevation model data included in the multimodal remote sensing data, terrain correction processing is performed on the filtered radar data to obtain corrected radar data.

4. The method of claim 2, wherein, The registered optical satellite image data includes a reflectance information set and a scene classification band information set. Each reflectance information in the reflectance information set includes: blue band surface reflectance, green band surface reflectance, red band surface reflectance, and near-infrared band surface reflectance. The registered radar data includes various registered radar sub-data, which include: a first backscattering coefficient, a second backscattering coefficient, and an equivalent number of views. The registered digital elevation model data includes various registered digital elevation model sub-data, which include: elevation values ​​and coherence coefficients. The process involves feature extraction and stacking of the registered optical satellite image data, the registered radar data, and the registered digital elevation model data to obtain registered multimodal remote sensing data, including: The reflectance information set in the registered optical satellite image data is extracted, including the surface reflectance of each blue band, each green band, each red band, and each near-infrared band, to obtain the blue band matrix, green band matrix, red band matrix, and near-infrared band matrix. The first backscattering coefficient and the second backscattering coefficient are extracted from each registration radar sub-data in the registration radar data to obtain the first backscattering coefficient matrix and the second backscattering coefficient matrix; Extract the elevation values ​​from each registered digital elevation model data in the registered digital elevation model data to obtain an elevation matrix; Based on the scene classification band information set included in the registered optical satellite image data, a quality level label matrix is ​​generated; The blue band matrix, green band matrix, red band matrix, near-infrared band matrix, first backscattering coefficient matrix, second backscattering coefficient matrix, elevation matrix, and quality grade marker matrix are arranged according to preset rules to obtain registered multimodal remote sensing data.

5. The method according to claim 1, wherein, The step of inputting the registered multimodal remote sensing data into a pre-trained visual encoder to obtain a visual feature map includes: The registered multimodal remote sensing data is converted into tensors; Based on the tensor, a three-dimensional feature map is generated; The three-dimensional feature map is subjected to image patch embedding processing to obtain an embedding vector set; Perform position embedding processing on each embedding vector in the set of embedding vectors to obtain a position embedding vector; The obtained position embedding vectors are defined as a set of position embedding vectors; Based on the aforementioned location embedding vector set, a context encoding information set is generated; The context encoding information in the set of context encoding information is reshaped and combined to obtain a visual feature map.

6. The method according to claim 1, wherein, The process of generating optimal prompt text information based on the registered multimodal remote sensing data includes: The registered multimodal remote sensing data is input into a preset scene classifier to obtain scene labels; Based on the scene tags, obtain prompt information that meets the preset conditions from the preset prompt information library; Get the category name sequence; Based on the category name sequence, a preset basic prompt template is populated to obtain a basic prompt template, wherein the category name sequence includes each category name; Based on the aforementioned prompt information and the preset structured prompt information, the basic prompt template is subjected to discrete category prior knowledge fusion processing to obtain the optimal prompt text information, wherein the optimal prompt text information includes the names of each category.

7. A land-sea image segmentation device, comprising: The acquisition unit is configured to acquire multimodal remote sensing data; The registration and fusion processing unit is configured to perform registration and fusion processing on the multimodal remote sensing data to obtain registered multimodal remote sensing data; The first input unit is configured to input the registered multimodal remote sensing data into a pre-trained visual encoder to obtain a visual feature map; The generation unit is configured to generate optimal prompt text information based on the registered multimodal remote sensing data; The micro-feature extraction processing unit is configured to perform micro-feature extraction processing on the optimal prompt text information to obtain a set of global text feature information and word-level text feature information. The multimodal interactive reasoning processing unit is configured to perform multimodal interactive reasoning processing on the visual feature map, the global text feature information, and the word-level text feature information set to generate a conditional feature information set. The second input unit is configured to input the set of conditional feature information into a pre-trained fine-grained land-sea segmentation model to obtain a land-sea segmentation image.

8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.