Multi-source remote sensing image data target extraction method, device, equipment and medium

By acquiring multi-source remote sensing image data and preprocessing, using pre-trained networks and SAM models to achieve rapid extraction of offshore oil film targets, solving the problems of poor oil film extraction accuracy and low degree of automation in the prior art, and achieving high accuracy automated oil film extraction.

CN120014249AActive Publication Date: 2025-05-16齐鲁空天信息研究院 +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510487710.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The prior art has problems such as large calculation volume, poor accuracy, and lack of uniformity in the verification of indicators between different source data in offshore oil film extraction, which is difficult to meet the needs of large-scale and automated oil film extraction.

Method used

By acquiring multi-source remote sensing image data and preprocessing, the target positioning box is obtained based on the pre-trained network, the boundary feature parameters are obtained, and the SAM model is input to achieve the target extraction result.

Benefits of technology

It realizes rapid extraction of offshore oil film targets under large batches and automated conditions, and the target extraction results are highly accurate, so it can effectively extract targets on non-training source data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014249A_ABST
    Figure CN120014249A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source remote sensing image data target extraction method and device, equipment and a medium, and can be applied to the technical field of image processing. The multi-source remote sensing image data target extraction method comprises the following steps: acquiring and preprocessing multi-source remote sensing image data; obtaining a target positioning frame based on the preprocessed data and a pre-trained network; obtaining boundary feature parameters based on the target positioning frame parameters; and inputting an SAM model based on the target positioning frame and the boundary feature parameters to obtain a target extraction result. The invention further provides a multi-source remote sensing image data target extraction device and equipment, a storage medium and a program product. The method has the advantages that rapid extraction of the offshore oil film target under the large-scale and automatic condition is achieved, and the accuracy of the target extraction result is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and more specifically, to a method, device, equipment, medium and program product for extracting targets from multi-source remote sensing image data. Background Art

[0002] Multi-source remote sensing images can provide rich information about land objects. By fusing and comprehensively analyzing remote sensing images from different sources, we can more comprehensively and accurately understand and comprehend various phenomena and changes on the earth's surface. In recent years, marine oil spills caused by ship collisions and drilling platform leaks have occurred frequently, causing water pollution, forming harmful algal blooms, and causing serious economic losses. Therefore, it is necessary to promptly determine the location and scope of the oil spill, identify the oil type, quantify the thickness of the oil film, and reduce the secondary harm caused by the accident.

[0003] Existing oil film extraction technologies can be summarized into two categories. The first category is the traditional oil film extraction method, which mainly includes the following steps: (1) dark spot detection, that is, locating the possible location of the oil film; (2) feature extraction, that is, manually or automatically screening the multi-dimensional features of the target, such as geometric features, physical features, texture features, and polarization features; (3) classifier classification, that is, implementing pixel-level classification of feature sets, such as support vector machines, decision trees, artificial neural networks, etc. The second category is based on deep learning methods, such as CNN, VGG, etc. Most of these models directly migrate models in the field of optical images to the field of SAR images, losing phase information. In addition, a large number of samples need to be manually labeled, which is labor-intensive. Summary of the invention

[0004] In view of the above problems, the present disclosure provides a multi-source remote sensing image data target extraction method, device, equipment, medium and program product with the ability to automatically predict target positioning parameters and achieve rapid extraction of offshore oil film targets under large-scale and automated conditions.

[0005] According to a first aspect of the present disclosure, a method for extracting targets from multi-source remote sensing image data is provided, including acquiring and preprocessing multi-source remote sensing image data; acquiring a target positioning frame based on the preprocessed data and a pre-trained network; acquiring boundary feature parameters based on the target positioning frame parameters; and obtaining a target extraction result by inputting a SAM model based on the target positioning frame and the boundary feature parameters.

[0006] According to an embodiment of the present disclosure, the acquiring a target positioning frame based on the preprocessed data and the pretrained model includes: extracting basic features of the preprocessed data; outputting multi-scale features based on the basic features; acquiring a target candidate region based on the multi-scale features and a pretrained region candidate network; and acquiring a target positioning frame based on the target candidate region and a pretrained classification regression network.

[0007] According to an embodiment of the present disclosure, the training method of the region candidate network includes: taking multiple anchor frames of different sizes and proportions based on each feature map of the multi-scale feature using training set data, and training the category probability and coordinate parameters of the anchor frames; the training loss function includes calculating the offset of the true value frame and the predicted frame relative to the anchor frame.

[0008] According to an embodiment of the present disclosure, the method includes: dividing the preprocessed image data into blocks, using a multi-head attention mechanism, through convolution and dimension adjustment to obtain image embedding information; the image embedding information is used to input the SAM model to obtain the extraction result.

[0009] According to an embodiment of the present disclosure, the extraction result is obtained by inputting the SAM model based on the target positioning box and the boundary feature parameters, including: based on the image embedding information, the target positioning box and the boundary feature parameters, after multi-head attention mechanism and transposed convolution processing, generating an extraction result after prediction of the image to be segmented.

[0010] According to an embodiment of the present disclosure, the preprocessing includes polarization feature extraction, and the polarization feature extraction includes: fusing all channel scattering contribution values ​​of the image data to extract span features, and the polarization features are used for boundary feature parameter extraction.

[0011] Another aspect of the disclosed embodiment provides a multi-source remote sensing image data target extraction device, including: a data set processing module, used to acquire and pre-process multi-source remote sensing image data; a target rapid positioning module, used to acquire a target positioning frame based on the pre-processed data and a pre-trained network; a SAM automatic segmentation module, used to obtain boundary feature parameters based on the parameters of the target positioning frame, and to obtain a target extraction result based on the target positioning frame and the boundary feature parameters.

[0012] Another aspect of an embodiment of the present disclosure provides an electronic device, including: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the method described above.

[0013] Another aspect of an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor executes the method as described above.

[0014] Another aspect of an embodiment of the present disclosure provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0015] One or more of the above embodiments have the following beneficial effects:

[0016] The method implemented in the present disclosure obtains and preprocesses multi-source remote sensing image data; obtains a target positioning frame based on the preprocessed data and a pre-trained network; obtains boundary feature parameters based on the target positioning frame parameters; and obtains a target extraction result by inputting the target positioning frame and the boundary feature parameters into a SAM model. The target positioning parameters are automatically predicted by the pre-trained model, and the boundary feature parameters are used to achieve rapid extraction of offshore oil film targets under large-scale and automated conditions, and the target extraction result has a high accuracy rate.

[0017] Based on the use of level 1 SLC data, the input data includes the original t matrix data, comprehensive polarization feature extraction, and boundary feature information acquisition, the full-process rapid target extraction of offshore oil film using multi-polarization and multi-phase SAR is realized. Based on the migration training of the produced sample set, the target model is obtained. The target positioning parameters and boundary feature parameters obtained by model prediction are supplemented with the modified SAM prediction results.

[0018] In addition, one or more embodiments of the present disclosure also have the following beneficial effects:

[0019] (1) Meet the needs of large-scale, automated oil film extraction.

[0020] The existing technology requires manual or image processing methods to locate the target area from the acquired image before subsequent feature extraction and classification can be performed. The method proposed in the embodiment of the present disclosure is based on a pre-trained model, which can automatically complete the entire process from data acquisition to target extraction, and can also extract targets in non-training source data sets.

[0021] (2) Solve the problem of lack of uniformity in indicator verification among data.

[0022] Most of the existing feature extraction methods do not have a unified standard, such as polarization features, texture features and statistical features, or these features are screened and combined using some methods and sent to the classifier for classification. The proposed method automatically extracts high-level, multi-scale features and fuses the multi-scales to avoid the phenomenon that the features have to change with the change of the data set.

[0023] (3) Reduce the manpower and energy consumed by large model labels.

[0024] When the existing technology uses a large model to implement segmentation tasks, a large number of manually labeled samples are required, and the accuracy of the model directly applied to the actual downstream branch tasks is very poor. The basic features of the method proposed in the embodiment of the present disclosure are extracted using an open source model, and only multi-scale high-level features need to be trained later, reducing the required sample size and training time. In addition, the SAM technology is used when labeling a small number of samples, with an average of 2 seconds per sample, which greatly improves the speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0026] Figure 1 The application scenario diagram of a multi-source remote sensing image data target extraction method, apparatus, device, medium and program product according to an embodiment of the present disclosure is schematically shown.

[0027] Figure 2 A flowchart of a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown.

[0028] Figure 3 The flowchart of obtaining a data sample set according to an embodiment of the present disclosure is schematically shown.

[0029] Figure 4 A flowchart of a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown.

[0030] Figure 5 A flowchart of another method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown.

[0031] Figure 6 The structure diagram of BottleNeck according to an embodiment of the present disclosure is schematically shown.

[0032] Figure 7 The figure schematically shows a multi-head intention mechanism diagram according to an embodiment of the present disclosure.

[0033] Figure 8 A diagram showing an implementation effect according to an embodiment of the present disclosure is shown.

[0034] Fig. 9 Another implementation effect diagram according to an embodiment of the present disclosure is shown.

[0035] Fig.10 The structural block diagram of a multi-source remote sensing image data target extraction device according to an embodiment of the present disclosure is schematically shown.

[0036] Fig.11 A block diagram of an electronic device suitable for implementing a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown.

[0037] It should be noted that, for the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the sizes of the overall / local structures or the overall / local areas may be enlarged or reduced, that is, these drawings are not drawn according to the actual scale. DETAILED DESCRIPTION

[0038] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0039] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0040] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0041] When using expressions such as "at least one of A, B, and C", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0042] Among the existing oil film extraction technologies, traditional oil spill extraction methods have problems such as large amount of calculation, poor accuracy, and lack of uniformity in indicator verification between different source data. They can no longer meet the needs of large-scale, automated oil film extraction. Although a large number of models based on deep learning have been successfully transplanted to oil film extraction, due to the lack of a large number of standard data sets for actual downstream branch tasks, there is still a defect of requiring true value labels as support, and data labels mostly rely on manual annotation, which is time-consuming, labor-intensive, and unreliable. Moreover, most of the above two types of extraction methods are to crop out some areas from the entire scene image for oil film extraction, and do not have the ability to quickly extract the entire process from data acquisition to oil film segmentation.

[0043] In view of the above problems existing in the prior art, the oil film extraction method proposed in the present invention is used to at least partially solve the above technical problems. The method is based on the SAM general segmentation model, automatically predicts the target positioning model parameters through training, and is supplemented by boundary feature parameters to achieve rapid extraction of offshore oil film targets under large-scale and automated conditions. The effectiveness of the method in practical applications is proved by multi-source data, verifying the great potential of the SAM model in oil film analysis.

[0044] An embodiment of the present disclosure provides a method for extracting targets from multi-source remote sensing image data, comprising the steps of: acquiring and preprocessing multi-source remote sensing image data; acquiring a target positioning frame based on the preprocessed data and a pre-trained network; acquiring boundary feature parameters based on the target positioning frame parameters; and inputting the target positioning frame and the boundary feature parameters into a SAM model to obtain a target extraction result.

[0045] Figure 1 The following schematically shows an application scenario diagram of the multi-source remote sensing image data target extraction method according to an embodiment of the present disclosure. It should be noted that: Figure 1 What is shown are merely examples to which the embodiments of the present disclosure can be applied, so as to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0046] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0047] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples). The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, etc.

[0048] The server 105 may be a server that provides various services, such as a background management server that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. For example, the server 105 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud computing, network services, and middleware services.

[0049] In applications such as client applications and web applications (abbreviated as APP in English), the client (i.e., front-end) and the server (i.e., back-end) can communicate data through network messages. For example, in the APP client, the parameters that the server needs to obtain from the client are assembled, and the network request method is called to send them to the server.

[0050] It should be noted that the multi-source remote sensing image data target extraction method provided in the embodiment of the present disclosure can generally be executed by at least one of a terminal device or a server. Accordingly, the multi-source remote sensing image data target extraction device provided in the embodiment of the present disclosure can generally be set in at least one of a terminal device or a server.

[0051] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.

[0052] The following will be based on Figure 1 The scene described by Figure 2-Figure 9 The multi-source remote sensing image data target extraction method of the disclosed embodiment is described in detail.

[0053] Figure 2 A flowchart of a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown.

[0054] like Figure 2 As shown, the multi-source remote sensing image data target extraction method of this embodiment includes:

[0055] In operation S100, multi-source remote sensing image data is acquired and pre-processed.

[0056] In operation S200, a target positioning frame is obtained based on the preprocessed data and the pre-trained network.

[0057] In operation S300, boundary feature parameters are obtained based on the parameters of the target positioning frame; and a SAM model is input based on the target positioning frame and the boundary feature parameters to obtain an extraction result.

[0058] In this embodiment, the acquired multi-source remote sensing image data is first used to realize image preprocessing and generate a standard data set, the preprocessed data is input into a pre-trained target rapid positioning network, the predicted target positioning frame parameter information is output, the boundary feature parameters are obtained based on the parameters of the target positioning frame, and finally the target positioning frame parameter information and the boundary feature parameter information are fused and input into the SAM automatic segmentation module, and fine-tuned to obtain the predicted target extraction result.

[0059] In this embodiment, multi-source remote sensing data refers to remote sensing data obtained by different types of sensors, different platforms or at different times, including optical remote sensing data, thermal infrared remote sensing data, microwave remote sensing data and other types. In this embodiment, the multi-source remote sensing image data can be obtained by the single-look complex (SLC) image data of the synthetic aperture radar (SAR). The single-look complex image data of SAR has the ability to work all day and all weather, is not limited by factors such as clouds, rain, fog, day and night, and has a certain penetration ability for vegetation, soil, etc., and can obtain information that optical remote sensing cannot obtain. The two can complement each other and provide more comprehensive ground object information. This embodiment specifically uses the level 1 SLC image data of Sentinel-1 and Gaofen-3, which contains phase and amplitude information. SLC data refers to SAR data with a look number of 1 in both the range and azimuth directions. It is the original high-resolution data obtained by directly processing the echo signal received by the radar, retaining the complete phase information and amplitude information of the signal, stored in complex form, and each pixel point consists of a real part and an imaginary part. SLC data has the characteristics of high resolution, complete phase information, large data volume and obvious noise.

[0060] In this embodiment, the preprocessing includes at least one processing technology such as orbit correction, radiation calibration, polarization feature extraction, speckle filtering and geocoding. According to the specific image data situation, one or more of the above preprocessing technologies can be selected for use alone or in combination. Among them, the orbit correction technology is implemented using a precise orbit file to correct the differences between different orbits. The radiation calibration technology converts the brightness value recorded by the sensor into a backscattering coefficient. For example, in this embodiment, the image data of Gaofen-3 can be calibrated by the following formula:

[0061]

[0062] In the above formula, is the radar backscatter coefficient (unit: dB); , I is the real part of the 1A-level product, Q is the imaginary part of the 1A-level product, and QV is the maximum value of the scene image before quantization, which can be obtained by parsing the metadata file.

[0063] Polarization feature extraction technology is mainly used to improve the accuracy of SAR image interpretation. This embodiment uses span features, and other multi-polarization features can also be used. The present invention extracts span features by fusing all channel scattering contribution values, as shown in the following formula:

[0064]

[0065] In the above formula, Represents the scattering amplitude value of different channels. Specifically, polarization correction is required. Suppose the original full polarization image before correction is M. Correct each group of pixels according to the following formula to obtain the relatively real ground object scattering matrix S after correction, where i and j represent the row and column numbers corresponding to the pixels. As shown in the following formula:

[0066]

[0067] The speckle filtering technique is to remove the speckle noise in the SAR feature map. The present invention uses the polarization LEE method with a window size of 5X5 to remove this effect. Specifically, edge template matching is first performed on the span feature map to select a direction window, and then a local statistical filter is applied within the direction window to achieve speckle suppression processing.

[0068] The geocoding technology mainly converts SAR data of different phases from the slant range coordinate system to the geographic coordinate system, including image positioning and projection conversion, which will not be described in detail.

[0069] In some embodiments of the present application, the training method of the region candidate network includes: taking a plurality of anchor boxes of different sizes and proportions based on each feature map of the multi-scale feature using training set data, and training the category probability and coordinate parameters of the anchor boxes;

[0070] The training loss function involves calculating the offset of the true box and the predicted box relative to the anchor box.

[0071] In this embodiment, the pre-trained network is trained based on a data sample set. The data sample set can be obtained based on the following method. Figure 3 The flowchart of obtaining a data sample set according to an embodiment of the present disclosure is schematically shown. The process of obtaining a data sample set specifically includes:

[0072] Operation S110 includes data acquisition, orbit correction, radiation calibration, polarization feature extraction, speckle filtering and geocoding preprocessing operations.

[0073] Operation S120 includes constructing a multi-scale pyramid model, implementing simulation, cropping and normalization, and enhancing data robustness by using rotation, folding and difference.

[0074] Operation S130 includes extracting embedded information of the data, implementing automatic labeling of training samples, and generating a data sample set in a standard format.

[0075] In this embodiment, the data set can be expanded by constructing a multi-scale pyramid model, and a series of images of different scales can be obtained by performing smoothing filtering and downsampling operations on the original image to different degrees. Since images of different scales can capture the features of targets of different sizes, they can better cope with the changes in target scales when performing feature matching, thereby improving the accuracy and robustness of matching. In this embodiment, the robustness of the data can also be enhanced by simulating, cropping and normalizing the image data, or by using methods such as rotation, folding and difference. When the data diversity meets the conditions, the relevant image data expansion step can be omitted, or one or more combinations of the above expansion methods can be selected.

[0076] In this embodiment, the embedded information of the data can be extracted by a data annotation tool to realize automatic annotation of training samples and generate a data sample set in a standard format. The data sample set contains data, categories and coordinate information, and the data sample set can be further divided into a training set, a validation set and a test set, for example, in a ratio of 6:3:1.

[0077] In some embodiments of the present disclosure, the acquiring a target positioning frame based on the preprocessed data and the pretrained model includes: extracting basic features of the preprocessed data; outputting multi-scale features based on the basic features; acquiring a target candidate region based on the multi-scale features and a pretrained region candidate network; and acquiring a target positioning frame based on the target candidate region and a pretrained classification regression network.

[0078] Figure 4 The flowchart of a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown. Figure 4 As shown, in this embodiment, the multi-source remote sensing image data target extraction method includes:

[0079] In operation S210, ResNet is used as a backbone model to extract basic data features and output the last layer feature map of each stage.

[0080] In operation S220, sampling and fusion processing of features at different levels are implemented based on the feature map of each stage, and a multi-scale feature map is output.

[0081] In operation S230, anchor boxes of different sizes and proportions are constructed based on the selected multi-scale feature map, and offsets are obtained through a pre-trained region candidate network to obtain a suitable candidate box, that is, a target candidate region.

[0082] In operation S240, the candidate box is assigned to a suitable feature layer to obtain the region of interest, and then the final target parameters are obtained through a pre-trained classification regression network, that is, the target positioning box is obtained.

[0083] In this embodiment, the residual network (ResNet) is used to extract the basic features of the data set. The disclosed embodiment uses ResNet as the backbone model to solve the problem of gradient disappearance. In this embodiment, there are specifically 5 stages, each of which is composed of multiple BottleNeck structures stacked together, as shown in the schematic diagram Figure 6 shown.

[0084] In this embodiment, multi-scale features are output based on the basic features. Specifically, the feature maps of different stages output from the network model initialization unit are further processed, that is, the features output by the last residual block layer are respectively extracted, which are recorded as {C 2 ,C 3 ,C 4 ,C 5 ,C 6}, and the last four features are convolved by 1*1 to obtain {M 3 ,M 4 ,M 5 ,M 6}, then M i The feature layers are respectively recorded as {N 3 ,N 4 ,N 5 ,N 6}, then N 3 、N 4 、N 5 The feature layers are respectively downsampled by 2x, 4x, fused, and convolved as {P 3 ,P 41 ,P 42 ,P 51 ,P 52 ,P 6}, P 41 and P 42 , P 51 and P 52 After weighted processing, {P 2 , P 3 , P 4 , P 5 , P 6}Multi-scale features. This embodiment extracts multi-scale features and fuses multiple scales to avoid the phenomenon that features change with changes in data sets.

[0085] In some embodiments of the present disclosure, the training method of the region candidate network includes: taking multiple anchor frames of different sizes and proportions based on each feature map of the multi-scale feature using training set data, and training the category probability and coordinate parameters of the anchor frames respectively; the training loss function is to calculate the offset of the true value frame and the predicted frame relative to the anchor frame.

[0086] The region candidate network is used to recommend target candidate regions. The training of the region candidate network first takes anchor frames of various sizes and proportions on each feature map output by the multi-scale feature extraction unit, and then trains the category probabilities and coordinate parameters of the anchor frames respectively. In this embodiment, specifically, four anchor frames of various sizes and proportions are taken on each feature map output by the multi-scale feature extraction unit. The Loss during training is calculated by calculating the offset of the true value frame and the predicted frame relative to the anchor frame. The specific calculation formula is as follows:

[0087] , .

[0088] , .

[0089] , .

[0090] , .

[0091]

[0092] In the above formula, , and Represent the center coordinates, width and height of the predicted box, anchor box and true value box respectively.

[0093] In some embodiments of the present disclosure, a pre-trained classification regression network is used to implement re-screening of the prediction box to obtain the final parameters of the target. Specifically, the feature map output from the multi-scale feature extraction unit and the prediction box output from the region candidate unit are combined to obtain ROIs, and the allocation of the prediction box to the feature layer is determined by the following formula.

[0094]

[0095] In the above formula, k is the feature layer of the multi-scale feature extraction unit, k 0The feature layer of the initialized model unit mapped to the input image size, w and h represent the size of the prediction box, and the symbol Represents the lower bound of the internal value.

[0096] The basic features of the method proposed in the embodiment of the present disclosure are extracted using an open source model. Subsequently, only multi-scale high-level features need to be trained, which reduces the required sample size and training time. In addition, the SAM technology is used when marking a small number of samples, with an average of one sample every 2 seconds, which greatly improves the speed.

[0097] In some embodiments of the present disclosure, the preprocessed image data is divided into blocks, each depth vector of each block is assigned a position code, and a multi-head attention mechanism is used to obtain image embedding information through convolution and dimension adjustment; the image embedding information is used to input the SAM model to obtain the extraction result.

[0098] In some embodiments of the present disclosure, the extraction result is obtained by inputting the SAM model based on the target positioning box and the boundary feature parameters, including: based on the image embedding information, the target positioning box and the boundary feature parameters are processed by a multi-head attention mechanism and a transposed convolution to generate an extraction result after prediction of the image to be segmented.

[0099] The SAM model is used to achieve pixel-level segmentation of objects in image data. The SAM model mainly includes an image encoder, a hint encoder, and a mask decoder.

[0100] Figure 5 The flowchart of a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown. Figure 5 As shown, the multi-source remote sensing image data target extraction method of this embodiment includes:

[0101] In operation S310, the preprocessed image data is convolved and positionally encoded, and the information is enhanced using a multi-head attention mechanism, and the image embedding information is output after the dimension is adjusted.

[0102] In operation S320, the image embedding information, ie, the object positioning box parameters and the boundary feature parameters are fused to form the input of the mask decoder.

[0103] In operation S330, after processing such as multi-head attention mechanism and transposed convolution, a target extraction result after prediction of the image to be segmented is generated.

[0104] In this embodiment, the image encoder part divides the image into multiple small blocks through the convolution kernel, then assigns a position code to each depth vector of each block, enhances the information of the feature vector using a multi-head attention mechanism, and finally adjusts the dimension through convolution and batch processing to obtain image embedding information.

[0105] Figure 7 The schematic diagram of the multi-head intention mechanism according to the embodiment of the present disclosure is shown schematically. Figure 7 As shown in the figure, the multiple blocks on the left side of the figure (labeled 1, 2, 3, ..., N) represent the image data of the input sequence. The multi-head attention mechanism inputs the input in parallel to multiple independent "heads" (represented as head1 to headN in the figure). Each head has its own independent weight matrix. For each head, the input elements are operated with multiple weight matrices separately. Here W q , W k It is used to calculate the query and key related representation, W v Used to calculate the representation related to the value. The normalized (Softmax) attention score is weighted and summed with the value vector to obtain the attention output of each head. The outputs of all heads (Attention Map1 to Attention MapN) are concatenated (Concat operation) to obtain a vector that integrates the information of multiple heads. The concatenated vector is then passed through a weight matrix W o Perform linear transformation to get the final output.

[0106] In this embodiment, the prompt encoder part uses points, boxes, masks and other methods to embed prompt information into the model. This embodiment uses a fusion method as the input form of the encoding unit, which is specifically composed of two parts. The first part is the output of the classification regression network, and the second part is the boundary feature parameter, which is intended to achieve supervision of the model prediction results.

[0107] In this embodiment, the boundary feature parameter extraction process is as follows: first, the positioning frame parameters output by the classification regression network are used as a window, and the span feature map in the window is segmented and edge detected to define the boundary buffer area. Several points are randomly selected in the buffer area to form the boundary feature parameters {Q i , T i}, the following formula is used to determine whether the two feature sets Q and T are separable. If they are not separable, the boundary feature parameters are empty. If they are separable, the boundary feature parameters are merged with the positioning frame and input into the SAM mask decoding unit, supplemented by the correction of the oil film extraction boundary. In this embodiment, the points containing different targets are obtained through the boundary feature parameters, and the point information and frame information are input into the SAM model together to extract the target.

[0108]

[0109] In the above formula, and are the means and variances of different sets respectively.

[0110] In this embodiment, the mask decoder part uses a lightweight mask decoder to realize the mask prediction of the image to be segmented. First, the output of the hint encoding unit and the output of the image encoding unit are processed together through a multi-head attention mechanism, a transposed convolution, etc., to generate an extraction result after the prediction of the image to be segmented.

[0111] In one embodiment of the present disclosure, a multi-source remote sensing image data target extraction method is used to quickly extract offshore oil films, and the implementation process and effects are as follows.

[0112] (1) Preprocessing stage: more than 50 multi-polarization images of the accident area were obtained. After the image data was preprocessed with orbit correction, radiation calibration, polarization feature extraction, speckle filtering and geocoding, the subsequent process was divided into two branches. The first branch randomly extracted a small number of images from the data set, and generated a standard data set through data set expansion and sample set generation operations, which was used as a support for subsequent training and verification of the model. The second branch served as a data pool for rapid prediction of the entire process after model training was completed, which was used to verify the automated extraction of offshore targets and to verify that targets could also be extracted from non-training source data sets.

[0113] (2) Training phase: Obtain the first branch standard data set in the preprocessing phase, extract the basic features and multi-scale features of the data set, and then use the region candidate network and classification regression network to screen the prompt boxes of the target candidate regions respectively. Through training and parameter adjustment, the position parameters of the final target positioning box are obtained. Specifically, the loss during network training is as follows:

[0114]

[0115]

[0116]

[0117] In the above formula, is the category loss, is the regression loss; is the classification probability predicted by the anchor box, is the category label, which is 1 for positive samples and 0 for negative samples; Represents the offset of the prediction box in the region proposal network unit or classification regression unit, Represents the offset of the two types of prediction boxes relative to the true value box; Represents the number of samples in a mini-batch, Represents the number of prediction boxes, is the weight balancing parameter.

[0118] After model training, sample data was used to verify the model. The verification accuracy is shown in the following table:

[0119] Table 1. Sea target extraction accuracy

[0120] Metric Oil Precision 0.960374 Recall 0.978448 F1 value 0.969326

[0121] (3) Prediction stage: The second branch data set in the preprocessing stage is input into the image encoder to obtain image embedding information, and the parameters obtained by training the prompt encoding unit and the boundary feature parameters together form the prompt information. The embedded information and the prompt information are input into the mask decoding unit to generate the predicted extraction result.

[0122] Figure 8 A diagram showing an implementation effect according to an embodiment of the present disclosure is shown. Fig. 9 Another implementation effect diagram according to an embodiment of the present disclosure is shown. Figure 8-Figure 9 It can be seen that the oil film (Oil) target positioning combination in the image is in line with expectations. This embodiment achieves rapid extraction of offshore oil film targets under large-scale and automated conditions.

[0123] Based on the above multi-source remote sensing image data target extraction method, the present disclosure also provides a multi-source remote sensing image data target extraction device. Fig.10 The device is described in detail.

[0124] like Fig.10 As shown, the multi-source remote sensing image data target extraction device 1200 of this embodiment includes a data set processing module 1201, a target rapid positioning module 1202, and a SAM automatic segmentation module 1203.

[0125] The data set processing module 1201 is used to obtain and pre-process multi-source remote sensing image data;

[0126] A target fast positioning module 1202 is used to obtain a target positioning frame based on the preprocessed data and the pre-trained network;

[0127] The SAM automatic segmentation module 1203 is used to obtain boundary feature parameters based on the parameters of the target positioning frame, and to obtain extraction results based on the target positioning frame and the boundary feature parameters.

[0128] In some embodiments of the present disclosure, the data set processing module 1201 may include: a preprocessing unit for implementing technical processing such as data acquisition, orbit correction, radiation calibration, polarization feature extraction, speckle filtering and geocoding. The data set processing module 1201 may also include: a data set expansion unit for expanding the training data set, which may be implemented by constructing a multi-scale pyramid model, simulating, cropping and normalizing the image, and enhancing the robustness of the data by using methods such as rotation, folding and difference; a sample set generation unit for implementing automatic annotation of training samples by extracting embedded information of the data, and generating a sample set in a standard format.

[0129] In some embodiments of the present disclosure, the preprocessing unit is used to extract polarization features of image data, including: fusing all channel scattering contribution values ​​of the image data to extract span features, which will not be repeated here.

[0130] In some embodiments of the present disclosure, the target rapid positioning module 1202 includes: a network model initialization unit, used to extract basic features of the preprocessed data; a multi-scale feature extraction unit, used to output multi-scale features based on the basic features; a region candidate network unit, used to obtain a target candidate region based on the multi-scale features and a pre-trained region candidate network; a classification regression unit, used to obtain a target positioning frame based on the target candidate region and a pre-trained classification regression network.

[0131] In some embodiments of the present disclosure, the SAM automatic segmentation module 1203 includes: an image encoding unit, which is used to obtain image embedding information based on preprocessed image data by using a multi-head attention mechanism, through convolution and dimension adjustment; a prompt encoding unit; which is used to embed prompt information into the SAM model by means of points, boxes, masks, etc., and obtain boundary feature parameters based on the parameters of the target positioning box; a mask decoding unit, which is used to generate an extraction result after prediction of the image to be segmented based on the image embedding information, the target positioning box and the boundary feature parameters after multi-head attention mechanism and transposed convolution processing.

[0132] For the parts not mentioned in the device part, they can be understood with reference to the various embodiments of the above method. That is, the device part includes modules for executing the various steps of any one of the method embodiments described above. In addition, the implementation methods, technical problems solved, functions implemented, and technical effects achieved of each module / unit / subunit, etc. in the device part embodiment are respectively the same or similar to the implementation methods, technical problems solved, functions implemented, and technical effects achieved of each corresponding step in the method part embodiment, and will not be repeated here.

[0133] According to an embodiment of the present disclosure, at least one of the data set processing module 1201, the target rapid positioning module 1202, and the SAM automatic segmentation module 1203 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the data set processing module 1201, the target rapid positioning module 1202, and the SAM automatic segmentation module 1203 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function can be performed.

[0134] Fig.11 A block diagram of an electronic device suitable for implementing a method for extracting targets from multi-source remote sensing image data according to an embodiment of the present disclosure is schematically shown.

[0135] like Fig.11 As shown, the electronic device 1300 according to an embodiment of the present disclosure includes a processor 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage part 1308 to a random access memory (RAM) 1303. The processor 1301 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 1301 may also include an onboard memory for caching purposes. The processor 1301 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0136] In RAM 1303, various programs and data required for the operation of electronic device 1300 are stored. Processor 1301, ROM 1302 and RAM 1303 are connected to each other through bus 1304. Processor 1301 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 1302 and / or RAM 1303. It should be noted that the program can also be stored in one or more memories other than ROM 1302 and RAM 1303. Processor 1301 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.

[0137] According to an embodiment of the present disclosure, the electronic device 1300 may further include an input / output (I / O) interface 1305, which is also connected to the bus 1304. The electronic device 1300 may further include one or more of the following components connected to the input / output (I / O) interface 1305: an input portion 1306 including a keyboard, a mouse, etc.; an output portion 1307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1308 including a hard disk, etc.; and a communication portion 1309 including a network interface card such as a LAN card, a modem, etc. The communication portion 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the input / output (I / O) interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed, so that a computer program read therefrom is installed into the storage portion 1308 as needed.

[0138] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present disclosure.

[0139] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 1301. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0140] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 1309, and / or installed from the removable medium 1311. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0141] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1309, and / or installed from the removable medium 1311. When the computer program is executed by the processor 1301, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, apparatus, module, unit, etc. described above can be implemented by a computer program module.

[0142] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).

[0143] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0144] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0145] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A method for extracting targets from multi-source remote sensing image data, characterized in that: Includes steps: Acquire multi-source remote sensing image data and preprocess it; Obtain the target positioning frame based on the preprocessed data and pre-trained network; Obtaining boundary feature parameters based on the target positioning frame parameters; The target extraction result is obtained by inputting the SAM model based on the target positioning frame and the boundary feature parameters.

2. The method according to claim 1, characterized in that The obtaining of the target positioning frame based on the preprocessed data and the pre-trained model includes: Extracting basic features of the preprocessed data; outputting multi-scale features based on the basic features; Obtaining a target candidate region based on the multi-scale features and a pre-trained region candidate network; A target positioning frame is obtained based on the target candidate region and a pre-trained classification regression network.

3. The method according to claim 2, characterized in that The training method of the region candidate network includes: Taking a plurality of anchor frames of different sizes and proportions based on each feature map of the multi-scale feature using training set data, and training the category probability and coordinate parameters of the anchor frames; The training loss function involves calculating the offset of the true box and the predicted box relative to the anchor box.

4. The method according to claim 1, characterized in that: The method comprises: The preprocessed image data is divided into blocks, and the multi-head attention mechanism is used to obtain image embedding information through convolution and dimension adjustment; the image embedding information is used to input the SAM model to obtain the extraction result.

5. The method according to claim 1, characterized in that The step of inputting the SAM model based on the target positioning frame and the boundary feature parameters to obtain the target extraction result includes: Based on the image embedding information, the target positioning frame and the boundary feature parameters, after multi-head attention mechanism and transposed convolution processing, an extraction result after prediction of the image to be segmented is generated.

6. The method according to claim 1, characterized in that The preprocessing includes polarization feature extraction, which includes: fusing all channel scattering contribution values ​​of the image data to extract span features, and the polarization features are used for boundary feature parameter extraction.

7. A multi-source remote sensing image data target extraction device, characterized in that: The device comprises: Dataset processing module, used to acquire and preprocess multi-source remote sensing image data; The target fast positioning module is used to obtain the target positioning frame based on the preprocessed data and pre-trained network; The SAM automatic segmentation module is used to obtain boundary feature parameters based on the parameters of the target positioning frame, and to obtain target extraction results based on the target positioning frame and the boundary feature parameters.

8. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Oil spilling area detection method and device, electronic equipment and medium

    CN117197120A

  • Large-area photovoltaic panel detection and extraction method based on deep convolutional neural network

    CN117496124A

  • Cross-scene multi-domain fusion small sample remote sensing target robust identification method

    CN118918476A

  • Remote sensing image segmentation method, system and equipment based on visual large model, and medium

    CN119006833A

  • Object-oriented high-resolution remote sensing image multi-scale segmentation method and system

    WO2022141145A1