Satellite remote sensing image zero sample automatic generation sample set method based on visual model and knowledge feature association
By using the method of correlation between visual models and knowledge features in satellite remote sensing images, the expert knowledge base is automatically segmented and dynamically adjusted, and the problem of zero sample annotation and automatic processing is solved, efficient and accurate sample set generation is achieved, and manual annotation costs are reduced.
Patent Information
- Application Number
- CN202510148768.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art has difficulties in zero-sample annotation and automatic processing in satellite remote sensing images, especially on massive data and large-size images, resulting in high cost of manual annotation and low automatic annotation efficiency.
Using a method based on the correlation between visual model and knowledge feature, the automatic segmentation of visual model and dynamic adjustment of expert knowledge base conditions is automatically generated to reduce the dependence on manual annotation.
It effectively reduces the time cost of manual satellite remote sensing image sample production tasks, improves the efficiency and accuracy of automatic labeling, and can process the data sets that generate new targets at the slice level.
Smart Images

Figure CN120088600A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite remote sensing image sample production, and particularly to a method for automatically generating a sample set of zero samples of satellite remote sensing images based on the association of visual models and knowledge features. Background Art
[0002] With the rapid development of commercial satellite remote sensing, the contradiction between the insufficient data processing ability of massive remote sensing images and the rapid increase in downstream applications has become increasingly prominent. Especially when facing new demands of downstream applications, it has become a huge challenge to label new target type samples from massive data. On the one hand, it takes a certain learning cost for annotators to re-recognize and interpret new targets; on the other hand, there are bottlenecks such as low retrieval efficiency and difficult boundary positioning in locating small targets on large-size satellite remote sensing images.
[0003] At present, to solve the above problems, some semi-automatic annotation applications have emerged to save the cost of data annotation. For example, Labelstudio-SAM and MMRotate-SAM first rely on a target detection model or manual pre-determination of the range, and then use SAM to automatically segment the range area to frame the target, solving the annotation problem of rotated targets.
[0004] However, the above methods are still greatly restricted in the use of satellite remote sensing images. First, annotating targets requires relying on manpower to complete target retrieval, and manually retrieving and annotating targets on remote sensing images is the main factor consuming time cost; second, the current integrated annotation software based on this method completes annotation at the slice level and cannot be directly applied to satellite remote sensing images; third, when facing zero-sample tasks, it is still necessary to pre-annotate samples before application; fourth, for actual zero-sample annotation applications, it is impossible to automatically process massive data. Summary of the Invention
[0005] The present invention aims to solve the technical problems existing in the prior art and provides a method for automatically generating a sample set of zero samples of satellite remote sensing images based on the association of visual models and knowledge features.
[0006] To solve the above technical problems, the technical solution of the present invention is specifically as follows:
[0007] A method for automatically generating a sample set of zero samples of satellite remote sensing images based on the association of visual models and knowledge features, comprising the following steps:
[0008] Step 1: Batch acquisition of remote sensing image data sources;
[0009] Step 2: Automatic segmentation by the visual model;
[0010] Step 3: Dynamic adjustment of the conditions of the expert knowledge base;
[0011] Step 4: Optimize and eliminate non-target results.
[0012] In the above technical solution, step 1 is specifically: automatically read the geographic information parameters and image size of the remote sensing image according to the library functions in the remote sensing image library GDAL, calculate the longitude and latitude coverage range of a single remote sensing image, and determine whether to skip the automatic segmentation of the model for this image by judging whether there is an overlap between the remote sensing image range and the guiding geographic area range.
[0013] In the above technical solution, in step 2, the processing flow of the image encoder in the visual model includes:
[0014] Step 1): Convolve the input image to extract it into a feature map with a size 16 times smaller and the number of channels increased from 3 to 768;
[0015] Step 2): Add position information encoding to the feature map of the input image, and its dimension remains unchanged;
[0016] Step 3): After the feature map completes the position information embedding process, it is processed through the core unit Transformer Block in the Transformer structure. Each Block contains MHSA for capturing the relationship between any positions in the input;
[0017] Step 4): After two layers of convolution, reduce the number of channels to 256 as the output of the image encoder.
[0018] In the above technical solution, step 3 is specifically: judge whether the target type of the currently executed target sample generation task exists in the sample library;
[0019] If the target type does not exist in the current sample library, it is determined that the current task type is a zero-shot generation task, and the target association matching method based on the target morphological features is used to dynamically adjust the conditions of the expert knowledge base;
[0020] If the target type exists in the current sample library, the target association matching method based on the target invariant moment features is used to dynamically adjust the conditions of the expert knowledge base.
[0021] In the above technical solution, in step 3, if the target type does not exist in the current sample library, it is determined that the current task type is a zero-shot generation task, and the target association matching method based on the target morphological features is used to dynamically adjust the conditions of the expert knowledge base, specifically:
[0022] In the zero-shot case, first, according to the conditions of the expert knowledge base, the shallow features of the target are preset; in the remote sensing satellite image, when the resolution is sufficient to interpret the target, the targets of the same size can be uniformly screened and processed at different resolutions.
[0023] In the above technical solution, in step 3, if the target type exists in the current sample library, an object correlation matching method based on the target invariant moment feature is used to dynamically adjust the conditions of the expert knowledge base. Specifically:
[0024] First, select the Daubechies wavelet basis decomposition for the image data;
[0025] Denoise the image after the Daubechies wavelet basis decomposition;
[0026] Based on the histogram of the denoised image and the wavelet transform at each scale, find the approximate image and the detail image at the corresponding scale. By calculating the maximum value of the approximate image at low resolution, determine the number of segmentation regions according to the independent peak width judgment criterion. Track back layer by layer from the threshold selected at low resolution to calculate the corresponding segmentation threshold at the highest resolution. Use this threshold to complete the binarization process of the image to obtain the final effective region, and calculate the region invariant moment feature for this region.
[0027] In the above technical solution, step 3 also includes: after calculating the invariant moment feature of the target effective region, to quantify the similarity degree of the image content, a method based on the vector space model is used to measure the similarity between images.
[0028] In the above technical solution, in step 3, using the method based on the vector space model to measure the similarity between images is specifically as follows:
[0029] Regard the image features as points in the vector space, and quantify the gap between images by calculating the proximity between two points;
[0030] Sample the Euclidean distance as the measurement basis, and its similarity measurement criterion calculation method is as follows:
[0031]
[0032] where, I s represents the invariant moment feature vector of the template, and I t represents the invariant moment feature vector of the current target region. Compare the calculated Euclidean distance with the preset threshold. Those that meet the threshold are the results of image correlation matching and should be retained during the automatic sample production process for the final manual verification of the results.
[0033] In the above technical solution, step 4 is specifically: use user decision-making to further improve the sample quality of the automatically produced sample set and screen out high-fidelity negative samples.
[0034] In the above technical solution, step 4 is specifically as follows: When a certain number of samples have not been extracted to complete the calculation and correlation matching of the invariant moment, the negative samples and difficult samples are eliminated by the method of user decision, and the remaining true and reliable samples are used as the templates for the invariant moment features.
[0035] The present invention has the following beneficial effects:
[0036] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention is a method for automatically generating a sample set for satellite remote sensing images with zero samples by combining a deep learning visual model with expert knowledge features, using the visual model as the basic technology for target positioning and the expert knowledge base as the basis for screening and classification.
[0037] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention, on the one hand, makes good use of the high-precision characteristics of the current visual model method, reducing the time for manual search of targets in large-scale remote sensing images; on the other hand, by combining the method of knowledge association to directly utilize expert knowledge, it reduces the difficulty of target confirmation caused by the arbitrariness of target directions and scales in satellite remote sensing images. By combining the two, the time cost of manual work in the task of satellite remote sensing image sample production is effectively reduced. A data set of new targets can be produced only by processing at the slice level.
[0038] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention effectively fills the gap in the automatic annotation and production of target samples under the condition of zero samples in the field of remote sensing image sample generation, and can be widely applied to the automatic recognition applications in current satellite remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0040] Figure 1 It is a general schematic diagram of the method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention.
[0041] Figure 2 It is a schematic diagram of the step flow of the method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention.
[0042] Figure 3 It is a general schematic diagram of the process of a large visual model for processing natural image data.
[0043] Figure 4Schematic diagram of the processing flow optimized according to the application scenario for the method of automatically generating a sample set with zero samples for satellite remote sensing images based on the association of visual models and knowledge features of the present invention.
[0044] Figure 5 Structural diagram of a network model based on the CNN structure.
[0045] Figure 6 Schematic diagram of the calculation and matching process for the target association and matching method based on target invariant moment features.
[0046] Figure 7 Schematic diagram of the accumulation result for the method of automatically generating a sample set with zero samples for satellite remote sensing images based on the association of visual models and knowledge features of the present invention, using remote sensing image data as the satellite data source and taking the oil depot samples in the port base as an example.
[0047] Figure 8 Schematic diagram of the accumulation result for the method of automatically generating a sample set with zero samples for satellite remote sensing images based on the association of visual models and knowledge features of the present invention, using remote sensing image data as the satellite data source and taking the oil depot samples in the airport base as an example. Detailed implementation manner
[0048] The inventive concept of the present invention is as follows:
[0049] With the continuous improvement of the resolution of satellite remote sensing images, the interpretability of each unit pixel in the remote sensing image is continuously improved, enhancing the boundary clarity between targets and enabling the regional bounding of smaller-sized targets in the remote sensing image. In response to the above situation improvement, according to different requirements of accuracy and time, the method of the present invention constructs a semantic segmentation function service based on a large visual model and a visual convolutional model; performs image segmentation processing on the input remote sensing image through the visual model, and saves the result information of the visual model processing; after the result information is output, with the help of the expert knowledge base system, verifies the feature information of the specified type of image area according to the composite expert knowledge, retains the qualified feature areas, and outputs them as slice result files; in the uniformly generated result files at the slice level, the method of the present invention supports user decision-making, can traverse and browse the coarsely verified slice set, and decides whether to retain the slice by simply judging whether the target correctly exists within the slice range, greatly reducing the search cost and annotation cost of manual annotation; on the determined target slices to be retained, templates can be automatically generated and compared with the newly coarsely screened slices to further improve the sample accuracy.
[0050] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of visual models and knowledge features of the present invention, an automatic segmentation service based on a visual model, can automatically perform zero-sample-level automatic annotation on a set of satellite remote sensing images input into the system through dynamically adjustable expert prior knowledge conditions, and finally output the annotation results in the form of target slices and annotations. The annotator only needs to judge the retention problem of the target at the slice level.
[0051] The present invention will be described in detail below with reference to the accompanying drawings.
[0052] As Figure 1 and 2 shown, the method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of visual models and knowledge features of the present invention includes the following steps:
[0053] Step 1: Batch obtain remote sensing image data sources.
[0054] In the data source part, to ensure the effect of subsequent algorithm segmentation of the target, it is required that the resolution of remote sensing images all reaches within sub-meter level. Usually, the geographical range area of the remote sensing image data set cannot be fully counted. When it contains a large number of geographical locations and is difficult to manually screen and distinguish, the system applicable to the method of the present invention provides a geographical area screening function, which can set specified geographical area conditions before the model automatic segmentation process. For a large number of remote sensing images, the images that do not meet the geographical area conditions can be automatically skipped during the execution process. The specific implementation is to automatically read the six geographical information parameters and image size of the remote sensing image according to the library functions in the GDAL of the remote sensing image library, calculate the longitude and latitude coverage range of a single remote sensing image, and determine whether to skip the model automatic segmentation of the image by judging whether there is an overlap between the remote sensing image range and the guided geographical area range.
[0055] According to actual applications, taking the data standard scene download platform of the domestic constellation "Jilin-1" as an example, the resolution of its constellation is generally better than 1 meter, and a set of coarse-screened satellite image data containing multiple target points can be obtained on the global base map by means of map point selection and custom area range. Since the method is mainly aimed at the case where the target scene is zero-sample, in order to ensure the quality of the initially screened samples, the side-sway angle during satellite image imaging is set to be better than 20°, and the image cloud cover is lower than 20%.
[0056] Step 2: Automatic segmentation by the visual model.
[0057] The size of satellite remote sensing images far exceeds that of ordinary natural images. In the method of visual model segmentation, the time cost is the most time-consuming part in the implementation process of the whole method. Considering the performance of hardware devices and aiming to balance efficiency and time, the method of the present invention integrates a semantic segmentation model based on a large vision model and a semantic segmentation model based on a CNN structure. Under the same conditions, the accuracies of the two methods are similar, but the processing efficiency of the semantic segmentation model based on the CNN structure is dozens of times that of the large model structure. Therefore, it is recommended to give priority to using the method based on the CNN structure to maximize work efficiency.
[0058] The following will separately introduce the characteristic structures of the two methods.
[0059] Generally speaking, the semantic segmentation model based on a large vision model consists of three parts: an image encoder, a prompt encoder, and a mask decoder. The image encoder part is responsible for mapping the remote sensing image to be segmented into the image feature space. The prompt encoder part is responsible for mapping the input prompt information into the prompt feature space. The mask decoder is responsible for integrating the outputs of the previous two decoders and decoding the final segmented mask result from the separately output discrete high-dimensional data. The general process of the large vision model for processing natural image data is as Figure 3 shown.
[0060] In the professional field, usually a prompt encoder is selected to provide key guiding information for the final mask decoder, and providing multiple prompt methods does help to improve the segmentation accuracy. However, in the context of zero-shot annotation of massive remote sensing data, the workload of completing manual feedback prompts for a single scene of remote sensing data far exceeds that of a single scene of natural image data, which will greatly increase the cumbersome human-computer interaction cost and deviate from its original usage scenario. Therefore, in the method of the present invention, it is innovatively composed of only the image encoder and the mask decoder to prioritize efficiency. To prevent the sudden increase of negative samples caused by direct segmentation, on the one hand, the method of the present invention uses a rich expert knowledge base to initially screen the segmentation result set and eliminate the results that do not conform to the expert data knowledge; on the other hand, after the initial screening, user decision intervention is added, which can retain the user decision to generate a target template for the result, complete subsequent template matching, and continuously iterate the method to improve the accuracy of data accumulation in the zero-shot scenario. After optimization according to the application scenario, its processing process is as Figure 4 shown.
[0061] In the large vision model of the method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association between visual models and knowledge features of the present invention, the processing process of the image encoder in the visual model is generally divided into four steps:
[0062] Step 1): The input image is convolved to extract a feature map with a size 16 times smaller and the number of channels increased from 3 to 768;
[0063] Step 2): Add position information encoding to the feature map of the input image, and its dimension remains unchanged. The position information encoding belongs to a given parameter matrix, which is initialized to 0 during the network training process, but the inference process is determined;
[0064] Step 3): After the feature map completes the position information embedding process, it will be processed by the core unit Transformer Block in the Transformer structure. Each Block contains MHSA (multi-head self-attention mechanism) to capture the relationship between any positions in the input;
[0065] Step 4): After two layers of convolution, the number of channels is reduced to 256 as the output of the image encoder.
[0066] In the existing large visual model structure, due to the simplification of the prompt encoder, the original mask decoder no longer repeatedly merges the output of the image encoder and the prompt encoder to update the generated result, but chooses to replace the prompt encoder with a fixed vector, and finally obtains the segmentation result based on the characteristics of the image itself. Completing the image segmentation task based on a large model has advantages in flexibility and versatility, but the Transformer-based architecture involves a large amount of computing resources and takes a long time to infer on high-resolution images. In the context of the need for efficient processing of massive remote sensing images, there is still room for improvement.
[0067] In order to greatly improve the reasoning speed, a semantic segmentation model based on deep learning convolutional structure is adopted, which is essentially an algorithm using CNN structure applied to segmentation tasks. However, unlike the end-to-end Transformer structure of the large model, the CNN-based method introduces certain artificial prior structure designs, such as local connections of convolution and object allocation strategies related to the receiving domain, which greatly reduces the computational redundancy of the original Transformer structure in the large model on general perception tasks, allowing the model to achieve segmentation effects close to those of the large model with fewer parameters.
[0068] The network model based on the CNN structure is divided into three parts, adopting the structure of the classic backbone layer, neck layer, and head layer. The backbone layer network is used to extract features from the input image. This layer network consists of three structural units: downsampling convolution, coupled fusion, and spatial pyramid pooling. The downsampling convolution will increase the number of channels of the feature map while decreasing the resolution of the feature map; the coupled fusion is a structural unit for deep feature extraction, which ensures that the sizes of the input and output feature maps remain unchanged; the spatial pyramid pooling uses multiple consecutive max-pooling operations to splice the input feature map and multi-scale features in the channel dimension. The neck layer network is used to fuse the features of the feature map output by the backbone network. Specifically, in the backbone layer network, three feature maps of different scales from the third layer to the fifth layer are spliced by upsampling. There are two differences between this layer and the backbone layer network in the coupled fusion structural unit. One is that the input and output channel numbers of the feature maps before and after this structure will change to keep the channel number of the same feature map output by the backbone layer network; the other is that the neck layer does not contain a residual structural unit, and the number of basic units in it is only one. The head layer network is responsible for outputting the final result, and this layer will output the feature map after the original segmentation.
[0069] The structure of the network model based on the CNN structure is as Figure 5 shown. In the structure, the downsampling convolution units of the backbone layer and the neck layer both represent a structural unit composed of a convolution, pooling, and activation unit, while the downsampling convolution unit of the head layer adds a convolution unit on this basis. As the final output, the result of the head layer contains the feature map of the predicted regional position, category, and corresponding mask coefficient.
[0070] Step 3: Dynamically adjust the conditions of the expert knowledge base.
[0071] After the visual preprocessing of satellite remote sensing images is completed based on the visual model, since it is a pixel-level segmentation process, a large number of segmentation results based on satellite remote sensing images will be generated. To eliminate a large number of pixel-level non-target negative samples from the segmentation result set, the method for automatically generating a sample set with zero samples of satellite remote sensing images based on the association between the visual model and knowledge features of the present invention provides two sample rough screening methods based on knowledge features with different granularities, namely the target association matching method based on target morphological features and the target association matching method based on target invariant moment features.
[0072] First, determine whether the target type of the currently executed target sample generation task exists in the sample library.
[0073] (1) If the target type does not exist in the current sample library, the current task type is identified as a zero-shot generation task. In the zero-shot case, it is necessary to first set the shallow features of the target according to the conditions of the expert knowledge base, including its length range, width range, color gamut range, etc. Among the pre-set shallow features, for example, the length and width in the size features are set according to the actual size of the target, which is different from the pixel scale features in natural images. In remote sensing satellite images, the pixel size can be converted into the actual size according to the geographical information of the image, and it is possible to effectively screen and process targets of the same size at different resolutions when the resolution is sufficient to interpret the target. Taking the oil depot in the port base as an example, its size is about 50 meters. There will be a large pixel scale difference at 0.5-meter resolution and 0.75-meter resolution, but through the combination of the target position and geographical information conversion, its true size can be accurately restored. The calculation method is as follows:
[0074] GeoTrans = (leftup_x, pixel x , rotate x , leftup y , pixel y , rotate y )
[0075] The six parameters in the affine geographic transformation parameters respectively represent the abscissa of the upper left corner, the horizontal resolution, the horizontal rotation parameter, the ordinate of the upper left corner, the vertical resolution, and the vertical rotation parameter.
[0076] The relevant formula for converting pixel coordinates to projection coordinates is as follows. In fact, dGeoTrans represents the affine parameters, X pixel and Y line respectively represent the horizontal and vertical coordinates of the pixel, and X p and Y p respectively represent the longitude and latitude information of the finally converted projection coordinates.
[0077] X p = dGeoTrans[0] + X pixel * dGeoTrans[1] + Y line * dGeoTrans[2]
[0078] Y p = dGeoTrans[3] + X pixel * dGeoTrans[4] + Y line * dGeoTrans[5]
[0079] After completing the conversion of pixel coordinates to projection coordinates according to the affine parameter matrix, the original projection coordinate system and the target projection coordinate system are constructed according to the projection coordinate system in the remote sensing image, and the final converted geographical coordinates of the pixel coordinates can be obtained. According to the geographical coordinates, the actual distance between two points in the real geographical space can be directly calculated using the encapsulated functions in the geodesic library to obtain the scale feature of the target. In addition to the scale feature, the color gamut range can also be used as a feature (typically including RGB channel pixel values, hue, saturation, etc.). Taking the oil depot in the port base as an example, the RGB pixel value range is between 220 and 255. After segmenting the target area, the average pixel value inside the area can be obtained. If it meets the pixel range, it can be retained. When screening the scale information in the shallow features, the negative samples that do not meet the scale information will be initially screened out.
[0080] (2) If the target type exists in the current sample library, that is, after completing a small number of sample checks, the sample library already contains a very small number of target type samples. For the situation where the samples are extremely few, in order to further improve the accuracy of target sample generation, a target association and matching method based on target invariant moment features can be adopted. Its calculation and matching process is as Figure 6 shown.
[0081] After the segmentation is initially completed, the targets will be temporarily retained in the form of the minimum bounding rectangle. Therefore, there may be some background information on non-rectangular objects. To ensure the accuracy of the calculated features, further processing will be carried out on this basis. First, select the Daubechies wavelet basis decomposition for the image data, and perform wavelet decomposition according to scales of 2, 4, and 8 respectively. For the decomposed image, use the non-linear soft threshold denoising in the wavelet domain to denoise the image. The histogram of the denoised image and the wavelet transform at each scale can be used to calculate the approximate image and the detail image at the corresponding scale. By calculating the maximum value of the approximate image at low resolution and determining the number of segmentation regions according to the independent peak width judgment criterion, track layer by layer from the threshold selected at low resolution to calculate the corresponding segmentation threshold at the highest resolution, and use this threshold to binarize the image to obtain the final effective region, and calculate the region invariant moment feature for this region.
[0082] Moments in statistics reflect the distribution of random variables. Regarding the gray level of an image as a multi-dimensional density distribution function, the moments extracted from it are a highly concentrated image feature, which has translational, gray level, scale, and rotational invariance within the region range. The invariant moment feature has good adaptability to the arbitrariness of the direction and scale of remote sensing targets under satellite remote sensing images. For an image, the definition formula of the moment is as follows:
[0083] For an image with a gray level distribution of f(x, y), the definition of its p+q order moment is as follows:
[0084] mpq = ∫∫ x p y q f(x, y) dxdy p, q = 0, 1, 2, …
[0085] The central moment of order p + q is defined as follows:
[0086] μ pq = ∫∫ (x - x 0 ) p (y - y 0 ) q f(x, y) dxdy
[0087] The normalized central moment is defined as follows:
[0088]
[0089] In the method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association between a visual model and knowledge features of the present invention, seven invariant moment features are constructed using the second - order and third - order central moments as a set of features:
[0090] I = (I 1 , I 2 , I 3 , I 4 , I 5 , I 6 , I 7 )
[0091] The calculation formulas for the features are as follows:
[0092] I 1 = h 20 + h 02
[0093]
[0094] I 3 = (h 30 - 3h 12 ) 2 +(3h 21 - h 03 ) 2
[0095] I 4 = (h 30 + h 12 ) 2 +(h 21 + h 03 ) 2
[0096] I 5 = (h 30 - 3h 12 )(h30 +h 12 )[(h 30 +h 12 ) 2 -3(h 21 -h 03 ) 2 +(3h 21 -h 03 )(h 21 +h 03 )[3(h 30 +h 12 ) 2 -(h 21 +h 03 ) 2
[0097] I 6 =(h 20 -h 02 )[(h 30 +h 12 ) 2 -(h 21 +h 03 ) 2 +4h 11 (h 30 +h 12 )(h 21 +h 03 )I 7 =(3h 21 -h 03 )(h 30 +h 12 )[(h 30 +h 12 ) 2 -3(h 21 +h 03 ) 2 -(h 30 -3h 12 )(h 21 +h 03 )[3(h 30 +h 12 ) 2 -(h 21 +h 03 ) 2
[0098] After completing the calculation of the invariant moment features of the target effective area, in order to quantify the similarity of the image content, the method of automatically generating sample sets based on the visual model and knowledge feature association of satellite remote sensing images using zero samples adopts a method based on the vector space model to measure the similarity between images. The image features are regarded as points in the vector space, and the distance between the images is quantified by calculating the proximity between the two points. The sampling Euclidean distance is used as the basis for measurement, and the similarity measurement criterion is calculated as follows:
[0099]
[0100] Among them, I s Represents the invariant moment eigenvector of the template, I t The invariant moment feature vector representing the current target area is used to compare the calculated Euclidean distance with the preset threshold. The one that meets the threshold is the result of image association matching and should be retained during the automatic sample production process for the final manual verification result.
[0101] Step 4: Optimize and eliminate non-target results.
[0102] The purpose of the non-target result optimization and screening step is to further improve the sample quality of the automatically produced sample set by user decision, and to screen out negative samples with high simulation degree. The method of the present invention determines the target area through high-precision segmentation based on the visual model. Only shallow features cannot guarantee that there are no wrong category samples in the sample library in the initial screening stage. In the early stage of the method, when a certain number of samples are not extracted to complete the calculation and association matching of the feature invariant moment, the negative samples and difficult samples can be eliminated through the user decision method, and the retained real and reliable samples can be used as templates for the invariant moment features. With the brief intervention of user decisions, the accuracy of the sample production method will be able to be continuously and reliably improved.
[0103] At the user decision level, the operation method of the present invention is to arrange the current preliminary screening sample set in the form of slices, and the user only needs to choose to keep or eliminate while switching different slices. User decision intervention does not need to refer to the traditional sample labeling task, which needs to be labeled to the specified workload before it can be completed. It only needs to confirm that a small number of samples are retained before the operation of optimizing and screening non-target results can be completed at any time. Subsequent processing will return to the dynamic adjustment of the expert knowledge base in step 3. With the continuous accumulation of real samples, the accuracy of the expert knowledge base will be dynamically and continuously improved.
[0104] The method for automatically generating sample sets based on zero samples of satellite remote sensing images based on visual models and knowledge feature association of the present invention selects remote sensing image data as the satellite data source, and the satellite resolution used is sub-meter level. The above method takes the port base oil depot samples and the airport base oil depot samples as examples, and the accumulated result samples are as follows: Figure 7and 8 as shown
[0105] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention uses the visual model as the basic technology for target positioning and the expert knowledge base as the basis for screening and classification, and realizes the method for automatically generating a sample set with zero samples for satellite remote sensing images by combining the deep learning visual model and expert knowledge features.
[0106] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention, on the one hand, makes good use of the high-precision characteristics of the current visual model method, reducing the time for manual search of targets in large-scale remote sensing images; on the other hand, by combining the method of knowledge association to directly utilize expert knowledge, it reduces the difficulty of target confirmation caused by the arbitrariness of target directions and scales in satellite remote sensing images. By combining the two, the time cost of manual work in the task of satellite remote sensing image sample production is effectively reduced. A data set of new targets can be produced only by processing at the slice level.
[0107] The method for automatically generating a sample set with zero samples for satellite remote sensing images based on the association of a visual model and knowledge features of the present invention effectively fills the blank of automatic annotation and production of target samples under zero-sample conditions in the field of remote sensing image sample generation, and can be widely applied to automatic recognition applications in current satellite remote sensing images.
[0108] Obviously, the above embodiments are only examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for automatically generating sample sets from zero-sample satellite remote sensing images based on visual models and knowledge feature association, characterized in that: The following steps are involved: Step 1: Batch obtain remote sensing image data sources; Step 2: Automatic segmentation of visual model; Step 3: Dynamic adjustment of expert knowledge base conditions; Step 4: Optimize and eliminate non-target results.
2. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 1 is characterized in that: Step 1 is as follows: the geographic information parameters and image size of the remote sensing image are automatically read according to the library function in the remote sensing image library GDAL, the longitude and latitude coverage of a single remote sensing image is calculated, and whether the model automatic segmentation skips the image is determined by judging whether the remote sensing image range overlaps with the guidance geographic area range.
3. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 1 is characterized in that: In step 2, the processing flow of the image encoder in the visual model includes: Step 1): The input image is convolved to extract a feature map with a size 16 times smaller and the number of channels increased from 3 to 768; Step 2): Add position information encoding to the feature map of the input image, and keep its dimension unchanged; Step 3): After the feature map completes the position information embedding process, it is processed by the core unit Transformer Block in the Transformer structure. Each Block contains MHSA to capture the relationship between any positions in the input; Step 4): After two layers of convolution, the number of channels is reduced to 256 as the output of the image encoder.
4. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 1 is characterized in that: Step 3 is as follows: Determine whether the target type of the target sample generation task currently being executed exists in the sample library; If the target type does not exist in the current sample library, the current task type is considered to be a zero-sample generation task, and the target association matching method based on target morphological features is used to dynamically adjust the expert knowledge base conditions; If the target type exists in the current sample library, the target association matching method based on the target invariant moment feature is used to dynamically adjust the expert knowledge base conditions.
5. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 4 is characterized in that: In step 3, if the target type does not exist in the current sample library, the current task type is considered to be a zero-sample generation task, and the target association matching method based on target morphological features is used to dynamically adjust the expert knowledge base conditions, specifically: In the case of zero samples, the shallow features of the target are pre-set according to the expert knowledge base conditions; Remote sensing satellite images can effectively screen and process targets of the same size at different resolutions when the resolution is sufficient to interpret the target.
6. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 4, characterized in that: In step 3, if the target type exists in the current sample library, the target association matching method based on the target invariant moment feature is used to dynamically adjust the expert knowledge base conditions, specifically: First, Daubechies wavelet basis decomposition is selected for image data; De-noising the image after Daubechies wavelet decomposition; Based on the histogram of the denoised image and the wavelet transform at each scale, the approximate image and detail image at the corresponding scale are obtained. The maximum value of the approximate image at low resolution is calculated, and the number of segmentation areas is determined according to the independent peak width judgment criterion. The threshold selected at low resolution is traced back layer by layer to calculate the corresponding segmentation threshold at the highest resolution. The image is binarized using this threshold to obtain the final valid area, and the regional invariant moment feature calculation is completed for this area.
7. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 6 is characterized in that: Step 3 also includes: after completing the calculation of the invariant moment features of the target effective area, in order to quantify the similarity of the image contents, a method based on a vector space model is used to measure the similarity between the images.
8. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 7 is characterized in that: In step 3, the method based on the vector space model to measure the similarity between images is as follows: Image features are considered as points in vector space, and the distance between images is quantified by calculating the proximity between two points; The sampling Euclidean distance is used as the measurement basis, and the similarity measurement criterion is calculated as follows: Among them, I s Represents the invariant moment eigenvector of the template, I t The invariant moment feature vector representing the current target area is used to compare the calculated Euclidean distance with the preset threshold. The one that meets the threshold is the result of image association matching and should be retained during the automatic sample production process for the final manual verification result.
9. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 1, characterized in that: Step 4 is as follows: using user decisions to further improve the quality of automatically generated sample sets and filter out negative samples with high simulation degree.
10. The method for automatically generating sample sets from zero samples of satellite remote sensing images based on visual model and knowledge feature association according to claim 9, characterized in that: Step 4 is specifically as follows: when a certain number of samples are not extracted to complete the calculation and association matching of the feature invariant moment, the negative samples and difficult samples are eliminated through the user decision method, and the remaining real and reliable samples are used as templates for the invariant moment features.