Space-time correlation domain auxiliary labeling and quality control method for interpreting remote sensing image
Through the combination of remote sensing image data standardization and multiple annotation methods, the problems of low labeling efficiency and low quality of large and wide remote sensing image data are solved, and the acquisition of fast and accurate labeling and high-quality labeling data are achieved, meeting the needs of intelligent interpretation models.
Patent Information
- Application Number
- CN202510189676.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology is difficult to quickly and efficiently label and quality control large wide remote sensing image data, resulting in low labeling efficiency and low quality, which cannot meet the demand of intelligent interpretation models for a large number of high-quality samples.
Through the standardization of remote sensing image data, a point package data set is constructed, and combined with latitude and longitude information and time series, interactive and automated annotation methods are provided, including manual annotation, interactive data annotation and automated data annotation. At the same time, the labeling quality control method is adopted for time-spatial data comparison and target domain data comparison to ensure the accuracy and consistency of the labeling data.
It realizes fast and accurate labeling of remote sensing image data, reduces the labeling threshold and learning cost, improves the labeling efficiency and quality, and meets the demand for high-quality samples of intelligent interpretation models.
Smart Images

Figure CN120107762A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite remote sensing image interpretation, and in particular to a spatiotemporal correlation domain auxiliary annotation and quality control method for interpreting remote sensing images. Background Art
[0002] With the development of satellite remote sensing technology and deep learning technology, the intelligent interpretation of satellite remote sensing images using deep learning technology has become a major research direction. In order to achieve the intelligent interpretation of remote sensing images, deep learning technology needs to use remote sensing image sample data to train the intelligent interpretation model, and the number and quality of samples will directly affect the training effect of the intelligent interpretation model. In practice, in order to ensure that the model has sufficient generalization ability and detection accuracy, it is necessary not only to use a large number of rich training samples, but also to require samples with sufficiently high annotation quality.
[0003] In the highly professional field of remote sensing interpretation, such as the classification of targets such as aircraft and ships, the labelers are usually required to have professional interpretation knowledge and be able to distinguish the subtle differences between targets. This is not friendly to labelers who are new to the field and requires a long period of guidance and training. However, with the rapid development of satellite remote sensing technology and the continuous increase in the number of remote sensing satellites in orbit, a huge amount of remote sensing image data is generated every day. How to make full use of these remote sensing image data, realize the rapid interpretation of remote sensing image data, lower the threshold for labeling remote sensing data, ensure the quality of remote sensing data labeling, and meet the needs of intelligent interpretation models for a large number of high-quality samples is an urgent problem to be solved.
[0004] At present, there are many data annotation methods and tools on the market, which can be divided into two categories. One is the data annotation tools for traditional images. Such tools do not have the ability to process large-width remote sensing images, cannot load remote sensing image data, and cannot perform remote sensing image data annotation work; the other is the interpretation tools for remote sensing image data. Such tools have powerful interpretation functions, but lack intelligence, have many redundant operations, and are not highly specialized. The learning cost of such tools is high, and the complex operations reduce the annotation efficiency and are more likely to produce incorrect annotations. At the same time, they do not have the quality control capabilities of the annotated data and cannot guarantee the accuracy of data annotation. In summary, the existing image data annotation methods have the following problems: (1) There is a lack of rapid annotation methods and tools for large-scale remote sensing image data. Existing annotation tools are mainly for small-scale images and cannot load and process large-scale remote sensing image data. Remote sensing image interpretation tools are complex to operate, lack intelligence, and are not highly specialized. (2) The labeling of remote sensing image data relies on human experience and knowledge, which places high demands on labelers. Labelers need a long time and energy to learn and practice, and cannot quickly get started; (3) Remote sensing image data annotation relies on manual operation, which is complex and prone to more erroneous annotations. There is also a lack of effective methods for annotation quality control. Summary of the invention Therefore, the technical problem to be solved by the present invention is to overcome the defects existing in the above-mentioned prior art, thereby providing a spatiotemporal correlation domain auxiliary annotation and quality control method for interpreting remote sensing images.
[0005] A spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images comprises the following steps: S1. Based on the remote sensing image data normalization method, the remote sensing image data is normalized, and the normalized remote sensing image data set is realized in the form of point packages; S2. Select any standardized original remote sensing image and further select the annotation form; S3. Select an appropriate annotation tool for annotation according to the annotation form selection result of step S2; wherein the annotation tools include: manual annotation and intelligent annotation method; S4. Fine-tune the annotation result of step S3; S5. Determine whether the labeling of all remote sensing images is completed; if so, complete the labeling; if not, repeat steps S2-S4 until the judgment result of step S5 is yes.
[0006] Preferably, the remote sensing image data normalization method is specifically as follows: combining remote sensing image data units belonging to the same point to construct a remote sensing image point package data set; combining multiple remote sensing image point package data sets to complete remote sensing image data normalization; Among them, the smallest unit of remote sensing image data annotation is the remote sensing image data unit; the remote sensing image point package dataset is constructed based on the remote sensing image data unit; the point package dataset uses the latitude and longitude information of the remote sensing image to screen multiple remote sensing images in the same longitude and latitude range, and form a spatially fixed and time-series related remote sensing image data set.
[0007] Preferably, the intelligent annotation method specifically includes an interactive data annotation method and an automatic data annotation method; Among them, the interactive data annotation method is a method of combining manual operation with the general interpretation model of the encoder-decoder architecture; The automatic data annotation method uses a target model constructed based on deep learning interpretation principles for a specific target to perform panoramic reasoning on the original remote sensing image to be processed, and converts the reasoning results into target annotations.
[0008] Preferably, the interactive data annotation method is a method of combining manual operation with a general interpretation model of an encoder-decoder architecture, specifically including the following process: Grid the original remote sensing images; Manual operation is used to extract the position of the target of interest in the original remote sensing image. The specific operations include: positive and negative point selection and positive and negative box selection. Among them, positive point selection and positive box selection both indicate the selection of the target of interest, and negative point selection and negative box selection both indicate the exclusion of the current area. An interactive data labeling task is completed using a universal interpretation model of the encoder-decoder architecture, which specifically includes the following processes: the remote sensing image to be processed is input into the encoder, and the encoder outputs the encoding result and caches it; the encoding result and the manually selected coordinates are input into the decoder, and the decoder extracts the target at the specified position, and the extraction result is used as the target boundary; according to the labeling task requirements, the extraction result is converted into the labeling style to complete an interactive data labeling task.
[0009] Preferably, the automated data annotation method specifically includes the following process: Select the corresponding target model according to the specific labeling task; The remote sensing image to be inferred is processed by partitioning to complete the detection and annotation transformation of the entire remote sensing image; the specific partitioning processing method used is: window sliding reasoning; After the window completes the whole scene reasoning, the detection results are converted into absolute positions on the original remote sensing image according to the relative position of the window, non-maximum suppression is performed, redundant detection results in overlapping areas are removed, and finally the detection results are converted into target annotations.
[0010] Preferably, step S4 specifically includes: after converting the detection result into the target annotation, verifying the automatically labeled result and making fine adjustments according to the verification result; Among them, fine-tuning includes: supplementing missing annotations, adjusting annotation ranges and modifying annotation properties.
[0011] Preferably, it also includes a target type auxiliary interpretation method for assisting the labeling personnel to quickly interpret the target type; specifically: Calculate the length and width of the target based on the latitude and longitude information of the remote sensing image, and obtain all suspected types of the labeled target based on the current annotation task type and point information: Among them, it is necessary to pre-build a corresponding target knowledge base according to the specific labeling task type; the target knowledge base integrates the target's own attributes and statistical characteristics, and further constructs a professional knowledge system of multi-source information fusion for specific targets, so that when creating a target labeling box, the labeling personnel can compare the partial attributes of the label with the information stored in the target knowledge base to extract all suspected types; the labeling personnel comprehensively judge the specific type of the target based on the remaining attributes of the label.
[0012] A method for auxiliary annotation and quality control of time-space correlation domain for interpreting remote sensing images, including the method for auxiliary annotation of time-space correlation domain for interpreting remote sensing images, and also including: S6. Implement labeling quality control methods; S7. Correct incorrect annotations; S8. Determine whether all the annotations are correct and meet the specifications. If so, archive the data; if not, repeat steps S6-S7 until the result of step S8 is yes.
[0013] Preferably, a method for quality control of annotation is performed, specifically including two methods: comparison of spatiotemporal domain data and comparison of target domain data; The spatiotemporal data comparison is a quality control method for comparing multi-phase annotation results of the same area, and utilizes the latitude and longitude information of remote sensing images and the temporal correlation of the previous and next phases to achieve the previous and next phase comparison of the local area.
[0014] The target domain data comparison is a quality control method implemented by using target differences. All target annotations are cut out and displayed in groups according to annotation type and attribute information, so as to realize intuitive and rapid inspection of all annotations.
[0015] The technical solution of the present invention has the following advantages: 1. The present invention standardizes the remote sensing image data annotation task, fully considers the actual annotation task requirements of remote sensing images, and constructs a standardized remote sensing image data set in the form of a point package. During the annotation process, the data can be quickly and accurately interpreted and annotated by combining point information, longitude and latitude coordinates, time series, etc.
[0016] 2. Aiming at the remote sensing image data annotation task, the present invention provides two types of annotation forms for regular targets and irregular targets: rectangular box and polygonal box annotation forms, which have all necessary annotation operations, including a series of convenient operations such as creation, modification, deletion, and movement; in order to further improve the efficiency of data annotation and make full use of the advantages of deep learning models, an intelligent data annotation method is provided, including interactive data annotation and automatic data annotation; 3. The present invention proposes a target assisted interpretation method. For specific task types, the target's own characteristics and statistical characteristics are integrated to build a professional knowledge base of multi-source information fusion about the target. The annotation process will extract information such as target size, location, task type, time dimension, etc., compare with the target knowledge base, extract all suspected types, and then manually identify the target type based on multi-angle views, key identification features and other information. This method significantly improves the speed and accuracy of interpretation and greatly reduces the threshold for target interpretation; 4. Aiming at the problem that the intelligent interpretation model of remote sensing images is easily affected by the target imaging quality, the present invention classifies the target imaging quality. Subsequently, target samples with different imaging qualities can be used for model training according to different task requirements. The detection effect of the model on different target imaging qualities can also be analyzed, and the model can be optimized in a targeted manner. 5. The present invention provides a method for labeling quality control, including two aspects: spatiotemporal data comparison and target domain data comparison. The spatiotemporal data comparison method can reduce the cases of missed labels and wrong labels, and can also effectively observe the changes of targets in time series, effectively improving the speed and accuracy of target interpretation; the spatiotemporal data comparison method will extract all target labels and display them in groups, so that the labeling personnel can quickly detect the target labels with differences, and then locate and correct the wrong labels. Through this labeling quality control method, the quality of data labeling is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a general schematic diagram of the method provided in the second embodiment of the present invention; Figure 2 Normalize the remote sensing image data of the present invention; Figure 3 Marking style for the object of the present invention; Figure 4 It is a schematic diagram of the process flow of the interactive target labeling method of the present invention; Figure 5 A schematic diagram of gridding a large-width remote sensing image of the present invention; Figure 6 A schematic diagram of the interactive operation and adaptive grid merging of the present invention; Figure 7 A schematic diagram of the automatic target labeling method of the present invention; Figure 8 It is a schematic diagram of the window sliding reasoning of the present invention; Fig. 9 This is a schematic diagram of the target auxiliary interpretation method of the present invention; Fig.10 It is a schematic diagram of the window locking effect of the time-space domain data comparison of the present invention; Fig.11 It is a schematic diagram of target domain data comparison of the present invention. DETAILED DESCRIPTION The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0020] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0021] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0022] Example 1 A spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images comprises the following steps: S1. Based on the remote sensing image data normalization method, the remote sensing image data is normalized, and the normalized remote sensing image data set is realized in the form of point packages; S2. Select any standardized original remote sensing image and further select the annotation form; S3. Select an appropriate annotation tool for annotation according to the annotation form selection result of step S2; wherein the annotation tools include: manual annotation and intelligent annotation method; S4. Fine-tune the annotation result of step S3; S5. Determine whether the labeling of all remote sensing images is completed; if so, complete the labeling; if not, repeat steps S2-S4 until the judgment result of step S5 is yes.
[0023] Specifically: For step S1: S1-1. Standardization of remote sensing image data; S1-2. Remote sensing image point package data set; S1-3. Remote sensing data loading and list display; It should be noted that the remote sensing image data normalization method is: combining remote sensing image data units belonging to the same point to construct a remote sensing image point package data set; combining multiple remote sensing image point package data sets to complete remote sensing image data normalization; The main purpose of normalizing remote sensing image data in this embodiment is to reasonably organize remote sensing image data sets and store remote sensing image data sets processed once in a unified form and specification. Remote sensing image data belongs to raster data and has longitude and latitude information. Using this feature, a normalized remote sensing image data set is constructed in this embodiment.
[0024] Specifically, constructing a remote sensing image point package data set enables the standardized remote sensing image data set to be implemented in the form of a point package. In this embodiment, the point package refers to using the longitude and latitude information of the remote sensing image to screen out multiple remote sensing images in the same longitude and latitude range, and forming a remote sensing image data set with fixed spatial time series correlation. Using this remote sensing image point package data set, when marking static targets, based on the characteristics that the target has fixed longitude and latitude information and the difference between the previous and next phases is small, the previous phase scene is quickly copied and marked to the current foreground to complete the target marking; when marking dynamic targets, the changes in the target have time series correlation, and it is easier to distinguish the target type based on the changes in the previous and next phases relative to the target, especially in the case of poor imaging, the target features are difficult to distinguish. At this time, the previous and next phases are used for auxiliary interpretation, which can judge the target faster and more accurately.
[0025] In this embodiment, the smallest unit of remote sensing image data annotation is called remote sensing image data unit, and the remote sensing image data unit is composed of remote sensing image TIF file, annotation vector SHP file, image metadata file, image quick view and other files. Among them, the remote sensing image TIF file is a remote sensing image, which belongs to raster data and has longitude and latitude information; the annotation vector SHP file is an annotation file, which stores the longitude and latitude, type, attribute and other information of the annotated target. The remote sensing image TIF file and the annotation vector SHP file are superimposed to realize the annotation and viewing of all targets; the image metadata file records various parameter information of the current remote sensing image imaging.
[0026] By combining remote sensing image data units belonging to the same point, a remote sensing image point package data set is constructed, and by combining multiple remote sensing image point package data sets, the remote sensing image data normalization is completed. The effect of remote sensing image data normalization is as follows: Figure 2 shown.
[0027] The vector annotation SHP file in the remote sensing image data unit stores the necessary attribute information to accurately describe the target. Its main attribute information and meaning are shown in the following table.
[0028] Table 1 Label vector attribute information table
[0029] For step S2: There are two types of target annotation. One is rectangular annotation, which is mainly used for regular targets such as aircraft and ships. These targets are densely distributed and have arbitrary directions. The traditional approach is to use horizontal rectangular boxes to represent them, but this will introduce too much background information and make it difficult to distinguish densely distributed targets. To address this problem, existing remote sensing target detection algorithms mainly use rotating rectangular boxes for representation. The other type of annotation is polygonal annotation, which is mainly used for irregular targets such as forests, fields, and water bodies. By annotating along the outer contour of the target, the target is distinguished from the background. It is widely used in semantic segmentation, instance segmentation, and other fields.
[0030] In view of these two main types of data annotation, this embodiment summarizes the target convenient annotation method and provides the annotation functions necessary for the data annotation process, including creating annotations, adjusting positions and boundaries, adjusting rotation angles, modifying categories and attributes, showing and hiding annotations, copying annotations, and deleting annotations. Figure 3 shown.
[0031] For step S3: The target convenient labeling method provides the necessary operations required for the target labeling process, but the labeling process mainly relies on manual operation, which is inefficient. In view of this problem, in this embodiment, the advantages of deep learning technology are brought into play, and the general interpretation model accumulated historically is fully utilized to construct an intelligent labeling method. Intelligent labeling methods include two categories, one is an interactive data labeling method, and the other is an automated data labeling method. Therefore, the labeling tools available in this embodiment include: manual labeling, interactive labeling, and automatic labeling.
[0032] Among them, the interactive data annotation method is a combination of manual operation and the general interpretation model of the encoder-decoder architecture. Figure 4 As shown; The interactive data annotation method is a method of combining manual operation with a general interpretation model of an encoder-decoder architecture, and specifically includes the following processes: Grid the original remote sensing images; Manual operation is used to extract the position of the target of interest in the original remote sensing image. The specific operations include: positive and negative point selection and positive and negative frame selection. Among them, positive point selection and positive frame selection both indicate the selection of the target of interest, and negative point selection and negative frame selection both indicate the exclusion of the current area. The positive and negative point selection and positive and negative frame selection operations are used to quickly locate the target of interest. The general interpretation model of the encoder-decoder architecture is used to complete an interactive data annotation task, which specifically includes the following processes: the remote sensing image to be processed is input into the encoder, and the encoder outputs the encoding result and caches it; the encoding result and the manually selected coordinates are input into the decoder, and the decoder extracts the target at the specified position, and the extraction result is used as the target boundary; according to the requirements of the annotation task, the extraction result is converted into a polygon, a rotated rectangle, or a horizontal rectangle annotation style to complete an interactive data annotation task. The general interpretation model of "encoder-decoder" is used to separate the encoding and decoding processes. The encoder is used to extract image features, which has a large amount of calculation and is relatively time-consuming, but it is only performed once; the decoder uses the encoding result and the manually selected coordinates to extract the target. The decoding operation has a small amount of calculation and can be repeated many times. The use of this "encoder-decoder" general interpretation model greatly improves the efficiency of interactive data annotation.
[0033] In practical applications, the width of remote sensing images is relatively large, and the horizontal and vertical pixels are usually more than 20,000 or 30,000. The existing deep learning model cannot be input at one time. In order to realize the interactive annotation of large-width remote sensing images, the remote sensing images are gridded, and the grid size is set to 1024. The visualization effect is as follows Figure 5 As shown in the figure, the combination of grid and point selection box is used to realize interactive labeling of objects within the grid and across the grid.
[0034] The gridding of remote sensing images reduces the amount of data that is input into the deep learning model at a time, and also causes some targets to be segmented by grids, especially those with larger sizes, which cannot be completely enclosed by a single grid, while remote sensing image data annotation requires that they be annotated as a whole. To simplify the problem, in this embodiment, the relative relationship between the target and the grid is divided into two cases: one is that the target is completely contained in the grid, and the other is that the target spans multiple grids.
[0035] For the first case: the target is completely contained in a single grid. This is the easiest case to handle in actual processing. It supports both point selection and box selection interactive data annotation methods. The target position information within a single grid is obtained by clicking the target body or by box selecting the envelope target. It is then combined with the remote sensing image slices covered by the current grid to input into the deep learning model to extract the target boundary.
[0036] For the second case: the target spans multiple grids, in actual processing, it is necessary to use a box selection method to completely enclose the target, the box selection spans multiple grids, and the multiple grids involved in the box selection are merged into one grid, and the corresponding remote sensing image slices are extracted as the input of the encoder. In addition, when the target spans multiple grids, it indicates that the target has a large size, so the corresponding remote sensing image slices need to be downsampled. While compressing the model input parameters, it retains the main features of the target without affecting the extraction effect of the target boundary. The implementation effect of point selection and box selection is as follows: Figure 6 shown.
[0037] The automatic data annotation method uses a target model built based on the deep learning interpretation principle for a specific target to perform panoramic reasoning on the original remote sensing image to be processed, and converts the reasoning results into target annotations, thereby reducing the workload of data annotation. The automatic data annotation method can also be used to evaluate the detection effect of the target model, determine the problems existing in the target model, and then supplement the sample data to iterate the target model in a targeted manner, gradually improve the detection effect of the target model, and realize two-way iteration of data and target model.
[0038] like Figure 7 The automated data annotation method specifically includes the following processes: Select the corresponding target model according to the specific labeling task, which includes but is not limited to aircraft models, ship models and vehicle models in actual applications, subject to actual needs; Since the deep learning model cannot infer a wide range of remote sensing images at one time, it is necessary to partition the remote sensing images. The remote sensing images to be inferred are partitioned to complete the detection and annotation transformation of the entire remote sensing image; the specific partitioning processing method used is: window sliding reasoning; window sliding reasoning uses a fixed-size window to slide horizontally and vertically through the entire remote sensing image data to complete the detection and annotation transformation of the entire remote sensing image. Figure 8 As shown in the figure, the window size is set to 1024. To ensure the integrity of the target, two adjacent windows retain the overlapping area. Based on the above rules, the windows slide from left to right and from top to bottom until the entire scene image is traversed. The points distributed horizontally and vertically in the figure are the center points where the windows stay each time.
[0039] After the window completes the whole scene reasoning, the detection results are converted to absolute positions on the original remote sensing image according to the relative position of the window, and non-maximum suppression processing is performed to remove redundant detection results in overlapping areas.
[0040] For step S4: Specifically: verify the labeling results of intelligent labeling and make fine adjustments based on the verification results; Among them, fine-tuning includes: supplementing missing annotations, adjusting annotation ranges and modifying annotation properties.
[0041] Step S4 also includes a target type auxiliary interpretation method such as Fig. 9 After obtaining all the suspected types of the labeled target, the labeler only needs to select the most similar type. In this way, the labeler can quickly judge the target type, which greatly reduces the possibility of misjudgment. Specifically: Take the rotation box annotation as the starting step as an example: Start - spin box annotation - target knowledge base comparison - suspected type list - manual screening and comparison - output true category - end.
[0042] Calculate the length and width of the target based on the latitude and longitude information of the remote sensing image, and obtain all suspected types of the labeled target based on the current annotation task type and point information: Among them, it is necessary to pre-build a corresponding target knowledge base according to the specific annotation task type; the target knowledge base integrates the target's own attributes and statistical characteristics, and further builds a professional knowledge system of multi-source information fusion for specific targets, so that when the annotation personnel create the target annotation box, they compare the partial attributes of the annotation with the information stored in the target knowledge base to extract all suspected types; the annotation personnel combine the remaining attributes of the annotation to comprehensively judge the specific type of the target. Specifically: in practice, when targeting regular targets such as aircraft, ships, and vehicles, the target knowledge base integrates the target's own characteristics, such as type, length and width, information introduction, performance, multi-angle view, key identification features, etc., and also integrates the target's statistical characteristics, such as quantity, country, historical appearance area, historical events, etc., to build a professional knowledge system of multi-source information fusion for specific targets. When creating the target annotation box, the annotation personnel will compare the length and width of the annotation, the location, task type, time information, etc. with the knowledge base to extract all suspected types, and the annotation personnel will combine the remaining attributes of the annotation, such as multi-angle view, key identification features, etc., to comprehensively judge the specific type of the target.
[0043] Using this method, the difficulty of identifying the target type is greatly reduced and the speed of identifying the target type is improved. For labelers who encounter such targets for the first time, the data labeling threshold is significantly lowered and the accuracy of data labeling is improved.
[0044] Example 2 like Figure 1 On the basis of Example 1, this example further discloses a method for auxiliary annotation and quality control of the spatiotemporal correlation domain of interpreting remote sensing images, including a method for quality control of the spatiotemporal correlation domain of interpreting remote sensing images, and also includes steps S6-S8: In this embodiment, before executing step S6, the target imaging quality grading method is further executed, specifically: The imaging process of remote sensing images is affected by objective conditions such as weather, lighting, and roll angle, resulting in different imaging effects of the target. In particular, in some cases, the imaging effect of the target is not good, and it is difficult to identify the target by image alone. For example, in the case of cloud cover and insufficient light, it will be difficult to identify the target type and the target boundary will be uncertain. Therefore, the imaging quality of the target in the remote sensing image is an issue that must be considered. Generally speaking, these target samples with poor imaging quality have great difficulties in the manual interpretation and annotation stage, and are prone to annotation errors. If they are directly introduced into the training data of the intelligent interpretation model, it is easy to cause poor model training convergence effect, increase false detection, and reduce the detection effect of the intelligent interpretation model.
[0045] Therefore, in order to better train and evaluate the general interpretation model, it is necessary to grade the target imaging quality in remote sensing images so that in subsequent use, target samples with different imaging qualities can be used to train the target model according to different task requirements. At the same time, the detection effect of the target model at different imaging qualities can be analyzed, thereby achieving targeted optimization of the target model.
[0046] In view of the requirements of the target model involved in deep learning on data quality, this embodiment divides the imaging quality of the target. The specific target imaging quality classification is shown in the following table.
[0047] Table 2 Target imaging quality grading table
[0048] For steps S6-S8: Implement annotation quality control methods to control the annotated remote sensing image data, including two methods: spatiotemporal data comparison and target domain data comparison; Since the quality of data annotation also affects the training effect of deep learning models, it is necessary to control the annotation quality of remote sensing image data to avoid missing labels, wrong labels, and irregular labels. The remote sensing image data annotation quality control method includes two parts: one is the spatiotemporal data comparison, which mainly solves the problems of missing labels and wrong labels, and the other is the target domain data comparison, which mainly solves the problems of type errors, attribute errors, and irregular labels.
[0049] The spatiotemporal data comparison method is a quality control method for comparing multiple phases of annotation results for the same area. The data annotation of remote sensing images is performed in the form of point packages. Since remote sensing images are raster data with longitude and latitude information, and there is a temporal correlation between the remote sensing images of the previous and next phases, the above characteristics can be used to achieve the previous and next phase comparison of the local area, thereby improving the speed and accuracy of target identification, especially for those remote sensing images with poor imaging quality. The previous and next phases can be used to more accurately identify the target type. Based on this idea, a window anchor mechanism is designed in this embodiment to lock the window in the area to be observed, and then switch the remote sensing images at different times to observe the changes in the remote sensing images and the data annotation. The window locking effect of this method is as follows: Fig.10 As shown in the figure, in actual use, the image is switched back and forth within a window. This method can reduce the omission of targets in the window lock area, and is also convenient for comparing the changes of targets before and after the time phase, ensuring the accuracy of target category and attribute information.
[0050] The target domain data comparison method is a quality control method that uses target differences. In the data annotation process, it is necessary to annotate the bounding box and category attribute information of the target. After completing the annotation of a batch of remote sensing data, a large number of target annotations will be obtained. In order to ensure the annotation quality, it is necessary to verify the accuracy of all annotations. The traditional approach requires checking each image and each target, which is time-consuming, cumbersome, and not very practical. This embodiment proposes a target domain data comparison method, which crops all target annotations and displays them in groups according to the annotation type and attribute information, so as to realize intuitive and rapid inspection of all annotations. The target domain data comparison method is as follows: Fig.11 As shown in the figure, the target annotations with wrong type, wrong attributes, and irregular annotations will be significantly distinguished from adjacent targets. The annotators can quickly locate and correct the erroneous annotations, thus ensuring the quality of data annotation.
[0051] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.
Claims
1. A spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images, characterized in that: The following steps are involved: S1. Based on the remote sensing image data normalization method, the remote sensing image data is normalized, and the normalized remote sensing image data set is realized in the form of point packages; S2. Select any standardized original remote sensing image and further select the annotation form; S3. Select an appropriate annotation tool for annotation according to the annotation form selection result of step S2; wherein the annotation tools include: manual annotation and intelligent annotation method; S4. Fine-tune the annotation result of step S3; S5. Determine whether the labeling of all remote sensing images is completed; if so, complete the labeling; if not, repeat steps S2-S4 until the judgment result of step S5 is yes.
2. The spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images according to claim 1, characterized in that: The remote sensing image data normalization method is specifically as follows: combining remote sensing image data units belonging to the same point to construct a remote sensing image point package data set; combining multiple remote sensing image point package data sets to complete the remote sensing image data normalization; Among them, the smallest unit of remote sensing image data annotation is the remote sensing image data unit; the remote sensing image point package dataset is constructed based on the remote sensing image data unit; the point package dataset uses the latitude and longitude information of the remote sensing image to screen multiple remote sensing images in the same longitude and latitude range, and form a spatially fixed and time-series related remote sensing image data set.
3. The spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images according to claim 2, characterized in that: Intelligent labeling methods, including interactive data labeling methods and automated data labeling methods; Among them, the interactive data annotation method is a method of combining manual operation with the general interpretation model of the encoder-decoder architecture; The automatic data annotation method uses a target model constructed based on deep learning interpretation principles for a specific target to perform panoramic reasoning on the original remote sensing image to be processed, and converts the reasoning results into target annotations.
4. The spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images according to claim 3, characterized in that: The interactive data annotation method is a method of combining manual operation with a general interpretation model of an encoder-decoder architecture. The following processes are included: Grid the original remote sensing images; Manual operation is used to extract the position of the target of interest in the original remote sensing image. The specific operations include: positive and negative point selection and positive and negative box selection. Among them, positive point selection and positive box selection both indicate the selection of the target of interest, and negative point selection and negative box selection both indicate the exclusion of the current area. An interactive data labeling task is completed using a universal interpretation model of the encoder-decoder architecture, which specifically includes the following processes: the remote sensing image to be processed is input into the encoder, and the encoder outputs the encoding result and caches it; the encoding result and the manually selected coordinates are input into the decoder, and the decoder extracts the target at the specified position, and the extraction result is used as the target boundary; according to the labeling task requirements, the extraction result is converted into the labeling style to complete an interactive data labeling task.
5. The spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images according to claim 4, characterized in that: The automated data annotation method specifically includes the following processes: Select the corresponding target model according to the specific labeling task; The remote sensing image to be inferred is processed by partitioning to complete the detection and annotation transformation of the entire remote sensing image; the specific partitioning processing method used is: window sliding reasoning; After the window completes the whole scene reasoning, the detection results are converted into absolute positions on the original remote sensing image according to the relative position of the window, non-maximum suppression is performed, redundant detection results in overlapping areas are removed, and finally the detection results are converted into target annotations.
6. The spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images according to claim 5, characterized in that: For step S4, specifically: after converting the detection result into the target annotation, verify the automatic annotation result and make fine adjustments according to the verification result; Among them, fine-tuning includes: supplementing missing annotations, adjusting annotation ranges and modifying annotation properties.
7. The spatiotemporal correlation domain auxiliary annotation method for interpreting remote sensing images according to claim 6, characterized in that: It also includes a target type auxiliary interpretation method to assist labelers in quickly interpreting target types; specifically: Calculate the length and width of the target based on the latitude and longitude information of the remote sensing image, and obtain all suspected types of the labeled target based on the current annotation task type and point information: Among them, it is necessary to pre-build a corresponding target knowledge base according to the specific labeling task type; the target knowledge base integrates the target's own attributes and statistical characteristics, and further constructs a professional knowledge system of multi-source information fusion for specific targets, so that when creating a target labeling box, the labeling personnel can compare the partial attributes of the label with the information stored in the target knowledge base to extract all suspected types; the labeling personnel comprehensively judge the specific type of the target based on the remaining attributes of the label.
8. A method for controlling the quality of spatial-temporal correlation domains in interpreting remote sensing images, characterized in that: The method for auxiliary annotation of spatiotemporal correlation domain for interpreting remote sensing images according to claim 7 further comprises: S6. Implement labeling quality control methods; S7. Correct incorrect annotations; S8. Determine whether all the annotations are correct and meet the specifications. If so, archive the data; if not, repeat steps S6-S7 until the result of step S8 is yes.
9. The method for controlling the quality of spatial-temporal correlation domain of remote sensing image interpretation according to claim 8, characterized in that: Implement annotation quality control methods, including two methods: spatiotemporal data comparison and target domain data comparison; The spatiotemporal data comparison is a quality control method for comparing multi-phase annotation results of the same area, and utilizes the latitude and longitude information of remote sensing images and the temporal correlation of the previous and next phases to achieve the previous and next phase comparison of the local area. The target domain data comparison is a quality control method implemented by using target differences. All target annotations are cut out and displayed in groups according to annotation type and attribute information, so as to realize intuitive and rapid inspection of all annotations.