Time sequence remote sensing image segmentation method and device for image-level label weak supervision training

Through image-level label weak supervision training and time-series-aware affinity propagation, the problems of under-activation and over-activation in time-series remote sensing images are solved, high-precision semantic segmentation is achieved, and segmentation accuracy is improved.

CN120673059APending Publication Date: 2025-09-19INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510727793.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing weakly supervised segmentation methods suffer from under-activation and over-activation phenomena in time-series remote sensing images, cannot effectively utilize image-level annotations for high-precision semantic segmentation, and temporal feature confusion leads to erroneous activation.

Method used

We adopt the image-level label weakly supervised training method, extract temporal and spatial features through visual transformer, and combine prototype learning and temporal-aware affinity propagation to generate high-quality pseudo labels for segmentation.

Benefits of technology

It significantly improves the segmentation accuracy of time-series remote sensing images, reaching 95% of the fully supervised accuracy, and outperforms existing methods on multiple datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673059A_ABST
    Figure CN120673059A_ABST
Patent Text Reader

Abstract

The invention provides a time sequence remote sensing image segmentation method and device for image-level label weak supervision training. The method comprises the following steps: taking a time sequence remote sensing image as training data; time sequence features of training data are extracted through a time encoder of the visual converter, space features of a time sequence feature graph are extracted through a space encoder of the visual converter, the space features are separated, dense features and global features are obtained, a classifier obtains a classification result according to the global features, and a segmentation decoder obtains a segmentation result according to the dense features; constructing a first loss function and a second loss function based on the differences between the overall category of the time sequence remote sensing image and the classification result and the segmentation result, so as to train a visual converter; and inputting a to-be-processed time sequence remote sensing image to be subjected to semantic segmentation into the trained visual converter, and inputting the obtained dense features of the to-be-processed time sequence remote sensing image into a segmentation decoder to obtain a semantic segmentation result of the to-be-processed time sequence remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image semantic segmentation technology, and more particularly to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for time-series remote sensing image segmentation using weakly supervised training with image-level labels. The present invention proposes a weakly supervised learning paradigm, using only image-level annotations to train the SITS semantic segmentation model, aiming to eliminate the SITS semantic segmentation task's reliance on cumbersome annotation. Background Art

[0002] Semantic segmentation of ground objects using Satellite Image Time Series (SITS) remote sensing imagery has become a key area of ​​remote sensing image interpretation and is widely used in fields such as smart agriculture, natural disaster monitoring, urban planning, and climate monitoring. Compared to semantic segmentation tasks for general natural images, SITS requires more pixel-by-pixel annotation due to its complex data structure. However, due to SITS's low spatial resolution and unclear ground object boundaries, annotating pixel-level masks in SITS is extremely complex and time-consuming.

[0003] The current mainstream weakly supervised segmentation methods are based on the Class Activation Map (CAM) to generate semantic segmentation pseudo-labels. In the field of natural scene images, due to the discreteness of the distribution within the category, the model is forced to learn an extremely sharp decision boundary. The weight of the classifier tends to interact with the more discriminative feature areas of the target object. This causes CAM to often highlight the local salient areas of the target object, such as Figure 1 As shown in the figure, it tends to interact with the dog's head area while ignoring the dog's body, resulting in only the head being classified as a dog, while ignoring the fact that the body should also belong to the dog category. This phenomenon is called under-activation. Therefore, previous weakly supervised segmentation methods in the field of natural scene images focused on how to expand the under-activated CAM from the local salient area to the entire object area to generate a complete segmentation pseudo-label for the object.

[0004] However, existing weakly supervised techniques face many challenges in the SITS semantic segmentation task.

[0005] On the one hand, from the spatial perspective of SITS data, objects within the same category usually have consistent appearance and color, showing strong neighborhood consistency. The compactness of local patterns within this category reduces the classifier's tolerance to noise perturbations, causing CAM to activate not only the entire object area, but also incorrectly activate in many background areas, such as Figure 1As shown in Figure 2, this phenomenon is called over-activation. This activation characteristic is completely different from that in the field of natural scene images, which makes it impossible for weakly supervised algorithms in the field of natural images to directly benefit the segmentation task in SITS.

[0006] On the other hand, from the perspective of the time series of SITS data, although different categories have different phenological cycles and change characteristics, they may show similar appearances in certain specific periods, e.g. Figure 2 This confusion introduces incorrect semantic biases into the model's learning process, activating some non-target semantic regions in the CAM. Those temporal segments that deviate from key semantics will affect the CAM's ability to perceive the correct category regions. Summary of the Invention

[0007] The purpose of this invention is to overcome the problem that existing weakly supervised semantic segmentation technology is not suitable for the SITS semantic segmentation field, and proposes a SITS weakly supervised semantic segmentation training method based on prototype learning and temporal-aware affinity propagation.

[0008] In view of the shortcomings of existing technologies, such as Figure 9 As shown, the present invention proposes a time series remote sensing image segmentation method based on image-level label weak supervision training, which includes:

[0009] In the initial step, the time series remote sensing images with their overall categories labeled are used as training data and input into the visual transformer;

[0010] A training step comprises extracting temporal features of the training data through a temporal encoder of the visual transformer, extracting spatial features of the temporal feature graph through a spatial encoder of the visual transformer, separating the spatial features to obtain dense features and global features, obtaining a classification result based on the global features by a classifier, and obtaining a segmentation result based on the dense features by a segmentation decoder, and constructing a first loss function and a second loss function based on the gap between the overall category of the temporal remote sensing image and the classification result and the segmentation result, respectively, to train the visual transformer;

[0011] In the image segmentation step, the time series remote sensing image to be processed for semantic segmentation is input into the trained visual transformer, and the dense features of the time series remote sensing image to be processed are input into the segmentation decoder to obtain the semantic segmentation result of the time series remote sensing image to be processed.

[0012] The time series remote sensing image segmentation method of image-level label weak supervision training, wherein

[0013] This initial step includes:

[0014] The input of the visual transformer is a time series remote sensing image X∈R T×C×H×W, where T represents the length of the time series, C represents the number of channels of the remote sensing image, and H×W represents the spatial dimension; by mapping the time series remote sensing image X into multiple image blocks, the time series remote sensing image X is reshaped into an image and time multi-class d is the feature dimension and K is the number of image categories, where h×w represents the spatial range of each image block; the time of the time series remote sensing image is added as the time position embedding P T ∈R T×d And concatenate the multiple types of tokens of the time to get the input of the time encoder

[0015]

[0016] The training steps include:

[0017] The temporal features output by the time encoder before extraction The first K tokens of Will The dimension exchange obtains the feature Based on the feature Z S , we get spatial multi-class and spatial position embedding Combine this feature Z in all K spatial representations S , spatial multi-class Spatial position embedding P S , and get the combined result

[0018]

[0019] Combine the results Input the spatial encoder and output the spatial features Separation,

[0020] This global feature Input the classifier to obtain the classification result; the dense feature Input the segmentation decoder to obtain the segmentation result;

[0021] The dense feature Through average pooling and the weight of the segmentation decoder w∈R K×D Multiply; weighted by the weight w and sum the dense features Each channel is generated:

[0022]

[0023] Will Normalize to the specified range and apply a global threshold to filter background pixels to obtain the segmentation result.

[0024] The method for temporal remote sensing image segmentation using weakly supervised training of image-level labels, wherein the training step includes the step of exploring spatial perception clues:

[0025] The overall category y∈[0,1] of the labeled time series remote sensing image K , through the weight w and Calculate its normalized fusion Use threshold μ l and μ h Filter foreground, background and uncertain areas;

[0026] Time-intensive embedding Perform spatial clustering;

[0027] Establish a set of positive prototypes of categories and negative prototype sets Given a mapping matrix C k , which represents the assignment between each pixel and its prototype and can be considered as an element of the transport prototype

[0028]

[0029] where N k Represents the number of pixels belonging to category k; 1 is a fully one vector, u and r are the mapping matrices C k Marginal projection of rows and columns; optimize the mapping matrix C by maximizing the following objective function k :

[0030]

[0031] is the corresponding time-dense embedding belonging to class k, and η controls the smoothness of the entropy regularization term;

[0032] According to the distribution matrix and embedding, the nth p Prototypes are updated with momentum:

[0033]

[0034] Here α∈[0,1] is the momentum coefficient;

[0035] Leveraging Fusion CAM Update the category prototype set P as a pseudo label pos and use 1- Update the negative prototype set P neg ;

[0036] For each time-dense embedding and its most similar prototype The cosine distance is used to measure their similarity:

[0037]

[0038] Where τ represents the temperature parameter; the following third loss function is constructed:

[0039]

[0040] in is an indicator function, which is 1 if category k appears in the overall category of the labeled time series remote sensing image, otherwise it is 0;

[0041] The visual transformer is trained using the first loss function, the second loss function, and the third loss function.

[0042] The time series remote sensing image segmentation method of image-level label weakly supervised training, wherein the training step includes a time series perception affinity propagation step:

[0043] In this temporal encoder, the input tokens are normalized and projected into the query matrix Q∈R (K+T)×d and bond matrix K∈R (K+T)×d ;according to Sequence calculation self-attention N h ·N w :

[0044]

[0045] From self-attention Explicitly extracting time-to-category attention The attention Indicates the contribution of high-level representation to category recognition in different time segments; weighted by the following formula

[0046]

[0047] High-level representation This increases the variation between different semantic categories, thereby modeling pixel relationships more accurately.

[0048] The temporal-aware pairwise affinity of a pixel at position i:

[0049]

[0050] here yes The standard deviation of Represents a local receptive field; it propagates time-aware representation through iteration The pairwise affinity is used to denoise the original CAM:

[0051]

[0052] Will and Align, and get the fourth loss function:

[0053]

[0054] The visual transformer is trained using the first loss function, the second loss function, and the fourth loss function.

[0055] like Figure 10 As shown, the present invention also proposes a time series remote sensing image segmentation device with image-level label weak supervision training, which includes:

[0056] The initial module takes the time series remote sensing images with the overall categories of the time series remote sensing images as training data and inputs them into the visual transformer;

[0057] A training module extracts temporal features of the training data through a temporal encoder of the visual transformer, extracts spatial features of the temporal feature graph through a spatial encoder of the visual transformer, separates the spatial features to obtain dense features and global features, a classifier obtains a classification result based on the global features, a segmentation decoder obtains a segmentation result based on the dense features, and constructs a first loss function and a second loss function based on the gap between the overall category of the temporal remote sensing image and the classification result and the segmentation result, respectively, to train the visual transformer;

[0058] The image segmentation module inputs the time series remote sensing image to be processed for semantic segmentation into the trained visual transformer, inputs the obtained dense features of the time series remote sensing image to be processed into the segmentation decoder, and obtains the semantic segmentation result of the time series remote sensing image to be processed.

[0059] The image-level label weakly supervised training time series remote sensing image segmentation device, wherein

[0060] This initial module includes:

[0061] The input of the visual transformer is a time series remote sensing image X∈R T×C×H×W , where T represents the length of the time series, C represents the number of channels of the remote sensing image, and H×W represents the spatial dimension; by mapping the time series remote sensing image X into multiple image blocks, the time series remote sensing image X is reshaped into an image and time multi-class d is the feature dimension and K is the number of image categories, where h×w represents the spatial range of each image block; the time of the time series remote sensing image is added as the time position embedding P T ∈R T×d And concatenate the multiple types of tokens of the time to get the input of the time encoder

[0062]

[0063] This training module includes:

[0064] The temporal features output by the time encoder before extraction The first K tokens of Will The dimension exchange obtains the feature Based on the feature Z S , we get spatial multi-class and spatial position embedding Combine this feature Z in all K spatial representations S , spatial multi-class Spatial position embedding P S , and get the combined result

[0065]

[0066] Combine the results Input the spatial encoder and output the spatial features Separation,

[0067] This global feature Input the classifier to obtain the classification result; the dense feature Input the segmentation decoder to obtain the segmentation result;

[0068] The dense feature Through average pooling and the weight of the segmentation decoder w∈R K×D Multiply; weighted by the weight w and sum the dense features Each channel is generated:

[0069]

[0070] Will Normalize to the specified range and apply a global threshold to filter background pixels to obtain the segmentation result.

[0071] The image-level label weakly supervised training temporal remote sensing image segmentation device, wherein the training module includes a spatial perception clue exploration module:

[0072] The overall category y∈[0,1] of the labeled time series remote sensing image K , through the weight w and Calculate its normalized fusion Use threshold μ l and μ h Filter foreground, background and uncertain areas;

[0073] Time-intensive embedding Perform spatial clustering;

[0074] Establish a set of positive prototypes of categories and negative prototype sets Given a mapping matrix C k , which represents the assignment between each pixel and its prototype and can be considered as an element of the transport prototype

[0075]

[0076] where N k Represents the number of pixels belonging to category k; 1 is a fully one vector, u and r are the mapping matrices C k Marginal projection of rows and columns; optimize the mapping matrix C by maximizing the following objective function k :

[0077]

[0078] is the corresponding time-dense embedding belonging to class k, and η controls the smoothness of the entropy regularization term;

[0079] According to the distribution matrix and embedding, the nth p Prototypes are updated with momentum:

[0080]

[0081] Here α∈[0,1] is the momentum coefficient;

[0082] Using fusion Update the category prototype set P as a pseudo label pos and use Update the negative prototype set P neg ;

[0083] For each time-dense embedding and its most similar prototype The cosine distance is used to measure their similarity:

[0084]

[0085] Where τ represents the temperature parameter; the following third loss function is constructed:

[0086]

[0087] in is an indicator function, which is 1 if category k appears in the overall category of the labeled time series remote sensing image, otherwise it is 0;

[0088] Training the visual transformer using the first loss function, the second loss function, and the third loss function;

[0089] The training module includes the time-aware affinity propagation module:

[0090] In this temporal encoder, the input tokens are normalized and projected into the query matrix Q∈R (K+T)×d and bond matrix K∈R (K+T)×d ;according to Sequence calculation self-attention N h ·N w :

[0091]

[0092] From self-attention Explicitly extracting time-to-category attention The attention Indicates the contribution of high-level representation to category recognition in different time segments; weighted by the following formula

[0093]

[0094] High-level representation This increases the variation between different semantic categories, thereby modeling pixel relationships more accurately.

[0095] The temporal-aware pairwise affinity of a pixel at position i:

[0096]

[0097] here yes The standard deviation of Represents a local receptive field; it propagates time-aware representation through iteration The pairwise affinity is used to denoise the original CAM:

[0098]

[0099] Will and Align, and get the fourth loss function:

[0100]

[0101] The visual transformer is trained using the first loss function, the second loss function, and the fourth loss function.

[0102] The present invention also proposes an electronic device, which includes the aforementioned time-series remote sensing image segmentation device with image-level label weakly supervised training. The electronic device may be connected to an information display device, which is used to display the semantic segmentation results using display parameters and attributes set by the user or through an artificial intelligence model.

[0103] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the temporal remote sensing image segmentation method using weakly supervised training of image-level labels.

[0104] The present invention also proposes a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the temporal remote sensing image segmentation method using image-level label weakly supervised training are implemented.

[0105] From the above scheme, it can be seen that the advantages of the present invention are:

[0106] In summary, the proposed semantic segmentation method for time-series remote sensing images based on weakly supervised training with image-level labels effectively adapts weakly supervised algorithms to the field of time-series remote sensing images, significantly reducing the labeling cost of time-series remote sensing images. Using only categorical labeling, it achieves 95% of the full supervision accuracy.

[0107] like Figure 4As shown, the present invention compared the overall accuracy (OA) and mean intersection over union (mIoU) of different methods in generating pseudo labels in one stage on the Germany and PASTIS datasets, and the method of the present invention achieved the best results. On the Germany dataset, the method of the present invention achieved an OA of 88.3 and an mIoU of 80.6, which were 5.8 and 6.5 higher than the baseline results, respectively, and 4.4 and 4.3 higher than the best other method PAMR results. On the PASTIS dataset, the method of the present invention achieved an OA of 84.1 and an mIoU of 75.6, which were 2.9 and 6.1 higher than the baseline results, respectively, and 2.1 and 4.4 higher than the best other method PAMR results.

[0108] like Figure 5 As shown in the figure, the present invention compares the OA and mIoU of the final semantic segmentation results of different methods on the PASTIS dataset, and the present invention achieves the best results. When using only image-level labels, the present invention achieves an OA of 80.2 and an mIoU of 62.0, which are improvements of 3.0 and 4.2 respectively compared to the baseline results, and 1.7 and 3.3 respectively compared to the best other method PAMR. Compared with the model trained with pixel-level labels, the present invention achieves 95% of its accuracy.

[0109] like Figure 6 As shown in the figure, the present invention conducts sufficient ablation experiments on the PASTIS dataset. The three modules designed by the present invention for the three key points can all improve the OA and mIoU of the final segmentation results.

[0110] Figure 7 Partial visualization results of Silly Dog on the PASTIS dataset are shown. The pseudo labels generated in the first stage and the final segmentation results of the method of the present invention are better than the baseline. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Figure 1 Schematic diagram of under-activation and over-activation;

[0112] Figure 2 The different diagrams show the different discrimination abilities of different time series;

[0113] Figure 3 It is a diagram of the method architecture;

[0114] Figure 4This is a pseudo-label accuracy comparison experiment (overall accuracy, average intersection-over-union ratio);

[0115] Figure 5 This is a diagram of the segmentation accuracy comparison experiment (overall accuracy, average intersection-over-union ratio);

[0116] Figure 6 This is the result of the ablation experiment;

[0117] Figure 7 Display graphs for visualizing results;

[0118] Figure 8 This is the generation process diagram of CB-CAM;

[0119] Figure 9 Flow chart of the method of the present invention;

[0120] Figure 10 This is a module diagram of the device of the present invention;

[0121] Figure 11 This is a schematic structural diagram of a first electronic device of the present invention;

[0122] Figure 12 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;

[0123] Figure 13 This is a schematic structural diagram of a second electronic device according to the present invention.

[0124] Reference numerals:

[0125] A-First electronic device;

[0126] B- image-level label weakly supervised training of temporal remote sensing image devices;

[0127] C-data acquisition equipment;

[0128] D-information display device;

[0129] 1000- second electronic device;

[0130] Ⅰ-computing unit;

[0131] II-ROM;

[0132] III-RAM;

[0133] IV-bus;

[0134] V-interface;

[0135] VI-input unit;

[0136] VII-output unit;

[0137] VIII-Storage medium;

[0138] IX-Communication unit. DETAILED DESCRIPTION

[0139] It should be noted that, in this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0140] Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.

[0141] The processor described in the present invention is the control center of an electronic device and can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).

[0142] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.

[0143] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include: servers, desktop computers, laptops, smartphones, tablet computers, embedded computers, etc., wherein the embedded computers include vehicles and robots, etc.

[0144] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0145] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto, and the actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or a combination of certain components, or a different arrangement of components.

[0146] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0147] It should also be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0148] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0149] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0150] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0151] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0152] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0153] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0154] Taking into account the shortcomings of the existing technologies, the present invention finds that the existing weakly supervised segmentation algorithms do not take into account the characteristics of small differences in spatial features within SITS data classes and easy confusion of temporal features between classes, resulting in poor performance of existing weakly supervised technologies on SITS data. Through research on SITS's CAM, it was found that SITS's CAM will have serious over-activation phenomenon. Therefore, the present invention abandons the pseudo-label expansion strategy commonly used by previous weakly supervised segmentation algorithms under the characteristics of under-activated CAM of natural image data, and adopts a pseudo-label generation strategy based on prototype vector similarity. In addition, the present invention also uses the temporal difference characteristics of SITS data to design a temporal-aware affinity propagation module to highlight the temporal sequence with category recognition to further correct the pseudo-label. Based on the above design, the present invention proposes the first weakly supervised segmentation framework for temporal remote sensing data, called Exact, such as Figure 3 shown.

[0155] In summary, in order to achieve the above technical effects, the present invention proposes the following key technical points:

[0156] Key Point 1: Based on the low intra-class feature variance in SITS data, this paper introduces a set of spatial cues to explicitly capture patterns across different classes. Using the filtered CAM as indicators, this paper spatially clusters the regions with the highest class relevance, thereby updating representative cues. These cues are used to regularize the feature space by optimizing contrast targets, thereby strengthening the model's decision boundary and mitigating interference from spurious patterns.

[0157] Key Point 2: To address the semantic bias caused by abnormal time segments, this paper proposes a time-aware affinity propagation method to highlight the importance of key time segments in class differentiation. Specifically, this paper extracts the time-to-class attention weights from the model and uses them to reweight the time series embedding. This adjusted representation is used to perform time-aware pairwise affinity propagation on the original CAM, effectively suppressing non-target semantic regions in a self-supervised manner.

[0158] Key Point 3: Unlike existing weakly supervised segmentation methods that rely on classifier weights to generate CAMs, this paper utilizes updated spatiotemporal cues to generate cue-based CAMs (CB-CAMs) as pseudo-labels for segmentation. Compared to the original CAM, CB-CAMs offer the following advantages: They significantly suppress spatial and temporal interference and more accurately delineate class regions, providing more reliable supervision for subsequent segmentation.

[0159] To illustrate the above-mentioned features and effects of the present invention more clearly and easily, the following embodiments are specifically described below with reference to the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are for illustrative purposes only. The scope of protection of the present invention is not limited to the disclosed embodiments; the present invention is defined by the appended claims.

[0160] The method of processing the temporal dimension first and then extracting spatial information shows significant superiority in processing SITS data. Following the design of TSViT (Temporo-Spatial Vision Transformer), the present invention uses the variant visual transformer ViT as the backbone network of the temporal encoder and the spatial encoder. The present invention considers a SITS input X∈R T×C×H×W , where T represents the length of the time series, C represents the number of channels of the remote sensing image, for example, color images have RGB 3 channels, and remote sensing images also have infrared, near-infrared and other channels, and H×W represents the spatial dimension, such as the resolution of the remote sensing image. The input X is mapped into a series of Patch Tokens and reshaped into d is the feature dimension, where h×w represents the spatial range of each Patch. Then, the present invention adds the encoding of the image shooting date as the time position embedded in P T ∈R T×d , and multi-classify time Splice to get the input of the time encoder

[0161]

[0162] Among them, K represents the number of categories, and P T Repeated N h ·N w times to match the spatial shape, the input image is mapped to and P T Need to repeat N h ·N w times, becomes an N h ·N w ×K×d and N h ·N w ×1×d vector can be spliced ​​together with Z in the third dimension to become

[0163] Output feature map of the temporal encoder The present invention extracts the first K And in order to do similarity calculation in the spatial dimension later, Swap the order of the first and second dimensions to get the input of the spatial encoder These inputs are then compared with the spatial category and spatial position embedding Combine all K spatial representations to obtain the combined result

[0164]

[0165] Among them, the spatial position embedding P S It is determined according to the position of each patch in the image.

[0166] After the spatial encoder stage, the present invention will output the features Separate to adapt to different downstream tasks. For classification tasks, the present invention will globally Input the classifier to obtain the classification Logits, which are usually crop categories such as wheat and corn. For the segmentation task, in order to generate a dense prediction mask, the dense Enter the segmentation decoder.

[0167] Class Activation Maps (CAMs) are widely used in weakly supervised semantic segmentation (WSSS) to provide weak annotations that rely on image-level labels. Given a natural image, its feature map F∈R H′×W′×D Extracted by the classification backbone. In order to obtain the classification score, the feature map is processed by average pooling and combined with the classifier weight w∈R K×D Multiplied together, the classifier weight is the weight of the classifier that processes the global token above. The CAM is generated by weighting each channel in the feature map with the classifier weight and summing it, as shown below:

[0168]

[0169] represents CAM, k represents category channel, xy represents position, RELU is activation function, D represents the number of feature channels, w is the weight of the above classifier, and F is and average.

[0170] Most WSSS methods will Normalize to the range of [0,1], and apply the threshold value that makes the pseudo label accuracy the highest by iteration on the data set to filter the background pixels to finally obtain the pseudo label. The normalized M can refer to the probability of each pixel belonging to each category. For example, close to 0 means it does not belong to the current category, and close to 1 means it belongs to the current category. In the spatiotemporal network, the present invention will densely label and Input into the classifier, output the category prediction of the entire image, and calculate the additional classification loss which also uses the overall image category as the supervision signal To generate the fused original CAM of SITS.

[0171] Exploring spatial perception cues

[0172] Given an input SITSX and its image-level label y∈[0,1] K The present invention first calculates the normalized fusion of the classifier weight and the output feature map F of the spatiotemporal encoder The present invention uses two thresholds μ l and μ h to filter out reliable foreground, background and uncertain areas.

[0173] In order to capture the compact patterns of different categories, the present invention establishes a set of category representative prototypes, which are later used as perceptual clues to generate high-quality pseudo labels. In the task of semantic segmentation of time-series remote sensing images, temporal features usually provide more information than spatial context, so the present invention chooses to embed temporal dense features on the entire dataset. Specifically, the present invention establishes a class prototype set by continuously updating during storage and training. and negative prototype sets If the kth class appears in the training batch, the present invention updates the prototype p by solving the optimal transportation problem k There are two purposes for updating the prototypes: 1. Use these prototypes to optimize the feature space of the model; 2. Use the updated prototypes to calculate the similarity distance on the feature map to obtain the pseudo label. If the prototype is not updated, the prototype will result in only the initial random value. On the contrary, if the kth class does not appear, it will not be considered. Given a mapping matrix C k , which represents the assignment between each pixel and its prototype and can be considered as an element of the transport prototype:

[0174]

[0175] where N k represents the number of pixels belonging to class k. 1 represents an all-one vector of appropriate dimension, u and r are the mapping matrices C k The marginal projection of rows and columns. The present invention can optimize the mapping matrix by maximizing the following objective function:

[0176]

[0177] Here is the corresponding time dense embedding belonging to category k, and η controls the smoothness of the entropy regularization term. The time dense embedding is the feature input of the spatial encoder, and the dimension of the feature vector at each position is N k ×d,N k is the number of categories k, and the kth channel belongs to category k.

[0178] The continuous approximate solution of the equation can be obtained by iteratively applying the Sinkhorn-Knopp algorithm. p Prototypes are updated with momentum:

[0179]

[0180] Here α∈[0,1] is the preset momentum coefficient. Belongs to C k , the subscript np represents the number of prototypes for each category k. Using the classifier weights and feature maps of the two spatiotemporal encoders, two CAMs can be obtained. The two CAMs are fused to obtain a fused CAM. Update the category prototype set P as a pseudo label pos and use Update the negative prototype set P neg It is worth noting that each prototype does not participate in the gradient back-propagation to avoid noise from the classifier.

[0181] Based on spatial perception cues, this paper introduces cue-based contrastive learning to normalize the embedding space. More specifically, for each time-dense embedding output by the temporal encoder and its most relevant prototypes The nth one represents the kth class p prototypes, the present invention uses cosine distance to measure their similarity:

[0182]

[0183] where τ represents the temperature parameter. Subsequently, the invention forces the pixel embedding to closely match its prototype and to be clearly distinguishable from other prototypes:

[0184]

[0185] in Is an indicator function that is 1 if category k appears in the image-level label and 0 otherwise. Minimizing the above objective can pull the pixel embedding closer to the semantic center and push it away from other hallucination patterns, thereby improving the clarity of the model decision boundary and alleviating noise interference.

[0186] Time-Aware Affinity Propagation

[0187] This paper proposes a time-aware affinity propagation method to deal with the incorrect semantic deviation caused by abnormal time segments. In the temporal encoder, the input tokens are normalized and projected into the query matrix Q∈R (K+T)×d and bond matrix K∈R (K+T)×d Then according to Sequence calculation self-attention N h ·N w , as shown below:

[0188]

[0189] Without additional computation or supervision, the present invention can be Explicitly extracting time-to-category attention The attention represents the contribution of high-level representation to category recognition in different time segments. Embedding time series To reweight:

[0190]

[0191] Intuitively, the high-level representation after modulation This increases the variation between different semantic categories, thereby modeling pixel relationships more accurately.

[0192] In obtaining time-aware representation Finally, the present invention performs affinity propagation to suppress the erroneous semantic regions on the original CAM.

[0193] In particular, the time-aware pairwise affinity of a pixel at position i can be estimated as follows:

[0194]

[0195] here yes The standard deviation of Represents the local receptive field (e.g., 8-way local neighbors). The present invention iterates the time-aware representation The pairwise affinity of denoising the original CAM:

[0196]

[0197] In obtaining improved activation maps Finally, the present invention aligns it with the original CAM to effectively guide the learning process:

[0198]

[0199] By minimizing Our proposed method can correct the erroneous activations in the original CAM and incorporate the temporal perception prior into the embedding space, which ultimately benefits the perceptual cues.

[0200] Clue-based CAM generation strategy

[0201] like Figure 3 As shown in Figure 2, the overall objective function for training a classification network consists of four parts:

[0202]

[0203] here and is the traditional binary cross entropy loss, λ i Represents the weight used to rescale the loss term.

[0204] After training, the present invention performs per-pixel perception based on the updated well-defined spatiotemporal cues in the spatiotemporal dense embedding space to generate a cue-based CAM (CB-CAM). Each embedding z in i , we measure the maximum similarity with the positive prototype and subtract the misleading activation with the negative prototype:

[0205]

[0206] If the class k does not appear in the training image, the present invention will k Set to all zeros. Figure 8 A visualization of this generation process is provided. Finally, the global background score is used to filter CB-CAMY to generate pseudo labels, which are then used to train the SITS semantic segmentation network.

[0207] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0208] like Figure 10 As shown, the present invention also proposes a time series remote sensing image segmentation device with image-level label weak supervision training, which includes:

[0209] The initial module takes the time series remote sensing images with the overall categories of the time series remote sensing images as training data and inputs them into the visual transformer;

[0210] A training module extracts temporal features of the training data through a temporal encoder of the visual transformer, extracts spatial features of the temporal feature graph through a spatial encoder of the visual transformer, separates the spatial features to obtain dense features and global features, a classifier obtains a classification result based on the global features, a segmentation decoder obtains a segmentation result based on the dense features, and constructs a first loss function and a second loss function based on the gap between the overall category of the temporal remote sensing image and the classification result and the segmentation result, respectively, to train the visual transformer;

[0211] The image segmentation module inputs the time series remote sensing image to be processed for semantic segmentation into the trained visual transformer, inputs the obtained dense features of the time series remote sensing image to be processed into the segmentation decoder, and obtains the semantic segmentation result of the time series remote sensing image to be processed.

[0212] The image-level label weakly supervised training time series remote sensing image segmentation device, wherein

[0213] This initial module includes:

[0214] The input of the visual transformer is a time series remote sensing image X∈R T×C×H×W , where T represents the length of the time series, C represents the number of channels of the remote sensing image, and H×W represents the spatial dimension; by mapping the time series remote sensing image X into multiple image blocks, the time series remote sensing image X is reshaped into an image and time multi-class d is the feature dimension and K is the number of image categories, where h×w represents the spatial range of each image block; the time of the time series remote sensing image is added as the time position embedding P T ∈R T×d And concatenate the multiple types of tokens of the time to get the input of the time encoder

[0215]

[0216] This training module includes:

[0217] The temporal features output by the time encoder before extraction The first K tokens of Will The dimension exchange obtains the feature Based on the feature Z S , we get spatial multi-class and spatial position embedding Combine this feature Z in all K spatial representations S , spatial multi-class Spatial position embedding P S , and get the combined result

[0218]

[0219] Combine the results Input the spatial encoder and output the spatial features Separation,

[0220] This global feature Input the classifier to obtain the classification result; the dense feature Input the segmentation decoder to obtain the segmentation result;

[0221] The dense feature Through average pooling and the weight of the segmentation decoder w∈R K×D Multiply; weight and sum the dense features by the weight w Each channel is generated:

[0222]

[0223] Will Normalize to the specified range and apply a global threshold to filter background pixels to obtain the segmentation result.

[0224] The image-level label weakly supervised training temporal remote sensing image segmentation device, wherein the training module includes a spatial perception clue exploration module:

[0225] The overall category y∈[0,1] of the labeled time series remote sensing image K , through the weight w and Calculate its normalized fusion Use threshold μ l and μ h Filter foreground, background and uncertain areas;

[0226] Time-intensive embedding Perform spatial clustering;

[0227] Establish a set of positive prototypes of categories and negative prototype sets Given a mapping matrix C k , which represents the assignment between each pixel and its prototype and can be considered as an element of the transport prototype

[0228]

[0229] where Nk Represents the number of pixels belonging to category k; 1 is a fully one vector, u and r are the mapping matrices C k Marginal projection of rows and columns; optimize the mapping matrix C by maximizing the following objective function k :

[0230]

[0231] is the corresponding time-dense embedding belonging to class k, and η controls the smoothness of the entropy regularization term;

[0232] According to the distribution matrix and embedding, the nth p Prototypes are updated with momentum:

[0233]

[0234] Here α∈[0,1] is the momentum coefficient;

[0235] Using fusion Update the category prototype set P as a pseudo label pos and use Update the negative prototype set P neg ;

[0236] For each time-dense embedding and its most similar prototype The cosine distance is used to measure their similarity:

[0237]

[0238] Where τ represents the temperature parameter; the following third loss function is constructed:

[0239]

[0240] in is an indicator function, which is 1 if category k appears in the overall category of the labeled time series remote sensing image, otherwise it is 0;

[0241] Training the visual transformer using the first loss function, the second loss function, and the third loss function;

[0242] The training module includes the time-aware affinity propagation module:

[0243] In this temporal encoder, the input tokens are normalized and projected into the query matrix Q∈R (K+T)×d and bond matrix K∈R (K+T)×d ;according to Sequence calculation self-attention N h ·Nw :

[0244]

[0245] From self-attention Explicitly extracting time-to-category attention The attention Indicates the contribution of high-level representation to category recognition in different time segments; weighted by the following formula

[0246]

[0247] High-level representation This increases the variation between different semantic categories, thereby modeling pixel relationships more accurately.

[0248] The temporal-aware pairwise affinity of a pixel at position i:

[0249]

[0250] here yes The standard deviation of Represents a local receptive field; it propagates time-aware representation through iteration The pairwise affinity of denoising the original CAM:

[0251]

[0252] Will and Align, and get the fourth loss function:

[0253]

[0254] The visual transformer is trained using the first loss function, the second loss function, and the fourth loss function.

[0255] like Figure 11 As shown, the present invention further proposes a first electronic device A in another embodiment, which includes the above-mentioned time series remote sensing image segmentation device with image-level label weak supervision training.

[0256] like Figure 12 As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect and obtain remote sensing images, and the information display device D is used to display the semantic segmentation results obtained by the analysis of the present invention.

[0257] The information display device D can organize and process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be manually preset, for example, the data output by the first electronic device A is visually displayed, which can be based on the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll, etc. The user is presented with the key information specified by the user, such as the cultivated land area, range, planting stage, etc. The user can understand this information more promptly without having to access the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key information based on the user's previous usage habits, such as viewing time, number of clicks, number of edits, etc., and then automatically present the user with rich and necessary key information.

[0258] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the time series remote sensing image segmentation method with image-level label weak supervision training provided by the above methods.

[0259] In another embodiment, the present invention further proposes a storage medium VIII for storing a computer program for performing the time series remote sensing image segmentation method of the image-level label weakly supervised training. It should be understood that the storage medium in the embodiment of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM).

[0260] Figure 13 A schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention is shown. The second electronic device 1000 electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0261] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. Various programs and data required for the operation of the device 1000 can also be stored in the RAM III. The computing unit I, ROM II, and RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0262] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard and mouse; an output unit VII, such as various types of displays and speakers; a storage medium VIII, such as a magnetic disk and optical disk; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0263] Computing unit I can be various general and / or special processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. Computing unit I performs the various methods and processes described above, such as method steps S1-S3. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the method described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to execute the method in any other appropriate manner (e.g., by means of firmware).

[0264] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A time series remote sensing image segmentation method based on image-level label weakly supervised training, characterized in that: include: In the initial step, the time series remote sensing images with their overall categories labeled are used as training data and input into the visual transformer; A training step comprises extracting temporal features of the training data through a temporal encoder of the visual transformer, extracting spatial features of the temporal feature graph through a spatial encoder of the visual transformer, separating the spatial features to obtain dense features and global features, obtaining a classification result based on the global features by a classifier, and obtaining a segmentation result based on the dense features by a segmentation decoder, and constructing a first loss function and a second loss function based on the gap between the overall category of the temporal remote sensing image and the classification result and the segmentation result, respectively, to train the visual transformer; In the image segmentation step, the time series remote sensing image to be processed for semantic segmentation is input into the trained visual transformer, and the dense features of the time series remote sensing image to be processed are input into the segmentation decoder to obtain the semantic segmentation result of the time series remote sensing image to be processed.

2. The time series remote sensing image segmentation method based on image-level label weakly supervised training according to claim 1, characterized in that: This initial step includes: The input of the visual transformer is a time series remote sensing image X∈R T×C×H×W , where T represents the length of the time series, C represents the number of channels of the remote sensing image, and H×W represents the spatial dimension; by mapping the time series remote sensing image X into multiple image blocks, the time series remote sensing image X is reshaped into an image and time-based multi-class tokens d is the feature dimension and K is the number of image categories, where h×w represents the spatial range of each image block; the time of the time series remote sensing image is added as the time position embedding P T ∈R T×d And concatenate the multiple types of tokens of the time to get the input of the time encoder The training steps include: The temporal features output by the time encoder before extraction The first K tokens of Will The dimension exchange obtains the feature Based on the feature Z S , we get spatial multi-class and spatial position embedding Combine this feature Z in all K spatial representations S , spatial multi-class Spatial position embedding P S , and get the combined result Combine the results Input the spatial encoder and output the spatial features Separation, This global feature Input the classifier to obtain the classification result; the dense feature Input the segmentation decoder to obtain the segmentation result; The dense feature Through average pooling and the weight of the segmentation decoder w∈R K×D Multiply; weighted by the weight w and sum the dense features Each channel is generated: Will Normalize to the specified range and apply a global threshold to filter background pixels to obtain the segmentation result.

3. The time series remote sensing image segmentation method based on image-level label weakly supervised training according to claim 2, characterized in that: The training steps include exploring spatial perception cues: The overall category y∈[0,1] of the labeled time series remote sensing image K , through the weight w and Calculate its normalized fused CAM Use threshold μ l and μ h Filter foreground, background and uncertain areas; Time-intensive embedding Perform spatial clustering; Establish a set of positive prototypes of categories and negative prototype sets Given a mapping matrix C k , which represents the assignment between each pixel and its prototype and can be considered as an element of the transport prototype where N k Represents the number of pixels belonging to category k; 1 is a fully one vector, u and r are the mapping matrices C k Marginal projection of rows and columns; optimize the mapping matrix C by maximizing the following objective function k : is the corresponding time-dense embedding belonging to class k, and η controls the smoothness of the entropy regularization term; According to the distribution matrix and embedding, the nth p Prototypes are updated with momentum: Here α∈[0,1] is the momentum coefficient; Leveraging Fusion CAM Update the category prototype set P as a pseudo label pos and use 1- Update the negative prototype set P neg ; For each time-dense embedding and its most similar prototype The cosine distance is used to measure their similarity: Where τ represents the temperature parameter; the following third loss function is constructed: in is an indicator function, which is 1 if category k appears in the overall category of the labeled time series remote sensing image, otherwise it is 0; The visual transformer is trained using the first loss function, the second loss function, and the third loss function.

4. The time series remote sensing image segmentation method based on image-level label weakly supervised training according to claim 2, characterized in that: This training step includes a time-aware affinity propagation step: In this temporal encoder, the input tokens are normalized and projected into the query matrix Q∈R (K+T)×d and bond matrix K∈R (K +T)×d ;according to Sequence calculation self-attention N h ·N w : From self-attention Explicitly extracting time-to-category attention The attention Indicates the contribution of high-level representation to category recognition in different time segments; weighted by the following formula High-level representation This increases the variation between different semantic categories, thereby modeling pixel relationships more accurately. The temporal-aware pairwise affinity of a pixel at position i: here yes The standard deviation of Represents a local receptive field; it propagates time-aware representation through iteration The pairwise affinity is used to denoise the original CAM: Will and Align, and get the fourth loss function: The visual transformer is trained using the first loss function, the second loss function, and the fourth loss function.

5. A time series remote sensing image segmentation device with image-level label weak supervision training, characterized in that: include: The initial module takes the time series remote sensing images with the overall categories of the time series remote sensing images as training data and inputs them into the visual transformer; A training module extracts temporal features of the training data through a temporal encoder of the visual transformer, extracts spatial features of the temporal feature graph through a spatial encoder of the visual transformer, separates the spatial features to obtain dense features and global features, a classifier obtains a classification result based on the global features, a segmentation decoder obtains a segmentation result based on the dense features, and constructs a first loss function and a second loss function based on the gap between the overall category of the temporal remote sensing image and the classification result and the segmentation result, respectively, to train the visual transformer; The image segmentation module inputs the time series remote sensing image to be processed for semantic segmentation into the trained visual transformer, inputs the obtained dense features of the time series remote sensing image to be processed into the segmentation decoder, and obtains the semantic segmentation result of the time series remote sensing image to be processed.

6. The time series remote sensing image segmentation device for image-level label weakly supervised training according to claim 5, characterized in that: This initial module includes: The input of the visual transformer is a time series remote sensing image X∈R T×C×H×W , where T represents the length of the time series, C represents the number of channels of the remote sensing image, and H×X represents the spatial dimension; by mapping the time series remote sensing image W into multiple image blocks, the time series remote sensing image X is reshaped into an image and time-based multi-class tokens d is the feature dimension and K is the number of image categories, where h×w represents the spatial range of each image block; the time of the time series remote sensing image is added as the time position embedding P T ∈R T×d And concatenate the multiple types of tokens of the time to get the input of the time encoder This training module includes: The temporal features output by the time encoder before extraction The first K tokens of Will The dimension exchange obtains the feature Based on the feature Z S , we get spatial multi-class and spatial position embedding Combine this feature Z in all K spatial representations S , spatial multi-class Spatial position embedding P S , and get the combined result Combine the results Input the spatial encoder and output the spatial features Separation, This global feature Input the classifier to obtain the classification result; the dense feature Input the segmentation decoder to obtain the segmentation result; The dense feature Through average pooling and the weight of the segmentation decoder w∈R K×D Multiply; weight and sum the dense features by the weight w Each channel is generated: Will Normalize to the specified range and apply a global threshold to filter background pixels to obtain the segmentation result.

7. The time series remote sensing image segmentation device for image-level label weakly supervised training according to claim 2, characterized in that: This training module includes modules for exploring spatial perception cues: The overall category y∈[0,1] of the labeled time series remote sensing image K , through the weight w and Calculate its normalized fused CAM Use threshold μ l and μ h Filter foreground, background and uncertain areas; Time-intensive embedding Perform spatial clustering; Establish a set of positive prototypes of categories and negative prototype sets Given a mapping matrix C k , which represents the assignment between each pixel and its prototype and can be considered as an element of the transport prototype where N k Represents the number of pixels belonging to category k; 1 is a fully one vector, u and r are the mapping matrices C k Marginal projection of rows and columns; optimize the mapping matrix C by maximizing the following objective function k : is the corresponding time-dense embedding belonging to class k, and η controls the smoothness of the entropy regularization term; According to the distribution matrix and embedding, the nth p Prototypes are updated with momentum: Here α∈[0,1] is the momentum coefficient; Leveraging Fusion CAM Update the category prototype set P as a pseudo label pos and use 1- Update the negative prototype set P neg ; For each time-dense embedding and its most similar prototype The cosine distance is used to measure their similarity: Where τ represents the temperature parameter; the following third loss function is constructed: in is an indicator function, which is 1 if category k appears in the overall category of the labeled time series remote sensing image, otherwise it is 0; Training the visual transformer using the first loss function, the second loss function, and the third loss function; The training module includes the time-aware affinity propagation module: In this temporal encoder, the input tokens are normalized and projected into the query matrix Q∈R (K+T)×d and bond matrix K∈R (K +T)×d ;according to Sequence calculation self-attention N h ·N w : From self-attention Explicitly extracting time-to-category attention The attention Indicates the contribution of high-level representation to category recognition in different time segments; weighted by the following formula High-level representation This increases the variation between different semantic categories, thereby modeling pixel relationships more accurately. The temporal-aware pairwise affinity of a pixel at position i: here yes The standard deviation of Represents a local receptive field; it propagates time-aware representation through iteration The pairwise affinity is used to denoise the original CAM: Will and Align, and get the fourth loss function: The visual transformer is trained using the first loss function, the second loss function, and the fourth loss function.

8. An electronic device, characterized in that: It includes a temporal remote sensing image segmentation device with image-level label weakly supervised training as described in claims 5-7, and the electronic device is connected to an information display device, which is used to display the semantic segmentation results with display parameters and attributes set by the user or through an artificial intelligence model.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the time series remote sensing image segmentation method using image-level label weakly supervised training as recited in any one of claims 1 to 4.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the temporal remote sensing image segmentation method using image-level label weakly supervised training described in any one of claims 1 to 4 are implemented.

Citation Information

Cited By

  • Time sequence remote sensing land coverage change detection method and device, equipment and medium

    CN121789080A