Methods, systems, apparatus, and media for processing a time series of images

CN116681905BActive Publication Date: 2026-09-25ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310651456.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2026-09-25
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

然而,现有的对图像时间序列的语义分割模型通常依赖于对序列中的每个图像的像素级标注,这需要消耗大量的人力等资源

Benefits of technology

[0021]允许在单帧标注下实现模型训练;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116681905B_ABST
    Figure CN116681905B_ABST
Patent Text Reader

Abstract

A method for processing a time sequence of images is disclosed, including identifying a plurality of neighboring frames and a reference frame of a query frame of the time sequence of images; obtaining a short-range temporal representation of the query frame; obtaining a long-range temporal representation of the query frame; and obtaining an enhanced representation of the query frame based on the short-range temporal representation and the long-range temporal representation of the query frame. A corresponding system, apparatus, and computer-readable storage medium are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image processing, and more particularly to methods, systems, apparatus, and media for processing image time series. Background Technology

[0002] Currently, semantic segmentation of images has been widely implemented in applications such as autonomous driving, medical imaging, and industrial inspection. Image semantic segmentation refers to assigning labels to pixels in an image, while semantic segmentation of image time series refers to assigning labels to pixels in multiple images (e.g., all images) within the sequence. Furthermore, semantic segmentation models for image time series (e.g., videos) have also been developed. However, existing semantic segmentation models for image time series typically rely on pixel-level annotation of each image in the sequence, which requires significant human and other resources. Moreover, the performance of these models needs improvement, especially in scenarios such as remote sensing image time series.

[0003] Therefore, there is a need for solutions that can improve the processing of image time series. Summary of the Invention

[0004] One or more embodiments of this specification achieve their above-mentioned objectives through the following technical solutions.

[0005] In one aspect, a method for processing an image time series is provided, comprising: identifying a plurality of neighboring frames and a reference frame of a query frame of the image time series; obtaining a short-range time representation of the query frame based on the query frame and the plurality of neighboring frames; obtaining a long-range time representation of the query frame based on the short-range time representation of the query frame and the reference frame; and obtaining an enhanced representation of the query frame based on the short-range time representation and the long-range time representation of the query frame.

[0006] Preferably, the image time series is a remote sensing image time series.

[0007] Preferably, obtaining the short-range temporal representation of the query frame based on the query frame and the plurality of neighboring frames includes: performing feature extraction on the query frame and the plurality of neighboring frames using a visual backbone network to obtain their visual features; and using a spatiotemporal Transformer module to obtain the short-range temporal representation of the query frame based on the visual features of the query frame and the plurality of neighboring frames.

[0008] Preferably, each of the query frame and the plurality of neighboring frames has multiple feature maps with different step sizes, wherein the spatiotemporal Transformer module includes spatiotemporal Transformer units corresponding to the step sizes respectively.

[0009] Preferably, obtaining the short-range temporal representation of the query frame based on the visual features of the query frame and the plurality of neighboring frames using the spatiotemporal Transformer module includes: stacking features from each frame for each step; processing the features for the corresponding step using the spatiotemporal Transformer unit; upsampling only the features of the query frame to the same scale from the plurality of outputs from the spatiotemporal Transformer unit; and performing concatenation on the upsampled features to obtain the short-range temporal representation of the query frame.

[0010] Preferably, each spatiotemporal Transformer unit includes a 3D window multi-head self-attention mechanism, a hybrid feedforward network, and two layer normalization layers.

[0011] Preferably, obtaining the long-range temporal representation of the query frame based on the short-range temporal representation of the query frame and the reference frame includes: performing feature extraction on the reference frame to obtain its pixel representation; using a shared object-context representation module to obtain category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively; and constructing the input of a multi-head self-attention mechanism based at least in part on the short-range temporal representation of the query frame and the category-level representation of the reference frame to obtain the long-range temporal representation of the query frame.

[0012] Preferably, the input for constructing the multi-head self-attention mechanism based at least in part on the short-range temporal representation of the query frame and the category-level representation of the reference frame includes: using the short-range temporal representation of the query frame as the query of the multi-head self-attention mechanism, and using the concatenation of the global category representation and the category-level representation of the reference frame as the key and value of the multi-head self-attention mechanism.

[0013] Preferably, the global category representation is a set of learnable cluster centers maintained for each category.

[0014] Preferably, the method further includes: inputting the enhanced representation of the query frame into a semantic segmentation header to perform semantic segmentation on the query frame.

[0015] On the other hand, a method for performing semantic segmentation on a remote sensing image time series is also disclosed, comprising: for a query frame to be processed in the remote sensing image time series, identifying a plurality of neighboring frames and a reference frame of the query frame; obtaining a short-range temporal representation of the query frame based on the query frame and the plurality of neighboring frames; obtaining a long-range temporal representation of the query frame based on the short-range temporal representation of the query frame and the reference frame; obtaining an enhanced representation of the query frame based on the short-range temporal representation and the long-range temporal representation of the query frame; and performing semantic segmentation on the enhanced representation of the query frame using a semantic segmentation head.

[0016] Preferably, the method further includes: performing feature extraction on the reference frame to obtain its pixel representation; using a shared object-context representation module to obtain the category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively; and using the short-range temporal representation of the query frame as the query for a multi-head self-attention mechanism, and using the concatenation of the global category representation and the category-level representation of the reference frame as the key and value of the multi-head self-attention mechanism to obtain the long-range temporal representation of the query frame.

[0017] In another aspect, a system for processing image time series is also disclosed, comprising: means for identifying a plurality of neighboring frames and a reference frame of a query frame of the image time series; means for obtaining a short-range time representation of the query frame based on the query frame and the plurality of neighboring frames; means for obtaining a long-range time representation of the query frame based on the short-range time representation of the query frame and the reference frame; and means for obtaining an enhanced representation of the query frame based on the short-range time representation and the long-range time representation of the query frame.

[0018] In another aspect, an apparatus for processing image time series is provided, comprising a processor; and a memory coupled to the processor, the memory storing processor-executable instructions that, when executed by the processor, cause the processor to perform the method as described above.

[0019] In another aspect, a non-transient processor-readable storage medium is provided, comprising processor-executable instructions that, when executed by a processor, cause the processor to perform the method described above.

[0020] One or more embodiments of this specification can achieve one or more of the following technical effects:

[0021] Allows model training to be performed with single-frame annotation;

[0022] Reduce the resources required for annotation;

[0023] Improve model performance. Attached Figure Description

[0024] The above-described invention and the following detailed embodiments will be better understood when read in conjunction with the accompanying drawings. It should be noted that the drawings are merely examples of the claimed invention. In the drawings, the same reference numerals represent the same or similar elements.

[0025] Figure 1 A flowchart illustrating an example method for processing image time series according to an embodiment of this specification is shown.

[0026] Figure 2 A schematic diagram illustrating an overview of an example architecture for processing image time series according to embodiments of this specification is shown.

[0027] Figure 3 A schematic diagram illustrating an example processing procedure of the spatiotemporal Transformer module according to an embodiment of this specification.

[0028] Figure 4 A schematic diagram illustrating an example process for generating a long-range time representation of a query frame according to an embodiment of this specification.

[0029] Figure 5 A schematic flowchart of an example method for performing semantic segmentation on a time series of remote sensing images according to embodiments of this specification is shown.

[0030] Figure 6 A schematic flowchart of an example system for performing semantic segmentation on image time series according to an embodiment of this specification is shown.

[0031] Figure 7 A table showing a comparison of the performance of the method according to the embodiments of this specification with other methods.

[0032] Figure 8 A table showing the effect of each module on the model performance according to embodiments of this specification is provided.

[0033] Figure 9 A schematic block diagram of an apparatus for implementing a system or method according to one or more embodiments of this specification is shown. Detailed Implementation

[0034] The following detailed description is sufficient to enable any person skilled in the art to understand the technical content of one or more embodiments of this specification and to implement them accordingly. Furthermore, based on the specification, claims, and drawings disclosed in this specification, those skilled in the art can easily understand the objectives and advantages associated with one or more embodiments of this specification.

[0035] As mentioned earlier, in existing technologies, training semantic segmentation models or other machine learning models for image time series typically requires labeling every frame (or at least a large number of frames) of the image time series. This is not only the current practice but also an intuitive one. Training based on only one labeled frame of the image time series is counterintuitive. However, labeling a large number of frames in an image time series consumes significant resources, such as human and time resources.

[0036] In the embodiments of this specification, a scheme is proposed that can still be trained when only one frame in the sample of the image time series is labeled, and can achieve a similar effect to the traditional full labeling.

[0037] See Figure 1 The flowchart illustrates an example method 100 for processing image time series according to an embodiment of this specification. Figure 1 The method can be referred to Figure 2 To understand, among them Figure 2 A schematic diagram of an example architecture overview 200 for processing image time series according to an embodiment of this specification is shown.

[0038] Method 100 may include, in operation 102, identifying a plurality of neighboring frames and a reference frame of the query frame of the image time series.

[0039] An image time series is a sequence of images arranged in chronological order. Preferably, the time interval between two adjacent images in the image time series is a fixed value. Alternatively, the time interval may not be a fixed value. In a preferred embodiment, in addition to being arranged in chronological order, some or all of the images in the sequence may also include corresponding timestamps. It can be seen that the image time series includes time information, which is helpful for the analysis of machine learning models.

[0040] The most common form of image time series is a video or video clip, where each image in the video is a frame of that video. Other forms of image time series can also be used, such as a collection of images arranged in chronological order.

[0041] For ease of understanding, in the following text, images in an image time series will sometimes be referred to as "frames," regardless of whether the image time series is in the form of a video or video clip.

[0042] In a preferred embodiment, the images in the image time series are remote sensing images. Therefore, the remote sensing image time series can reflect the changes of ground features at a location over time. Due to the existence of the phenomenon of "same object, different spectrum; same spectrum, different object" in remote sensing images, performing semantic segmentation on remote sensing image time series is challenging. The embodiments of this specification, through the techniques described in detail below, can better utilize the temporal information (e.g., long-term and short-term temporal information) in remote sensing image time series, thereby improving the performance of remote sensing image semantic segmentation. Furthermore, the solutions of the embodiments of this specification also utilize the characteristics of ground features changing over time in remote sensing images (e.g., they usually change gradually rather than abruptly), exhibiting good performance when performing semantic segmentation on remote sensing image time series.

[0043] The image can also be any image other than remote sensing imagery.

[0044] During the prediction phase, the query frame is the frame in the image time series for which prediction is to be performed. During the training phase, the query frame is the frame in the image time series currently used for training, which may be labeled. By means of the embodiments described in this specification, even if only one frame in the image time series is labeled, that labeled frame can be used as the query frame, and the model can still be trained, achieving good technical results.

[0045] Neighboring frames are frames that are temporally close to the query frame in an image time series, such as multiple frames that are closest to the query frame.

[0046] A reference frame is a frame in the image time series that is temporally distant from the query frame, and it can be used to contribute long-term temporal information. Preferably, during the prediction phase, the frame in the image time series that is farthest from the query frame can be selected as the reference frame. During the training phase, frames in the image time series can be randomly selected as reference frames. By randomly selecting reference frames from the entire image time series during the training phase, the model in the embodiments of this specification potentially utilizes all data for training, not just the query frame and its neighboring frames. In this way, overfitting can be prevented to some extent.

[0047] Frames in the image time series that are not adjacent to the query frame can be selected as reference frames in any other manner that can be conceived by those skilled in the art.

[0048] During the prediction period, the query frame, multiple neighboring frames, and the reference frame together constitute the set of frames to be processed: Therefore, the frame set has n+2 frames, where n is the number of neighboring frames.

[0049] During training, the above set of frames, together with the labels of the query frames, constitutes the training samples:

[0050] Among them, I q Indicates a query frame, I ref Indicates reference frame, I adj (For example ) indicates a neighboring frame, L q Indicates the label of the query frame. (See reference) Figure 2 It shows query frame I according to an embodiment of this specification. q Reference Frame I ref and n neighboring frames I adj A schematic diagram.

[0051] In one example, the number of neighboring frames, n = 3. Preferably, the maximum distance between a neighboring frame and the query frame does not exceed 10 frames. In a preferred example, the time distance (e.g., the number of interval frames) between a neighboring frame and the query frame can be 3, 6, and 9. For example, it can be the -3rd, -6th, or -9th frame of the query frame (all before the query frame), or the 3rd, 6th, or 9th frame (all after the query frame), or frames selected both before and after the query frame. In other examples, the number and selection of neighboring frames can be set according to actual needs.

[0052] Compared to images, image time-series processing (e.g., semantic segmentation) can utilize inter-frame correlations. To effectively utilize both short-term and long-term inter-frame correlations simultaneously, two modules are proposed in the embodiments of this specification: a spatiotemporal Transformer module (e.g., ... Figure 2 The spacetime Transformer module 204, for details, can be found in [reference needed]. Figure 3 ), used to model short-term correlations; Reference Frame Context Enhancement (RFCE) module (e.g. Figure 2 The reference frame context enhancement module 208 is used to model long-term relevance and is optimized for query frames and reference frames.

[0053] In a more preferred embodiment, the system may further include a Global Category Context (GCC) module (e.g. Figure 2 The global category context module (206) is used to further supplement the RFCE module with category information.

[0054] The specific implementations of the Spatiotemporal Transformer module, the Reference Context Enhancement module, and the Global Category Context module will be described in detail below.

[0055] Method 100 may include, in operation 104, obtaining a short-range time representation of the query frame based on the query frame and the plurality of neighboring frames. The short-range time representation models the short-term correlation between the query frame and its neighboring frames.

[0056] Preferably, obtaining the short-range temporal representation of the query frame based on the query frame and the plurality of neighboring frames includes: performing feature extraction on the query frame and the plurality of neighboring frames using a visual backbone network to obtain their visual features; and using a spatiotemporal Transformer module to obtain the short-range temporal representation of the query frame based on the visual features of the query frame and the plurality of neighboring frames.

[0057] like Figure 2 As shown, the query frame and its multiple neighboring frames can be input into the feature extraction module 202 to generate a feature map for each frame. The feature extraction module 202 can be, for example, a visual backbone network.

[0058] The visual backbone network can be based on, for example, the ResNet-101 model. For instance, the visual backbone network can employ ResNet+FPN, with PPM as the neck. The dimension of the FPN can be, for example, 256 or other suitable values ​​conceived by those skilled in the art. Preferably, the backbone ResNet modules can be initialized using ImageNet pre-trained weights, and other parts of the modules can take random values.

[0059] This visual backbone network can typically generate representations with different strides. For example, strides can be 4, 8, 16, and 32. In this way, each frame can have multiple feature maps (e.g., 4 feature maps, each corresponding to one stride). Feature maps are sometimes simply referred to as features in this paper. For example, Figure 3 The features shown for step length 4, step length 8, step length 16, and step length 32 can be generated by the feature extraction module.

[0060] Compared to semantic segmentation of a single image, the key to semantic segmentation of time-series images lies in leveraging the temporal correlation between frames. In the embodiments of this specification, a Spatial-Temporal Transformer (STT) module is proposed to facilitate temporal modeling. This STT module comprises multiple parallel STT units. Each frame has multiple feature maps with different strides. The number of STT units is the same as the number of feature maps. Parallel STT units are used to process the corresponding strides separately, such as... Figure 3 As shown.

[0061] Specifically, first, for each step size, stack the feature maps of the query frame and its n neighboring frames. For example, see... Figure 2 The feature extraction module 202 outputs a stacked feature map 201 of four steps of the query frame and a stacked feature map 203 of four compensated neighboring frames.

[0062] Then, the feature maps of the stacked query frames and their n neighboring frames are input into the corresponding spatiotemporal Transformer units in the spatiotemporal Transformer module, so that the spatiotemporal Transformer units can process the features of the corresponding step size.

[0063] See Figure 3 This illustrates a schematic diagram of an example processing procedure 300 of the spatiotemporal Transformer module according to an embodiment of this specification. Figure 3As shown, the spatiotemporal Transformer module mainly includes multiple spatiotemporal Transformer units 302, 304, 306, and 308, which are used to process features 301 with a step size of 4, features 303 with a step size of 8, features 305 with a step size of 16, and features 307 with a step size of 32, respectively, and generate outputs 309, 311, 313, and 315. The dimensions of the spatiotemporal Transformer units 302, 304, 306, and 308 can be, for example, 32, 64, 128, and 256, or other suitable values.

[0064] exist Figure 3 The box below shows an example structure of a preferred embodiment of the spatiotemporal Transformer unit. Each spatiotemporal Transformer unit may include a 3D window multi-head self-attention (DW-MSA) mechanism, a hybrid feedforward network (Mix-FFN), and multiple (e.g., two) layer normalization (LN) layers.

[0065] like Figure 3 As shown, the Layer Normalization (LN) layer can be set before 3D W-MSA and Mix-FFN, respectively.

[0066] 3D W-MSA can uniformly segment a 3D input feature map to form a set of non-overlapping cubes and apply MSA to them. The window size of 3D W-MSA can be, for example, 4x7x7. Further details about 3D W-MSA can be found in the paper "Video Swin Transformer" published by Ze Liu et al. in 2021.

[0067] Mix-FFN introduces a depthwise 3×3 convolution between two MLPs to connect these non-overlapping cubes. Further details about Mix-FFN can be found in the 2021 paper "Segformer: Simple and efficient design for semantic segmentation with transformers" by Enze Xie et al. and the 2019 paper "Billion-scale semi-supervised learning for image classification" by I. Zeki Yalniz et al.

[0068] Other structures that can be conceived by those skilled in the art can be used to implement the spatiotemporal Transformer unit.

[0069] like Figure 3 As shown, for the outputs 309, 311, 313, and 315 of the spatiotemporal Transformer unit, only the features of the query frames in the stacked outputs (such as...) can be taken. Figure 3 The lighter-colored features in the stacked output (such as features 317, 319, 321, and 323) are not directly concatenated because features 317, 319, 321, and 323 correspond to inputs of different lengths and therefore have different scales.

[0070] Therefore, to concatenate them, the features of each step size can be upsampled to have the same scale as the features of the largest step size, and only the features of the query frames in the stack can be taken. For example, as Figure 3 As shown, features 319, 321 and 323 can be upsampled so that the upsampled features have the same scale as feature 317.

[0071] Subsequently, the features of the upsampled query frames from each spatiotemporal Transformer unit are concatenated to obtain the short-range time representation 325 of the query frame, which can be represented as: As shown in the following formula:

[0072]

[0073] The operation of the spacetime Transformer module is represented as f. STT The operation of the feature extraction module is represented as f FE .symbol This represents the composition operator. It can be seen that through the combined operation of the feature extraction module and the spatiotemporal Transformer module, it is possible to extract information from the query frame I. q and adjacent frame I adj Obtain the short-range time representation of the query frame.

[0074] Method 100 may include: in operation 106, obtaining a long-range time representation of the query frame based on the short-range time representation of the query frame and the reference frame.

[0075] The spatiotemporal Transformer can use a fixed-size window to fuse temporal information within a small spatial region. However, when there are objects with significant variations in the time series of images, it may be difficult to capture their long-term correlations using only the spatiotemporal Transformer.

[0076] In the embodiments described in this specification, one alternative approach is to provide feature maps of distantly spaced frames to the spatiotemporal Transformer model in the long-term model and significantly expand the pane size to capture moving objects. However, this could ultimately lead to prohibitively high computational costs.

[0077] In the preferred embodiments of this specification, a more preferred approach is proposed, namely, using Reference Frame Context Enhancement (RFCE) to model long-term relationships to generate long-range temporal representations of query frames, such as... Figure 2 As shown on the right.

[0078] See Figure 4 This illustrates a schematic diagram of an example process 400 for generating a long-range time representation of a query frame according to an embodiment of this specification. This is achieved through the following process:

[0079] In operation 402, feature extraction can be performed on the reference frame to obtain its pixel representation.

[0080] Feature extraction of the reference frame can be performed using the feature extraction model 202 described above. For example, it can be performed using a visual backbone network as described above. Preferably, feature extraction of the reference frame can be performed in parallel with feature extraction of the query frame and the plurality of neighboring frames. Feature extraction of the query frame, neighboring frames, and reference frame can be performed independently, sequentially, interleaved, in parallel, or in any other suitable manner.

[0081] Alternatively, feature extraction of the reference frame can be performed using a separate feature extraction module.

[0082] Specifically, refer to frame I ref pixel representation It can be determined by the following formula:

[0083]

[0084] Where f FE Indicates the feature extraction module (e.g.) Figure 2 The feature extraction module 202 (such as a visual backbone network, etc.) is shown above.

[0085] In operation 404, a shared Object-Contextual Representation (OCR) module can be used to obtain category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively. Further details on object-contextual representations can be found in the 2020 paper "Object-contextual representations for semantic segmentation" by Yuhui Yuan et al.

[0086] Specifically, a shared object-context representation module can be used to extract short-range temporal representations from the query frame. Pixel representation of the reference frame Extract query frame I q and reference frame I ref Category-level representation and

[0087] In operation 406, the input to the multi-head self-attention mechanism can be constructed at least in part based on the short-range temporal representation of the query frame and the category-level representation of the reference frame.

[0088] In a preferred embodiment, the input to the multi-head self-attention mechanism can be constructed based on the short-range temporal representation of the query frame, the global category representation, and the category-level representation of the reference frame. The self-attention mechanism may fail when the query frame contains a category not present in the reference frame, because information about the corresponding category is missing from the reference frame. This problem can be mitigated by using a global category representation. The details of the global category representation will be described in detail below.

[0089] For example, the short-range temporal representation of the query frame can be used as the query, and the concatenation of the global category representation and the category-level representation of the reference frame can be used as the key and value input into a multi-head self-attention mechanism. Through the multi-head self-attention mechanism, the long-range temporal representation of the query frame can be obtained.

[0090] Specifically, the short-range time representation of the query frame can be used. As a query, the global category representation G and the category-level representation of the reference frame are used. The concatenation of the query, key, and value is used as the key and value. Subsequently, a multi-head self-attention mechanism is run on the above query, key, and value to generate a long-range temporal representation of the query frame. As shown in the following formula:

[0091]

[0092]

[0093] Concat represents the concatenation operation, while Softmax represents the Softmax operation.

[0094] In an alternative embodiment, instead of using a global category representation, the input to the multi-head self-attention mechanism can be constructed solely based on the short-range temporal representation of the query frame and the category-level representation of the reference frame. For example, the short-range temporal representation of the query frame can be used as the query, and the category-level representation of the reference frame can be directly used as the key and value inputs to the multi-head self-attention mechanism.

[0095] In this scenario, when the query frame contains categories not included in the reference frame, the system may perform worse compared to embodiments using global category representations. However, avoiding global category representations can reduce system complexity, improve system speed, and reduce resource consumption.

[0096] Method 100 may include, in operation 108, obtaining an enhanced representation of the query frame based on the short-range time representation and the long-range time representation of the query frame.

[0097] Specifically, the query frame I q Short-range time representation and long-range time representation Concatenate them together to obtain an enhanced representation of the query frame. As shown in the following formula:

[0098]

[0099] After obtaining the enhanced representation of the query frame, this enhanced representation can be used to perform operations, such as image recognition, object detection, semantic segmentation, and so on. The following section will use semantic segmentation as an example.

[0100] Method 100 may include: optionally, inputting an enhanced representation of the query frame into a semantic segmentation header to produce a semantic segmentation result for the query frame.

[0101] During the prediction phase, this semantic segmentation result serves as the semantic segmentation result for the query frame in the image time series.

[0102] During the training phase, the semantic segmentation result and the tag of the query frame (e.g., tag L) can be used. q The loss is calculated using the cross-entropy loss function, and training is performed based on this loss. For example, the cross-entropy loss function can be used to calculate the loss, as shown in the following equation:

[0103]

[0104] in Represents the cross-entropy loss, CE is the cross-entropy loss function, and φ Seq This indicates a semantic segmentation operation. This represents an enhanced representation of the query frame, while L q The label represents the query frame.

[0105] Other suitable loss calculation methods that can be conceived by those skilled in the art can be used to calculate the loss.

[0106] In this way, even if only one image in the image time series is labeled (i.e., has a label), that labeled image can be used as the query image, and training can be performed using the method described above. In other words, using the scheme of this specification's embodiments, even with only a single labeled image, it is possible to achieve performance similar to when every image in the image time series is labeled. Therefore, the embodiments of this specification can reduce the requirements for training samples and reduce the resources required for labeling.

[0107] The solutions in the embodiments of this specification mainly utilize the following principles to achieve better performance:

[0108] First, by using the reference frame context enhancement module (regardless of whether a global category representation is used), long-term information from the reference frame can effectively improve segmentation accuracy by providing richer prior context and implicitly mitigating overfitting.

[0109] Second, by using a global category representation, category information can be taken into account, which in particular improves performance when there is no relevant category in the reference frame.

[0110] Third, for the short-range time representation of the query frame. Pixel representation of the reference frame Pixel-level correlation modeling is computationally very expensive. In the embodiments of this specification, a compact representation of correlation modeling can be generated by region (e.g., pooling) or category.

[0111] Fourth, by using an object-context representation module, category-level compact representations typically exhibit performance advantages compared to region-based representations. Therefore, an object-context representation module is utilized in the RFCE module of the embodiments in this specification. The object-context representation module requires an auxiliary segmentation head φAux to predict the coarse segmentation map. In this work, a short-range temporal representation of the query frame can be used. and the tag L of the query frame q To train the auxiliary segmentation head φAux. The training loss... It can be expressed by the following formula:

[0112]

[0113] The following describes an example process for obtaining a global category representation according to a preferred embodiment of this specification.

[0114] As mentioned above, the self-attention mechanism may fail when the query frame contains a category that is not included in the reference frame, because the reference frame lacks information about the corresponding category.

[0115] To address this issue, this specification proposes a Global Category Context (GCC) module to simulate a global category representation in its embodiments. The basic idea is to maintain a set of learnable cluster centers for each category, i.e., a tensor of shape T×C×D, where T, C, and D represent the number of clusters, the number of categories, and the dimension of the features in each category, respectively. Figure 2 As shown. This tensor is the global category representation. Global category representation The number of clusters T can be set to 3 or other suitable values ​​that can be conceived by those skilled in the art.

[0116] like Figure 2 As shown, during the training phase, we first obtain the query frame. The category representation. Then, for each category j, from the tensor Extract from all clusters of this category The nearest cluster center v j 'i' represents the cluster index. This represents the i-th cluster center of category j. In the embodiments of this specification, a ground truth value can be used to filter invalid categories, or to select categories. It is the tag L q A collection of categories.

[0117]

[0118] in

[0119] Then, for each category The category representation of the query frame and the nearest cluster center of the corresponding category. j The mean squared error (MSE) loss is calculated as follows:

[0120]

[0121] It is important to note that Used only for optimizing global category representation Without query frame The category represents the propagation gradient. For example... Figure 2 As shown in the figure, SG represents the stopping gradient.

[0122] It should be noted that this only runs during the training phase. Figure 2 The portion within the dashed box is used to update the global category representation. During the prediction phase, the global category representation... It is fixed, therefore no operation is required. Figure 2 The part within the dashed box in the image is not used; instead, a fixed, trained algorithm is used. Use the value instead.

[0123] Global category contexts can be viewed as a complement to RFCEs, so we concatenate their outputs together before feeding them into multi-head attention, such as... Figure 2 As shown.

[0124] The final training loss is shown in the following formula:

[0125]

[0126] Here, α and β are hyperparameters. In the preferred example, hyperparameters α and β can take values ​​of 1.0 and 0.5, respectively. In other examples, hyperparameters α and β can take other values.

[0127] Preferably, the method described in the embodiments of this specification can be combined with a pseudo-label based method. In the pseudo-label based method, pseudo-labels are first generated, and then used to perform subsequent processing. Therefore, pseudo-labels generated by the pseudo-label based method can be used as labels for query frames, and processed using the method according to the embodiments of this specification.

[0128] See Figure 5 The diagram illustrates a schematic flowchart of an example method 500 for performing semantic segmentation on a time series of remote sensing images according to an embodiment of this specification.

[0129] Method 500 may include, in operation 502, identifying multiple neighboring frames and a reference frame of a query frame of a remote sensing image time series. For details, please refer to the description of operation 102 above.

[0130] Method 500 may further include: in operation 504, obtaining a short-range temporal representation of the query frame based on the query frame and multiple neighboring frames. Specifically, a visual backbone network can be used to perform feature extraction on the query frame and multiple neighboring frames to obtain their visual features. Furthermore, a spatiotemporal Transformer module can be used to obtain the short-range temporal representation of the query frame based on the visual features of the query frame and multiple neighboring frames. Each frame in the query frame and multiple neighboring frames contains multiple feature maps with different strides, and the spatiotemporal Transformer module includes spatiotemporal Transformer units corresponding to each stride. Specifically, this can be done by: stacking features from each frame for each stride; processing the features corresponding to the stride using the spatiotemporal Transformer unit; upsampling only the features of the query frame to the same scale from the multiple outputs from the spatiotemporal Transformer unit; and concatenating the upsampled features to obtain the short-range temporal representation of the query frame. Each spatiotemporal Transformer unit includes a 3D window multi-head self-attention mechanism, a hybrid feedforward network, and two layers of normalization. For details, refer to the description of operation 104 above.

[0131] Method 500 may further include: in operation 506, obtaining a long-range temporal representation of the query frame based on the short-range temporal representation of the query frame and the reference frame. Specifically, this may include the following operations: performing feature extraction on the reference frame to obtain its pixel representation; using a shared object-context representation module to obtain category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively; and constructing the input of a multi-head self-attention mechanism based at least in part on the short-range temporal representation of the query frame and the category-level representation of the reference frame to obtain the long-range temporal representation of the query frame. Preferably, the short-range temporal representation of the query frame can be used as the query for the multi-head self-attention mechanism, and the concatenation of the global category representation and the category-level representation of the reference frame can be used as the key and value of the multi-head self-attention mechanism. Alternatively, the short-range temporal representation of the query frame can be used as the query for the multi-head self-attention mechanism, and the category-level representation of the reference frame can be used as the key and value of the multi-head self-attention mechanism. The global category representation is a set of learnable cluster centers maintained for each category. The global category representation may be fixed during the prediction phase. For specific details, please refer to the description of operation 106 above.

[0132] Method 500 may further include: in operation 508, obtaining an enhanced representation of the query frame based on the short-range time representation and the long-range time representation of the query frame. For example, the enhanced representation can be obtained by concatenating the short-range time representation and the long-range time representation. For details, please refer to the description of operation 108 above.

[0133] Method 500 may further include, in operation 510, inputting an enhanced representation of the query frame into a semantic segmentation header to perform semantic segmentation on the query frame. See the description above for details.

[0134] See Figure 6 The diagram shows a schematic flowchart of an example system 600 for performing semantic segmentation on image time series according to an embodiment of this specification.

[0135] System 600 may include: a frame identification module 602, which can be used to identify multiple neighboring frames and a reference frame of a query frame in a remote sensing image time series. For details, please refer to the description of operation 102 above.

[0136] System 600 may further include: a short-range time representation generation module 604, which can be used to obtain a short-range time representation of the query frame based on the query frame and multiple neighboring frames. For details, please refer to the description of operation 104 above.

[0137] System 600 may further include a long-range time representation generation module 606, which can be used to obtain the long-range time representation of the query frame based on the short-range time representation of the query frame and the reference frame. For details, please refer to the description of operation 106 above.

[0138] System 600 may further include: an enhanced representation generation module 608, which can be used to obtain an enhanced representation of the query frame based on the short-range time representation and the long-range time representation of the query frame. For details, please refer to the description of operation 108 above.

[0139] See Figure 7 Table 700 shows a comparison of the performance of the method according to embodiments of this specification with other methods. As shown in the table, all methods use the same training and testing settings: a batch size of 8, a random cropping size of 512 × 1024, 100 training epochs, and a single-scale evaluation strategy. It can be seen that this method achieves better mIoU compared to other methods.

[0140] See Figure 8 The table illustrates the impact of the Spatiotemporal Transformer module (STT), the Reference Frame Context Enhancement (RFCE) module, and the Global Class Context (GCC) module on model performance according to embodiments of this specification. Figure 8 As can be seen, the STT, RFCE, and GCC modules all improve the performance of this method.

[0141] Figure 9A schematic block diagram is shown of an apparatus 900 for implementing a system (such as system 1000 above) according to one or more embodiments of this specification or performing a method (such as method 800 above) as described in one or more embodiments of this specification. The apparatus may include a processor 910 and a memory 915, the processor being configured to perform operations of any of the methods described above. The memory may store, for example, acquired data, algorithms and / or models used, and intermediate data generated during operation, etc.

[0142] The device 900 may include a network connectivity element 925, such as a network connectivity device that can connect to other devices via a wired or wireless connection. The wireless connection may be, for example, a WiFi connection, a Bluetooth connection, or a 3G / 4G / 5G network connection. The device may also receive user input from other devices or transmit data to other devices for display via the network connectivity element.

[0143] The device may also optionally include other peripheral components 920, such as input devices (e.g., keyboard, mouse) and output devices (e.g., monitor). It can also output relevant information to the user via the output devices.

[0144] Each of these modules can communicate with each other directly or indirectly, for example, via one or more buses (e.g., bus 905).

[0145] Furthermore, embodiments of this specification also disclose an apparatus including a processor; and a memory coupled to the processor, the memory storing processor-executable instructions that, when executed by the processor, cause the processor to perform operations implementing the methods of one or more embodiments of this specification.

[0146] Furthermore, embodiments of this specification also disclose a system including means for implementing various operations of the methods of one or more embodiments of this specification.

[0147] It is understood that the methods according to one or more embodiments of this specification can be implemented in software, firmware, or a combination thereof.

[0148] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. In particular, for the apparatus and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0149] It should be understood that the foregoing description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0150] It should be understood that the use of a singular form to describe an element or to show only one element in the accompanying drawings does not imply that the number of such element is limited to one. Furthermore, modules or elements described or shown as separate herein may be combined into a single module or element, and modules or elements described or shown as single herein may be broken down into multiple modules or elements.

[0151] In this specification, unless otherwise specified, “nearly,” “almost,” or “approximately” (if used) means a deviation of no more than 10%; preferably, a deviation of no more than 5%; more preferably, a deviation of no more than 1%.

[0152] It should also be understood that the terminology and expressions used herein are for descriptive purposes only, and one or more embodiments described herein should not be limited to these terms and expressions. The use of these terms and expressions does not exclude any illustrative and descriptive equivalent features (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be considered to cover all such equivalents.

[0153] Similarly, it should be noted that although specific embodiments have been described with reference to the present invention, those skilled in the art should recognize that the above embodiments are merely illustrative of one or more embodiments of this specification, and various equivalent changes or substitutions can be made without departing from the spirit of the invention. Therefore, any changes or modifications to the above embodiments within the scope of the essential spirit of the invention will fall within the scope of the claims of this application.

Claims

1. A method for processing image time series, comprising: A plurality of neighboring frames and a reference frame are used to identify a query frame in the image time series, wherein the neighboring frames are frames that are temporally close to the query frame in the image time series, and the reference frame is a frame that is not temporally close to the query frame in the image time series. The short-range time representation of the query frame is obtained based on the query frame and the plurality of neighboring frames, wherein the short-range time representation models the short-term correlation between the query frame and its neighboring frames. Feature extraction is performed on the reference frame to obtain its pixel representation; The shared object-context representation module is used to obtain the category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively; The input of the multi-head self-attention mechanism is constructed at least in part based on the short-range temporal representation of the query frame and the category-level representation of the reference frame to obtain the long-range temporal representation of the query frame, wherein the long-range temporal representation models the long-term relationship between the query frame and the reference frame. as well as An enhanced representation of the query frame is obtained based on the short-range time representation and the long-range time representation of the query frame.

2. The method as described in claim 1, wherein the image time series is a remote sensing image time series.

3. The method of claim 1, wherein obtaining the short-range time representation of the query frame based on the query frame and the plurality of neighboring frames comprises: The visual backbone network is used to perform feature extraction on the query frame and the multiple neighboring frames to obtain their visual features; as well as The Spatiotemporal Transformer module is used to obtain the short-range temporal representation of the query frame based on the visual features of the query frame and the multiple neighboring frames.

4. The method of claim 3, wherein each of the query frame and the plurality of neighboring frames has a plurality of feature maps with different step sizes, wherein the spatiotemporal Transformer module includes spatiotemporal Transformer units corresponding to the step sizes respectively.

5. The method of claim 4, wherein obtaining the short-range temporal representation of the query frame using the spatiotemporal Transformer module based on the visual features of the query frame and the plurality of neighboring frames includes: For each step, features from each frame are stacked; Use the spatiotemporal Transformer unit to process features corresponding to the step size; In the multiple outputs from the spatiotemporal Transformer unit, only the features of the query frame are upsampled to the same scale; as well as The upsampled features are concatenated to obtain a short-range time representation of the query frame.

6. The method of claim 5, wherein each spatiotemporal Transformer unit comprises a 3D window multi-head self-attention mechanism, a hybrid feedforward network, and two layer normalization layers.

7. The method of claim 1, wherein the input for constructing the multi-head self-attention mechanism based at least in part on the short-range temporal representation of the query frame and the category-level representation of the reference frame comprises: The short-range time representation of the query frame is used as the query for the multi-head self-attention mechanism, and the concatenation of the global category representation and the category-level representation of the reference frame is used as the key and value for the multi-head self-attention mechanism.

8. The method of claim 7, wherein the global category representation is a set of learnable cluster centers maintained for each category.

9. The method of claim 1, further comprising: The enhanced representation of the query frame is input into the semantic segmentation header to perform semantic segmentation on the query frame.

10. A method for performing semantic segmentation on a time series of remote sensing images, comprising: For a query frame to be processed in the remote sensing image time series, identify multiple neighboring frames and a reference frame of the query frame, wherein the neighboring frames are frames that are temporally close to the query frame in the remote sensing image time series, and the reference frame is a frame that is not temporally close to the query frame in the remote sensing image time series. The short-range time representation of the query frame is obtained based on the query frame and the plurality of neighboring frames, wherein the short-range time representation models the short-term correlation between the query frame and its neighboring frames. Feature extraction is performed on the reference frame to obtain its pixel representation; The shared object-context representation module is used to obtain the category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively; The input of the multi-head self-attention mechanism is constructed at least in part based on the short-range temporal representation of the query frame and the category-level representation of the reference frame to obtain the long-range temporal representation of the query frame, wherein the long-range temporal representation models the long-term relationship between the query frame and the reference frame. An enhanced representation of the query frame is obtained based on the short-range time representation and the long-range time representation of the query frame; as well as Semantic segmentation is performed on the enhanced representation of the query frame using a semantic segmentation header.

11. The method of claim 10, wherein the input for constructing the multi-head self-attention mechanism based at least in part on the short-range temporal representation of the query frame and the category-level representation of the reference frame comprises: The short-range time representation of the query frame is used as the query for the multi-head self-attention mechanism, and the concatenation of the global category representation and the category-level representation of the reference frame is used as the key and value for the multi-head self-attention mechanism.

12. A system for processing image time series, comprising: A means for identifying a plurality of neighboring frames and a reference frame of a query frame in an image time series, wherein the neighboring frames are frames that are temporally adjacent to the query frame in the image time series, and the reference frame is a frame that is not temporally adjacent to the query frame in the image time series. An apparatus for obtaining a short-range time representation of a query frame based on the query frame and the plurality of neighboring frames, wherein the short-range time representation models the short-range correlation between the query frame and its neighboring frames. A means for performing feature extraction on the reference frame to obtain its pixel representation; An apparatus for using a shared object-context representation module to obtain category-level representations of the query frame and the reference frame from the short-range temporal representation of the query frame and the pixel representation of the reference frame, respectively; An apparatus for constructing a multi-head self-attention mechanism based at least in part on the input of the short-range temporal representation of the query frame and the category-level representation of the reference frame to obtain a long-range temporal representation of the query frame, wherein the long-range temporal representation models the long-term relationship between the query frame and the reference frame. as well as An apparatus for obtaining an enhanced representation of the query frame based on the short-range time representation and the long-range time representation of the query frame.

13. An apparatus for processing image time series, comprising: processor; as well as A memory coupled to the processor stores processor-executable instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1-9.

14. A non-transient processor-readable storage medium comprising processor-executable instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Feature extraction method and device, equipment, storage medium and program product

    CN112861830A

  • Continuous sign language recognition method and device

    CN115393949A

  • Semantic segmentation system for medical image sequence

    CN115861616A