Remote sensing image multi-task joint prediction method and device and medium
By introducing a task cross-fusion module based on state space model and a two-stage decoding optimization strategy in remote sensing image multitask learning, the problem of difficulty in capturing synergistic relationships between tasks and limited convolution operations is solved, and efficient and accurate multitasking prediction effect is achieved.
Patent Information
- Application Number
- CN202510057514.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-02
AI Technical Summary
When the existing multi-task learning method processes remote sensing images, it is difficult to effectively capture the synergistic relationship between tasks, and the convolution operation is limited by the receptive field, making it difficult to capture local details and global features at the same time, resulting in a degradation of model performance.
A multi-task joint learning framework based on state space model is designed, and the correlation characteristics between tasks are explicitly modeled through the task cross-fusion module, and a two-stage decoding optimization strategy is adopted, combined with an auxiliary loss mechanism to balance the performance between multi-tasks.
It significantly improves the information interaction and collaborative learning ability between multitasks, improves the generalization ability and robustness of the model, and can efficiently and accurately complete multitask prediction of remote sensing images.
Smart Images

Figure CN119919665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image analysis and processing, and in particular to a remote sensing image multi-task joint prediction method, device and medium. Background Art
[0002] Remote sensing image analysis, as an important technology in the fields of geographic information, environmental monitoring and urban planning, has developed rapidly in recent years. Traditional single-task learning methods usually focus on solving a single task, such as semantic segmentation, height estimation or boundary detection. Although these methods have made some progress in specific tasks, they have failed to fully utilize the potential correlation between tasks. Multi-task learning can significantly improve the overall performance and computational efficiency of the model by sharing the feature extraction network in a single model and optimizing multiple tasks at the same time. However, existing multi-task learning methods still face many technical bottlenecks when processing remote sensing images. Most methods use convolutional neural networks (CNNs) or encoder-decoder structures to extract common features in a shared encoder, but lack explicit modeling of the interdependencies and correlation information between tasks. This implicit feature sharing method leads to insufficient information exchange between tasks, making it difficult to effectively capture the synergistic relationship between semantic segmentation and height estimation. In addition, the convolution operation is limited by the receptive field, making it difficult to simultaneously capture local details and global features in remote sensing images. This limitation significantly reduces model performance, especially when dealing with complex terrain or large-scale targets.
[0003] The optimization strategies of existing methods usually define loss functions in a task-independent manner, fail to balance the optimization requirements among multiple tasks, and are prone to task competition, that is, the performance improvement of some tasks may be at the expense of other tasks. This imbalance in optimization objectives further limits the performance of multi-task learning methods in practical applications. At the same time, since the feature extraction and decoding modules of each task are often designed independently, the computing resource overhead of existing models is significantly increased. Especially in large-scale remote sensing data processing scenarios, existing technologies are difficult to achieve a balance between performance and efficiency. Therefore, the existing multi-task learning framework still has a lot of room for improvement in modeling the relationship between tasks, optimizing global and local features, and improving resource utilization. An innovative method is urgently needed to overcome these technical obstacles to meet the actual needs of multi-task learning of remote sensing images. Summary of the invention
[0004] The purpose of the present invention is to provide a remote sensing image multi-task joint prediction method, device and medium, which designs a multi-task joint learning framework based on a state space model, fully integrates the correlation features between tasks, enhances task interaction, and thus improves the multi-task collaborative learning ability of the model; secondly, by designing a two-stage decoding optimization strategy and combining an auxiliary loss mechanism, the performance between multiple tasks is effectively balanced, and the generalization ability and robustness of the model are significantly improved; the present invention can efficiently and accurately complete the multi-task prediction of remote sensing images, and meet the actual needs of large-scale remote sensing data processing.
[0005] In order to achieve the above object, the present invention provides the following technical solutions: In a first aspect, the present invention provides a remote sensing image multi-task joint prediction method, wherein the multi-task includes a semantic segmentation task, a height estimation task and an edge detection task, and the method includes: Input remote sensing image data set, perform pixel value normalization and size adjustment on the input image; Design a shared encoder to extract common features from input images; Design an initial decoder to recover high-resolution predictions from features extracted by the shared encoder; Construct the initial loss function; The feature tensors obtained from the initial decoders of each task are input into the task cross-fusion module based on the state-space model for cross-task fusion; Based on the features output by the task cross-fusion module, an optimized decoder is designed to further refine the task features and generate the final prediction results; Construct the final loss function.
[0006] As an embodiment, the input remote sensing image data set performs pixel value normalization and size adjustment on the input image in sequence, including: The remote sensing image contains four channels, RGB and near infrared, as well as corresponding semantic segmentation and elevation maps. The data format is required to be a standard image format. Each image has four channels and a high resolution. Scale the value of each pixel to between 0 and 1 using the formula: ,in, is the original pixel value, are the minimum and maximum pixel values in the image respectively; then all images are resized to a uniform size to ensure the size consistency of the input images; the images are scaled by bilinear interpolation to keep the main structure of the image unchanged.
[0007] As an embodiment, the design shares an encoder to extract common features from an input image, including: The pre-trained convolutional neural network is used as a shared encoder, the first few layers of the network extract low-level features, and the later layers extract high-level features; The resized remote sensing image is input into the shared encoder to extract common features; the output feature tensor shape is N×1024×16×16, where N is the batch size, 1024 is the number of channels, and 16 is the size of the spatial dimension; Batch Normalization is applied after each convolution operation to increase training speed, reduce the model's dependence on initialization, and ensure the stability of the training process; The activation function after each convolution layer uses ReLU (rectifier linear unit) to introduce nonlinearity and enhance the representation ability of the network.
[0008] As an embodiment, the designing of the initial decoder to recover the high-resolution prediction result from the features extracted by the shared encoder includes: An independent initial decoder is designed for each task. The decoder adopts a convolutional neural network (CNN) structure, which consists of two layers of 3×3 convolution layers and an upsampling layer. The number of convolution kernels is set to 128 to maintain a high feature expression capability. The upsampling operation uses bilinear interpolation to gradually restore the low-resolution feature map to the size of the original image. The feature map output by each decoder is passed through a task-specific prediction head to generate the final task output; for semantic segmentation, the Softmax activation function is used to generate the category prediction of each pixel; for height estimation, the linear activation function is used to output the height value of each pixel; for boundary detection, the Sigmoid activation function is used to output the probability of whether each pixel is a boundary point.
[0009] As an embodiment, constructing the initial loss function includes: Based on the output of the initial decoder , construct the initial loss function; The semantic segmentation task uses the cross entropy loss function: , Among them, N is the number of pixels in the image, C is the number of categories of semantic segmentation, is the true category label for semantic segmentation; The height estimation task uses the mean squared error (MSE) loss function: , in, Represents the true label value of the height estimation task; Edge detection uses a gradient-based loss function: , in, represents the gradient of the real label edge detection task, The gradient of the initial prediction result; Based on the loss functions of the above subtasks, the initial loss function is established, and the formula is expressed as: , in, They represent the adjustment parameters of each task, which are used to ensure that the loss amount of each task remains in the same value range.
[0010] As an embodiment, the step of inputting the feature tensors obtained from the initial decoders of each task into the task cross-fusion module based on the state space model for cross-task fusion includes: The feature tensor shape is N×128×32×32; the features from the shared encoder and the multi-task features from the initial decoder Input to the task cross-fusion module, where They represent the initial decoder features of semantic segmentation, height estimation, and edge detection respectively; after convolution and interpolation operations, the above features enter the spectral dimension and spatial dimension state space fusion module; The state space model calculation process can be expressed as: , , Where A, B, C are the parameter matrices of the state space model, is the input feature, is the hidden state, is the output feature.
[0011] As an embodiment, based on the features output by the task cross-fusion module, designing an optimized decoder to further refine the task features and generate the final prediction result includes: The characteristics of the shared encoder , the characteristics of the task cross-fusion module and the characteristics of the initial encoder , are input into the optimized decoder; Characteristics of a shared encoder , the characteristics of the task cross-fusion module After convolution, the feature is input into the state space model, and the output feature is convolved and interpolated again, which is consistent with the feature of the initial encoder. Feature concatenation is used to output the final task prediction results.
[0012] As an embodiment, constructing the final loss function includes: The loss function definition of each task of the optimized decoder is the same as the initial loss function, denoted as , the final loss function includes the initial loss function, and the formula is expressed as: , in, It is an adjustment parameter used to control the proportion of the initial loss function in the final loss function.
[0013] In a second aspect, the present invention provides a remote sensing image multi-task joint prediction device, comprising: Image preprocessing module: used to perform pixel value normalization and size adjustment on the input image; Shared encoder module: used to extract common features from the input image; Initial decoder module: used to recover high-resolution prediction results from the features extracted by the shared encoder; The first analysis module is used to perform the calculation of the initial loss function and output the initial results of multi-task prediction; Task cross-fusion module: used to fuse the feature tensors obtained from the initial decoders of each task across tasks to generate shared features across tasks; Optimized decoder module: used to further refine the task features and generate the final prediction results based on the features output by the task cross-fusion module; The second analysis module is used to perform the calculation of the final loss function and output the final result of multi-task prediction.
[0014] In a third aspect, the present invention provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the remote sensing image multi-task joint prediction method described in the first aspect is executed.
[0015] In summary, the beneficial effects of the present invention are as follows: the present application designs a multi-task joint learning framework based on the state-space model. In the encoding-decoding process of the multi-task model, a task cross-fusion module based on the state-space model is innovatively introduced to extract multi-task cross-features from the spatial and spectral dimensions of remote sensing images, explicitly model the correlation characteristics between modeling tasks, and improve the information interaction and collaborative learning capabilities among multiple tasks; in addition, the present invention designs a two-stage decoding optimization strategy, combines the preliminary prediction results output by the initial decoder, and introduces an auxiliary loss mechanism to further optimize the multi-task prediction performance of the decoder, effectively balance the performance among multiple tasks, and significantly enhance the generalization ability and robustness of the model; the present invention can efficiently and accurately complete the multi-task prediction of remote sensing images, is suitable for multi-task collaborative processing of large-scale remote sensing data, and the results obtained can be widely used in fields such as geographic information extraction, environmental monitoring, and urban planning.
[0016] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, some of the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A schematic diagram of a flow chart of a multi-task joint prediction method for remote sensing images provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a state space model in a multi-task joint prediction method for remote sensing images provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a task cross-fusion module based on a state space model in a remote sensing image multi-task joint prediction method provided by an embodiment of the present invention; Figure 4 A schematic diagram of the structure of an optimized encoder in a multi-task joint prediction method for remote sensing images provided by an embodiment of the present invention; Figure 5 A schematic diagram of the structure of a remote sensing image multi-task joint prediction device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0020] See also Figure 1 - 4. An embodiment of the present application discloses a remote sensing image multi-task joint prediction method, wherein the multi-tasks include a semantic segmentation task, a height estimation task, and an edge detection task, and the method includes steps S10 to S80.
[0021] S10. Input remote sensing image dataset, the remote sensing image contains four channels of RGB and near infrared, as well as corresponding semantic segmentation and elevation map. The data format is required to be a standard image format, such as TIFF or JPEG. Each image has four channels (RGB and NIR) and a high resolution, for example, 1024×1024 pixels. Perform pixel value normalization on the input image, scaling the value of each pixel to between 0 and 1. The formula is: ,in, is the original pixel value, are the minimum and maximum pixel values in the image respectively; then resize all images to a uniform size to ensure the size consistency of the input images and facilitate model training; scale the images through bilinear interpolation to keep the main structure of the image unchanged.
[0022] S20. Use a pre-trained convolutional neural network (such as ResNet or ViT) as a shared encoder. The first few layers of the network extract low-level features, such as edges and textures, while the later layers extract high-level features, such as the shape and structure of objects. Input the resized remote sensing image into the shared encoder to extract common features. The output feature tensor has a shape of N×1024×16×16, where N is the batch size, 1024 is the number of channels, and 16 is the size of the spatial dimension. After each layer of convolution operation, batch normalization is applied to increase the training speed, reduce the model's dependence on initialization, and ensure the stability of the training process. The activation function after each layer of convolution uses ReLU (rectifier linear unit) to introduce nonlinearity and enhance the representation ability of the network.
[0023] S30. Design an independent initial decoder for each task. The decoder adopts a convolutional neural network (CNN) structure to recover high-resolution prediction results from the features extracted by the shared encoder. The decoder consists of two layers of 3×3 convolutional layers and an upsampling layer. The number of convolution kernels is set to 128 to maintain a high feature expression capability. The upsampling operation uses bilinear interpolation to gradually restore the low-resolution feature map to the size of the original image. The feature map output by each decoder is finally used to generate the final task output through a task-specific prediction head. For semantic segmentation, the Softmax activation function is used to generate the category prediction of each pixel. For height estimation, the linear activation function is used to output the height value of each pixel. For boundary detection, the Sigmoid activation function is used to output the probability of whether each pixel is a boundary point.
[0024] S40, based on the output result of the initial decoder , construct the initial loss function; Specifically, the semantic segmentation task uses the cross entropy loss function: , Among them, N is the number of pixels in the image, C is the number of categories of semantic segmentation, is the true category label for semantic segmentation; The height estimation task uses the mean squared error (MSE) loss function: , in, Represents the true label value of the height estimation task; Edge detection uses a gradient-based loss function: , in, represents the gradient of the real label edge detection task, The gradient of the initial prediction result; Based on the loss functions of the above subtasks, they are merged into a unified initial loss function, which is expressed as follows: , in, They represent the adjustment parameters of each task, which are used to ensure that the loss amount of each task remains in the same value range.
[0025] S50, such as Figure 2 As shown in the figure, the feature tensor obtained from the initial decoder of each task is input into the task cross fusion module based on the state space model for cross-task fusion, where the feature tensor shape is N×128×32×32; the multi-task features are modeled using the state space model to combine local and global information to generate shared features across tasks; the calculation process of the task cross fusion module based on the state space model is as follows Figure 3 As shown, the features from the shared encoder and the multi-task features from the initial decoder Input to the task cross-fusion module, where They represent the initial decoder features of semantic segmentation, height estimation, and edge detection respectively; after convolution and interpolation operations, the above features enter the spectral dimension and spatial dimension state space fusion module; The state space model calculation process can be expressed as: , , Where A, B, C are the parameter matrices of the state space model, is the input feature, is the hidden state, is the output feature.
[0026] In the specific implementation process, in order to further combine the spectral and spatial dimension characteristics of remote sensing images, spectral dimension pooling and spatial dimension pooling are introduced into the state-space fusion model respectively, and finally a state-space fusion module of spectral and spatial dimensions is formed.
[0027] S60, based on the features output by the multi-task cross-fusion module, an optimized decoder is designed to further refine the task features and generate the final prediction results, such as Figure 4 As shown; the characteristics of the shared encoder , the characteristics of the task cross-fusion module and the characteristics of the initial encoder , are input to the optimized decoder; specifically, the features of the shared encoder , the characteristics of the task cross-fusion module After convolution, the feature is input into the state space model, and the output feature is convolved and interpolated again, which is consistent with the feature of the initial encoder. Feature concatenation is used to output the final task prediction results.
[0028] S70, construct the final loss function, and optimize the loss function definition of each task of the decoder, which is the same as the initial loss function, denoted as , the final loss function includes the initial loss function, and the formula is expressed as: , in, It is an adjustment parameter used to control the proportion of the initial loss function in the final loss function.
[0029] S80, use Adam optimizer for training, with a learning rate of 0.001, a batch size of 32, and 50 training rounds. The learning rate is adjusted every 10 rounds and reduced to the original 0.1 to ensure the stability of the training process. After the model is trained, it is inferred on the test data set. The semantic segmentation results are output in the form of images, with each pixel corresponding to a category. The height estimation results are output as grayscale images, with grayscale values representing height. The boundary detection results are output as binary images, with white pixels representing boundary positions.
[0030] See also Figure 5 The embodiment of the present application also discloses a remote sensing image multi-task joint prediction device, including: Image preprocessing module: used to perform pixel value normalization and size adjustment on the input image; Shared encoder module: used to extract common features from the input image; Initial decoder module: used to recover high-resolution prediction results from the features extracted by the shared encoder; The first analysis module is used to perform the calculation of the initial loss function and output the initial results of multi-task prediction; Task cross-fusion module: used to fuse the feature tensors obtained from the initial decoders of each task across tasks to generate shared features across tasks; Optimized decoder module: used to further refine the task features and generate the final prediction results based on the features output by the task cross-fusion module; The second analysis module is used to perform the calculation of the final loss function and output the final result of multi-task prediction.
[0031] The embodiment of the present application further discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned remote sensing image multi-task joint prediction method are executed.
[0032] In summary, the present invention designs a multi-task joint learning framework based on the state-space model. In the encoding-decoding process of the multi-task model, a task cross-fusion module based on the state-space model is innovatively introduced to extract multi-task cross-features from the spatial dimension and spectral dimension of the remote sensing image, explicitly model the correlation characteristics between modeling tasks, and improve the information interaction and collaborative learning capabilities among multiple tasks. In addition, the present invention designs a two-stage decoding optimization strategy, combines the preliminary prediction results output by the initial decoder, and introduces an auxiliary loss mechanism to further optimize the multi-task prediction performance of the decoder, effectively balance the performance among multiple tasks, and significantly enhance the generalization ability and robustness of the model. The present invention can efficiently and accurately complete the multi-task prediction of remote sensing images, is suitable for multi-task collaborative processing of large-scale remote sensing data, and the results obtained can be widely used in geographic information extraction, environmental monitoring, urban planning and other fields.
[0033] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0034] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A remote sensing image multi-task joint prediction method, characterized in that: The multi-tasks include a semantic segmentation task, a height estimation task and an edge detection task, and the method includes: Input remote sensing image data set, perform pixel value normalization and size adjustment on the input image; Design a shared encoder to extract common features from input images; Design an initial decoder to recover high-resolution predictions from features extracted by the shared encoder; Construct the initial loss function; The feature tensors obtained from the initial decoders of each task are input into the task cross-fusion module based on the state-space model for cross-task fusion; Based on the features output by the task cross-fusion module, an optimized decoder is designed to further refine the task features and generate the final prediction results; Construct the final loss function.
2. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: The input remote sensing image dataset performs pixel value normalization and size adjustment on the input image, including: The remote sensing image contains four channels, RGB and near infrared, as well as corresponding semantic segmentation and elevation maps. The data format is required to be a standard image format. Each image has four channels and a high resolution. Scale the value of each pixel to between 0 and 1 using the formula: ,in, is the original pixel value, are the minimum and maximum pixel values in the image respectively; then all images are resized to a uniform size to ensure the size consistency of the input images; the images are scaled by bilinear interpolation to keep the main structure of the image unchanged.
3. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: The proposed design shares an encoder that extracts common features from the input image, including: The pre-trained convolutional neural network is used as a shared encoder, the first few layers of the network extract low-level features, and the later layers extract high-level features; The resized remote sensing image is input into the shared encoder to extract common features; the output feature tensor shape is N×1024×16×16, where N is the batch size, 1024 is the number of channels, and 16 is the size of the spatial dimension; Batch Normalization is applied after each convolution operation to increase training speed, reduce the model's dependence on initialization, and ensure the stability of the training process; The activation function after each convolution layer uses ReLU (rectifier linear unit) to introduce nonlinearity and enhance the representation ability of the network.
4. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: The initial decoder is designed to recover high-resolution prediction results from the features extracted by the shared encoder, including: An independent initial decoder is designed for each task. The decoder adopts a convolutional neural network (CNN) structure, which consists of two layers of 3×3 convolution layers and an upsampling layer. The number of convolution kernels is set to 128 to maintain a high feature expression capability. The upsampling operation uses bilinear interpolation to gradually restore the low-resolution feature map to the size of the original image. The feature map output by each decoder is passed through a task-specific prediction head to generate the final task output; for semantic segmentation, the Softmax activation function is used to generate the category prediction of each pixel; for height estimation, the linear activation function is used to output the height value of each pixel; for boundary detection, the Sigmoid activation function is used to output the probability of whether each pixel is a boundary point.
5. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: The constructing of the initial loss function comprises: Based on the output of the initial decoder , construct the initial loss function; The semantic segmentation task uses the cross entropy loss function: , Among them, N is the number of pixels in the image, C is the number of categories of semantic segmentation, is the true category label for semantic segmentation; The height estimation task uses the mean squared error (MSE) loss function: , in, Represents the true label value of the height estimation task; Edge detection uses a gradient-based loss function: , in, represents the gradient of the real label edge detection task, The gradient of the initial prediction result; Based on the loss functions of the above subtasks, the initial loss function is established, and the formula is expressed as: , in, They represent the adjustment parameters of each task, which are used to ensure that the loss amount of each task remains in the same value range.
6. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: The feature tensors obtained from the initial decoders of each task are input into the task cross-fusion module based on the state space model for cross-task fusion, including: The feature tensor shape is N×128×32×32; the features from the shared encoder and the multi-task features from the initial decoder Input to the task cross-fusion module, where They represent the initial decoder features of semantic segmentation, height estimation, and edge detection respectively; after convolution and interpolation operations, the above features enter the spectral dimension and spatial dimension state space fusion module; The state space model calculation process can be expressed as: , , Where A, B, C are the parameter matrices of the state space model, is the input feature, is the hidden state, is the output feature.
7. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: Based on the features output by the task cross-fusion module, an optimized decoder is designed to further refine the task features and generate the final prediction results, including: The characteristics of the shared encoder , the characteristics of the task cross-fusion module and the characteristics of the initial encoder , are input into the optimized decoder; Characteristics of a shared encoder , the characteristics of the task cross-fusion module After convolution, the feature is input into the state space model, and the output feature is convolved and interpolated again, which is consistent with the feature of the initial encoder. Feature concatenation is used to output the final task prediction results.
8. The remote sensing image multi-task joint prediction method according to claim 1, characterized in that: The final loss function is constructed as follows: include: The loss function definition of each task of the optimized decoder is the same as the initial loss function, denoted as , the final loss function includes the initial loss function, and the formula is expressed as: , in, It is an adjustment parameter used to control the proportion of the initial loss function in the final loss function.
9. A remote sensing image multi-task joint prediction device, characterized in that: include: Image preprocessing module: used to perform pixel value normalization and size adjustment on the input image; Shared encoder module: used to extract common features from the input image; Initial decoder module: used to recover high-resolution prediction results from the features extracted by the shared encoder; The first analysis module is used to perform the calculation of the initial loss function and output the initial results of multi-task prediction; Task cross-fusion module: used to fuse the feature tensors obtained from the initial decoders of each task across tasks to generate shared features across tasks; Optimized decoder module: used to further refine the task features and generate the final prediction results based on the features output by the task cross-fusion module; The second analysis module is used to perform the calculation of the final loss function and output the final result of multi-task prediction.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the remote sensing image multi-task joint prediction method according to any one of claims 1 to 8 is executed.
Citation Information
Cited By
Geometry and topology collaborative guidance medical image segmentation method
CN120997227A
Ground and satellite observation constraint combined GPP multi-task learning estimation method, system, medium and equipment
CN121072810A
A method, system, medium and device for estimating GPP multi-task learning with combined ground and satellite observation constraints
CN121072810B