A multi-task image processing method, system and device

By building a multi-task network model and combining the loss function optimization weight, the competition and consistency problems of feature extraction in multi-task image processing are solved, efficient fine-grained feature extraction and global consistency optimization are achieved, and the efficiency and quality of image processing are improved.

CN119445287BActive Publication Date: 2025-07-25CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202411516628.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-07-25
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

In traditional multi-task image processing methods, there are competition and conflicts in feature extraction between tasks, and it is difficult to take into account global consistency and fine-grained feature requirements, which affects the processing effect.

Method used

A multi-task network model is built, including an encoder, a multi-grained heterogeneous convolution module and an adaptive fine structure enhancement module. Combined with details, the global consistency loss, L1 loss and structural similarity loss are understood, and the weight is optimized through multiple training iterations to achieve fine-grained feature extraction and global consistency optimization.

Benefits of technology

It improves the efficiency and image quality of multitasking, ensures coordinated optimization and visual consistency between tasks, and is suitable for field such as object detection, image classification and segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445287B_ABST
    Figure CN119445287B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task image processing method, system and device, relating to the technical field of image processing, and comprising the following steps: constructing a multi-task network model; inputting an image data set into the multi-task network model to respectively obtain an infrared enhanced image, an infrared super-resolution image, a fused image of infrared and visible light images, and a colorized infrared image; calculating loss functions between different images output by the multi-task network model and the infrared image and visible light image in the input image data set, and performing cyclic training on the multi-task network model. The technical solution of the present invention not only effectively extracts local details through the unified training of the multi-task model, but also maintains global consistency, greatly improves the processing efficiency and image quality, reduces redundant steps in the processing process, and enhances the robustness and adaptability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a multi-task image processing method, system and device. Background Art

[0002] In recent years, multi-task image processing technology has been widely studied and applied. Its core idea is to simultaneously process multiple related tasks by sharing the feature extraction layer in the network model, thereby improving computational efficiency and reducing resource consumption. However, traditional multi-task processing methods usually face the problems of competition and conflict between tasks, that is, different tasks may have different requirements for feature extraction, resulting in the shared features being difficult to meet the needs of all tasks, thus affecting the overall processing effect. In addition, there is also a global consistency problem in feature extraction in multi-task processing, that is, the feature expressions of each task need to maintain a certain consistency globally to ensure information complementarity and collaborative processing between tasks.

[0003] Current research mainly focuses on how to better achieve the distinction of fine-grained features and the maintenance of global consistency in feature extraction. Some methods attempt to regulate the mutual influence between different tasks by introducing specific loss functions. Based on this, the present application proposes a multi-task image processing method combining fine-grained feature extraction and global consistency loss, which can share efficient and accurate feature representations between tasks and ensure the consistency of global information, thereby improving the comprehensive performance of multi-task processing and being applicable to multiple fields such as object detection, image classification, and image segmentation. Summary of the Invention

[0004] The technical solution of the present invention to solve the above technical problems is to provide a multi-task image processing method, including the following steps:

[0005] Construct a multi-task network model, including an encoder and an image enhancement decoder, an image super-resolution decoder, an image fusion decoder, and an image colorization enhancement decoder; the encoder includes a multi-granularity heterogeneous convolution module and an adaptive fine structure reinforcement module;

[0006] Input the image data set into the multi-task network model to obtain an infrared enhanced image, an infrared super-resolution image, an infrared and visible light image fusion image, and an infrared image colorization image respectively;

[0007] Calculate the loss functions between the different images output by the multi-task network model and the infrared images and visible light images in the input image data set, and perform cyclic training on the multi-task network model. Calculate the loss function values through multiple training iterations until the number of training times reaches the set threshold or the value of the loss function reaches the set range;

[0008] Among them, the combination of the detail-aware global consistency loss, L1 loss, and structural similarity loss is used as the loss function. The definitions of the detail-aware global consistency loss, L1 loss, and structural similarity loss functions are as follows:

[0009]

[0010] Among them, represents the detail-aware global consistency loss function, represents the L1 loss function, represents the structural similarity loss function, represents the local detail retention loss, represents the global consistency loss, represents the detail consistency loss, I out represents the output images of different decoders, I gt represents the original input image, || ||1 represents calculating the L1 norm;

[0011] The definitions of the local detail retention loss, global consistency loss, and detail consistency loss functions are as follows:

[0012]

[0013] Among them, represents the Sobel gradient operator, H represents the function for calculating the histogram of the image, N is the number of gray levels of the histogram, j represents the j-th gray level, patch i represents the i-th local region of the image, M is the number of local regions, Autocorr represents the autocorrelation function;

[0014] The definition of the total loss function is:

[0015]

[0016] Among them, α, β, γ are weight parameters designed to balance the contributions of each loss term; the weight update method is:

[0017] Calculate the loss change rate,

[0018]

[0019] Among them, represents the derivative of each loss with respect to time (training steps), represents calculating the derivative, t represents time (training steps);

[0020] Adaptive weight adjustment, adjusting the weights based on the loss change rate, the weight update formula is as follows:

[0021]

[0022] Among them, η is the learning rate hyperparameter, which is used to control the step size of weight adjustment. is the normalization factor of the total loss change rate, which is the sum of the change rates of all loss terms. t + 1 represents the next moment in time.

[0023] Normalized weights, which normalize the weights after each update:

[0024]

[0025] Furthermore, the multi-granularity heterogeneous convolution module consists of convolution blocks 1-1, 1-2, 1-3, 2-1, 2-2, 2-3, 3-1, 3-2, and 3-3. The convolution kernels of convolution blocks 1-1, 2-1, and 3-1 are of size n1×n1 and use the HardSwish activation function. The convolution kernels of convolution blocks 1-2, 2-2, and 3-2 are of size n2×n2 and use the Sigmoid activation function. The convolution kernels of convolution blocks 1-3, 2-3, and 3-3 are of size n3×n3 and use the Sigmoid activation function.

[0026] The adaptive fine-structure reinforcement module consists of detail reinforcement blocks 1, 2, 3, and 4. Detail reinforcement blocks 1 and 2 receive the feature information output by the multi-granularity heterogeneous convolution module, and after cross-splicing the outputs of detail reinforcement blocks 1 and 2, they are output to detail reinforcement blocks 3 and 4. The size of the feature map is the same as the size of the input image.

[0027] Furthermore, the image enhancement decoder consists of 2 convolutional layers, 2 activation functions, and 1 Gabor filter.

[0028] The image super-resolution decoder consists of 2 convolutional layers, 3 activation functions, pixel shuffling, and 1 Gabor filter.

[0029] The image fusion decoder consists of 3 convolutional layers and 3 activation functions.

[0030] The image colorization decoder consists of 6 convolutional layers, 4 dilated convolutional layers, 8 activation functions, and 5 splicing blocks.

[0031] Furthermore, the infrared images and visible light images in the image dataset come from five datasets: CVC, MSRS, M3FD, KAIST, and Urban100. 15,000 pairs of high-quality infrared images and visible light images are selected for multi-task network training.

[0032] To solve the above technical problems, the present application also proposes a multi-task image processing system for implementing the multi-task image processing method as described above, including:

[0033] An image acquisition module for acquiring training data; the training data being infrared images and visible light images;

[0034] A feature extraction module for using an encoder to extract features from each image in the training data to obtain a set of feature maps for each image; the set of feature maps including multiple feature maps of different sizes;

[0035] An infrared image enhancement module for enhancing the feature maps of a specified size in the set of feature maps of each image to obtain an infrared enhanced image;

[0036] An infrared image super-resolution module for performing super-resolution on the feature maps of a specified size in the set of feature maps of each image to obtain an infrared super-resolution image;

[0037] An infrared and visible light image fusion module for performing infrared and visible light image fusion on the feature maps of a specified size in the set of feature maps of each image to obtain an infrared and visible light image fusion image;

[0038] An infrared image colorization module for performing infrared image colorization on the feature maps of a specified size in the set of feature maps of each image to obtain an infrared image colorization image;

[0039] A model training module for performing model training with multi-task cyclic supervision on the infrared enhanced image, infrared super-resolution image, infrared and visible light image fusion image, and infrared image colorization image obtained by each module to obtain a multi-task image processing model.

[0040] To solve the above technical problems, the present application also proposes a multi-task image processing device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the multi-task image processing method as described above.

[0041] Compared with the prior art solutions, the present invention has the following beneficial effects:

[0042] 1. Multi - task collaborative processing to improve system efficiency: The multi - task image processing method proposed by the present invention can simultaneously execute multiple processing tasks such as image enhancement, image fusion, image super - resolution, and image colorization. Under the framework of multi - task processing, fine - grained image features extracted by the encoder are shared among tasks, reducing redundant calculations and significantly improving the efficiency of the overall processing system. At the same time, through the global consistency loss function, the collaborative optimization among tasks is ensured, guaranteeing the visual quality consistency of the processed images and the linkage between tasks.

[0043] 2. Fine - grained feature extraction to enhance detail performance: By extracting fine - grained features in the image through the encoder, the present invention can capture rich detail information in infrared and visible light images. Whether in image enhancement or super - resolution tasks, the precise extraction of local details can significantly improve the detail fidelity of the output image, being applicable to high - demand image processing scenarios such as low - light and noisy environments.

[0044] 3. Global consistency optimization to ensure visual quality: Through the self - designed global consistency loss function, the present invention effectively solves the contradiction between local feature optimization and global consistency that may occur during multi - task processing. By combining the detail - aware global consistency loss with the L1 loss and the structural similarity loss, both the precise retention of image details and the maintenance of the visual quality and consistency of the overall image are ensured, thus obtaining high - quality processing results in various tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0046] Figure 1 is the flowchart of the steps of the multi - task image processing method described in the present invention;

[0047] Figure 2 is the structural diagram of the multi - task network model described in the present invention;

[0048] Figure 3 is the structural diagram of the encoder network of the present invention;

[0049] Figure 4 is the structural diagram of each convolution block from convolution block 1 - 1 to convolution block 3 - 3 in the multi - granularity heterogeneous convolution module of the present invention;

[0050] Figure 5This is the structural diagram of the detail enhancement block in the encoder network of the present invention;

[0051] Figure 6 This is the structural diagram of the multi-layer perception module in the detail enhancement block of the present invention;

[0052] Figure 7 This is the structural diagram of the multi-granularity attention mechanism of the present invention;

[0053] Figure 8 This is the structural diagram of the image enhancement decoder of the present invention;

[0054] Figure 9 This is the structural diagram of the image super-resolution decoder of the present invention;

[0055] Figure 10 This is the structural diagram of the image fusion decoder of the present invention;

[0056] Figure 11 This is the structural diagram of the image colorization decoder of the present invention;

[0057] Figure 12 This is the structural schematic diagram of a multi-task image processing system of the present invention;

[0058] Figure 13 This is the structural schematic diagram of a multi-task image processing device of the present invention. Detailed implementation manners

[0059] The present invention provides a multi-task image processing method, system and device, aiming to design a multi-task image processing method capable of simultaneously performing fine-grained feature extraction and global consistency optimization.

[0060] The multi-task image processing method proposed by the present invention will be described in the following specific embodiments:

[0061] Embodiment 1:

[0062] As Figure 1 、 Figure 2 shown, a multi-task image processing method includes the following steps:

[0063] S10: Construct a multi-task network model, including an encoder and an image enhancement decoder, an image super-resolution decoder, an image fusion decoder, and an image colorization enhancement decoder; the encoder includes a multi-granularity heterogeneous convolution module and an adaptive fine-structure enhancement module;

[0064] S20: Input an image data set into the multi-task network model to obtain an infrared enhanced image, an infrared super-resolution image, an infrared and visible light image fusion image, and an infrared image colorization image respectively;

[0065] S30: Calculate the loss function between different images output by the multi-task network model and the infrared images and visible light images in the input image dataset, and perform iterative training on the multi-task network model. Calculate the loss function value through multiple training iterations until the number of training times reaches the set threshold or the value of the loss function reaches the set range;

[0066] Among them, the combination of the detail-aware global consistency loss, L1 loss, and structural similarity loss is the loss function. The definitions of the detail-aware global consistency loss, L1 loss, and structural similarity loss functions are as follows:

[0067]

[0068] Among them, represents the detail-aware global consistency loss function, represents the L1 loss function, represents the structural similarity loss function, represents the local detail retention loss, represents the global consistency loss, represents the detail consistency loss, I out represents the output images of different decoders, I gt represents the original input image, || ||1 represents calculating the L1 norm;

[0069] The definitions of the local detail retention loss, global consistency loss, and detail consistency loss functions are as follows:

[0070]

[0071] Among them, represents the Sobel gradient operator, H represents the function for calculating the histogram of the image, N is the number of gray levels of the histogram, j represents the j-th gray level, patch i represents the i-th local region of the image, M is the number of local regions, and Autocorr represents the autocorrelation function;

[0072] The definition of the total loss function is:

[0073]

[0074] Among them, α, β, and γ are weight parameters designed to balance the contributions of each loss term. An innovative weight update method is designed, using an adaptive optimization method to make α, β, and γ dynamically adjusted according to the importance of each loss term during the training process.

[0075] Specifically, an adaptive adjustment strategy based on the gradient change rate of each loss term is used, and the design scheme is as follows:

[0076] Calculate the loss change rate,

[0077]

[0078] Among them, represents the derivative of each loss with respect to time (training steps), represents calculating the derivative, t represents time (training steps);

[0079] Adaptive weight adjustment, adjusting the weights based on the loss change rate, and the weight update follows the following strategy: If the change rate of a certain loss term is large, it means that this loss term has a greater impact on the optimization of the model and should be given a higher weight; on the contrary, if the change rate is small, the weight should be reduced. The weight update formula is as follows:

[0080]

[0081]

[0082] Among them, η is the learning rate hyperparameter used to control the step size of weight adjustment, is the normalization factor of the total loss change rate, which is the sum of the change rates of all loss terms, t + 1 represents the next moment of time (training steps plus one),

[0083] Normalized weights. To ensure the stability of the weights and the sum to be 1, it is necessary to normalize the weights after each update:

[0084]

[0085] Furthermore, the encoder module is composed of a multi-granularity heterogeneous convolution module and an adaptive fine-structure reinforcement module. The specific composition structure is as Figure 3 shown; the multi-granularity heterogeneous convolution module is composed of convolution block 1-1, convolution block 1-2, convolution block 1-3, convolution block 2-1, convolution block 2-2, convolution block 2-3, convolution block 3-1, convolution block 3-2, and convolution block 3-3. The convolution kernel sizes of convolution block 1-1, convolution block 2-1, and convolution block 3-1 are n1×n1, and the HardSwish activation function is used. The convolution kernel sizes of convolution block 1-2, convolution block 2-2, and convolution block 3-2 are n2×n2, and the S-shaped activation function is used. The convolution kernel sizes of convolution block 1-3, convolution block 2-3, and convolution block 3-3 are n3×n3, and the S-shaped activation function is used;

[0086] The adaptive fine-structure reinforcement module is composed of detail reinforcement block 1, detail reinforcement block 2, detail reinforcement block 3, and detail reinforcement block 4. Detail reinforcement block 1 and detail reinforcement block 2 receive the feature information output by the multi-granularity heterogeneous convolution module, and after cross-stitching the outputs of detail reinforcement block 1 and detail reinforcement block 2, they are output to detail reinforcement block 3 and detail reinforcement block 4. The size of the feature map is the same as the size of the input image.

[0087] Specifically, each convolutional block consists of two convolutional layers, two activation functions, skip connections, and a splicing block. The specific composition structure is as follows Figure 4 shown; among them, in the multi-granularity heterogeneous convolutional module, the convolutional kernels of convolutional blocks 1-1, 2-1, and 3-1 have a size of 17×17 and use the HardSwish activation function. The convolutional kernels of convolutional blocks 1-2, 2-2, and 3-2 have a size of 13×13 and use the Sigmoid activation function. The convolutional kernels of convolutional blocks 1-3, 2-3, and 3-3 have a size of 9×9 and use the Sigmoid activation function. The definitions of the HardSwish activation function and the Sigmoid activation function are as follows

[0088]

[0089] Specifically, the structure of the detail enhancement block is as follows Figure 5 shown, including a multi-layer perceptron module. The multi-layer perceptron module consists of 5 convolutional layers, 6 activation functions, and 1 splicing block. The convolutional kernels of convolutional layers 1, 2, and 3 have a size of 5×5 and use the ReLU activation function. The specific composition structure is as follows Figure 6 shown; the adaptive fine-structure enhancement module consists of detail enhancement block 1, detail enhancement block 2, detail enhancement block 3, and detail enhancement block 4. Each detail enhancement block consists of layer normalization, a multi-granularity attention mechanism, and a multi-layer perceptron module. The multi-granularity attention mechanism consists of a 1×1 convolutional layer, a query matrix (Q), a key matrix (K), a value matrix (V), and a splicing block. The specific composition structure is as follows Figure 7 shown; the definitions of layer normalization (LayerNorm) and the ReLU activation function are as follows

[0090]

[0091] where x is the input feature vector, μ is the mean of the feature vector, σ is the standard deviation of the feature vector, γ and β are the learned scaling coefficient and translation coefficient, and ∈ is a very small number used to stabilize the calculation

[0092] Furthermore, as shown in Figure 8As shown in the figure, the image enhancement decoder consists of 2 convolutional layers, 2 activation functions, and 1 Gabor filter. The visible light image is input into the Gabor filter to obtain high-frequency features and low-frequency features. The high-frequency features are added to the features output by the encoder as the input of convolutional layer 1, and then passed through activation function 1. The output of activation function 1 is added to the low-frequency features as the input of convolutional layer 2, and then passed through activation function 2. Finally, an infrared enhanced image is obtained. The convolutional kernels of the 2 convolutional layers have a size of n3×n3. The 2 activation functions use sigmoid activation functions. The kernel function of the Gabor filter has an inclination angle of θ1. Four filters of different sizes are set, and each image obtains 4 different high-frequency information. The original image is subtracted from the 4 high-frequency information respectively to obtain 4 different low-frequency information;

[0093] Specifically, the convolutional kernels of the 2 convolutional layers have a size of 3×3. The 2 activation functions use sigmoid activation functions. The kernel function of the Gabor filter has an inclination angle of 180°. Four filters of different sizes are set, and the size of each filter is 7×7, 11×11, 15×15, and 17×17.

[0094] As Figure 9 shown in the figure, the image super-resolution decoder consists of 2 convolutional layers, 3 activation functions, pixel shuffle, and 1 Gabor filter. First, the visible light image is input into the Gabor filter to obtain high-frequency features. The high-frequency features are added to the features output by the encoder as the input of convolutional layer 1, and then passed through activation function 1. Then, the output of activation function 1 is added to the high-frequency features as the input of convolutional layer 2, and then passed through activation function 2. Again, the output of activation function 2 is used as the input of pixel shuffle, and then passed through activation function 3. Finally, an infrared super-resolution image is obtained. The convolutional kernels of the 2 convolutional layers have a size of n4×n4. The 3 activation functions use sigmoid activation functions. The kernel function of the Gabor filter has an inclination angle of θ2. Four filters of different sizes are set, and each image obtains 4 different high-frequency information;

[0095] Specifically, the convolutional kernels of the 2 convolutional layers have a size of 7×7. The 3 activation functions use sigmoid activation functions. The kernel function of the Gabor filter has an inclination angle of 90°. Four filters of different sizes are set, and the size of each filter is 9×9, 13×13, 17×17, and 21×21. The pixel shuffle convolutional kernel has a size of 7×7, a stride of 2, a padding method of SAME, and an output channel number of 1.

[0096] As Figure 10As shown in the figure, the image fusion decoder consists of 3 convolutional layers and 3 activation functions. First, the features output by the encoder are used as the input of convolutional layer 1, and then pass through activation function 1. Second, the output of activation function 1 is used as the input of convolutional layer 2, and then pass through activation function 2. Then, the output of activation function 2 is used as the input of convolutional layer 3, and then pass through activation function 3. Finally, the infrared and visible light fusion image is obtained. The convolutional kernels of the 3 convolutional layers are of size n5×n5, and the 3 activation functions use sigmoid activation functions.

[0097] Specifically, the convolutional kernels of the 3 convolutional layers are of size 3×3, and the 3 activation functions use sigmoid activation functions.

[0098] As Figure 11 shown in the figure, the image colorization decoder consists of 6 convolutional layers, 4 dilated convolutional layers, 8 activation functions and 5 splicing blocks. First, the output of the encoder is respectively input into convolutional layer 1 and dilated convolutional layer 1, and then pass through activation functions 1 and 2 respectively. The visible light image is respectively input into convolutional layer 2 and dilated convolutional layer 2, and then pass through activation functions 3 and 4 respectively. Second, the outputs of activation functions 1 and 2 are used as the input of splicing block 1, and the output of splicing block 1 is respectively used as the input of convolutional layer 3 and dilated convolutional layer 3, and then pass through activation functions 5 and 6 respectively. The outputs of activation functions 3 and 4 are used as the input of splicing block 2, and the output of splicing block 2 is respectively used as the input of convolutional layer 4 and dilated convolutional layer 4, and then pass through activation functions 7 and 8 respectively. At the same time, the outputs of splicing block 1 and splicing block 2 are multiplied, and the result of multiplication is input into convolutional layer 5. Then, the outputs of activation functions 5 and 6 are used as the input of splicing block 3, and the outputs of activation functions 7 and 8 are used as the input of splicing block 4. The outputs of splicing block 3, convolutional layer 5 and splicing block 4 are used as the input of splicing block 5, and the output of splicing block 5 is used as the input of convolutional layer 6. Finally, the infrared colorized image is obtained from the output of convolutional layer 6. The convolutional kernels of convolutional layer 1, convolutional layer 2, convolutional layer 3 and convolutional layer 4 are of size n6×n6. The convolutional kernels of dilated convolutional layer 1, dilated convolutional layer 2, dilated convolutional layer 3 and dilated convolutional layer 4 are of size n7×n7, and the dilation rate is m. Activation functions 1 to 8 use rectified linear unit activation functions. The convolutional kernels of convolutional layer 5 and convolutional layer 6 are of size n8×n8.

[0099] Specifically, the convolutional kernels of convolutional layer 1, convolutional layer 2, convolutional layer 3 and convolutional layer 4 are of size 3×3. The convolutional kernels of dilated convolutional layer 1, dilated convolutional layer 2, dilated convolutional layer 3 and dilated convolutional layer 4 are of size 3×3, and the dilation rate is 4. Activation functions 1 to 8 use rectified linear unit activation functions. The convolutional kernels of convolutional layer 5 and convolutional layer 6 are of size 5×5.

[0100] Furthermore, the infrared images and visible light images in the image dataset are from five datasets, namely CVC, MSRS, M3FD, KAIST, and Urban100. 15,000 pairs of high-quality infrared images and visible light images are selected for multi-task network training.

[0101] A multi-task image processing method constructed in this application can directly generate enhanced, super-resolution, fused, and colorized output images from the input infrared images and visible light images, avoiding the complexity of manually designing and adjusting each task step in traditional image processing methods. Through the unified training of the multi-task model, this method not only effectively extracts local details but also maintains global consistency, greatly improving the processing efficiency and image quality, reducing redundant steps in the processing process, and enhancing the robustness and adaptability of the system.

[0102] By calculating the relevant metrics of the images obtained by the existing method, the feasibility and superiority of this method are further verified. The existing method is: Chinese Patent Publication No.: "CN112132258B", titled "A Multi-Task Image Processing Method Based on Deformable Convolution". The comparison of the relevant metrics between the existing technology and the method proposed in the present invention is shown in Table 1:

[0103] Table 1 Comparison of relevant metrics between the existing technology and the method proposed in the present invention ("↑" indicates the higher the better, "↓" indicates the lower the better):

[0104]

[0105] As can be seen from the table, the method proposed in the present invention has higher image spatial frequency, edge intensity, information entropy, visual information fidelity, structural similarity, peak signal-to-noise ratio, learnable perceptual image similarity, and Fréchet inception distance. These metrics further illustrate that the method proposed in the present invention has richer image details, maintains the visual quality and consistency of the overall image, and obtains high-quality image processing results in various tasks.

[0106] Example 2:

[0107] A multi-task image processing system, as Figure 12 shown, is used to implement the multi-task image processing method described in Example 1 and includes:

[0108] An image acquisition module for acquiring training data; the training data is infrared images and visible light images;

[0109] A feature extraction module for using an encoder to extract features from each image in the training data to obtain a set of feature maps for each image; the set of feature maps includes multiple feature maps of different sizes;

[0110] An infrared image enhancement module, which is used to enhance the feature maps of a specified size in the feature map set of each of the images to obtain infrared enhanced images;

[0111] An infrared image super-resolution module, which is used to perform super-resolution on the feature maps of a specified size in the feature map set of each of the images to obtain infrared super-resolution images;

[0112] An infrared and visible light image fusion module, which is used to perform infrared and visible light image fusion on the feature maps of a specified size in the feature map set of each of the images to obtain infrared and visible light image fusion images;

[0113] An infrared image colorization module, which is used to perform infrared image colorization on the feature maps of a specified size in the feature map set of each of the images to obtain infrared image colorization images;

[0114] A model training module, which is used to perform model training with multi-task loop supervision on the infrared enhanced images, infrared super-resolution images, infrared and visible light image fusion images, and infrared image colorization images obtained by each module to obtain a multi-task image processing model.

[0115] Specifically, it further includes a storage medium for storing a computer program, and running the computer program can execute the multi-task image processing method provided in Embodiment 1.

[0116] Embodiment 3:

[0117] A multi-task image processing device, as Figure 13 shown, includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the multi-task image processing method as described in Embodiment 1.

[0118] Specifically, the memory can be a ROM, a static storage device, a dynamic storage device or a RAM; the memory can store a program, and when the program stored in the memory is executed by the processor, the processor and the communication interface are used to execute each step of the training method of the infrared and visible light image fusion network in the embodiments of the present invention;

[0119] The processor can adopt a CPU, a microprocessor, an ASIC, a GPU or one or more integrated circuits, and is used to execute relevant programs to implement the functions required to be executed by the units in the infrared and visible light image fusion training system of the present invention, or to execute the infrared and visible light image fusion training method of the present invention;

[0120] The processor can also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the infrared and visible light image fusion training method of the present invention can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The above-mentioned processor can also be a general-purpose processor, DSP, ASIC, FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the infrared and visible light image fusion method, steps and logic block diagrams of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the infrared and visible light image fusion method of the present invention can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module can be located in the random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register and other mature storage media in the art. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the functions required to be executed by the units included in the infrared and visible light image fusion training system of the present invention, or execute the infrared and visible light image fusion training method of the present invention.

[0121] The communication interface uses a transceiver system such as, but not limited to, a transceiver to implement communication between the system and other devices or communication networks. For example, the to-be-processed image or the initial feature map of the to-be-processed image can be obtained through the communication interface.

[0122] The bus can include a path for transmitting information between various components of the system (such as the memory, processor, communication interface).

[0123] Embodiment 4:

[0124] A computer-readable storage medium for multi-task image processing. The computer-readable storage medium can be the computer-readable storage medium included in the system described in the above embodiments, or can exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the present invention.

[0125] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multi-task image processing method, characterized in that, It includes the following steps: Construct a multi-task network model, including an encoder and an image enhancement decoder, an image super-resolution decoder, an image fusion decoder, and an image colorization enhancement decoder; the encoder includes a multi-granularity heterogeneous convolution module and an adaptive fine-structure reinforcement module; Input the image dataset into the multi-task network model to obtain an infrared enhanced image, an infrared super-resolution image, a fused image of infrared and visible light images, and a colorized infrared image respectively; Calculate the loss functions between the different images output by the multi-task network model and the infrared images and visible light images in the input image dataset, perform cyclic training on the multi-task network model, calculate the loss function values through multiple training iterations until the number of training times reaches the set threshold or the value of the loss function reaches the set range; Among them, the combination of the detail-aware global consistency loss, the L1 loss, and the structural similarity loss is the loss function, and the definitions of the detail-aware global consistency loss, the L1 loss, and the structural similarity loss function are: Among them, represents the detail-aware global consistency loss function, represents the L1 loss function, represents the structural similarity loss function, represents the local detail preservation loss, represents the global consistency loss, represents the detail consistency loss, I out represents the output images of different decoders, I gt represents the original input image, || ||1 represents calculating the L1 norm, and SSIM is to calculate the structural similarity between two images; The definitions of the local detail preservation loss, the global consistency loss, and the detail consistency loss function are: Among them, represents the Sobel gradient operator, H represents the function for calculating the histogram of the image, N is the number of gray levels of the histogram, j represents the j-th gray level, and patch i represents the i-th local region of the image, M is the number of local regions, and Autocorr represents the autocorrelation function; The definition of the total loss function is: Among them, α, β, and γ are weight parameters designed to balance the contributions of each loss term; the weight update method is: Calculate the loss change rate, Among them, represents the derivative of each loss with respect to time, represents calculating the derivative, and t represents time; Adaptive weight adjustment, adjust the weights based on the loss change rate, and the weight update formula is as follows: where η is a learning rate hyperparameter used to control the step size of weight adjustment, is the normalization factor of the total loss change rate, which is the sum of the change rates of all loss terms, and t + 1 represents the next moment in time, Normalize the weights, and normalize the weights after each update: The multi-granularity heterogeneous convolution module is composed of convolution block 1-1, convolution block 1-2, convolution block 1-3, convolution block 2-1, convolution block 2-2, convolution block 2-3, convolution block 3-1, convolution block 3-2, and convolution block 3-3. The convolution kernels of convolution block 1-1, convolution block 2-1, and convolution block 3-1 are n1×n1, and the HardSwish activation function is used. The convolution kernels of convolution block 1-2, convolution block 2-2, and convolution block 3-2 are n2×n2, and the Sigmoid activation function is used. The convolution kernels of convolution block 1-3, convolution block 2-3, and convolution block 3-3 are n3×n3, and the Sigmoid activation function is used; The adaptive fine-structure reinforcement module is composed of detail reinforcement block 1, detail reinforcement block 2, detail reinforcement block 3, and detail reinforcement block 4. Detail reinforcement block 1 and detail reinforcement block 2 receive the feature information output by the multi-granularity heterogeneous convolution module, and after cross-splicing the outputs of detail reinforcement block 1 and detail reinforcement block 2, output them to detail reinforcement block 3 and detail reinforcement block 4. The size of the feature map is the same as the size of the input image.

2. The multi-task image processing method according to claim 1, wherein The image enhancement decoder is composed of 2 convolutional layers, 2 activation functions, and 1 Gabor filter, The image super-resolution decoder is composed of 2 convolutional layers, 3 activation functions, pixel recombination, and 1 Gabor filter, The image fusion decoder is composed of 3 convolutional layers and 3 activation functions; The image colorization enhancement decoder is composed of 6 convolutional layers, 4 dilated convolutional layers, 8 activation functions, and 5 splicing blocks.

3. The multi-task image processing method according to claim 1, wherein The infrared images and visible light images in the image dataset are from five datasets, namely CVC, MSRS, M3FD, KAIST, and Urban100. 15,000 pairs of high-quality infrared images and visible light images are selected for multi-task network training.

4. A multi-task image processing system for implementing the multi-task image processing method according to any one of claims 1 to 3, characterized in that, Including: An image acquisition module for acquiring training data; The training data is infrared images and visible light images; A feature extraction module for extracting features of each image in the training data using an encoder to obtain a set of feature maps for each image; the set of feature maps includes multiple feature maps of different sizes; An infrared image enhancement module for enhancing the feature maps of a specified size in the set of feature maps of each image to obtain infrared enhanced images; An infrared image super-resolution module for super-resolving the feature maps of a specified size in the set of feature maps of each image to obtain infrared super-resolution images; An infrared and visible light image fusion module for fusing infrared and visible light images of the feature maps of a specified size in the set of feature maps of each image to obtain infrared and visible light image fusion images; An infrared image colorization module for colorizing the infrared images of the feature maps of a specified size in the set of feature maps of each image to obtain infrared image colorization images; A model training module for performing model training with multi-task cyclic supervision on the infrared enhanced images, infrared super-resolution images, infrared and visible light image fusion images, and infrared image colorization images obtained by each module to obtain a multi-task image processing model.

5. A multi-task image processing device, characterized in that, Including: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-task image processing method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • A Multi-Task Image Processing Method Based on Deformable Convolution

    CN112132258B

  • Weak supervision fine-grained image recognition method based on visual self-attention mechanism

    CN111539469A

  • Multi-task medical image enhancement method based on generative adversarial network

    CN114596285A

Cited By

  • Multi-task document image enhancement method and system based on low-rank adaptation

    CN120725893A

  • A multi-task document image enhancement method and system based on low-rank adaptation

    CN120725893B