A Multi-View Industrial Image Fusion Method Based on Cross-Attention Mechanism
By introducing cross attention mechanisms and multi-objective optimization strategies in image fusion, the problem that traditional methods are difficult to make full use of image information at different perspectives is solved, and high-quality and high-precision image fusion is achieved to meet the needs of complex industrial scenarios.
Patent Information
- Application Number
- CN202510219101.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional image fusion methods are difficult to make full use of complementary information of images from different perspectives, resulting in limited quality and accuracy of generated fusion images, which is difficult to meet the needs of complex industrial scenarios.
A multi-view industrial image fusion method based on the cross attention mechanism is adopted, features are extracted through an encoder with weight sharing, feature fusion is performed by combining self-attention and cross attention mechanism, and the fusion effect is optimized through the calculation of pixel reconstruction loss and structural loss.
It improves the quality and semantic consistency of image fusion, enhances the effect of feature extraction and fusion, and meets the needs of high-quality image fusion in complex industrial scenarios.
Smart Images

Figure CN119693764B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly relates to a multi-view industrial image fusion method, device, equipment and storage medium based on a cross-attention mechanism. Background Art
[0002] In recent years, with the continuous development of industrial automation and intelligent manufacturing technologies, machine vision has been widely applied in the industrial field. However, due to the complexity and variability of the industrial environment, traditional image fusion methods often struggle to fully utilize the complementary information of images from different perspectives, resulting in limited quality and accuracy of the generated fused images and difficulty in meeting the requirements of practical applications.
[0003] To address the above problems, the attention mechanism is usually utilized to improve the fusion effect. Specifically, by learning the correlation and importance between image regions, the weights of the feature maps are adaptively adjusted, enhancing the quality and semantic consistency of the fused images. This image fusion method based on the attention mechanism has improved the fusion effect to a certain extent and can better handle complex industrial scenarios. However, most of the above image fusion methods adopt unidirectional or symmetric attention calculations, lacking the full utilization of the interaction information between images from different perspectives, resulting in insufficient detail retention and structural consistency of the fused images and inability to meet the requirements of high-quality image fusion in complex industrial scenarios. Summary of the Invention
[0004] The present application provides a multi-view industrial image fusion method, device, equipment and storage medium based on a cross-attention mechanism to meet the requirements of high-quality image fusion in complex industrial scenarios.
[0005] In a first aspect, the present application provides a multi-view industrial image fusion method based on a cross-attention mechanism. The method includes: respectively inputting two industrial images with different views after normalization processing into a first encoder and a second encoder, and correspondingly obtaining a first feature image and a second feature image, where the first encoder and the second encoder are encoders with shared weights; inputting the first feature image and the second feature image into a decoder to obtain a first fusion feature and a second fusion feature; fusing the first fusion feature and the second fusion feature to generate a target fusion feature, and mapping the target fusion feature to the image space through a convolutional layer to generate a fused image; calculating the pixel reconstruction loss value and the structural loss value between the fused image and the first feature image and the second feature image, and performing a weighted sum of the pixel reconstruction loss value and the structural loss value to generate a total loss function; adjusting the parameters in the first encoder, the second encoder, and the decoder according to the total loss function to correspondingly obtain a first target encoder, a second target encoder, and a target decoder; respectively inputting the two industrial images with different views into the first target encoder and the second target encoder, and correspondingly obtaining a first target feature image and a second target feature image, and inputting the first target feature image and the second target feature image into the target decoder to obtain a target fused image.
[0006] By adopting the above technical solution, by using encoders with shared weights to process industrial images with different views, the number of model parameters is effectively reduced, and the calculation efficiency is improved. The decoder used in the method combines the self-attention and cross-attention mechanisms, which can fully capture the internal features of the image and the interaction information between images with different views, thereby enhancing the effect of feature extraction and fusion. By designing a feature fusion formula to generate a target fusion feature and using a convolutional layer to map it to the image space, high-quality image fusion is achieved. The calculation of pixel reconstruction loss and structural loss is also introduced, and a total loss function is generated through weighted summation. This multi-objective optimization strategy not only ensures the pixel-level accuracy of the fused image but also maintains the overall structure of the image, thereby improving the visual quality and semantic consistency of the fused image. By optimizing the model parameters through the backpropagation algorithm, the obtained target encoders and decoders can better adapt to complex industrial scenarios, and the finally generated target fused image has significant improvements in both detail retention and structural consistency, meeting the requirements of industrial applications for high-quality image fusion.
[0007] Optionally, the step of inputting the first feature image and the second feature image into the decoder to obtain a first fused feature and a second fused feature includes: respectively inputting the first feature image and the second feature image into the self-attention module in the decoder, and enhancing the features of the first feature image and the second feature image through the self-attention mechanism to obtain a first enhanced feature image and a second enhanced feature image; inputting the first enhanced feature image and the second enhanced feature image into the cross-attention module in the decoder, and performing feature interaction and fusion on the first enhanced feature image and the second enhanced feature image through the cross-attention mechanism to obtain the first fused feature and the second fused feature.
[0008] By adopting the above technical solution, the self-attention module respectively enhances the features of the first feature image and the second feature image to generate a first enhanced feature image and a second enhanced feature image, which can effectively capture the long-range dependence relationship within a single image and improve the feature expression ability. Subsequently, the cross-attention module performs feature interaction and fusion on these two enhanced feature images to obtain a first fused feature and a second fused feature. This process fully utilizes the complementary information between images from different perspectives and enhances the comprehensiveness and robustness of the features. This design combining self-attention and cross-attention not only improves the model's understanding ability of the internal structure of a single image but also enhances the ability to capture the relationship between multi-perspective images, thereby realizing more comprehensive and accurate feature fusion while retaining the unique information of each perspective.
[0009] Optionally, the step of fusing the first fused feature and the second fused feature to generate a target fused feature includes: substituting the first fused feature and the second fused feature into a feature fusion formula to generate the target fused feature; where the feature fusion formula is: F fusion = αG1 + (1 - α)G2; in the formula, F fusion is the target fused feature, α is the weight coefficient, G1 is the first fused feature, and G2 is the second fused feature.
[0010] By adopting the above technical solution, the adaptive fusion of the first fusion feature and the second fusion feature is achieved by introducing a feature fusion formula to generate the target fusion feature. This weighted fusion method allows the model to dynamically adjust the weight coefficients according to the characteristics and importance of the images from different perspectives, so as to more flexibly balance the contributions of the two features during the fusion process. This design can not only make full use of the complementary information of the two perspectives, but also adaptively adjust the fusion strategy according to the requirements of the specific scenario. By adjusting the weight coefficients, a balance can be achieved between retaining the unique information of each perspective and enhancing the common features, thereby generating a more comprehensive and accurate target fusion feature. This adaptive fusion mechanism is particularly suitable for processing complex industrial scenarios, where the importance of different perspectives may change with the scene. Therefore, this method can generate higher-quality fusion features, thereby improving the visual quality and semantic consistency of the final fused image, and better meeting the diverse requirements for image fusion in industrial applications.
[0011] Optionally, calculating the pixel reconstruction loss values of the fused image, the first feature image, and the second feature image includes: obtaining the width and height of the fused image, the pixel values of the fused image at each position, the pixel values of the first feature image at each position, and the pixel values of the second feature image at each position;; substituting the width and height of the fused image, the pixel values of the fused image at each position, the pixel values of the first feature image at each position, and the pixel values of the second feature image at each position into the pixel reconstruction loss formula to calculate the pixel reconstruction loss values of the fused image, the first feature image, and the second feature image; where the pixel reconstruction loss formula is:
[0012] In the formula, L pixel is the pixel reconstruction loss value, W is the width of the fused image, H is the height of the fused image, I fusion (i, j) is the pixel value of the fused image at the position (i, j), I1(i, j) is the pixel value of the first feature image at the position (i, j), and I2(i, j) is the pixel value of the second feature image at the position (i, j).
[0013] By adopting the above technical solution, the width and height of the fused image, as well as the pixel values of the fused image, the first feature image, and the second feature image at each position are obtained, and then the loss value is calculated through the pixel reconstruction loss formula. This design takes into account all the pixels of the entire image. By calculating the mean square error between the pixel values of the fused image and the average of the pixel values of the original feature images, the fusion effect is comprehensively evaluated. By dividing the error value by the total number of pixels in the image, a normalized loss value is obtained, enabling fair comparison of images of different sizes. This pixel-level evaluation method can capture subtle image differences, helping to maintain the consistency between the fused image and the original image at the pixel level. At the same time, it also provides an effective feedback signal for model optimization, prompting the fused result to better fuse multi-view features while retaining the information of the original image. This fine-grained loss calculation mechanism helps to improve the quality and accuracy of the fused image, especially suitable for processing industrial image fusion tasks that require high precision, thus better meeting the strict requirements for image fusion in industrial applications.
[0014] Optionally, calculating the structural loss value of the fused image with respect to the first feature image and the second feature image includes: obtaining the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image; substituting the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image into the structural loss formula to calculate the structural loss value of the fused image with respect to the first feature image and the second feature image; the structural loss formula is:
[0015]
[0016] In the formula, L structure is the structural loss value, W is the width of the fused image, H is the height of the fused image, S(I fusion )(i, j) is the value of the edge detection result of the fused image at the position (i, j), S(I1)(i, j) is the value of the edge detection result of the first feature image at the position (i, j), and S(I2)(i, j) is the value of the edge detection result of the second feature image at the position (i, j).
[0017] By adopting the above technical solution, the width and height of the fused image, as well as the edge detection results of the fused image, the first feature image, and the second feature image are obtained, and then the loss value is calculated through the structure loss formula. This design focuses on comparing the edge information of the images. By calculating the difference between the edge detection result of the fused image and the average value of the edge detection results of the original feature images, the retention degree of the structural information during the fusion process is comprehensively evaluated. By dividing the error value by the total number of pixels of the image, a normalized loss value is obtained, enabling fair comparison of images of different sizes. This edge-based evaluation method can effectively capture the structural features of the images, helping to maintain the consistency between the fused image and the original image at the structural level. At the same time, it provides a feedback signal that focuses on the image structure for model optimization, prompting the fused result to better fuse the features from multiple perspectives while retaining the structure of the original image.
[0018] Optionally, the weighted summation of the pixel reconstruction loss value and the structure loss value to generate the total loss function includes: substituting the pixel reconstruction loss value and the structure loss value into the loss function calculation formula to generate the total loss function; where the loss function is: L total = λ1L pixel + λ2L structure ; In the formula, L total is the total loss function, λ1 is the first preset parameter, L pixel is the pixel reconstruction loss value, λ2 is the second preset parameter, and L structure is the structure loss value.
[0019] By adopting the above technical solution, the calculation mechanism of the total loss function is introduced to achieve the weighted fusion of the pixel reconstruction loss value and the structure loss value. The method substitutes the pixel reconstruction loss value and the structure loss value into the loss function calculation formula. By adjusting the values of the first preset parameter and the second preset parameter, the relative importance of pixel-level reconstruction and structural consistency in the overall optimization goal can be flexibly controlled. This multi-objective optimization strategy can ensure the pixel-level accuracy of the fused image while fully considering the overall structure of the image, thus achieving a more comprehensive and balanced image fusion effect.
[0020] Optionally, adjusting the parameters in the first encoder, the second encoder, and the decoder according to the total loss function to obtain a first target encoder, a second target encoder, and a target decoder respectively, includes: obtaining gradient values of the total loss function with respect to each parameter in the first encoder, the second encoder, and the decoder; according to the gradient values, using the backpropagation algorithm to iteratively optimize the parameters in the first encoder, the second encoder, and the decoder to obtain the trained parameters of the first encoder, the second encoder, and the decoder; assigning the trained parameters of the first encoder to the first target encoder, assigning the trained parameters of the second encoder to the second target encoder, and assigning the trained parameters of the decoder to the target decoder to obtain the trained first target encoder, second target encoder, and target decoder.
[0021] By adopting the above technical solution, gradient values of the total loss function with respect to each parameter in the first encoder, the second encoder, and the decoder are obtained. These gradient values reflect the influence degree of each parameter on the overall performance. Then, the backpropagation algorithm is used to iteratively optimize the model parameters according to these gradient values. This process can effectively adjust the model parameters to minimize the total loss function. Through multiple rounds of iteration, the model gradually learns better feature extraction and fusion strategies, thereby improving the overall performance. Finally, the trained parameters are respectively assigned to the first target encoder, the second target encoder, and the target decoder to obtain the trained model. This optimization method can consider both pixel reconstruction and structural consistency, enabling the model to better capture and fuse key features when processing industrial images from different perspectives. Through this end-to-end training process, the model can adaptively learn the features of complex industrial scenarios, improving its ability to understand and fuse multi-perspective images. This optimization strategy not only improves the generalization ability of the model but also enhances its robustness in processing various industrial scenarios. Finally, the trained model can generate higher-quality fused images, maintaining good structural consistency while retaining detailed information, thus better meeting the strict requirements for image fusion in industrial applications.
[0022] Second aspect, the present application provides a multi-view industrial image fusion device based on cross-attention mechanism. The device includes: a first input module, a second input module, a fusion module, a calculation module, an adjustment module and a generation module; wherein, the first input module is configured to respectively input two industrial images with different views after normalization processing into a first encoder and a second encoder, and correspondingly obtain a first feature image and a second feature image, and the first encoder and the second encoder are encoders with shared weights; the second input module is configured to input the first feature image and the second feature image into a decoder to obtain a first fusion feature and a second fusion feature; the fusion module is configured to fuse the first fusion feature and the second fusion feature to generate a target fusion feature, and map the target fusion feature to the image space through a convolutional layer to generate a fusion image; the calculation module is configured to calculate the pixel reconstruction loss value and the structure loss value of the fusion image with the first feature image and the second feature image, and perform weighted summation on the pixel reconstruction loss value and the structure loss value to generate a total loss function; the adjustment module is configured to adjust the parameters in the first encoder, the second encoder and the decoder according to the total loss function, and correspondingly obtain a first target encoder, a second target encoder and a target decoder; the generation module is configured to respectively input the two industrial images with different views into the first target encoder and the second target encoder, and correspondingly obtain a first target feature image and a second target feature image, and input the first target feature image and the second target feature image into the target decoder to obtain a target fusion image.
[0023] Third aspect, the present application provides an electronic device, adopting the following technical solution: including a processor, a memory, a user interface and a network interface, the memory is used for storing instructions, the user interface and the network interface are used for communicating with other devices, and the processor is used for executing the instructions stored in the memory, so that the electronic device executes a computer program of any one of the above multi-view industrial image fusion methods based on cross-attention mechanism.
[0024] Fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: storing a computer program that can be loaded and executed by a processor for any one of the above multi-view industrial image fusion methods based on cross-attention mechanism.
[0025] In summary, the present application includes at least one of the following beneficial technical effects:
[0026] By using an encoder with weight sharing to process industrial images from different perspectives, the number of model parameters is effectively reduced, and the computational efficiency is improved. The decoder used in the method combines self-attention and cross-attention mechanisms, which can fully capture the internal features of the image and the interaction information between images from different perspectives, thereby enhancing the effect of feature extraction and fusion. By designing a feature fusion formula to generate target fusion features and using a convolutional layer to map them to the image space, high-quality image fusion is achieved. The calculation of pixel reconstruction loss and structural loss is also introduced, and the total loss function is generated through weighted summation. This multi-objective optimization strategy not only ensures the pixel-level accuracy of the fused image but also maintains the overall structure of the image, thereby improving the visual quality and semantic consistency of the fused image. By optimizing the model parameters through the backpropagation algorithm, the obtained target encoder and decoder can better adapt to complex industrial scenarios, and the finally generated target fusion image has significant improvements in terms of detail retention and structural consistency, meeting the requirements of industrial applications for high-quality image fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 FIG. is a schematic flowchart of a multi-perspective industrial image fusion method based on a cross-attention mechanism provided by an embodiment of the present application;
[0028] Figure 2 FIG. is a schematic structural diagram of a multi-perspective industrial image fusion device based on a cross-attention mechanism provided by an embodiment of the present application;
[0029] Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0030] DESCRIPTION OF REFERENCE NUMERALS: 1000, electronic device; 1001, processor; 1002, communication bus; 1003, user interface; 1004, network interface; 1005, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0032] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.
[0033] Figure 1 This is a schematic flowchart of a multi-view industrial image fusion based on a cross-attention mechanism provided by an embodiment of the present application. The present application discloses a multi-view industrial image fusion method based on a cross-attention mechanism. As Figure 1 shown, this method includes S101-S106.
[0034] S101, input two industrial images with different views after normalization processing into a first encoder and a second encoder respectively, and correspondingly obtain a first feature image and a second feature image. The first encoder and the second encoder are encoders with shared weights.
[0035] When implementing the image fusion method of the present invention, first, two industrial images with different views need to be obtained as inputs. In order to eliminate differences in image size, brightness, etc., and facilitate subsequent feature extraction and fusion processing, these two images are subjected to normalization processing. Normalization can unify the size of the images to a specified size and scale the pixel values to the range of [0, 1]. The two normalized images are respectively denoted as I1 and I2.
[0036] Then, the normalized images I1 and I2 are respectively input into the first encoder and the second encoder to extract features from the images. The encoder can adopt VIT (VISION TRANSFORMER). Before inputting the normalized images I1 and I2 into the VIT encoder, the images first need to be divided into blocks of a fixed size. For example, if the size of the input image is 384×384, it can be divided into 16×16 blocks, and the size of each block is 24×24. Then, each block is flattened into a vector, and position encoding (POSITION EMBEDDING) is added to retain the spatial position information of the blocks. Then, this sequence of vectors is input into the VIT encoder.
[0037] The VIT encoder is composed of multiple TRANSFORMER blocks, and each block includes two sub-layers: multi-head self-attention (MULTI-HEADSELF-ATTENTION) and feed-forward neural network (FEED-FORWARD NETWORK). In the self-attention mechanism, each block interacts with all other blocks to calculate the attention weights between them, thereby capturing the dependencies between blocks. Through multi-head self-attention, the model can learn different feature representations in different subspaces. After being processed by multiple TRANSFORMER blocks, the VIT encoder can extract high-level image features.
[0038] The first encoder and the second encoder adopt the same ViT structure and share weight parameters. For the first encoder, the divided image patch sequence I1 is input into it, and through the self-attention mechanism and the feed-forward neural network of multiple Transformer blocks, the first feature image F1 is obtained. Similarly, for the second encoder, the image patch sequence I2 is input into it, and through the exactly same Transformer blocks, the second feature image F2 is obtained. These two feature images are high-level feature representations extracted by the ViT encoder.
[0039] S102, input the first feature image and the second feature image into the decoder to obtain the first fused feature and the second fused feature.
[0040] In one example, the role of the decoder is to convert the feature images extracted by the encoder back to the original image space and fuse the feature information from different perspectives. In this solution, the decoder adopts a structure that combines a self-attention module and a cross-attention module. First, input the first feature image into the self-attention module of the decoder. The self-attention module is similar to the self-attention mechanism in the encoder. By calculating the attention weights between different positions in the feature image, it models the dependencies between features. After being processed by the self-attention module, the representation ability of the first feature image can be enhanced, highlighting the key information in it. Denote the enhanced first feature image as F1 ′ .
[0041] Similarly, input the second feature image into another self-attention module of the decoder. After the same processing process, the enhanced second feature image F2 is obtained ′ . The role of the self-attention module is to enhance the features within the feature image of each perspective, enabling the subsequent fusion process to more effectively extract and utilize the complementary information between perspectives.
[0042] Next, input the enhanced feature images F1 ′ and F2 ′ into the cross-attention module of the decoder. The purpose of the cross-attention module is to establish the association between features from different perspectives. By calculating the attention weights between the feature images, it realizes the interactive fusion of features. Specifically, first calculate the attention weights between each position in F1 ′ and all positions in F2 ′ , and then perform weighted summation on F2 ′ according to these weights to obtain the fused first feature image G1. Similarly, calculate the attention weights between each position in F2 ′ and all positions in F1 ′ , and perform weighted summation on F1 ′Perform weighted summation to obtain the fused second feature image G2. This process can be regarded as the mutual reference and complement of features from two perspectives in space, enabling the fused feature image to take into account the advantages of different perspectives and contain more comprehensive and accurate information.
[0043] Based on the above embodiments, as an alternative implementation, in S102, inputting the first feature image and the second feature image into the decoder to obtain the first fused feature and the second fused feature specifically includes S201 - S202:
[0044] S201, input the first feature image and the second feature image into the self - attention module in the decoder respectively, and enhance the features of the first feature image and the second feature image through the self - attention mechanism to obtain the first enhanced feature image and the second enhanced feature image.
[0045] In an example, input the first feature image and the second feature image into the self - attention module in the decoder respectively. The self - attention module performs weighted fusion on each position in the feature image through the self - attention mechanism to generate an enhanced feature representation. For the first feature image, the self - attention module first calculates the attention weights between each position and other positions, and these weights reflect the correlation and importance between different positions. Then, by multiplying these attention weights with the feature vectors at the corresponding positions and summing them up, the self - attention module obtains the enhanced first feature image. Similarly, for the second feature image, the self - attention module obtains the enhanced second feature image through the same process.
[0046] S202, input the first enhanced feature image and the second enhanced feature image into the cross - attention module in the decoder, and perform feature interaction and fusion on the first enhanced feature image and the second enhanced feature image through the cross - attention mechanism to obtain the first fused feature and the second fused feature.
[0047] In one example, the cross-attention module receives two inputs: the first enhanced feature image and the second enhanced feature image. Similar to the self-attention module, the cross-attention module also realizes the interaction and fusion of features through the attention mechanism. However, the difference is that the attention weights of the cross-attention module are calculated between two different feature images. For the first enhanced feature image, the cross-attention module regards it as the "query", regards the second enhanced feature image as the "key" and "value", and fuses the information of the second enhanced feature image into the first enhanced feature image by calculating the attention weights between each position in the first enhanced feature image and all positions in the second enhanced feature image, obtaining the first fused feature. Similarly, for the second enhanced feature image, the cross-attention module regards it as the "query", regards the first enhanced feature image as the "key" and "value", and fuses the information of the first enhanced feature image into the second enhanced feature image by calculating the attention weights between each position in the second enhanced feature image and all positions in the first enhanced feature image, obtaining the second fused feature.
[0048] S103, fuse the first fused feature and the second fused feature to generate a target fused feature, and map the target fused feature to the image space through a convolutional layer to generate a fused image.
[0049] In one example, the first fused feature and the second fused feature are added element-wise to obtain the target fused feature. This operation can be regarded as a feature-level fusion strategy, which equally superimposes the fused features from two perspectives, so that the fused feature contains comprehensive information from both perspectives. The advantage of additive fusion is that it is simple and efficient, and can retain all the features in G1 and G2 without introducing additional information loss.
[0050] Next, in order to transform the target fused feature into the final fused image, it is necessary to map it back to the image space from the feature space. This step is usually implemented using a convolutional layer. Specifically, a mapping network composed of multiple convolutional layers is designed. Taking the target fused feature as the input, after a series of convolutional operations, a fused image with the same size as the original input image is generated.
[0051] In the design of the mapping network, a residual connection structure is adopted, that is, the input feature is directly added to the output feature through a shortcut connection. This residual structure can help the network better learn the identity mapping and alleviate the difficulty of training deep networks. At the same time, batch normalization and activation functions (such as ReLU) are inserted between convolutional layers to accelerate network convergence and improve feature representation ability.
[0052] For example, assuming that the size of the target fusion feature is 48×48×512, a mapping network containing 3 convolutional layers can be designed. The first convolutional layer reduces the feature dimension from 512 to 256 and introduces non-linearity through the ReLU activation function. The second convolutional layer reduces the feature dimension from 256 to 128 and also uses ReLU activation. The last convolutional layer reduces the feature dimension from 128 to 3, obtaining an output with the same number of channels as the original image. Through this progressive feature dimension reduction and non-linear transformation, the mapping network can gradually transform the target fusion feature H into a visually clear and natural fusion image.
[0053] Based on the above embodiments, as an alternative implementation, in S103, fusing the first fusion feature and the second fusion feature to generate the target fusion feature specifically includes:
[0054] Substituting the first fusion feature and the second fusion feature into the feature fusion formula to generate the target fusion feature; where the feature fusion formula is: F fusion = αG1+(1-α)G2; in the formula, F fusion is the target fusion feature, α is the weight coefficient, G1 is the first fusion feature, and G2 is the second fusion feature.
[0055] Specifically, performing weighted summation on the first fusion feature and the second fusion feature to obtain a target fusion feature that combines the information from both perspectives. The value range of the weight coefficient is [0, 1], which controls the relative importance of the first fusion feature and the second fusion feature in the final fusion result.
[0056] The selection of the weight coefficient can be determined according to specific application scenarios and requirements. For example, if it is known in advance that the quality or information content of the first perspective image is higher than that of the second perspective image, the weight coefficient can be set to a larger value to increase the weight of the first perspective feature in the fusion result. Conversely, if the quality or information content of the second perspective image is higher, the weight coefficient can be set to a smaller value. In practical applications, methods such as cross-validation or grid search can also be used to automatically select the optimal weight coefficient from a candidate set to obtain the best fusion performance.
[0057] S104, calculating the pixel reconstruction loss value and the structure loss value between the fusion image and the first feature image and the second feature image, and performing weighted summation on the pixel reconstruction loss value and the structure loss value to generate the total loss function.
[0058] Calculating the pixel reconstruction loss value between the fusion image and the first feature image and the second feature image specifically includes: obtaining the width and height of the fusion image, the pixel values of the fusion image at each position, the pixel values of the first feature image at each position, and the pixel values of the second feature image at each position;
[0059] Substitute the width and height of the fused image, the pixel values of the fused image at each position, the pixel values of the first feature image at each position, and the pixel values of the second feature image at each position into the pixel reconstruction loss formula to calculate the pixel reconstruction loss value between the fused image and the first and second feature images; where,
[0060] The pixel reconstruction loss formula is:
[0061]
[0062] In the formula, L pixel is the pixel reconstruction loss value, W is the width of the fused image, H is the height of the fused image, I fusion (i, j) is the pixel value of the fused image at position (i, j), I1(i, j) is the pixel value of the first feature image at position (i, j), and I2(i, j) is the pixel value of the second feature image at position (i, j).
[0063] In one example, assume the size of the fused image is W×H, where W is the width and H is the height. Use two nested loops to traverse each pixel position (i, j) of the fused image, where i ranges from [1, W] and j ranges from [1, H]. For each position (i, j), we respectively read the pixel values of the fused image, the first feature image, and the second feature image at this position, denoted as I fusion (i, j), I1(i, j), and I2(i, j). These pixel values are usually a three-dimensional vector representing the RGB color value at that position.
[0064] After obtaining the required information, substitute the width W of the fused image, the height H, the pixel values I fusion (i, j) of the fused image at each position, the pixel values I1(i, j) of the first feature image at each position, and the pixel values I2(i, j) of the second feature image at each position into the pixel reconstruction loss formula to calculate the pixel reconstruction loss value between the fused image and the input images. The pixel reconstruction loss formula is defined as:
[0065] where, ||2 represents the L2 norm, i.e., the Euclidean distance. The physical meaning of this formula is that first, calculate the difference between the pixel value I fusion (i, j) of the fused image at each position (i, j) and the average value (I1(i, j) + I2(i, j)) / 2 of the input image pixel values, then square and sum the differences at all positions, and finally divide by the total number of pixels W×H to obtain the average pixel reconstruction loss. The value range of L pixel is [0, +∞), and the smaller its value, the closer the fused image is to the input image at the pixel level, and the higher the fusion quality.
[0066] For example, assume that the size of the fused image is 1024×768, and the sizes of the first feature image and the second feature image are the same as that of the fused image. For the position (100, 200), I fusion (100, 200) has RGB values of (0.8, 0.5, 0.2), I1(100, 200) has RGB values of (0.9, 0.6, 0.1), and I2(100, 200) has RGB values of (0.7, 0.4, 0.3). According to the pixel reconstruction loss formula, the loss value at this position = (0.8 - 0.8) 2 +(0.5 - 0.5) 2 +(0.2 - 0.2) 2 = 0. By analogy, sum up the loss values at each position and divide by the total number of pixels 1024×768 to obtain the final pixel reconstruction loss value.
[0067] Calculate the structural loss values of the fused image with respect to the first feature image and the second feature image, including:
[0068] Obtain the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image; substitute the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image into the structural loss formula to calculate the structural loss values of the fused image with respect to the first feature image and the second feature image;
[0069] The structural loss formula is:
[0070]
[0071] In the formula, L structure is the structural loss value, W is the width of the fused image, H is the height of the fused image, S(I fusion )(i, j) is the value of the edge detection result of the fused image at the position (i, j), S(I1)(i, j) is the value of the edge detection result of the first feature image at the position (i, j), and S(I2)(i, j) is the value of the edge detection result of the second feature image at the position (i, j).
[0072] In one example, to calculate the structural loss, it is first necessary to obtain the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image. The width and height of the fused image can be obtained by accessing its size attributes. The edge detection results require the use of specialized edge detection algorithms, such as the Sobel operator, Canny operator, etc., to perform edge extraction on the fused image, the first feature image, and the second feature image respectively. Edge detection algorithms can highlight the structural contours and texture details in the image, providing important feature information for the calculation of the structural loss.
[0073] Assume that the size of the fused image is W×H. Applying the edge detection algorithm to it, the edge detection result S(I fusion ) is obtained, where S(·) represents the edge detection operation. The size of S(I fusion ) is the same as that of I fusion , but the pixel values change from the original RGB values to grayscale values representing edge intensity. Similarly, edge detection is also performed on the first feature image and the second feature image to obtain their edge detection results S(I1) and S(I2).
[0074] After obtaining the required information, substitute the width W, height H of the fused image, the edge detection result S(I fusion ) of the fused image, the edge detection result S(I1) of the first feature image, and the edge detection result S(I2) of the second feature image into the structural loss formula to calculate the structural loss value L structure between the fused image and the input image. The structural loss formula is defined as:
[0075]
[0076] where S(I fusion )(i, j) represents the value of the edge detection result of the fused image at the position (i, j), and S(I1)(i, j) and S(I2)(i, j) represent the values of the edge detection results of the first feature image and the second feature image at the position (i, j) respectively. The physical meaning of this formula is that first, calculate the difference between the edge intensity value S(I fusion )(i, j) of the fused image at each position (i, j) and the average edge intensity of the input image, then square and sum the differences at all positions, and finally divide by the total number of pixels W×H to obtain the average structural loss. The value range of L structure is [0, +∞). The smaller its value, the closer the fused image and the input image are at the structural level, and the higher the fusion quality.
[0077] For example, assume the size of the fused image is 1024×768. At the position (100, 200), the value of S(I_fusion)(100, 200) is 0.8, the value of S(I1)(100, 200) is 0.9, and the value of S(I2)(100, 200) is 0.7. According to the structure loss formula, the loss value at this position = (0.8 - 0.8) 2 = 0; and so on. Summing up the loss values at each position and dividing by the total number of pixels 1024×768, we get the final structure loss value.
[0078] The pixel reconstruction loss value and the structure loss value are weighted and summed to generate the total loss function, including:
[0079] Substitute the pixel reconstruction loss value and the structure loss value into the loss function calculation formula to generate the total loss function; where the loss function is: L total = λ1L pixel + λ2L structure ;
[0080] In the formula, L total is the total loss function, λ1 is the first preset parameter, L pixel is the pixel reconstruction loss value, λ2 is the second preset parameter, and L structure is the structure loss value.
[0081] In one example, the pixel reconstruction loss value L pixel and the structure loss value L structure are combined into the total loss function L total by weighted summation. Specifically, we substitute L pixel and L structure into the loss function calculation formula to generate the total loss function: L total = λ1L pixel + λ2L structure where λ1 and λ2 are two preset weight parameters used to adjust the proportion of the pixel reconstruction loss and the structure loss in the total loss function. By adjusting the values of λ1 and λ2, we can flexibly control the model's emphasis on pixel-level details and structural semantic information, thereby generating a more balanced and high-quality fusion result.
[0082] For example, during the training process, the calculated pixel reconstruction loss value L pixel = 0.05 and the structure loss value L structure = 0.02. Set λ1 to 1.0 and λ2 to 0.8, then the value of the total loss function = 0.05 + 0.016 = 0.066.
[0083] This total loss value combines the pixel reconstruction loss and the structure loss, reflecting the reconstruction quality of the fused image at multiple levels. Substitute Ltotal As an optimization objective during the model training process, the model parameters are adjusted by minimizing the total loss function to generate a fusion result that is closer to the real image.
[0084] S105. Adjust the parameters in the first encoder, the second encoder, and the decoder according to the total loss function, and correspondingly obtain the first target encoder, the second target encoder, and the target decoder.
[0085] In one example, the backpropagation algorithm is adopted to perform gradient descent optimization on the parameters in the first encoder, the second encoder, and the decoder using the total loss function. First, a batch of input image pairs (I1, I2) are fed into the current encoder and decoder to generate corresponding fused images. Then, the pixel reconstruction loss and the structure loss are calculated based on the fused images and the input images, and the value of the total loss function is calculated according to the preset weight coefficients.
[0086] Next, use an automatic differentiation tool (such as PYTORCH or TENSORFLOW) to calculate the gradients of the total loss function with respect to each parameter in the first encoder, the second encoder, and the decoder. These gradients indicate the importance of each parameter in reducing the total loss function. The larger the gradient value of a parameter, the more effectively adjusting this parameter can reduce the loss function. According to the calculated gradients, use an optimizer (such as ADAM or SGD) to update each parameter, and adjust them along the direction of the negative gradient to minimize the total loss function.
[0087] Through multiple rounds of iterative optimization, the parameters of the model are continuously updated, making the value of the total loss function gradually decrease. In this process, the first encoder and the second encoder learn how to extract the key features of the input images, and the decoder learns how to fuse these features and map them back to high-quality images. After sufficient training, the first target encoder, the second target encoder, and the target decoder with optimized performance are finally obtained.
[0088] Based on the above embodiments, as an alternative implementation, in S105, adjusting the parameters in the first encoder, the second encoder, and the decoder according to the total loss function to correspondingly obtain the first target encoder, the second target encoder, and the target decoder specifically includes S501 - S503:
[0089] S501. Obtain the gradient values of the total loss function with respect to each parameter in the first encoder, the second encoder, and the decoder.
[0090] S502. According to the gradient values, use the backpropagation algorithm to perform iterative optimization on the parameters in the first encoder, the second encoder, and the decoder to obtain the trained parameters of the first encoder, the second encoder, and the decoder.
[0091] S503 assigns the trained first encoder parameters to the first target encoder, the trained second encoder parameters to the second target encoder, and the trained decoder parameters to the target decoder, obtaining the trained first target encoder, second target encoder, and target decoder.
[0092] In one example, first obtain the gradient values of the total loss function with respect to each parameter in the first encoder, second encoder, and decoder. The gradient value represents the rate of change of the total loss function in the parameter space, indicating the direction and magnitude of parameter adjustment to minimize the value of the total loss function. The gradient value of each parameter can be calculated by an automatic differentiation tool or manual derivation to obtain a gradient vector.
[0093] After obtaining the gradient values, use the backpropagation algorithm to iteratively optimize the parameters in the first encoder, second encoder, and decoder. The backpropagation algorithm is a commonly used neural network training method. By backpropagating the gradient of the loss function from the output layer to the input layer, the parameters of each layer are adjusted to make the prediction result of the model closer to the true value. In each iteration, according to the current parameter values and gradient values, update the parameters along the negative gradient direction to gradually reduce the value of the total loss function. Common parameter update rules include Stochastic Gradient Descent (SGD), ADAM, RMSPROP, etc. By setting appropriate learning rates and momentum factors, the step size and stability of parameter adjustment can be controlled.
[0094] After multiple rounds of iterative optimization, the trained first encoder, second encoder, and decoder parameters are obtained. These parameters have been adjusted according to the total loss function and can generate higher-quality and more realistic fused images. To apply the optimized parameters to the actual image fusion task, assign the trained first encoder parameters to the first target encoder, the trained second encoder parameters to the second target encoder, and the trained decoder parameters to the target decoder. In this way, the trained first target encoder, second target encoder, and target decoder are obtained, which already have stronger feature extraction and image reconstruction capabilities and can be applied to the subsequent image fusion process.
[0095] S106, input two industrial images with different perspectives into the first target encoder and the second target encoder respectively, obtaining the first target feature image and the second target feature image correspondingly, and input the first target feature image and the second target feature image into the target decoder to obtain the target fused image.
[0096] In one example, after completing the parameter optimization of the first encoder, second encoder, and decoder, the first target encoder, second target encoder, and target decoder with more excellent performance are obtained.
[0097] Given two industrial images I1 ′ and I2 ′ from different perspectives, first, input them into the first target encoder and the second target encoder respectively. The first target encoder extracts features from I1 ′ , and through the self-attention mechanism and feed-forward neural network of VIT, extracts the high-level feature representation of I1 ′ to obtain the first target feature image F1 ′ . Similarly, the second target encoder extracts features from I2 ′ to obtain the second target feature image F2 ′ .
[0098] Next, input the first target feature image F1 ′ and the second target feature image F2 ′ into the target decoder. The target decoder first enhances the features of F1 ′ and F2 ′ respectively through the self-attention module to obtain the enhanced feature images F1 ″ and F2 ″ . Then, through the cross-attention module, the target decoder deeply fuses F1 ″ and F2 ″ so that each feature in F1 ″ incorporates the information of F2 ″ to obtain the first fusion feature G1 ′ . Similarly, each feature in F2 ″ also incorporates the information of F1 ″ to obtain the second fusion feature G2 ′ .
[0099] Finally, the target decoder adds the first fusion feature G1 ′ and the second fusion feature G2 ′ element-wise to obtain the target fusion feature H'. Through the mapping network, the target decoder maps the target fusion feature H' from the feature space back to the image space to generate the final target fusion image.
[0100] Based on the above method, the present application also discloses a multi-perspective industrial image fusion device based on the cross-attention mechanism, as shown in Figure 2 , Figure 2 is a schematic structural diagram of a multi-perspective industrial image fusion device based on the cross-attention mechanism provided by an embodiment of the present application. The device includes: a first input module, a second input module, a fusion module, a calculation module, an adjustment module, and a generation module; wherein,
[0101] The first input module is used to input two industrial images with different perspectives after normalization processing into the first encoder and the second encoder respectively, and correspondingly obtain a first feature image and a second feature image. The first encoder and the second encoder are encoders with shared weights. The second input module is used to input the first feature image and the second feature image into the decoder to obtain a first fusion feature and a second fusion feature. The fusion module is used to fuse the first fusion feature and the second fusion feature to generate a target fusion feature, and map the target fusion feature to the image space through a convolutional layer to generate a fusion image. The calculation module is used to calculate the pixel reconstruction loss value and the structure loss value between the fusion image and the first feature image and the second feature image, and perform weighted summation on the pixel reconstruction loss value and the structure loss value to generate a total loss function. The adjustment module is used to adjust the parameters in the first encoder, the second encoder and the decoder according to the total loss function, and correspondingly obtain a first target encoder, a second target encoder and a target decoder. The generation module is used to input two industrial images with different perspectives into the first target encoder and the second target encoder respectively, and correspondingly obtain a first target feature image and a second target feature image, and input the first target feature image and the second target feature image into the target decoder to obtain a target fusion image.
[0102] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.
[0103] Please refer to Figure 3 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0104] Among them, the communication bus 1002 is used to realize the connection and communication between these components.
[0105] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.
[0106] Among them, the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0107] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling the data stored in the memory 1005, it performs various functions of the server and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and may be implemented separately by a single chip.
[0108] Among them, the memory 1005 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 3 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program of a multi-view industrial image fusion method based on a cross-attention mechanism.
[0109] In Figure 3In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 1001 can be used to call an application program stored in the memory 1005 for a multi-view industrial image fusion method based on the cross-attention mechanism. When executed by one or more processors, the electronic device is enabled to execute one or more of the methods as described in the above embodiments.
[0110] An electronic device-readable storage medium stores instructions. When executed by one or more processors, the electronic device is enabled to execute one or more of the methods as described in the above embodiments.
[0111] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0112] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0113] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.
[0114] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0115] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0116] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0117] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and practicing the present disclosure, those skilled in the art will easily think of other embodiments of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A multi-view industrial image fusion method based on cross-attention mechanism, characterized in that: The method comprises: Inputting the two normalized industrial images of different viewing angles into a first encoder and a second encoder respectively, and obtaining a first feature image and a second feature image correspondingly, wherein the first encoder and the second encoder are weight-sharing encoders; Inputting the first feature image and the second feature image into a decoder to obtain a first fused feature and a second fused feature; the inputting the first feature image and the second feature image into the decoder to obtain the first fused feature and the second fused feature comprises: inputting the first feature image and the second feature image into a self-attention module in the decoder respectively, and performing feature enhancement on the first feature image and the second feature image through a self-attention mechanism to obtain a first enhanced feature image and a second enhanced feature image; inputting the first enhanced feature image and the second enhanced feature image into a cross-attention module in the decoder, and performing feature interactive fusion on the first enhanced feature image and the second enhanced feature image through a cross-attention mechanism to obtain the first fused feature and the second fused feature; Fusing the first fusion feature and the second fusion feature to generate a target fusion feature, and mapping the target fusion feature to an image space through a convolution layer to generate a fused image; Calculating a pixel reconstruction loss value and a structure loss value of the fused image, the first feature image, and the second feature image, and performing a weighted summation of the pixel reconstruction loss value and the structure loss value to generate a total loss function; Adjusting parameters in the first encoder, the second encoder, and the decoder according to the total loss function to obtain a first target encoder, a second target encoder, and a target decoder accordingly; The two industrial images of different viewing angles are respectively input into the first target encoder and the second target encoder to obtain a first target feature image and a second target feature image respectively, and the first target feature image and the second target feature image are input into the target decoder to obtain a target fusion image.
2. The multi-view industrial image fusion method based on the cross attention mechanism according to claim 1 is characterized in that: The fusing the first fusion feature and the second fusion feature to generate a target fusion feature includes: Substitute the first fusion feature and the second fusion feature into the feature fusion formula to generate the target fusion feature; wherein, The feature fusion formula is: fusion =αG1+(1-α)G2; In the formula, F fusion is the target fusion feature, α is the weight coefficient, G1 is the first fusion feature, and G2 is the second fusion feature.
3. The multi-view industrial image fusion method based on cross attention mechanism according to claim 1 is characterized in that: The calculating the pixel reconstruction loss value of the fused image, the first feature image and the second feature image comprises: Acquire the width and height of the fused image, the pixel value of the fused image at each position, the pixel value of the first feature image at each position, and the pixel value of the second feature image at each position; Substituting the width and height of the fused image, the pixel values of the fused image at each position, the pixel values of the first feature image at each position, and the pixel values of the second feature image at each position into the pixel reconstruction loss formula, the pixel reconstruction loss values of the fused image, the first feature image, and the second feature image are calculated; wherein, The pixel reconstruction loss formula is: Where, L pixel is the pixel reconstruction loss value, W is the width of the fused image, H is the height of the fused image, and I fusion (i, j) is the pixel value of the fused image at position (i, j), I1(i, j) is the pixel value of the first feature image at position (i, j), and I2(i, j) is the pixel value of the second feature image at position (i, j).
4. The multi-view industrial image fusion method based on cross attention mechanism according to claim 1, characterized in that: The calculating the structural loss value of the fused image, the first feature image and the second feature image comprises: Acquire the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image; Substituting the width and height of the fused image, the edge detection result of the fused image, the edge detection result of the first feature image, and the edge detection result of the second feature image into a structural loss formula, and calculating the structural loss value of the fused image, the first feature image, and the second feature image; The structural loss formula is: Where, L structure is the structural loss value, W is the width of the fused image, H is the height of the fused image, S(I fusion )(i, j) is the value of the edge detection result of the fused image at position (i, j), S(I1)(i, j) is the value of the edge detection result of the first feature image at position (i, j), and S(I2)(i, j) is the value of the edge detection result of the second feature image at position (i, j).
5. The multi-view industrial image fusion method based on cross attention mechanism according to claim 1, characterized in that: The weighted summing of the pixel reconstruction loss value and the structure loss value to generate a total loss function includes: Substitute the pixel reconstruction loss value and the structure loss value into the loss function calculation formula to generate a total loss function; wherein the loss function is: L total =λ1L pixel +λ2L structure ; Where, L total is the total loss function, λ1 is the first preset parameter, L pixel is the pixel reconstruction loss value, λ2 is the second preset parameter, L structure is the structural loss value.
6. The multi-view industrial image fusion method based on cross attention mechanism according to claim 1, characterized in that: The adjusting the parameters in the first encoder, the second encoder and the decoder according to the total loss function to obtain a first target encoder, a second target encoder and a target decoder accordingly comprises: Obtaining the gradient value of the total loss function with respect to each parameter in the first encoder, the second encoder, and the decoder; Iteratively optimizing the parameters of the first encoder, the second encoder, and the decoder using a back propagation algorithm according to the gradient value to obtain trained parameters of the first encoder, the second encoder, and the decoder; The trained first encoder parameters are assigned to the first target encoder, the trained second encoder parameters are assigned to the second target encoder, and the trained decoder parameters are assigned to the target decoder to obtain the trained first target encoder, second target encoder and target decoder.
7. A multi-view industrial image fusion device based on a cross attention mechanism, characterized in that: The device comprises: a first input module, a second input module, a fusion module, a calculation module, an adjustment module and a generation module; wherein, The first input module is used to input the two normalized industrial images of different viewing angles into the first encoder and the second encoder respectively, to obtain the first feature image and the second feature image correspondingly, and the first encoder and the second encoder are weight-sharing encoders; The second input module is used to input the first feature image and the second feature image into the decoder to obtain the first fusion feature and the second fusion feature; the inputting the first feature image and the second feature image into the decoder to obtain the first fusion feature and the second fusion feature includes: inputting the first feature image and the second feature image into the self-attention module in the decoder respectively, and performing feature enhancement on the first feature image and the second feature image through the self-attention mechanism to obtain the first enhanced feature image and the second enhanced feature image; inputting the first enhanced feature image and the second enhanced feature image into the cross-attention module in the decoder, and performing feature interactive fusion on the first enhanced feature image and the second enhanced feature image through the cross-attention mechanism to obtain the first fusion feature and the second fusion feature; The fusion module is used to fuse the first fusion feature and the second fusion feature to generate a target fusion feature, and map the target fusion feature to the image space through a convolution layer to generate a fused image; The calculation module is used to calculate the pixel reconstruction loss value and the structure loss value of the fused image, the first feature image and the second feature image, and perform weighted summation of the pixel reconstruction loss value and the structure loss value to generate a total loss function; The adjustment module is used to adjust the parameters in the first encoder, the second encoder and the decoder according to the total loss function, and obtain a first target encoder, a second target encoder and a target decoder accordingly; The generation module is used to input the two industrial images of different perspectives into the first target encoder and the second target encoder respectively, to obtain a first target feature image and a second target feature image correspondingly, and to input the first target feature image and the second target feature image into the target decoder to obtain a target fusion image.
8. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Picture fusion method and system based on deep learning
CN119339200A
System and method of cross-modulated dense local fusion for few-shot image generation
US20240161360A1