Image processing method, electronic device, storage medium and program product

Through the self-attention mechanism of multi-scale sampling and adaptive blocking of images, the problems of redundancy and insufficient context association in traditional methods are solved, and the comprehensive and robust representation of image features is achieved, which improves the effect of image processing.

CN120387944AActive Publication Date: 2025-07-29INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510878671.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing image processing methods are difficult to effectively capture and integrate context information of different granularity, resulting in insufficient feature representation. Traditional methods have problems such as redundancy in information, loss of key details and insufficient context association when extracting and fusion of multi-scale feature.

Method used

By sampling the original image in multiple scales, using an adaptive blocking strategy and self-attention mechanism, an enhanced feature sequence is generated and converted into a feature map of preset scales, and finally generating an enhanced image through pixel-level fusion, realizing the organic combination of multi-scale information.

Benefits of technology

It significantly improves the feature expression ability and processing effect of the image, can better understand the image information in complex scenes, and enhances the robustness and refined analysis capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387944A_ABST
    Figure CN120387944A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method which can be applied to the technical field of artificial intelligence, and particularly relates to the field of digital image processing. The method comprises the following steps: carrying out multi-scale sampling on an original image to obtain a plurality of sampling images with different scales; each sampling image is subjected to block processing, a block sequence of each sampling image is obtained, and the scales of blocks of the sampling images with different scales are the same; performing self-attention calculation on the block sequence of each sampling image to obtain an enhanced feature sequence of each sampling image; converting the enhanced feature sequence of each sampling image into a feature map of a preset scale to obtain a plurality of feature maps of the preset scale; and generating an enhanced image according to the feature maps of the plurality of preset scales. The invention further provides electronic equipment, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to the field of digital image processing, and more specifically to an image processing method, an electronic device, a storage medium, and a program product. Background Art

[0002] In the fields of computer vision and image processing, effectively integrating feature information from different sources is one of the key challenges in improving algorithm performance. Due to the diversity of target objects in terms of size, shape, and texture, a single feature extraction method often fails to comprehensively cover their key attributes. When dealing with multi-scale information, related image processing methods are difficult to fully capture and integrate context information at different granularities, resulting in an incomplete feature representation. Summary of the Invention

[0003] In view of the above problems, this application provides an image processing method, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of this application, an image processing method is provided, including: performing multi-scale sampling on an original image to obtain multiple sampling images at different scales; performing block processing on each sampling image to obtain a block sequence of each sampling image, where the scales of the blocks of sampling images at different scales are the same; performing self-attention calculation on the block sequences of each sampling image to obtain an enhanced feature sequence of each sampling image; converting the enhanced feature sequences of each sampling image into feature maps of a preset scale to obtain multiple feature maps of the preset scale; and generating an enhanced image based on the multiple feature maps of the preset scale.

[0005] According to another aspect of this application, an electronic device is provided, including: one or more processors; a memory for storing one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.

[0006] According to another aspect of this application, a computer-readable storage medium is provided, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0007] According to another aspect of this application, a computer program product is provided, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented. Brief Description of the Drawings

[0008] Through the following description of the embodiments of this application with reference to the drawings, the above content and other objects, features, and advantages of this application will become clearer. In the drawings:

[0009] Figure 1Shows an application scenario diagram of an image processing method according to an embodiment of the present application;

[0010] Figure 2 Shows a flowchart of an image processing method according to an embodiment of the present application;

[0011] Figure 3 Shows a schematic diagram of an image processing method according to another embodiment of the present application;

[0012] Figure 4 Shows a visualization schematic diagram of an image processing method according to an embodiment of the present application;

[0013] Figure 5 Shows a structural block diagram of an image processing apparatus according to an embodiment of the present application;

[0014] Figure 6 Shows a block diagram of an electronic device suitable for implementing an image processing method according to an embodiment of the present application. Detailed implementation manners

[0015] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present application.

[0016] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0017] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0018] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0019] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, etc., all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0020] In the field of digital image processing, efficiently and comprehensively understanding image content is a core challenge for many advanced applications (such as computer vision, medical image analysis, intelligent driving, etc.). The information contained in images has significant multi-scale characteristics. From fine-grained textures to macroscopic structures, features at different scales are crucial for comprehensively depicting image semantics. However, when dealing with multi-scale information, related image processing methods often struggle to effectively capture and integrate context information at different granularities.

[0021] For example, although traditional Convolutional Neural Networks (CNNs) perform well in extracting local features, their fixed receptive fields limit their ability to capture global dependencies and long-range interactions in images. Simply expanding the receptive field by stacking more layers or using large-sized convolutional kernels often increases computational costs or leads to information redundancy, and it is difficult to adaptively focus on the truly important regions in the image. In addition, for some multi-scale processing strategies, such as the feature pyramid structure, although they can generate feature maps at different resolutions, how to effectively align and fuse these heterogeneous features from different scales to avoid information loss or conflict remains an urgent problem to be solved. In practical applications, the target size, texture details, and spatial distribution in images are often variable, which requires the processing system to be able to flexibly adapt and extract key information at different scales.

[0022] In recent years, the attention mechanism has received extensive attention due to its ability to model long-range dependencies, and the hierarchical feature extraction strategy provides an effective way to capture multi-scale information. However, how to specifically design a framework that can organically combine the multi-scale hierarchical feature extraction of images with global context modeling (through the attention mechanism) and ultimately achieve effective information fusion, so as to fully explore the joint features of images at different abstraction levels, thereby addressing the challenges of image understanding in complex scenarios and significantly enhancing the feature representation ability, remains a key issue in this field.

[0023] In related technologies, for the task of multi-scale feature extraction and fusion of images, the following key bottlenecks are often faced:

[0024] On the one hand, in the traditional hierarchical feature extraction process, there are often problems of information redundancy or loss of key details between different resolution levels. This makes it difficult for the features of each layer to effectively complement each other in the subsequent fusion process, thereby affecting the model's accurate understanding and refined analysis of the overall semantics of the image. For example, simple downsampling operations may irreversibly lose key detail information in high-resolution images, and relevant methods lack effective mechanisms to dynamically capture long-range dependencies and filter highly discriminative features during intra-layer feature extraction, thus being unable to effectively suppress the interference of redundant information on the final image representation and difficult to ensure that the model can focus on significant regions at different scales.

[0025] On the other hand, when relevant image processing methods integrate features from different resolution levels, it is often difficult to effectively model the hierarchical spatial associations of the image, resulting in the gradual attenuation of cross-scale semantic information during the fusion process, making the final image representation lack context consistency and robustness. Specifically, although some methods attempt to introduce attention mechanisms or Transformer structures, their tokenization strategies between different resolution levels usually lack unity or adaptability, and fail to achieve a consistent and fixed size for the tokens generated at all levels in the spatial dimension. This makes it difficult for intra-layer feature extraction to fully capture long-range dependencies at a unified granularity and limits the model's discriminative ability for complex texture and structural relationships.

[0026] In addition, relevant fusion methods, such as simple feature concatenation or global average pooling, do not fully consider the unique contributions and interactions of the features of each layer, resulting in information redundancy or dilution of key information, failing to achieve the complementary advantages of features at different scales, and unable to achieve fine-grained alignment and non-linear correlation mining of multi-resolution features through pixel-level addition, thus being difficult to bridge the "gap" between different-scale information, ultimately limiting the depth and robustness of the overall information representation of the image and difficult to achieve optimal performance in complex image processing tasks.

[0027] Figure 1 The application scenario diagram of the image processing method according to the embodiment of the present application is shown.

[0028] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0029] Users can interact with the server 105 through the network 104 using the first terminal device 101, the second terminal device 102, and the third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0030] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and so on.

[0031] The server 105 can be a server providing various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server can analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0032] It should be noted that the image processing method provided by the embodiments of the present application can generally be executed by the server 105. Correspondingly, the image processing device provided by the embodiments of the present application can generally be set in the server 105. The image processing method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image processing device provided by the embodiments of the present application can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0033] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0034] Figure 2 shows a flowchart of the image processing method according to an embodiment of the present application.

[0035] As Figure 2As shown, the image processing method of this embodiment includes operations S210 to S250, and this image processing method can be executed by at least one of the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105.

[0036] In operation S210, multi-scale sampling is performed on the original image to obtain multiple sampled images of different scales.

[0037] According to an embodiment of the present application, the scale of an image can be the resolution of the image. Multi-scale sampling can include upsampling and downsampling, which are not limited herein. Multi-scale sampling is preferably downsampling. To obtain sampled images of different scales, the original image can be downsampled based on different strides to obtain sampled images of different scales. In this case, the stride for each downsampling can be taken as 2, 4, 6, or 2, 5, 7, etc. in sequence, and is not limited thereto. It is also possible to downsample the original image to obtain a downsampled image, and then continue to downsample the downsampled image to obtain the next layer of downsampled image, and so on, to also obtain sampled images of different scales. In this case, the stride for each layer of downsampling can be the same or different.

[0038] For example, let the original image be X0, and its scale be H×W×C. Wherein, H is the height of the original image, W is the width of the original image, and C is the number of channels of the original image. If the stride for each layer of downsampling is 2, then the sampled image X k after the k-th layer of downsampling can be determined based on the following formula (1).

[0039] X k = Downsample(X k-1 ) = (X k-1 ) 2, k = 1, 2,..., N (1)

[0040] In formula (1), Downsample(·) represents the downsampling algorithm, (·) represents that the bilinear interpolation algorithm is specifically used for downsampling, 2 is the downsampling operation with a stride of 2. Through this operation, a total of N layers of downsampled images X1, X2,..., X N are generated, where X N is the downsampled image with the smallest scale, and its spatial resolution is H / 2 N × W / 2 N .

[0041] Based on the above method, the original image can be downsampled at multiple scales, gradually reducing the spatial resolution layer by layer to generate a feature pyramid of different scales. Each layer of features in the feature pyramid can represent a sampled image at a certain scale. The sampled image can be regarded as a low-resolution version of the original image while still maintaining the visual structure of the image.

[0042] In operation S220, each sampled image is block-processed to obtain a block sequence of each sampled image, where the scales of the blocks of sampled images at different scales are the same.

[0043] According to an embodiment of the present application, the block method for performing block processing is not limited to various block methods, as long as it is ensured that the scales of the blocks obtained by block-processing the sampled images at different scales are all the same. After obtaining one or more blocks by block-processing each sampled image, by sequentially splicing all the blocks of each sampled image, the block sequence of the sampled image can be obtained.

[0044] In operation S230, self-attention calculation is performed on the block sequences of each sampled image to obtain an enhanced feature sequence of each sampled image.

[0045] According to an embodiment of the present application, the enhanced feature sequence can represent abstract and attention-enhanced image features and no longer has the visual structure of the image.

[0046] In operation S240, the enhanced feature sequences of each sampled image are converted into feature maps of a preset scale to obtain multiple feature maps of the preset scale.

[0047] According to an embodiment of the present application, the preset scale can be custom-determined according to service requirements and is not limited herein. By restructuring the enhanced feature sequence back to its original spatial structure, a feature map can be obtained. The feature map can represent the deep understanding of the image inside the model and is used for subsequent fusion.

[0048] In operation S250, an enhanced image is generated based on the multiple feature maps of the preset scale.

[0049] According to an embodiment of the present application, corresponding to this operation, the multiple feature maps of the preset scale can be first fused to obtain a fused feature. Then, the fused feature is decoded to obtain an enhanced image that fuses the multi-scale features of the original image.

[0050] Through the above embodiments of the present application, by keeping the scales of the blocks obtained by block-processing images at different scales consistent, the spatial receptive fields represented by the blocks on the original image are made consistent, so that after the features of different scale layers are converted to the preset scale, their information can be "aligned" spatially, which ultimately helps to improve the fusion effect.

[0051] Figure 3A schematic diagram of an image processing method according to another embodiment of the present application is shown.

[0052] As Figure 3 shown, for the original image 310 with a scale of H×W×C, the original image is downsampled layer by layer with a preset step size to obtain multiple downsampled images, so that the downsampled images of different layers have different scales. The step size can be 2 n , where n is a positive integer. For example, by performing a hierarchical downsampling operation with a step size of 2, a first downsampled image 311 with a scale of H / 2×W / 2×C, a second downsampled image 312 with a scale of H / 4×W / 4×C, a third downsampled image 313 with a scale of H / 8×W / 8×C, etc. can be obtained.

[0053] Next, for each downsampled image, a blocking operation based on the aforementioned adaptive blocking strategy is performed respectively, and a first image block sequence 321, a second image block sequence 322, a third image block sequence 323, etc. can be correspondingly obtained.

[0054] According to an embodiment of the present application, the so-called adaptive blocking strategy means that for each sampled image, the sampled image is evenly grid-blocked according to the layer where the sampled image is located to obtain multiple blocks, so that the scales of the blocks of the sampled images of different scales are the same. For example, according to the layer where the sampled image is located and the scale of the sampled image, the number of blocks and the block scale of the sampled image can be determined; the sampled image is evenly grid-blocked according to the number of blocks and the block scale to obtain multiple blocks, such as Tokens. The scales of the multiple blocks are all the block scale, and the number of the multiple blocks is the number of blocks. According to the multiple blocks, a block sequence of the sampled image, such as a Token sequence, can be generated.

[0055] For example, for the N-layer downsampled images X1, X2,..., X N , for the sake of easy description, X N is regarded as the first layer (the smallest image layer), X N-1 is regarded as the second layer, and so on, X1 is regarded as the Nth layer (the largest downsampled image layer). Then the jth layer downsampled image can correspond to X N-j+1 in the N-layer downsampled images. Let the spatial resolution of the jth layer sampled image be H j ×W j = (H / 2 N-j+1 )×(W / 2 N-j+1 ). The jth layer feature map can be divided into a 2 j ×2 j grid, and M j = 4 jtokens. The scale of each token is (H j / 2 j ) × (W j / 2 j ). It should be noted that through this chunking strategy, the tokens generated by all layers have the same scale in the spatial dimension, i.e., (H / 2 N+1 ) × (W / 2 N+1 ). Recombine all M j tokens of the j-th layer into a token sequence T j . The dimension of the token sequence T j is M j × (P T 2 × C), where P T = H / 2 N+1 or W / 2 N+1 , representing the spatial side length of the token.

[0056] Through the above embodiments of the present application, by introducing a preset step size and combining the number of layers and scale of the sampled images, an adaptive chunking strategy for sampled images of different scales is determined, such that the scales of the chunks of sampled images of each scale are the same chunk scale, which can offset the sampling multiples between sampled images of different scales, thereby ensuring that the image patches generated for the sampled images of each scale finally have the same scale in the spatial dimension.

[0057] Next, continue to refer to Figure 3 . For each image patch sequence, self-attention calculation operations based on the aforementioned self-attention calculation method are performed respectively, and the first enhanced feature sequence 331, the second enhanced feature sequence 332, the third enhanced feature sequence 333, etc. can be obtained correspondingly. For example, hierarchical attention feature extraction can be achieved by combining multi-head self-attention (abbreviated as MHA) calculation, or hierarchical attention feature extraction can be achieved by combining self-attention (abbreviated as SA) calculation, which is not limited herein.

[0058] Taking MHA calculation as an example, based on the aforementioned embodiments, the token sequence T j generated by the downsampled image of each layer can be independently subjected to multi-head self-attention calculation.

[0059] Let the query, key, and value matrices corresponding to the token sequence T j of the j-th layer be .

[0060] The calculation principle of the single-head attention output is shown in formula (4).

[0061] (4)

[0062] In formula (4), is the dimension of the key vector, represents the matrix transpose.

[0063] Based on the single-head attention calculation principle shown in formula (4), after parallel calculation by multiple attention heads, the outputs of each head can be concatenated and linearly projected according to formula (5) to obtain the Token sequence T of the j-th layer j The enhanced feature sequence T output after self-attention calculation j out .

[0064] T j out = Concat(head1, ..., head h ) W o (5)

[0065] where h is the number of attention heads, and W o is the output projection matrix.

[0066] By obtaining blocks of the same scale from the differently-scaled sampled images based on the foregoing method, in the process of performing self-attention calculation, the same self-attention module can be reused for all layers, which helps to simplify the model architecture and reduce the number of parameters.

[0067] Next, the enhanced feature sequences of each sampled image can be converted into initial feature maps. Among them, the scale of the initial feature map is the same as that of the sampled image used to obtain the initial feature map. Then, the initial feature maps of each sampled image are resampled to obtain the feature maps of the preset scale of each sampled image.

[0068] According to the embodiments of the present application, the above preset scale can be the scale of a certain sampled image or a scale set by the user, which is not limited herein.

[0069] Based on the foregoing embodiments, the enhanced feature sequence T output by self-attention for each layer of Token sequence j out can be recombined back into its original image space structure to form the processed initial feature map F j processed , whose scale is the same as that of the downsampled image X of the j-th layer N-j+1 i.e., (H / 2 N-j+1 )×(W / 2 N-j+1 )×C. Subsequently, according to the preset scale, for each layer of the initial feature map Fj processed Resampling is performed. For example, bilinear interpolation can be used to restore the scale of each layer of the initial feature map to a preset scale to obtain the feature map of the preset scale of each sampled image.

[0070] According to an embodiment of the present application, the above preset scale can be the scale of the original image.

[0071] In this case, based on the foregoing embodiment, since each layer of sampled images is obtained by downsampling the original image, corresponding to this scenario, for each layer of the processed initial feature map F j processed Upsampling is performed. For example, bilinear interpolation can be used to restore its scale to the scale H×W×C of the original image. Let the upsampling operation be (·), then the feature map after upsampling the j-th layer is F j upsampled = (F j processed ).

[0072] For example, continuing to refer to Figure 3 , recombination operations for recombining each enhanced feature sequence back to the original image space result can be performed respectively. Correspondingly, a first initial feature map 341 with a scale of H / 2×W / 2×C, a second initial feature map 342 with a scale of H / 4×W / 4×C, a third initial feature map 343 with a scale of H / 8×W / 8×C, etc. are obtained. For each initial feature map, an upsampling operation of upsampling to the scale of the original image 310 can be performed respectively, and correspondingly, a first feature map 351, a second feature map 352, a third feature map 353, etc. with a scale of H×W×C can be obtained.

[0073] By determining the preset scale as the scale of the original image and restoring the feature map based on the scale of the original image, it is beneficial to more accurately represent the features of the feature map in the original image space, thereby facilitating the improvement of the accuracy of the subsequent enhanced image.

[0074] Next, pixel-level addition can be performed on the feature maps of multiple preset scales to obtain an enhanced feature map. Then, the enhanced feature map is decoded to obtain an enhanced image.

[0075] Based on the foregoing embodiment, taking the preset scale as the scale of the original image as an example, in combination with formula (6), all N layers of initial feature maps can be upsampled to the feature maps F1 upsampled , F2 upsampled ,..., F N upsampledPerforming pixel - level addition can generate a fusion result R that contains multi - scale attention fusion information. Among them, the dimension of R is H×W×C. This fusion result can be output as an enhanced image or used as an input for subsequent tasks (such as image recognition, segmentation, reconstruction, etc.), providing richer and more robust feature representations for downstream applications.

[0076] (6)

[0077] According to an embodiment of the present application, the enhanced image determined based on the fusion result R can fuse local features of the original image at different scales.

[0078] For example, continuing to refer to Figure 3 , an enhanced image 360 with a scale of H×W×C can be obtained by performing pixel - level fusion and decoding on the first feature map 351, the second feature map 352, the third feature map 353, etc.

[0079] Through the above - mentioned embodiments of the present application, it is possible to effectively utilize hierarchical downsampling to obtain multi - scale information of the image, and capture local features of the image at different resolution levels through adaptive block division and self - attention mechanism. Finally, through pixel - level fusion, the fine features and macroscopic features at each scale are organically combined, thereby significantly improving the feature expression ability and processing effect of the image.

[0080] According to an embodiment of the present application, during the process of executing the above - mentioned image processing method, the original image can also be block - processed to obtain a block sequence of the original image, where the scale of the blocks of the original image is the same as the scale of the blocks of each sampled image. Performing self - attention calculation on the block sequence of the original image to obtain an enhanced feature sequence of the original image. Converting the enhanced feature sequence of the original image into a feature map of a preset scale. For the process of performing pixel - level addition on the feature maps of multiple preset scales to obtain an enhanced feature map, it can also be expressed as: performing pixel - level addition on the feature map of the preset scale of the original image and the feature maps of the preset scales of multiple sampled images to obtain an enhanced feature map.

[0081] It should be noted that the specific implementation manners of block - processing, self - attention calculation, and restoration of the original image can refer to the description of the embodiments of block - processing, self - attention calculation, and restoration of the sampled image in the foregoing embodiments, and will not be elaborated here.

[0082] During the process of performing pixel - level fusion, the feature map of the original image and the feature maps of each sampled image can be combined for pixel - level addition, and the obtained enhanced image can include both the global features and local features of the original image.

[0083] Through the above embodiments of the present application, it is possible to effectively utilize hierarchical downsampling to obtain multi-scale information of an image, and through adaptive block division and self-attention mechanism, capture local and global context dependencies at different resolution levels. Finally, through pixel-level fusion, the fine features and macroscopic features at each scale are organically combined, thereby significantly improving the feature expression ability and processing effect of the image.

[0084] Figure 4 Fig. shows a visualization diagram of an image processing method according to an embodiment of the present application.

[0085] As Figure 4 shown, for the original image 400, after performing hierarchical downsampling operations, for example, three downsampled images 410 can be obtained. After respectively dividing each downsampled image into blocks, three block sequences 420 can be obtained. After respectively performing self-attention calculations on each block sequence, three enhanced feature sequences 430 can be obtained. According to the mapping space having the same scale as the scale of the downsampled image corresponding to each of the three enhanced feature sequences 430, the three enhanced feature sequences 430 can be respectively recombined back to their original image space structure to obtain three initial feature maps 440. After respectively performing upsampling on each initial feature map, three feature maps 450 can be obtained. After performing pixel-level fusion on the three feature maps 450, an enhanced feature 460 can be obtained. After decoding the enhanced feature map 460, an enhanced image ( Figure 4 not shown in the figure) can be obtained.

[0086] Through the above embodiments of the present application, an image processing method based on multi-scale hierarchical attention fusion is realized. By performing hierarchical downsampling on the input image to obtain multi-scale views, using an adaptive block division strategy to generate blocks at each layer and applying the self-attention mechanism to capture long-range dependencies within the layer, and finally upsampling the features processed at each layer to the original resolution and performing pixel-level fusion, it is possible to effectively extract and fuse multi-scale features of the image at different resolutions, fully explore the joint features of the image at different resolutions, enhance the representation ability and information utilization efficiency of the image, thereby coping with the challenges of multi-scale information processing and significantly improving the feature expression ability of the image, which has important theoretical and practical significance.

[0087] According to an embodiment of the present application, the original image may include a first image and a second image. The first image may be an image generated based on detection data of a first modality of an object. The second image may be an image generated based on detection data of a second modality of the object.

[0088] For example, the first modality may include spectroscopy and the second modality may include sonar. Thus, the first image may include a spectral image, and the second image may include a sonar image. After acquiring a spectral image and a sonar image of the same object, the spectral image and the sonar image may be processed separately based on operations S210 through S240 described above to obtain spectral feature maps at various preset scales for the spectral image and sonar feature maps at various preset scales for the sonar image. The spectral feature maps at various preset scales and the sonar feature maps at various preset scales may then be fused to obtain an enhanced image that incorporates multimodal information, including spectral and sonar information.

[0089] It should be noted that, for implementation details of the image processing method in this embodiment, reference may be made to the description in the aforementioned embodiments, which will not be repeated here.

[0090] According to the embodiments of the present application, through the collaborative design of the hierarchical attention mechanism and cross-modal convolutional fusion, multi-scale complementary feature extraction of spectral images and sonar images can be achieved, significantly improving the classification accuracy of complex scenes. The hierarchical attention module dynamically calculates feature sequences for different resolution levels, which can effectively suppress noise interference and focus on cross-modal key areas, so that the model can still maintain robust feature representation capabilities under low signal-to-noise ratio conditions. By modeling nonlinear associations through local receptive fields, the complementary information of spectral reflectivity and sonar echo intensity is fully explored. The joint optimization of the two can avoid the problems of feature over-smoothing or local information loss in traditional methods, while ensuring computational efficiency, and effectively improve the accuracy of multi-source data classification.

[0091] In addition, the architectural design based on layered downsampling and attention feature superposition breaks through the reliance of traditional multimodal fusion methods on single-resolution features and achieves full-level integration of multi-granularity semantic information. By generating a multi-scale feature pyramid through layer-by-layer downsampling and combining it with the process of backsampling to restore resolution, it can not only preserve the detailed texture of high-resolution data but also capture the global contextual associations of low-resolution layers. The attention superposition strategy splices feature maps of different levels in the channel dimension to construct a high-dimensional feature matrix with joint spatial-feature expression. It can effectively improve the model's detection sensitivity for small target features in tasks such as farmland edge recognition and underwater terrain classification, and alleviate the problems of false detection and missed detection caused by the semantic gap of heterogeneous data.

[0092] According to the embodiments of this application, the multi-scale hierarchical attention fusion framework proposed based on the above image processing method has the core idea of effectively integrating local and global information at different granularities. This framework can be further expanded to process more complex and multi-dimensional visual data formats.

[0093] According to an embodiment of the present application, the above-mentioned original image may also include a voxel image determined based on three-dimensional volume data. The three-dimensional volume data may include data of a three-dimensional volume formed by stacking multiple video frames along the time axis, and the like, and is not limited thereto. The voxel image may include at least one of the following: three-dimensional scan data in industrial non-destructive testing, computed tomography (CT), magnetic resonance imaging (MRI) in medical imaging, and the like, and is not limited thereto.

[0094] In the case where the original image is a voxel image, the above-mentioned image processing method may be expressed as: performing multi-scale sampling on the original voxel image to obtain multiple sampled voxel images at different scales. Performing a block processing on each sampled voxel image to obtain a voxel block sequence of each sampled voxel image, wherein the scales of the voxel blocks of the sampled voxel images at different scales are the same. Performing self-attention calculation on the voxel block sequence of each sampled voxel image to obtain a voxel enhanced feature sequence of each sampled voxel image. Converting the voxel enhanced feature sequence of each sampled voxel image into a voxel feature map of a preset scale to obtain multiple voxel feature maps of the preset scale. Generating an enhanced voxel image according to the multiple voxel feature maps of the preset scale.

[0095] It should be noted that the process of processing the original voxel image to generate an enhanced voxel image has the same or corresponding implementation manners as the process of processing the original image to generate an enhanced image in the foregoing embodiments. For the specific implementation manner of processing the original voxel image to generate an enhanced voxel image in this embodiment, reference may be made to the foregoing embodiments, and details are not described herein again.

[0096] Through the above embodiments of the present application, by performing multi-scale voxel sampling on the voxel image, generating voxel image blocks at different voxel resolutions and calculating attention, the fusion of multi-scale three-dimensional features is finally realized, which can effectively improve the accuracy of tasks such as lesion recognition and defect detection.

[0097] In the case where the original image is three-dimensional volume data formed by stacking multiple video frames along the time axis, in addition to including the width, height, and channel information of the video frames, the three-dimensional volume data may also include time information. The image processing method thereof may be the same as the method for processing the original voxel image to generate an enhanced voxel image described above, and details are not described herein again.

[0098] Through the above embodiments of the present application, by incorporating the time dimension into the consideration of hierarchical sampling of video images, forming a spatio-temporal pyramid, extracting video image blocks at different spatio-temporal scales and applying self-attention, the dynamic changes and temporal correlations in the video can be effectively captured, and better applications can be realized in fields such as behavior recognition and event detection.

[0099] Based on the above concept, this application proposes an innovative image processing method based on multi-scale hierarchical attention fusion, aiming to overcome the many bottlenecks existing in the prior art when dealing with multi-scale information such as images. The key technical points are as follows: First, by performing hierarchical downsampling on the original image, a multi-resolution feature pyramid is systematically constructed to capture the spatial information of the image at different granularities. Second, an adaptive block strategy is innovatively introduced: for each feature map after downsampling at each layer, starting from the layer with the smallest scale (the first layer), it is evenly divided into j 4 identical-sized Tokens. This design ensures that the Tokens generated at different resolution levels have a unified and fixed scale in the spatial dimension, providing a standardized input for subsequent attention calculation. Then, the independent Token sequence of each layer is input into the multi-head self-attention mechanism for processing, so as to effectively capture long-range dependencies and global context information within their respective resolution levels, and adaptively select highly discriminative features. Finally, all the feature maps of each layer after attention processing are upsampled to the resolution of the original image, and deep fusion is performed by pixel-level addition. This fusion method can organically combine the fine details and macroscopic context information from different scales to form a comprehensive and robust image feature representation.

[0100] Through the above technical solutions, this application effectively solves the problems of information loss, insufficient context association, and low fusion efficiency in multi-scale feature fusion of traditional methods, and significantly improves the image feature representation ability and the understanding ability of complex scenes.

[0101] According to the embodiments of this application, considering challenges such as high data acquisition costs or incomplete data in practical applications, it is possible to explore combining the above processing method of this application with advanced machine learning paradigms to improve its generalization ability and robustness. For example, it can be combined with self-supervised learning (SSL) or contrastive learning (CL), and a large amount of unlabeled multi-scale data is used for pre-training. By designing cross-scale or hierarchical self-supervised tasks, such as predicting missing Tokens, reconstructing high-resolution features, etc., the algorithm can learn powerful feature representations, and then fine-tuned with a small amount of labeled data to handle small-sample learning scenarios. In addition, it is also possible to explore combining with a knowledge distillation (KD) model to compress a large and complex multi-scale fusion model into a smaller and more efficient model for deployment on resource-constrained edge devices.

[0102] Figure 5 The structural block diagram of an image processing device according to an embodiment of this application is shown.

[0103] As shown Figure 5 in FIG. 500, the image processing apparatus includes a sampling module 510, a sampled image block division module 520, a sampled image self-attention calculation module 530, a sampled image conversion module 540, and an image enhancement module 550.

[0104] The sampling module 510 is configured to perform multi-scale sampling on the original image to obtain multiple sampled images with different scales.

[0105] The sampled image block division module 520 is configured to perform block division on each sampled image to obtain a block sequence of each sampled image, where the scales of the blocks of the sampled images with different scales are the same.

[0106] The sampled image self-attention calculation module 530 is configured to perform self-attention calculation on the block sequence of each sampled image to obtain an enhanced feature sequence of each sampled image.

[0107] The sampled image conversion module 540 is configured to convert the enhanced feature sequence of each sampled image into a feature map with a preset scale to obtain multiple feature maps with the preset scale.

[0108] The image enhancement module 550 is configured to generate an enhanced image based on the multiple feature maps with the preset scale.

[0109] According to an embodiment of the present application, the sampling module includes a downsampling sub-module.

[0110] The downsampling sub-module is configured to perform layer-by-layer downsampling on the original image with a preset step size to obtain multiple layers of sampled images, so that the sampled images of different layers have different scales.

[0111] According to an embodiment of the present application, the sampled image block division module includes a block division sub-module and a block sequence generation sub-module.

[0112] The block division sub-module is configured to perform uniform grid block division on the sampled image according to the layer where the sampled image is located to obtain multiple blocks, so that the scales of the blocks of the sampled images with different scales are the same.

[0113] The block sequence generation sub-module is configured to generate a block sequence of the sampled image based on the multiple blocks.

[0114] According to an embodiment of the present application, the block division sub-module includes a block information determination unit and a block division unit.

[0115] The block information determination unit is configured to determine the number of blocks and the block scale of the sampled image according to the layer where the sampled image is located and the scale of the sampled image.

[0116] A chunking unit, configured to perform uniform grid chunking on a sampled image according to the number of chunks and the chunk scale, to obtain a plurality of chunks, where the scale of the plurality of chunks is the chunk scale, and the number of the plurality of chunks is the number of chunks.

[0117] According to an embodiment of the present application, the sampled image conversion module includes an initial conversion sub-module and a resampling sub-module.

[0118] The initial conversion sub-module is configured to convert the enhanced feature sequence of each sampled image into an initial feature map, where the scale of the initial feature map is the same as the scale of the sampled image used to obtain the initial feature map.

[0119] The resampling sub-module is configured to resample the initial feature map of each sampled image to obtain a feature map of a preset scale of each sampled image.

[0120] According to an embodiment of the present application, the image enhancement module includes a pixel-level addition sub-module and a decoding sub-module.

[0121] The pixel-level addition sub-module is configured to perform pixel-level addition on a plurality of feature maps of a preset scale to obtain an enhanced feature map.

[0122] The decoding sub-module is configured to decode the enhanced feature map to obtain an enhanced image.

[0123] According to an embodiment of the present application, the image processing device further includes an original image chunking module, an original image self-attention calculation module, and an original image restoration module.

[0124] The original image chunking module is configured to perform chunking processing on the original image to obtain a chunk sequence of the original image, where the scale of the chunks of the original image is the same as the scale of the chunks of each sampled image.

[0125] The original image self-attention calculation module is configured to perform self-attention calculation on the chunk sequence of the original image to obtain an enhanced feature sequence of the original image.

[0126] The original image conversion module is configured to convert the enhanced feature sequence of the original image into a feature map of a preset scale.

[0127] According to an embodiment of the present application, the pixel-level addition sub-module includes a pixel-level addition unit.

[0128] The pixel-level addition unit is configured to perform pixel-level addition on the feature map of the preset scale of the original image and the feature maps of the preset scale of a plurality of sampled images to obtain an enhanced feature map.

[0129] According to an embodiment of the present application, any multiple of the sampling module 510, the sampled image block module 520, the sampled image self-attention calculation module 530, the sampled image conversion module 540, and the image enhancement module 550 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present application, at least one of the sampling module 510, the sampled image block module 520, the sampled image self-attention calculation module 530, the sampled image conversion module 540, and the image enhancement module 550 can be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application-specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, at least one of the sampling module 510, the sampled image block module 520, the sampled image self-attention calculation module 530, the sampled image conversion module 540, and the image enhancement module 550 can be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.

[0130] Figure 6 The block diagram of an electronic device suitable for implementing the image processing method according to an embodiment of the present application is shown.

[0131] As Figure 6 shown, the electronic device 600 according to an embodiment of the present application includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. The processor 601 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0132] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.

[0133] According to an embodiment of the present application, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.

[0134] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, an image processing method according to the embodiments of the present application is implemented.

[0135] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present application, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.

[0136] An embodiment of the present application also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the image processing method provided by the embodiment of the present application.

[0137] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0138] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0139] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0140] Those skilled in the art will understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.

[0141] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.

Claims

1. An image processing method, characterized in that, The method includes: Performing multi-scale sampling on the original image to obtain multiple sampled images of different scales; Performing block processing on each sampled image to obtain a block sequence of each sampled image, where the scales of the blocks of the sampled images of different scales are the same; Performing self-attention calculation on the block sequence of each sampled image to obtain an enhanced feature sequence of each sampled image; Converting the enhanced feature sequence of each sampled image into a feature map of a preset scale to obtain multiple feature maps of the preset scale; Generating an enhanced image based on the multiple feature maps of the preset scale.

2. The method according to claim 1, wherein the performing multi-scale sampling on the original image to obtain multiple sampled images of different scales includes: performing layer-by-layer downsampling on the original image with a preset step size to obtain multiple layers of sampled images, such that the sampled images of different layers have different scales; the performing block processing on each sampled image to obtain a block sequence of each sampled image includes: for each sampled image, performing uniform grid partitioning on the sampled image according to the layer where the sampled image is located to obtain multiple blocks, such that the scales of the blocks of the sampled images of different scales are the same; Generating a block sequence of the sampled image based on the multiple blocks.

3. The method according to claim 2, wherein The performing uniform grid partitioning on the sampled image according to the layer where the sampled image is located includes: determining the number of blocks and the block scale of the sampled image according to the layer where the sampled image is located and the scale of the sampled image; 4. The method according to claim 2, wherein The preset step size is 2 n , where n is a positive integer.

5. The method according to claim 1, wherein performing uniform grid partitioning on the sampled image according to the number of blocks and the block scale to obtain multiple blocks, the scales of the multiple blocks are all the block scale, and the number of the multiple blocks is the number of blocks. The converting the enhanced feature sequence of each sampled image into a feature map of a preset scale includes: converting the enhanced feature sequence of each sampled image into an initial feature map, where the scale of the initial feature map is the same as the scale of the sampled image used to obtain the initial feature map; 6. The method according to claim 1, wherein Performing resampling on the initial feature map of each sampled image to obtain a feature map of the preset scale of each sampled image. The generating an enhanced image based on the multiple feature maps of the preset scale includes: performing pixel-level addition on the multiple feature maps of the preset scale to obtain an enhanced feature map; 7. The method according to claim 6, characterized in that, decoding the enhanced feature map to obtain the enhanced image.

8. The method according to claim 7, wherein The preset scale is the scale of the original image. The method further includes: performing block processing on the original image to obtain a block sequence of the original image, where the scale of the blocks of the original image is the same as the scale of the blocks of each sampled image; performing self-attention calculation on the block sequence of the original image to obtain an enhanced feature sequence of the original image; converting the enhanced feature sequence of the original image into a feature map of a preset scale; The performing pixel-level addition on the multiple feature maps of the preset scale to obtain an enhanced feature map includes: performing pixel-level addition on the feature map of the preset scale of the original image and the feature maps of the preset scale of the multiple sampled images to obtain the enhanced feature map.

9. The method according to claim 1, wherein The original image includes a first image and a second image. The first image is an image generated based on detection data of a first modality of an object, and the second image is an image generated based on detection data of a second modality of the object.

10. The method according to claim 1, characterized in that, The self-attention calculation is a multi-head self-attention calculation, and the scale is the resolution.

11. The method according to claim 1, wherein The original image includes a voxel image determined based on three-dimensional volume data.

12. The method according to claim 11, wherein The three-dimensional volume data includes data of a three-dimensional volume formed by stacking a plurality of video frames along a time axis.

13. An electronic device, comprising: One or more processors; A memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.

14. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • CT and MRI image fusion method based on multi-branch multi-scale features and semantic information

    CN117788314A

  • Multi-modal feature fusion image classification method and application in humanoid robot

    CN118628802A

  • Medical image analysis method and system based on multi-modal image fusion

    CN119693365A

  • Low-field magnetic resonance image enhancement method and device, electronic equipment and storage medium

    CN119850762A

  • Medical image segmentation model and method based on efficient multi-scale feature self-attention decoder and enhanced jump connection

    CN119850956A

Cited By

  • Image-based defect detection method, electronic equipment, medium and program product

    CN120782774A