Multi-dimensional spatial transformation self-sensing attention mechanism image processing method and application thereof
Through the multi-dimensional spatial transformation self-perception attention mechanism, the problem of the spatial characteristics of the tilted target cannot be effectively processed in the prior art, and the features correction and alignment of the tilted target in the image are realized, and the detection accuracy is improved.
Patent Information
- Application Number
- CN202510275845.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing deep learning object detection methods have shortcomings in processing the spatial characteristics of skewed targets, and cannot effectively correct and align the spatial characteristics of skewed targets in the image, resulting in low detection accuracy.
The multi-dimensional spatial transformation self-perception attention mechanism is adopted. By decomposing the initial feature map into a multi-channel feature map, and performing clockwise and counterclockwise spatial offset operations, combining multi-layer perceptron for channel stitching and compression, corrected spatial attention calculation, and finally generating the output feature map through channel shuffling and multiplication operations.
Effectively correcting and aligning the spatial characteristics of tilted targets in the image improves the detection accuracy of computer vision processing models for surface defects of complex industrial products and drone inspections of specific targets.
Smart Images

Figure CN120107617A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-dimensional space transformation self-perception attention mechanism image processing method and application thereof, belonging to visual detection image processing technology. Background Art
[0002] Object detection methods based on deep learning are widely used in visual reasoning tasks such as surface defect detection of complex industrial products or detection of specific ground targets by drones. Due to the distortion of the image acquisition lens or the changing perspective of the drone movement, the defect features in the acquired image and the specific ground targets waiting to be detected will inevitably be tilted, which makes it easy for object detection methods based on deep learning to miss these tilted and deformed targets.
[0003] Introducing an attention mechanism into deep learning-based object detection methods can guide the network to focus on the target area, which helps improve the network's detection performance. However, the existing spatial attention mechanism still has shortcomings in processing the spatial features of tilted targets, and its ability to process the spatial features of tilted targets in images is not strong: it cannot correct and align the spatial features of tilted targets in images. Therefore, the deep learning object detection method that introduces the conventional spatial attention mechanism cannot meet the high-precision detection requirements of tilted targets in images. Summary of the invention
[0004] The technical problem solved by the present invention is: to address the problem that existing attention methods cannot align and correct the features of tilted targets in images, and to provide an image processing method and application of a multi-dimensional space transformation self-perception attention mechanism.
[0005] The present invention is implemented by the following technical solutions:
[0006] The present invention first provides a multi-dimensional space transformation self-perception attention mechanism image processing method, which automatically extracts spatial features from the input initial feature map, and expands the initial feature map X with C channels into a feature map X with 3C channels. MLP , transform the feature map X from the channel dimension MLP Decompose into three partial feature maps with the same dimensions, perform spatial offset operations in opposite directions on the first and second partial feature maps, then perform channel splicing on the third partial feature map and the first and second partial feature maps after the offset, and compress the channel splicing map into a single-channel intermediate feature map X MLP ', for the intermediate feature map X MLP 'Perform corrected spatial attention calculation to obtain the corrected feature map X MLP ", for the correction feature map X MLP "Shuffle the channels and multiply them with the initial feature map X to get the output feature map F.
[0007] In the multi-dimensional space transformation self-perception attention mechanism image processing method of the present invention, further, the initial feature map is acquired by a visual sensor, which is a visible light image or an infrared image, and the data dimensions of the initial feature map include batch size, channel, height and width.
[0008] In the multi-dimensional space transformation self-perception attention mechanism image processing method of the present invention, further, the expansion of the initial feature map and the compression of the channel splicing map are both processed using a multi-layer perceptron MLP.
[0009] In the multi-dimensional space transformation self-perception attention mechanism image processing method of the present invention, further, the spatial offset operation includes clockwise offset and counterclockwise offset;
[0010] The calculation process of the clockwise offset is as follows:
[0011]
[0012] The calculation process of the counterclockwise offset is as follows:
[0013]
[0014] In the formula, X MLP_1 is the first part of the feature map, is the first part of the feature map after the shift, X MLP_2 is the feature map of the second part, is the second part of the feature map after the shift, b 1 、c 1 、h 1 、w 1 The batch size, channels, height and width of the feature map corresponding to the first part, b 2 、c 2 、h 2 、w 2 Corresponding to the batch size, channel, height and width of the second part feature map, C is the channel value of the initial feature map.
[0015] In the multi-dimensional space transformation self-perception attention mechanism image processing method of the present invention, further, the intermediate feature map X MLP 'Correct spatial attention calculation by the following formula:
[0016]
[0017] In the formula, X MLP " is the output correction feature map, is a 7×7 convolution, F BN is the batch normalization layer (BN), F ReLU is the ReLU activation function, and σ is the Sigmoid function.
[0018] The present invention also provides a computer vision processing method, in which the multi-dimensional space transformation self-perception attention mechanism image processing method mentioned above is arranged in the first stage of the feature extraction network in the computer vision processing method.
[0019] The present invention also provides a computer-readable storage medium based on the above-mentioned computer vision processing method, wherein the computer-readable storage medium stores a computer program, and the computer program is called by a processor to implement the above-mentioned computer vision processing method of the present invention.
[0020] The above-mentioned computer vision processing method of the present invention, which includes the multi-dimensional space transformation self-perception attention mechanism image processing method, can be applied to industrial product surface defect visual inspection equipment.
[0021] The above-mentioned computer vision processing method of the present invention, which includes the multi-dimensional space transformation self-perception attention mechanism image processing method, can also be applied to drone inspection visual detection equipment.
[0022] The present invention adopts the above technical solution to achieve the following beneficial effects:
[0023] (1) The present invention decomposes the initial feature map into a multi-channel feature map according to the channel dimension, and divides the feature map into three parts. Then, the first part of the feature map is rotated to the left to correct the tilt and twist features of the image in the right direction; the second part of the feature map is rotated to the right to correct the tilt and twist features of the image in the left direction; the third part of the feature map remains unchanged to prevent the normal features in the initial feature map from being incorrectly corrected, thereby achieving comprehensive correction of the tilt and twist features in all directions of the image.
[0024] (2) The present invention establishes a spatial attention mechanism to perceive the tilt and distortion characteristics of the target in the spatial dimension, and aligns and corrects its characteristics, thereby improving the detection accuracy of computer vision processing models for complex industrial product surface defects and tilted targets in the process of drone inspection of specific targets.
[0025] In summary, the multi-dimensional space transformation self-perception attention mechanism image processing method provided by the present invention can well correct and align the tilted and distorted features in the feature image processed by machine vision, and is particularly suitable for the high-precision detection requirements of computer vision processing technology in surface defect detection of complex industrial products and drone inspections.
[0026] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flowchart of the image processing method of the multi-dimensional space transformation self-perception attention mechanism of the present invention.
[0028] Figure 2 This is an example diagram of an infrared thermal image feature diagram obtained for a 110KV substation substation equipment input in the embodiment.
[0029] Figure 3 For Figure 2 An example of the infrared thermal image feature map obtained from the output of the substation equipment of a 110KV substation. DETAILED DESCRIPTION
[0030] Example
[0031] Taking the visual inspection task of detecting a substation equipment from infrared images obtained from drones as an example, a total of 1,861 infrared images of substation equipment were generated, including seven types of substation equipment: lightning arrester 1, lightning arrester 2, current transformer 1, current transformer 2, voltage transformer, disconnector, and supporting porcelain bottle. Since the drone takes a bird's-eye view of the substation equipment, the initial image obtained includes many inclined substation equipment targets.
[0032] In this embodiment, YOLOv5 is used as the computer vision processing algorithm model for the above-mentioned visual detection task, and the multi-dimensional space transformation self-perception attention mechanism image processing method of the present invention is placed in the first stage of the feature extraction network of YOLOv5, and the initial feature map X∈R extracted in this stage is B×C×H×W Correction processing is performed, where B, C, H, and W represent the batch size, channel, height, and width of the initial feature map, respectively.
[0033] like Figure 1 As shown, in this embodiment, the initial feature map X∈R B×C×H×W The corrective processing process includes the following steps:
[0034] S100, the input initial feature map X is expanded to 3C by using a multi-layer perceptron (MLP) to obtain a three-channel feature map X MLP ∈R B×3C×H×W .
[0035] S200, transform the three-channel feature map X from the channel dimension MLP Decomposed into three partial feature maps with the same batch size, channel, height and width, respectively recorded as the first partial feature map X MLP_1 , the second part feature map X MLP_2 And the third part feature map X MLP_3 .
[0036] S300, for the first part of the feature map X MLP_1 Perform a clockwise spatial offset operation to obtain the first part of the feature map after offset For the second part feature map X MLP_2Perform a counterclockwise spatial offset operation to obtain the second part of the feature map after offset For the third part feature map X MLP_3 Remain unchanged.
[0037] Among them, the specific operation of the clockwise spatial offset operation is: the first part of the feature map X MLP_1 The first channel moves down one row, the second channel moves up one row, the third channel moves right one column, and the fourth channel moves left one column. The calculation process is as follows:
[0038]
[0039] The specific operation of the counterclockwise spatial offset operation is as follows: the second part of the feature map X MLP_2 The first channel moves one column to the right, the second channel moves one column to the left, the third channel moves one row down, and the fourth channel moves one row up. The calculation process is as follows:
[0040]
[0041] In the formula, X MLP_1 is the first part of the feature map, is the first part of the feature map after the shift, X MLP_2 is the feature map of the second part, is the second part of the feature map after the shift, b 1 、c 1 、h 1 、w 1 The batch size, channels, height and width of the feature map corresponding to the first part, b 2 、c 2 、h 2 、w 2 Corresponding to the batch size, channel, height and width of the second part feature map, C is the channel value of the initial feature map.
[0042] To better understand the above spatial offset transformation operation, assume that there is a 2×2 image block The clockwise spatial transformation operation is: A moves down, B moves up, D moves right, and C moves left, so the transformed image block is It can be seen that the image block has been rotated 45° clockwise to correct the left-leaning feature. The image block of the counterclockwise spatial transformation is also divided into four layers of channels. The first layer of channels moves one column to the right, which is equivalent to the local image being translated to the right; the second layer of channels moves one column to the left, which is equivalent to the local image being translated to the left; the third layer of channels moves one row downward, which is equivalent to the local image being translated upward; the fourth layer of channels moves one row upward, which is equivalent to the local image being translated upward. Similarly, using a 2×2 image block Demonstration, A moves to the right, B moves to the left, D moves down, and C moves up, and the transformed image block is It can be seen that the image is rotated 45° counterclockwise, correcting the feature tilted to the right. In this embodiment, the third partial feature map retains the features of the channel that are not tilted, preventing the normal features in the original feature map from being incorrectly corrected by the other two partial feature maps.
[0043] 400, to X MLP_3 , Perform the concat operation to concatenate the third feature map with the first feature map after the shift and the second feature map after the shift, and then compress the concatenated feature map into an intermediate feature map X of a single channel C through a multi-layer perceptron (MLP). MLP ′.
[0044] S500, for the intermediate feature map X MLP ′According to the following formula, the corrected spatial attention calculation is performed to obtain the corrected feature map X MLP ″.
[0045]
[0046] In the formula, is a 7×7 convolution, F BN is the batch normalization layer (BN), F ReLU is the ReLU activation function, and σ is the Sigmoid function.
[0047] S600, calibrating characteristic graph X MLP ″Perform a channel shuffle operation and multiply it with the input feature map X, that is, X×X MLP ”, due to X MLP What we get is the weight. By multiplying, we select the content in X according to the weight and get the output feature map F.
[0048] In some embodiments, a computer-readable storage medium based on the above-mentioned computer vision processing method is also provided, storing a computer program, and the computer program is called by the processor to implement the YOLOv5 computer vision processing algorithm mentioned above in this embodiment. YOLOv5 is a mature computer image vision processing algorithm, and this embodiment does not describe its specific data processing.
[0049] The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the software and hardware device described in any of the aforementioned embodiments, such as a hard disk or memory of a controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the controller. Furthermore, the readable storage medium may also include both an internal storage unit of the controller and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0050] Based on this understanding, the technical solution of the present invention, in essence or in other words, the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned readable storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., various media that can store program codes.
[0051] The multi-dimensional space transformation self-perception attention mechanism image processing method introduced in this embodiment can also be used in other mature computer vision processing algorithms. It is set in the first stage of the feature extraction network to correct and align the tilted and distorted spatial features in the feature image processed by machine vision.
[0052] The application of the present invention also includes the application to visual inspection equipment. For example, a visual inspection equipment for high-voltage transmission lines is used to inspect the high-voltage transmission lines by drones, and the visual images of the high-voltage transmission lines and the insulators and towers installed thereon are collected by drones. The transmitted image data is processed by the multi-dimensional space transformation self-perception attention mechanism image processing method of the above embodiment to align and correct the tilt features in the image data, and then the computer vision processing method is used for subsequent image recognition and analysis.
[0053] like Figure 2 and Figure 3The image shown is an infrared thermal image of a 110KV substation substation equipment captured by a drone carrying an infrared imager. The image includes voltage transformers and lightning arresters. Due to the random imaging angle of the drone, the substation equipment is tilted in the acquired infrared thermal image, which poses a huge challenge to the deep learning target detection model to detect and identify the substation equipment in the image. Figure 2 The basic YOLOv5 model directly inputs the arrester and mistakenly identifies it as a voltage transformer. Figure 2 Correction processing output is performed by the present invention Figure 3 After that, Figure 3 When input into the YOLOv5 model, the lightning arrester and voltage transformer can be correctly identified.
[0054] The present invention can also be applied to visual inspection equipment for industrial product surface defects. The visual inspection equipment for industrial product surface defects uses the computer vision processing method described in the present embodiment to perform alignment correction on the tilt in the visual image of the industrial product, and then uses the computer vision processing method to identify and analyze product surface defects.
[0055] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present application is a device for realizing the function specified in one flow or multiple flows of the flow chart and / or one or multiple blocks of the block diagram with reference to the instructions executed by the processor according to the method, device (system) and computer program product of the embodiment of the present application. These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which realizes the function specified in one flow or multiple flows of the flow chart and / or one or multiple blocks of the block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0056] In this document, the directions or positional relationships indicated by terms such as "up", "down", "front", "back", "left", "right", "top", "bottom", "inside", "outside", "vertical", and "horizontal" are based on the directions or positional relationships shown in the accompanying drawings and are only for the clarity of expressing the technical solutions and the convenience of description, and therefore should not be understood as limitations of the present invention.
[0057] In this document, the terms "comprises," "comprising," or any other variations thereof, are intended to cover a non-exclusive inclusion of elements other than those listed and may also include additional elements not expressly listed.
[0058] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A multi-dimensional spatial transformation self-perception attention mechanism image processing method that automatically extracts spatial features from the input initial feature map, characterized by: Expand the initial feature map X with channel C to a feature map X with channel 3C MLP , transform the feature map X from the channel dimension MLP Decompose into three partial feature maps with the same dimensions, perform spatial offset operations in opposite directions on the first and second partial feature maps, then perform channel splicing on the third partial feature map and the first and second partial feature maps after the offset, and compress the channel splicing map into a single-channel intermediate feature map X MLP ', for the intermediate feature map X MLP 'Perform corrected spatial attention calculation to obtain the corrected feature map X MLP ", for the correction feature map X MLP "Shuffle the channels and multiply them with the initial feature map X to get the output feature map F.
2. The multi-dimensional space transformation self-perception attention mechanism image processing method according to claim 1, characterized in that: The initial feature map is acquired by a visual sensor and is a visible light image or an infrared image. The data dimensions of the initial feature map include batch size, channel, height and width.
3. The multi-dimensional space transformation self-perception attention mechanism image processing method according to claim 1, characterized in that: The expansion of the initial feature map and the compression of the channel splicing map are both processed using a multi-layer perceptron MLP.
4. The multi-dimensional space transformation self-perception attention mechanism image processing method according to claim 1, characterized in that: The spatial offset operation includes clockwise offset and counterclockwise offset; The calculation process of the clockwise offset is as follows: The calculation process of the counterclockwise offset is as follows: Where, X MLP_1 is the first part of the feature map, is the first part of the feature map after the shift, X MLP_2 is the feature map of the second part, is the second part of the feature map after offset, b1, c1, h1, w1 correspond to the batch size, channel, height and width of the first part of the feature map, b2, c2, h2, w2 correspond to the batch size, channel, height and width of the second part of the feature map, and C is the channel value of the initial feature map.
5. The multi-dimensional space transformation self-perception attention mechanism image processing method according to claim 1, characterized in that: The intermediate feature map X MLP 'Correct spatial attention calculation by the following formula: Where, X MLP " is the output correction feature map, is a 7×7 convolution, F BN is the batch normalization layer (BN), F ReLU is the ReLU activation function, and σ is the Sigmoid function.
6. A computer vision processing method, characterized in that: The multi-dimensional space transformation self-perception attention mechanism image processing method of any one of claims 1-5 is arranged in the first stage of the feature extraction network in the computer vision processing method.
7. A computer-readable storage medium based on the computer vision processing method according to claim 6, characterized in that: A computer program is stored, and the computer program is called by a processor to implement the computer vision processing method according to claim 6.
8. Visual inspection equipment for surface defects of industrial products, characterized by: The visual inspection device uses the computer vision processing method of claim 6.
9. UAV inspection visual inspection equipment, characterized by: The visual inspection device uses the computer vision processing method of claim 6.
Citation Information
Patent Citations
Priori modulated dynamic visual self-attention model strip steel defect detection method
CN116385386A
Deep learning-based multi-aggregation feature pyramid remote sensing image detection method
CN118247670A
Optical image foreign object detection method and device based on end-to-end deep network
CN118628486A
Three-dimensional target detection method based on multi-modal voxel image feature fusion attention
CN119049008A
Ship image classification method based on local and global multi-head attention
CN119339224A