CAD (Computer Aided Design) image super-resolution enhancement method and system based on dual-polymerization Transformer

Through a network architecture based on dual-aggregation Transformer, combining self-attention and convolutional calculations, the global and local features of CAD images are extracted, and the problem of unable to effectively restore texture details in the existing technology is solved, and the generation of high-resolution images is achieved, which improves the readability and design accuracy of the image.

CN120339074APending Publication Date: 2025-07-18TAIZHOU TECHNICIAN COLLEGE (IN PREPARATION)
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510494052.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing deep convolutional neural network-based methods cannot effectively capture global information in the super-resolution enhancement of CAD images, resulting in insufficient texture details restoration, affecting the readability and design accuracy of the image.

Method used

Using a network architecture based on dual-aggregation Transformer, the local and global features of the image are captured through parallel self-attention and convolution calculations, and feature aggregation is performed. The high-frequency texture details of the CAD image are extracted by combining high-frequency filtering and adaptive residual blocks. The gated feedforward network is used to filter redundant information to generate clear high-resolution images.

Benefits of technology

It improves the readability and design accuracy of CAD images, provides more reliable data support, and is suitable for design work in the fields of construction, machinery, electronics and civil engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a CAD (Computer Aided Design) image super-resolution enhancement method and a CAD image super-resolution enhancement system based on a dual-polymerization Transform, and relates to the technical field of image super-resolution. The method comprises the following steps of: designing a resolution enhancement model based on a dual-polymerization Transform network to carry out resolution enhancement on a low-resolution CAD (Computer Aided Design) image; wherein self-attention branch calculation and convolution branch calculation are carried out in parallel in the dual-aggregation Transform module to jointly capture global and local features of an image, and spatial dimension and channel dimension self-attention are respectively adopted to realize inter-block feature fusion; performing adaptive fusion on the features of the two branches to realize intra-block feature aggregation; in the convolution calculation, high-frequency filtering and an adaptive residual block are adopted to perform local detail extraction, and a differentiable high-frequency filtering module can effectively extract high-frequency texture lines in the CAD image. Readability can be improved by utilizing the method to reconstruct the low-quality CAD image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image super-resolution, and particularly to a CAD image super-resolution enhancement method and system based on a dual-aggregation Transformer. Background Art

[0002] Computer-aided design (CAD) drawings are digital design drawings created by CAD software (such as AutoCAD, SolidWorks, Revit, etc.), and are usually used in fields such as architecture, machinery, electronics, and civil engineering. These drawings represent various design schemes in precise two-dimensional (2D) or three-dimensional (3D) formats, and support detailed information such as dimensions, geometric structures, and material annotations. However, in the actual application process, the resolution and clarity of CAD images often become important factors affecting work efficiency and design accuracy. For example, CAD images are usually stored in vector formats such as DWG and DXF, but during transmission or sharing, they are often converted into raster formats such as JPG, PNG, and TIFF, resulting in a decrease in image quality and affecting the clarity of lines and annotations. In order to save storage space or improve transmission speed, CAD images may also be compressed for storage, resulting in blurred details and making it difficult to identify small-sized text, symbols, or thin lines. In addition, many old CAD drawings are scanned and archived, but due to insufficient scanning accuracy, paper aging, or noise pollution, these images often have problems such as blurring, distortion, and low contrast, affecting subsequent editing and utilization. Therefore, CAD image super-resolution technology is extremely important.

[0003] At present, methods based on deep convolutional neural networks (CNNs) have shown remarkable effectiveness in the field of image super-resolution. For example, the literature [Dong C, Loy C C, He K, et al. Learning a deep convolutional network for image super-resolution[C] / / Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part IV 13. Springer International Publishing, 2014: 184-199] first introduced deep CNNs into image super-resolution and achieved obvious performance improvements; the literature [Zhang Y, Li K, Li K, et al. Image super-resolution using very deep residual channel attention networks[C] / / Proceedings of the European conference on computer vision (ECCV). 2018: 286-301] designed a deep residual network with a RIR (Residual-in-Residual) structure and a channel attention mechanism and constructed a model with more than 400 layers; the literature [Lim B, Son S, Kim H, et al. Enhanced deep residual networks for single image super-resolution[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition workshops. 2017: 136-144] optimized the residual blocks by removing unnecessary operations and expanding the model size. Although these models have achieved competitive results, they are pure CNN-based models. This means that they can only extract local features and cannot learn global information, which is not conducive to the restoration of texture details.

[0004] Due to the large number of similar or identical lines and structures in CAD images, these similar contents can be used as reference images for each other to restore the texture details of specific blocks. Transformer has a strong feature expression ability and can simulate long-distance dependencies in images. Therefore, the Transformer framework can be introduced into the super-resolution enhancement of CAD images to reconstruct low-quality CAD images, thereby improving readability and providing more reliable data support for work in related fields. The core idea of Transformer is "self-attention", which can capture long-term information between sequence elements. Transformer has been widely used in computer vision tasks and has been successfully applied to image recognition, object detection, and low-level image processing.

[0005] Although existing methods based on deep CNN or Transformer can effectively enhance image resolution, most of these methods are for natural images. CAD images usually contain precise lines, symbols, and texts. Therefore, how to combine these characteristics to achieve super-resolution enhancement of CAD images is worthy of research. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method and system for super-resolution enhancement of CAD images based on dual-aggregation Transformer.

[0007] According to one aspect of the present invention, a method for super-resolution enhancement of CAD images based on dual-aggregation Transformer is proposed. The method includes: Obtaining a training image set; the training image set includes low-resolution CAD images and corresponding high-resolution CAD images; Training a resolution enhancement model based on a dual-aggregation Transformer network based on the training image set; Inputting the CAD image to be enhanced into the trained resolution enhancement model to obtain a super-resolution CAD image.

[0008] Further, the resolution enhancement model based on the dual-aggregation Transformer network includes: a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; the process of training the resolution enhancement model based on the dual-aggregation Transformer network based on the training image set includes: for the input low-resolution CAD image, first extracting shallow features through a convolution layer with a convolution kernel size of 3×3 in the shallow feature extraction module; then extracting deep features through a deep feature extraction module composed of residual groups, where residual connections are used between each residual group, and each residual group contains a series of A dual-aggregation Transformer module and a convolutional layer with a convolutional kernel size of 3×3; then, a convolutional layer with a convolutional kernel size of 3×3 is used to further capture deep features; then, a high-resolution CAD image is restored through an image reconstruction module, including: using a convolutional layer with a convolutional kernel size of 1×1 to process the deep features; using bilinear interpolation for upsampling; and then using a convolutional layer with a convolutional kernel size of 1×1 for processing.

[0009] Further, the deep feature extraction through the deep feature extraction module composed of residual groups includes:

[0010] Each dual-aggregation Transformer module contains a first Transformer module and a second Transformer module in series; the first Transformer module first performs layer normalization on the input features, then jointly captures global and local context features using a spatial dimension self-attention branch and a convolutional neural network branch, and then adds channel attention and spatial attention after the parallel spatial dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches; then, a gated feed-forward network is used to filter the interactively fused features; the filtered features enter the second Transformer module, the second Transformer module first performs layer normalization on the features output by the first Transformer module, then jointly captures global and local context features using a channel dimension self-attention branch and a convolutional neural network branch, and then adds channel attention and spatial attention after the parallel channel dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches; then, a gated feed-forward network is used to filter the interactively fused features.

[0010] Further, the spatial dimension self-attention is calculated in local image patches in a sliding window manner, specifically including: for the features after layer normalization, first generate a query matrix , a key matrix , and a value matrix through linear mapping; then, perform shape transformation on and to obtain , ; divide , into heads respectively, denoted as , , and ; then, perform self-attention calculation as follows: , where represents Activation function represents the encoding at the corresponding position represents the number of channels for each head i; subsequently, the obtained is subjected to shape adjustment and channel connection, and then through a linear transformation, the final output feature is obtained; The channel - dimensional self - attention divides the feature after layer normalization along the channel dimension into multiple heads, and applies attention to each head respectively, specifically including: first, generating a query matrix through a linear mapping and a key matrix and a value matrix ; subsequently, and are subjected to shape transformation to obtain ; is divided along the channel dimension into heads, denoted as , and ; subsequently, the self - attention calculation is as follows: , where is a learnable parameter; subsequently, the obtained is subjected to shape adjustment and channel connection, and then through a linear transformation, the final output feature is obtained.

[0011] Furthermore, in the convolutional neural network branch, high - frequency filtering and multiple adaptive residual blocks are used to extract high - frequency features in CAD images, specifically including: first, using adaptive residual blocks for feature extraction; subsequently, inputting the extracted features into a high - frequency filtering module to extract high - frequency texture features of CAD images; again using adaptive residual blocks to extract deep features, and using global residual connections to add the original features.

[0012] Furthermore, the adaptive residual block includes two convolutions with kernel sizes of 1×1 and 3×3 respectively and a skip connection, expressed as the following formula:

[0013] where, and respectively represent the input and output of the adaptive residual block in the convolutional neural network; and respectively represent the weights of the residual path and the input feature path; represents the output of the residual path, that is, the output of the 3×3 convolution in the adaptive residual block.

[0014] Further, the step of inputting the extracted features into a high-frequency filtering module to extract the high-frequency texture features of the CAD image includes: performing average pooling on the input features using a pooling layer with a kernel size of k ; then, performing upsampling on the features after average pooling to obtain a new tensor with the restored shape ; finally, subtracting the new tensor element-wise from the input features to obtain the high-frequency texture features of the CAD image .

[0015] Further, the step of adding channel attention and spatial attention after the parallel spatial dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches includes: assuming that is the output of the spatial dimension self-attention, and is the output of the convolutional neural network branch; performing spatial attention calculation on to obtain the spatial attention weight map , and then multiplying pixel-wise with to obtain ; at the same time, performing channel attention calculation on to obtain the channel attention weight map , and then multiplying pixel-wise with to obtain ; adding and to obtain the final interactively fused features; The step of adding channel attention and spatial attention after the parallel channel dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches includes: assuming that is the output of the channel dimension self-attention, and is the output of the convolutional neural network branch; performing channel attention calculation on to obtain the channel attention weight map , and then multiplying pixel-wise with to obtain ; at the same time, performing spatial attention calculation on to obtain the spatial attention weight map , and then multiplying pixel-wise with to obtain ; adding and to obtain the final interactively fused features.

[0016] Further, the filtering of the features after interaction and fusion by using the gated feed-forward network includes: normalizing the input features by using layer normalization; expanding the feature channels by using a convolution with a kernel size of 1×1; performing feature mapping by using a depthwise separable convolution with a kernel size of 3×3; then dividing the output after feature mapping into two parallel branches and performing element-wise multiplication, where one branch is non-linearly activated by GELU; then processing the features after the multiplication operation by using a convolution with a kernel size of 1×1, and adding them to the input features of the gated feed-forward network through a residual connection to obtain the output features.

[0017] According to another aspect of the present invention, a CAD image super-resolution enhancement system based on a dual-aggregation Transformer is proposed, and the system includes: An image acquisition module configured to acquire a training image set; the training image set includes low-resolution CAD images and corresponding high-resolution CAD images; A model training module configured to train a resolution enhancement model based on a dual-aggregation Transformer network based on the training image set; A resolution enhancement module configured to input a CAD image to be enhanced into the trained resolution enhancement model to obtain a super-resolution CAD image.

[0018] The present invention has the following technical effects: The present invention proposes a CAD image super-resolution enhancement method and system based on a dual-aggregation Transformer. The designed resolution enhancement model uses a dual-aggregation Transformer architecture to enhance the resolution of CAD images. Self-attention and convolution calculations are performed in parallel within the Transformer module to jointly capture local and global features of the image to obtain a powerful feature expression ability; and the features of the two branches are adaptively fused to achieve intra-block feature aggregation; different Transformer modules alternately use spatial and channel dimension attention calculations to achieve inter-block spatial and channel feature aggregation. The spatial window can focus on rich feature space expressions, which helps to establish channel dependencies. The channel dimension self-attention provides global information for the spatial dimension self-attention and expands the receptive field of the window; high-frequency filtering and adaptive residual blocks are used in the convolutional neural network for local detail extraction. The differentiable high-frequency filtering module can effectively extract high-frequency texture lines in CAD images and filter out redundant noise information. Using the present invention to reconstruct low-quality CAD images can improve readability and provide more reliable data support for work in related fields. Description of the Drawings

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of a CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to an embodiment of the present invention.

[0021] Figure 2 It is an example diagram of an image in the ModelNet dataset according to an embodiment of the present invention.

[0022] Figure 3 It is a schematic diagram of the overall structure of a resolution enhancement model based on a dual-aggregation Transformer network according to an embodiment of the present invention.

[0023] Figure 4 It is a schematic diagram of the structure of a dual-aggregation Transformer module according to an embodiment of the present invention.

[0024] Figure 5 It is a schematic diagram of the structure of a convolutional neural network branch according to an embodiment of the present invention.

[0025] Figure 6 It is a schematic diagram of the structure of a gated feed-forward network according to an embodiment of the present invention.

[0026] Figure 7 It is a schematic diagram of the structure of a CAD image super-resolution enhancement system based on a dual-aggregation Transformer according to an embodiment of the present invention. Specific Embodiments

[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0028] An embodiment of the present invention proposes a CAD image super-resolution enhancement method based on a dual-aggregation Transformer. As Figure 1 shown, the method includes: S1. Obtain a training image set; S2. Train a resolution enhancement model based on a dual-aggregation Transformer network based on the training image set; S3. Input the CAD image to be enhanced into the trained resolution enhancement model to obtain a super-resolution CAD image.

[0029] The method starts from S1. In S1, a training image set is obtained.

[0030] According to an embodiment of the present invention, the original data of the training image set uses the ModelNet dataset. The ModelNet dataset is an authoritative benchmark dataset in the field of 3D shape recognition, containing 127,915 3D CAD models from 662 categories. The ModelNet dataset was released by the Computer Science and Artificial Intelligence Laboratory of the Massachusetts Institute of Technology in 2013 and has two versions, ModelNet10 and ModelNet40. In this embodiment, the ModelNet40 version is adopted. ModelNet40 covers 40 categories of daily objects (such as chairs, monitors, bookshelves, etc.), contains 12,311 3D models, is divided into 9,843 training samples and 2,468 test samples, and the models are stored in the OFF format, recording vertex coordinates, patch topology, and normal information. The data structure of ModelNet adopts a hierarchical directory design, with subdirectories for the training set and test set under each category, facilitating direct invocation for machine learning tasks. Example images of the dataset are shown as Figure 2 shown.

[0031] Use the rendering tool Blender to generate high-resolution two-dimensional multi-view images corresponding to the 3D CAD models in the ModelNet dataset. Specifically, set the output resolution of the high-resolution two-dimensional multi-view images to 2048×2048 pixels, use the Cycles physical rendering engine, and combine the icosahedron vertex distribution algorithm to generate 80 uniform viewpoints, covering a 360° omnidirectional view. The model surface simulates the real industrial texture (metallic 0.3, roughness 0.6) through the Principled BSDF material, and loads the HDR environment map and multi-directional point light source array to enhance the lighting details (such as shadow transitions and specular reflections). After the generated high-resolution images are denoised, they are downsampled to 512×512 pixels or a lower resolution (such as 256×256) through bicubic interpolation to generate corresponding low-resolution images, constructing low-resolution - high-resolution image pairs. The low-resolution images are used as the model input, and the high-resolution images are used as the ground truth images.

[0032] Then, S2 is executed. In S2, a resolution enhancement model based on the dual-aggregation Transformer network is trained based on the training image set.

[0033] According to an embodiment of the present invention, as Figure 3As shown, the overall model framework includes three modules: shallow feature extraction, deep feature extraction, and image reconstruction. Assume that the input low-resolution CAD image is First, it is processed by a convolutional layer with a kernel size of 3×3 to extract shallow features The deep feature extraction module consists of residual groups, and each residual group contains double-aggregation Transformer modules; a convolutional layer with a kernel size of 3×3 is introduced at the end of the residual group to refine the features extracted from the Transformer blocks; for each residual group, the residual connection method is adopted to effectively ensure the stability of training; subsequently, a 3×3 convolutional layer is added to capture more refined deep features. Finally, the image reconstruction module restores the high-resolution output image including: first, it is processed by a convolutional layer with a kernel size of 1×1, then upsampled by bilinear interpolation, and then processed by a convolutional layer with a kernel size of 1×1.

[0034] In the structure of the double-aggregation Transformer module, intra-block feature aggregation and inter-block feature aggregation are respectively performed on the features extracted by each Transformer module, and the feature maps are sequentially passed through the Transformer for multi-level feature transformation. The "double-aggregation" of the double-aggregation Transformer module is reflected in the feature aggregation between two cascaded blocks, as well as the aggregation of local and global features within each of the two blocks. Since CAD images contain a large number of lines and block diagrams, and there is a high degree of structural similarity between these texture edges, the Transformer model can establish long-distance dependencies, enabling it to find references at different positions in the same image and better restore blurred or damaged textures. However, the local texture details in CAD images are also crucial. The self-attention mechanism focuses on modeling global information. Therefore, a combination of convolution and self-attention is used to jointly capture global and local features.

[0035] Specifically, as Figure 4As shown, after shallow feature extraction of the input image, each Transformer module first performs layer normalization on the input features, and jointly captures local and global context features using a parallel convolutional neural network and self-attention module. Among them, different Transformer modules alternately use self-attention in the spatial dimension and the channel dimension to achieve inter-block feature aggregation; high-frequency filtering and multiple adaptive residual blocks are used in the convolutional neural network to extract high-frequency information in the CAD image; then channel attention and spatial attention are added after the parallel convolutional neural network and self-attention module to interact the features of the two branches and achieve intra-block feature adaptive aggregation; then a gated feed-forward network is used to filter redundant information, introduce non-linear transformation, and enhance the feature expression ability to generate a clear high-resolution CAD image.

[0036] 1) Parallel self-attention module and convolutional neural network 11) Self-attention branch: Inter-block aggregation is completed by alternately using self-attention in the spatial dimension and the channel dimension. Specifically, the self-attention of the first Transformer module in the dual-aggregation Transformer module uses spatial dimension self-attention, and the self-attention of the second Transformer module uses channel dimension self-attention. As Figure 4 shown, the upper and lower Transformer blocks alternately use self-attention in the spatial dimension and the channel dimension: the upper block - spatial dimension self-attention, the lower block - channel dimension self-attention, which are connected in series to form a dual-aggregation Transformer module. Generally, extracting spatial information and capturing channel context are crucial for the performance of Transformer in image super-resolution. Spatial dimension self-attention can model the fine spatial relationship between pixels, and channel dimension self-attention can simulate the relationship between feature maps, thus using global image information.

[0037] Spatial dimension self-attention is calculated in local image patches in a sliding window manner. Assume the input feature is , and first generate a query matrix , a key matrix , and a value matrix through linear mapping, respectively expressed as: , , , where represents linear projection; then and are shape-transformed to obtain , , where represents the number of pixels in non-overlapping image patches. The calculation is performed in a multi-head attention manner, that is, they are divided into number of heads, denoted as , and , and the number of channels for each head is .

[0038] Channel dimension self-attention divides the features along the channel dimension into multiple heads and applies attention to each head separately. First, still perform a linear mapping on the input features to generate a query matrix , a key matrix and a value matrix , and then transform their shapes into ; still divide the projection vector along the channel dimension into heads, and each head has its own independent query, key, and value mappings, so as to calculate different attention outputs. The self-attention calculation methods for each head in the spatial dimension and the channel dimension are as follows:

[0039]

[0040] Among them, represents the position encoding, is a learnable parameter used to adjust the size of the inner product before the softmax function.

[0041] Finally, perform shape adjustment and channel connection on the obtained and , and then through a linear transformation, obtain the final output, denoted as :

[0042]

[0043] Among them, represents the linear transformation matrix.

[0044] Spatial dimension self-attention is calculated within a local spatial window and adopts a sliding window mechanism, enabling information interaction between windows, not limited to the local area only; channel dimension attention focuses on the weight distribution of different feature channels. By assigning different attention weights to different channels, the model can focus on more important feature channels and suppress unimportant channels.

[0045] 12) Convolutional neural network branch: This branch processes the same input features in parallel with the self-attention branch, aiming to jointly capture local and global information. As Figure 5As shown, first, an adaptive residual block is used for feature extraction, and then the extracted features are input into a high-frequency filtering module to extract the high-frequency texture features of the CAD image. Again, an adaptive residual block is used to extract deep features, and a global residual connection is used to add the original features.

[0046] Specifically, the adaptive residual block includes two convolutions with kernel sizes of 1×1 and 3×3 respectively, and a skip connection. When using the skip connection, a residual scaling with adaptive weights is adopted to dynamically adjust the importance of the residual path and the original input path. Compared with fixed residual scaling, this strategy can improve the gradient flow and automatically adjust the residual feature map content of the input feature map.

[0047] In the high-frequency filtering module, first, a pooling layer with a kernel size of k is used to perform average pooling on the input features , and the feature shape changes from to . Each value in this feature can be regarded as the average intensity of each specified area in the input features. Then, the feature is upsampled to restore the shape to , obtaining a new tensor , can be regarded as an expression of the average smoothness information compared with the original . Finally, is subtracted element-wise from to obtain the high-frequency texture in the CAD image . The specific expressions of the adaptive residual block and the high-frequency filtering module are as follows:

[0048]

[0049] Among them, and represent the input and output of the adaptive residual block in the convolutional neural network respectively; and represent the weights of the residual path and the input feature path respectively. By training these two parameters, the model can automatically adjust the residual content of the input feature map and improve the gradient flow; represents the output of the residual path, that is, the output of the 3×3 convolution in the adaptive residual block. represents the CAD image features output by the adaptive residual block, is the high-frequency feature obtained after being processed by the high-frequency filtering module; represents the kernel size of the pooling layer, represents upsampling the features after average pooling; represents average pooling; finally, from Subtract the smoothed information element by element to obtain the high-frequency information 。

[0050] 2) Intra-block aggregation is completed by an adaptive interaction fusion module, that is, adaptively fuse the features extracted by the self-attention branch and the convolutional branch.

[0051] After the above parallel calculations of the self-attention branch and the convolutional branch, although global and local features can be obtained simultaneously, the features extracted by the two branches lack effective coupling. In addition, although alternating the self-attention in the spatial dimension and the channel dimension can capture spatial and channel features, the information in different dimensions still cannot be effectively utilized in a single self-attention. Therefore, an adaptive interaction fusion module is used to adaptively weight the features of the two branches from the spatial or channel dimension according to the type of self-attention mechanism.

[0052] Such as Figure 4 shown, in the Transformer module with parallel processing of spatial dimension self-attention and convolutional neural network, assume is the output of the spatial dimension self-attention,[[]] is the output of the convolutional layer, and the feature interaction is to perform spatial attention calculation on to obtain ; then multiply with pixel by pixel to obtain ; at the same time, perform channel attention calculation on to obtain the attention weight map , and multiply with pixel by pixel to obtain ; finally, add and to obtain the final interaction fusion feature.

[0053] Similarly, in the Transformer module with parallel processing of channel dimension self-attention and convolutional neural network, assume is the output of the channel dimension self-attention,[[]] is the output of the convolutional layer, perform channel attention calculation on to obtain the channel attention weight map , then multiply with pixel by pixel to obtain ; at the same time, perform spatial attention calculation on to obtain the spatial attention weight map , then multiply with pixel by pixel to obtain ; Add and to obtain the final interactive fusion feature. The calculation formulas for the spatial attention and channel attention weight maps are as follows:

[0054]

[0055] Among them, and respectively represent the attention in the spatial and channel dimensions. and are both convolutions of represents the RELU activation function, represents the sigmoid normalization operation; represents the output of the self-attention in the spatial dimension or the output of the self-attention in the channel dimension or the output of the convolutional neural network.

[0056] 3) The gated feed-forward network can filter redundant information in the fused features. As shown in Figure 6 , the gating mechanism is completed by multiplying the elements of two parallel paths of the linear transformation layer, where one path performs GELU non-linear activation; the features after multiplication are used to extract features through a convolution calculation, and finally added to the original input features to obtain the filtered information. The gating mechanism adds non-linear feature expression ability to the model by multiplying the elements of two parallel branches.

[0057] Specifically, assume that the input feature of the gated feed-forward network is . This gated feed-forward network first normalizes the input feature using layer normalization, then expands the feature channels using a 1×1 convolutional kernel, and then performs feature mapping using a depthwise separable convolution with a 3×3 convolutional kernel to learn the local image structure; the gating mechanism is reflected in dividing the output after feature mapping into two parallel branches and performing element multiplication, where one branch is activated by GELU non-linearly; the features after the multiplication operation are refined using a 1×1 convolutional layer and added to the input feature of the gated feed-forward network through a residual connection to obtain the output feature. The specific implementation method of the gated feed-forward network is as follows:

[0058]

[0059] Among them, and both represent convolutions with a 1×1 convolutional kernel, ​​​​​​Represents a depthwise separable convolution with a convolution kernel size of 3×3, represents element-wise multiplication, represents the GELU non-linear activation function, represents layer normalization.

[0060] During the model training process, the reconstructed image output by the model is calculated with the mean squared error loss against the real high-resolution CAD image in the training set to optimize the model training:

[0061] where n is the number of training samples, represents the enhanced high-resolution CAD image, represents the input low-resolution CAD image, are the trainable parameters in the model; represents the real high-resolution CAD image. The mean squared error calculates the squared error between the predicted value and the real value, is more sensitive to large errors, can more strongly penalize samples with larger prediction errors, and helps the model to perform more refined optimization.

[0062] Finally, execute S3. In S3, the CAD image to be enhanced is input into the trained resolution enhancement model to obtain the super-resolution CAD image.

[0063] In summary, the present invention proposes a CAD image super-resolution enhancement method based on dual-aggregation Transformer, in which a resolution enhancement model based on a dual-aggregation Transformer network is constructed. In the model, each Transformer block jointly captures local and global context features by using a parallel convolutional neural network and a self-attention module; different Transformer modules alternately use self-attention in the spatial dimension and the channel dimension to achieve inter-block feature aggregation; high-frequency filtering and multiple adaptive residual blocks are used in the convolutional neural network to extract high-frequency information in the CAD image; channel attention and spatial attention are added after the parallel convolutional neural network and the self-attention module to interact the features of the two branches, achieving adaptive aggregation of block features; a feed-forward network is used to filter redundant information, introducing non-linear transformation to enhance the feature expression ability to generate clear high-resolution CAD images.

[0064] Another embodiment of the present invention proposes a CAD image super-resolution enhancement system based on dual-aggregation Transformer, as Figure 7 shown. The system includes: An image acquisition module 710 configured to acquire a training image set; the training image set includes low-resolution CAD images and corresponding high-resolution CAD images; A model training module 720 configured to train a resolution enhancement model based on a dual-aggregation Transformer network using a training image set; A resolution enhancement module 730 configured to input a CAD image to be enhanced into the trained resolution enhancement model to obtain a super-resolution CAD image.

[0065] For the unspecified parts of a CAD image super-resolution enhancement system based on a dual-aggregation Transformer according to an embodiment of the present invention, please also refer to the specific description of the method embodiment above.

[0066] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A CAD image super-resolution enhancement method based on dual-aggregation Transformer, characterized in that Comprising: Obtain a training image set; The training image set includes low-resolution CAD images and corresponding high-resolution CAD images; Train a resolution enhancement model based on a dual-aggregation Transformer network based on the training image set; Input the CAD image to be enhanced into the trained resolution enhancement model to obtain a super-resolution CAD image.

2. The CAD image super-resolution enhancement method based on double aggregation Transformer according to claim 1, wherein, The resolution enhancement model based on the dual-aggregation Transformer network includes: a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; the process of training the resolution enhancement model based on the dual-aggregation Transformer network using a training image set includes: for the input low-resolution CAD image, first, the shallow feature extraction module extracts shallow features from it using a convolutional layer with a kernel size of 3×3; subsequently, deep features are extracted through a deep feature extraction module composed of residual groups, where residual connections are used between each residual group, and each residual group contains dual-aggregation Transformer modules connected in series and a convolutional layer with a kernel size of 3×3; subsequently, a convolutional layer with a kernel size of 3×3 is used to further capture deep features; subsequently, the high-resolution CAD image is restored through the image reconstruction module, including: processing the deep features using a convolutional layer with a kernel size of 1×1; performing upsampling using bilinear interpolation; and then processing using a convolutional layer with a kernel size of 1×1.

3. A CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 2, characterized in that, The deep feature extraction by means of a deep feature extraction module composed of residual groups includes: Each dual-aggregation Transformer module includes a first Transformer module and a second Transformer module connected in series; the first Transformer module first performs layer normalization on the input features, and then jointly captures global and local context features using a spatial dimension self-attention branch and a convolutional neural network branch. Subsequently, channel attention and spatial attention are added after the parallel spatial dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches. Subsequently, a gated feed-forward network is used to filter the interactively fused features; the filtered features enter the second Transformer module, and the second Transformer module first performs layer normalization on the features output by the first Transformer module, and then jointly captures global and local context features using a channel dimension self-attention branch and a convolutional neural network branch. Subsequently, channel attention and spatial attention are added after the parallel channel dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches. Subsequently, a gated feed-forward network is used to filter the interactively fused features.

4. A CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 3, characterized in that The spatial dimension self-attention is calculated in local image patches in a sliding window manner, specifically including: for the features after layer normalization, first generate a query matrix through a linear mapping , a key matrix and a value matrix ; subsequently, perform shape transformations on and to obtain , ; divide , into heads respectively, denoted as , and ; then perform self-attention calculation as follows: , where represents activation function, represents position encoding, represents the number of channels for each head i; then perform shape adjustment and channel connection on the obtained , and then through a linear transformation, obtain the final output features; The channel dimension self-attention divides the features after layer normalization along the channel dimension into multiple heads, and applies attention to each head respectively, which specifically includes: first generating a query matrix through a linear mapping , a key matrix , and a value matrix ; then performing shape transformations on and to obtain ; dividing along the channel dimension into heads, denoted as , and ; then performing self-attention calculation as follows: , where is a learnable parameter; then performing shape adjustment and channel connection on the obtained , and then through a linear transformation to obtain the final output features.

5. A CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 3, wherein In the convolutional neural network branch, high-frequency filtering and multiple adaptive residual blocks are used to extract high-frequency features in the CAD image, specifically including: first using an adaptive residual block for feature extraction; then inputting the extracted features into a high-frequency filtering module to extract the high-frequency texture features of the CAD image; using an adaptive residual block again to extract deep features, and using a global residual connection to add the original features.

6. The CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 5, wherein The adaptive residual block includes two convolutions with kernel sizes of 1×1 and 3×3 respectively and a skip connection, expressed as the following formula: ; Among them, and respectively represent the input and output of the adaptive residual block in the convolutional neural network; and respectively represent the weights of the residual path and the input feature path; represents the output of the residual path, that is, the output of the 3×3 convolution in the adaptive residual block.

7. A CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 6, characterized in that, The step of inputting the extracted features into a high-frequency filtering module to extract the high-frequency texture features of the CAD image includes: performing average pooling on the input features using a pooling layer with a kernel size of k ; then performing upsampling on the features after average pooling to obtain a new tensor with the restored shape ; finally, subtracting the new tensor from the input features element by element to obtain the high-frequency texture features of the CAD image .

8. A CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 3, characterized in that, Adding channel attention and spatial attention after the parallel spatial dimension self-attention branch and convolutional neural network branch to interactively fuse the features output by the two branches includes: Assume is the output of the spatial dimension self-attention, is the output of the convolutional neural network branch; perform spatial attention calculation on to obtain the spatial attention weight map , and then multiply and pixel by pixel to obtain ; at the same time, perform channel attention calculation on to obtain the channel attention weight map , and then multiply and pixel by pixel to obtain ; add and to obtain the final interactively fused feature; Adding channel attention and spatial attention after the parallel channel - dimension self - attention branch and the convolutional neural network branch to perform interactive fusion on the features output by the two branches includes: Assume is the output of the channel - dimension self - attention, and is the output of the convolutional neural network branch; perform channel attention calculation on to obtain the channel attention weight map , and then multiply and pixel - by - pixel to get ; at the same time, perform spatial attention calculation on to obtain the spatial attention weight map , and then multiply and pixel - by - pixel to get ; add to get the final interactive fusion feature.

9. A CAD image super-resolution enhancement method based on a dual-aggregation Transformer according to claim 3, characterized in that, The filtering of the interactively fused features using the gated feed-forward network includes: normalizing the input features using layer normalization; expanding the feature channels using a convolution with a kernel size of 1×1; performing feature mapping using a depthwise separable convolution with a kernel size of 3×3; then dividing the output after feature mapping into two parallel branches and performing element-wise multiplication, where one branch is activated by GELU non-linearity; subsequently, the features after the multiplication operation are processed using a convolution with a kernel size of 1×1, and the output features are obtained by adding them to the input features of the gated feed-forward network through a residual connection.

10. A CAD image super-resolution enhancement system based on a dual-aggregation Transformer, characterized in that, Comprising: An image acquisition module configured to obtain a training image set; The training image set includes low-resolution CAD images and corresponding high-resolution CAD images; A model training module configured to train a resolution enhancement model based on a dual-aggregation Transformer network based on the training image set; A resolution enhancement module configured to input a CAD image to be enhanced into a trained resolution enhancement model to obtain a super-resolution CAD image.

Citation Information

Cited By

  • Image super-resolution reconstruction method based on channel perception aggregation Transform

    CN120976024A

  • Oral cavity contour modeling method based on context enhanced mixed vision

    CN121259217A

  • Image super-resolution reconstruction method and system based on dual feature clustering

    CN122390970A