Pattern multi-scale feature fusion detection method based on light convolution and deformation attention
Through the multi-scale feature fusion detection method of patterns based on light convolution and deformation attention, the problems of computational redundancy, adaptability and accuracy in pattern detection are solved, and efficient and accurate pattern target detection is achieved, which is suitable for the digital protection of cultural heritage.
Patent Information
- Application Number
- CN202510679961.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Existing pattern detection technologies suffer from computational redundancy in high-resolution image processing, are difficult to adapt to the diversity of pattern topological structures, suffer from the loss of subtle features, lack sensitivity to low-contrast patterns, are prone to misjudgment in complex backgrounds, and have low efficiency in mobile terminal deployment.
A multi-scale feature fusion detection method for patterns is adopted based on light convolution and deformable attention. Through the cross-stage partial convolution module, deformable attention module, dynamic adaptive spatial feature fusion network and multi-scale feature fusion, combined with the image quality assessment mechanism, feature extraction and fusion are optimized.
It improves the accuracy and efficiency of pattern detection, reduces false detections and missed detections, adapts to complex backgrounds, supports efficient mobile deployment, and provides image quality assessment and re-capture prompts.
Smart Images

Figure CN120198763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pattern detection technology, and in particular to a pattern multi-scale feature fusion detection method based on light convolution and deformation attention. Background Art
[0002] With the increasing demand for digital protection of cultural heritage, automated detection of traditional patterns has become a core technology for cultural relic restoration, style dating, and intelligent pattern analysis. In museum digitization scenarios, ultra-high-precision scanning requires detection algorithms adapted to 80-megapixel-level image processing of cultural relics. It is necessary to identify the topological structure of patterns such as the Taotie pattern on bronze artifacts and the entwined lotus pattern on porcelain under complex background interference. Three-dimensional reconstruction of archaeological sites requires distinguishing subtle brushstroke differences between lacquerware cloud patterns and fabric linked bead patterns with submillimeter accuracy. In the protection of intangible cultural heritage, it is required to maintain a feature integrity of more than 90% for low-contrast patterns such as faded murals and damaged wood carvings, and accurately identify the topological variation characteristics of geometric patterns such as the zigzag pattern and the swastika pattern. In addition, the stringent requirements of mobile pattern retrieval systems for lightweight model deployment, as well as the challenges of cross-domain feature fusion in the task of migrating pattern styles from multiple dynasties, highlight the urgency of high-precision and high-efficiency pattern detection technology.
[0003] Although YOLOv8 has basic detection capabilities in traditional pattern detection scenarios, its native architecture faces significant limitations under the needs of cultural heritage digitization: the multi-level feature fusion mechanism of the C2f module generates computational redundancy when processing high-resolution cultural relic images, resulting in limited mobile deployment efficiency; the fixed-pattern convolution kernels of the backbone network are difficult to adapt to the morphological diversity of the pattern topology, resulting in a lack of continuous expression of the pattern; the excessive downsampling of deep feature extraction causes a serious loss of fine features such as micro-scratches and fine fabric patterns, and the standard detection head is not sensitive enough to low-contrast faded patterns and small patterns; the full-image computing paradigm of the traditional attention mechanism produces significant delays in multi-material mixed scenarios such as porcelain crackle patterns and wood carving three-dimensional patterns; the static fusion strategy of the feature pyramid is difficult to balance the micro-details of the pattern with the macro-cultural semantic association, especially under complex background interference, which is prone to misjudgment of pattern types and blurred boundaries, restricting the high-precision requirements of digital protection of cultural heritage.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a pattern multi-scale feature fusion detection method based on light convolution and deformation attention to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The multi-scale feature fusion detection method of patterns based on light convolution and deformation attention includes the following steps:
[0008] Step 1: Input the image to be detected into the improved CDDA-YOLOv8 model, perform lightweight feature extraction through the cross-stage partial convolution CSPPC module, and generate a CSPPC feature map;
[0009] Step 2: Input the obtained CSPPC feature map into the deformable attention DAT module for feature enhancement to generate the DAT enhanced feature map;
[0010] Step 3: Based on the ASFF spatial adaptive fusion network, the DWR expandable residual module is introduced to form an improved D-ASFF network. The DAT enhanced feature map is input into the D-ASFF network for multi-scale feature fusion to obtain the D-ASFF fused feature map. The DWR expandable residual module includes a regional residual branch and a semantic residual branch.
[0011] Step 4: Collect key parameters in the image processing process and determine the quality evaluation index of the image to be detected based on the key parameters, wherein the key parameters include feature map clarity index, attention focus index and feature consistency index;
[0012] Step 5: Add a 4x downsampled P2 detection head on the basis of the original detection head, generate a high-resolution feature map through the P2 detection head, and perform cross-scale fusion with the D-ASFF fusion feature map to obtain a cross-scale fused feature map. The cross-scale fusion adopts channel splicing and It is achieved by convolution dimensionality reduction;
[0013] Step 6: Input the cross-scale fused feature map into the detection head module, and output the final pattern target detection result and quality assessment report. The pattern target detection result includes category and location information, and the quality assessment report includes a quality assessment index and the corresponding reliability level. When the quality assessment index is lower than the preset quality assessment threshold, a prompt message suggesting re-capturing the image is output synchronously.
[0014] Furthermore, the CDDA-YOLOv8 model is an improved detection model based on the YOLOv8 architecture. Its improvements include: replacing the original C2f module with the cross-stage partial convolution CSPPC module, integrating the deformable attention DAT module in the backbone network, replacing the original feature pyramid network with the dynamic adaptive spatial feature fusion D-ASFF network, and adding a P2 small target detection head.
[0015] The specific logic for generating the CSPPC feature map is as follows: after the input feature map is adjusted for the number of channels by the convolution layer, it is divided into the first part of features and the second part of features along the channel dimension. The first part of features is directly transmitted through the jump connection, and the second part of features enters the container module containing the partial convolution operation for processing. The first part of features is spliced with the processed second part of features, and the CSPPC feature map is output through convolution layer integration. The input feature map refers to the first feature representation obtained after the image to be detected is processed by the initial convolution layer.
[0016] Furthermore, the DAT enhanced feature map is generated based on the following logic: first, multiple evenly distributed reference points are selected on the CSPPC feature map, a dynamic sampling offset based on the query vector is calculated, the reference point positions are adjusted according to the dynamic sampling offset to obtain deformed sampling points, and the deformed point feature values are extracted at the deformed sampling point positions through bilinear interpolation. The deformed point feature values are weightedly fused in combination with the multi-head attention mechanism to output the DAT enhanced feature map;
[0017] The dynamic sampling offset based on the query vector is calculated based on the following logic: The corresponding query vector is obtained by transformation. The offset network takes the query vector of each reference point as input, learns the relationship between the query vector and the offset vector, and obtains the dynamic sampling offset based on the query vector.
[0018] According to the dynamic sampling offset calculated by the offset network, the initial reference point is offset and moved to the new position to obtain the deformed sampling point;
[0019] For each deformed sampling point obtained, bilinear interpolation is performed to obtain a set of eigenvalues corresponding to the deformed sampling point on the feature map. The eigenvalues of the deformed points are weightedly fused in combination with the multi-head attention mechanism to output the DAT enhanced feature map.
[0020] The offset network consists of two convolutional modules, which capture local features through the deep convolution layer, pass the GELU activation function, and then obtain the final dynamic sampling offset through 1×1 convolution. The transformation is implemented through a 1×1 convolution layer, whose convolution kernel weight matrix It is learned during training that for dimensions The input feature map is first converted into The matrix, and then Matrix multiplication is finally reconstructed into The query vector, 、 and are the height of the input feature map, the width of the input feature map, and the number of channels of the input feature map, respectively. is the query vector dimension.
[0021] Furthermore, the DAT enhanced feature map is input into the D-ASFF network for multi-scale feature fusion. The specific logic is as follows: multi-scale context information is extracted through the expandable residual DWR module, and then an adaptive weight map is generated based on the adaptive spatial feature fusion ASFF network, and features at different levels are weightedly fused to output the D-ASFF fused feature map. The DWR expandable residual module includes a regional residual branch and a semantic residual branch. The output of the DWR module is connected to the adaptive weight map generation module of the ASFF network.
[0022] Furthermore, the regional residual branch refers to the use of dilated convolutions with dilation rates of 1, 3, and 5 to extract local fine-grained features in parallel; the semantic residual branch refers to the use of point-by-point convolution to fuse multi-scale features and retain global semantic information.
[0023] Furthermore, the feature map clarity index is calculated by performing Sobel operator gradient calculation on the feature map of each downsampling level in the D-ASFF network, and taking the standard deviation of the gradient amplitude of all pixels as the feature map clarity index. The formula is as follows:
[0024] ;
[0025] Where, is the feature map clarity index, is the total number of pixels in the feature map of each downsampling level in the D-ASFF network, is the index of the pixel, is the mean of all pixel gradient magnitudes, Represents the gradient amplitude of the i-th pixel in the feature map, which is calculated by the Sobel operator:
[0026] ;
[0027] Where, and Represents the gradient map of the feature map in the x and y directions respectively;
[0028] ;
[0029] ;
[0030] Where, is the horizontal gradient kernel, is the vertical gradient kernel, Represents the transpose operation of the matrix;
[0031] The method for calculating the attention focus index is: counting the attention weight distribution entropy of all attention heads in the DAT module, and taking the inverse characteristic number of the entropy value as the attention focus index. The formula is as follows:
[0032] ;
[0033] Where, is an indicator of attention focus. is the total number of attention heads in the DAT module, is the index of the attention head, It is The entropy of the weight distribution of attention heads:
[0034] ;
[0035] Where, It is In the attention head The normalized attention weights of the query vector, Represents the total number of query vectors of the attention head, where the query vector refers to the coordinates of the deformed sampling points generated by dynamic offset in the DAT module;
[0036] The feature consistency index is calculated by calculating the mutual information of two adjacent scale feature maps in the D-ASFF network and taking the average mutual information of all scale pairs as the feature consistency index. The formula is as follows:
[0037] ;
[0038] Where, is the feature consistency index, represents the total number of feature map levels in the D-ASFF network, is the adjacent scale feature map and The mutual information of The index of the feature map level:
[0039] ;
[0040] Where, represents the joint probability distribution, estimated by the eigenvalue histogram, and They represent the eigenvalues of the feature maps of adjacent levels in the D-ASFF network at the spatially aligned positions, and is the marginal probability distribution.
[0041] Furthermore, the quality evaluation index of the image to be detected is determined based on the feature map clarity index, attention focus index and feature consistency index, and the formula is as follows:
[0042] ;
[0043] Where, is the quality assessment index, is the feature map clarity index, is an indicator of attention focus. is the feature consistency index, 、 and is the preset weight value, , and satisfies .
[0044] Furthermore, a high-resolution feature map is generated by the P2 detection head and fused with the D-ASFF fusion feature map across scales. The specific logic is as follows:
[0045] The 4-fold downsampled feature map output by the backbone network is processed by 3×3 convolution and 1×1 convolution in sequence to generate The high-resolution feature map of resolution is obtained by bilinearly upsampling the D-ASFF fusion feature map by 2 times. The high-resolution feature map and the up-sampled feature map are spliced along the channel dimension to form The fusion feature map is finally compressed to the number of channels through 1×1 convolution. , output the cross-scale fused feature map, where is the number of feature map channels. The backbone network is the improved feature extraction backbone in the CDDA-YOLOv8 model, and its Stage 2 outputs a 4x downsampled feature map.
[0046] Furthermore, the method for determining the quality assessment threshold is: collecting a variety of pattern image data sets as samples, using the same method to calculate the quality assessment index of each image, selecting the 75% quantile of the sample quality assessment index as the dividing threshold between "clear" and "blurred", that is, the quality assessment threshold, and when the quality assessment index of the image to be detected is lower than the preset quality assessment threshold, synchronously outputting a prompt message suggesting re-capturing the image.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] The combination of the improved CSPPC module and the DAT module in this invention makes the feature extraction process more efficient and enables adaptive adjustment of feature fusion strategies, significantly improving detection accuracy. Furthermore, combined with an image quality assessment mechanism, it can identify low-quality images during real-time detection and provide recommendations for re-acquisition. This not only optimizes the detection process but also effectively reduces false detections and missed detections, improving the reliability and practicality of the overall system. Through these technological innovations, this solution has significant application value and market potential in the field of traditional pattern object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the overall method flow of the present invention;
[0050] Figure 2 、 Figure 3 and Figure 4 They are the fitting curves of the feature map clarity index and the quality evaluation index, the fitting curve of the attention focus index and the quality evaluation index, and the fitting curve of the feature consistency index and the quality evaluation index;
[0051] Figure 5 This is the improved model structure diagram. DETAILED DESCRIPTION
[0052] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.
[0053] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0054] Example:
[0055] See also Figure 1 and Figure 5 , the present invention provides a technical solution:
[0056] A multi-scale feature fusion detection method for patterns based on light convolution and deformation attention, the specific steps include:
[0057] Step 1: Input the image to be detected into the improved CDDA-YOLOv8 model, perform lightweight feature extraction through the cross-stage partial convolution CSPPC module, and generate a CSPPC feature map;
[0058] In this embodiment, the CDDA-YOLOv8 model is a detection model improved based on the YOLOv8 architecture. Its improvements include: replacing the original C2f module with a cross-stage partial convolution CSPPC module, integrating a deformable attention DAT module in the backbone network, replacing the original feature pyramid network with a dynamic adaptive spatial feature fusion D-ASFF network, and adding a P2 small target detection head.
[0059] The specific logic for generating the CSPPC feature map is as follows: after the input feature map is adjusted for the number of channels by the convolution layer, it is divided into the first part of features and the second part of features along the channel dimension. The first part of features is directly transmitted through the jump connection, and the second part of features enters the container module containing the partial convolution operation for processing. The first part of features is spliced with the processed second part of features, and the CSPPC feature map is output through convolution layer integration. The input feature map refers to the first feature representation obtained after the image to be detected is processed by the initial convolution layer.
[0060] Step 1 introduces the Cross-Stage Partial Convolution (CSPPC) module through the improved CDDA-YOLOv8 model, making the feature extraction process more lightweight and efficient. By splitting the feature map into two parts along the channel dimension, this module effectively maintains information transfer and feature richness, reducing computational complexity and improving the model's real-time performance. Furthermore, the use of skip connections ensures that important information is not degraded, helping to improve detection accuracy. Compared to traditional convolutional structures, this method not only optimizes feature extraction efficiency but also enhances the ability to capture details of traditional patterns.
[0061] Compared with existing technologies, this lightweight convolution-based design method is significantly superior in processing speed and accuracy. Traditional convolutional neural networks often face the problem of high computing resource consumption, while the introduction of the CSPPC module effectively reduces the number of model parameters and computational complexity, making it suitable for execution in resource-constrained environments. At the same time, due to its sophisticated feature extraction capabilities, it can better capture pattern details at multiple scales, improving the overall detection effect. This efficient feature extraction and fusion mechanism enables the model to demonstrate greater adaptability and accuracy in traditional pattern target detection tasks. In this solution, the feature extraction process in step 1 provides a solid foundation for subsequent feature enhancement and fusion. The high-quality feature map generated by the CSPPC module can provide rich contextual information for the subsequent deformable attention DAT module, promoting the effect of feature enhancement. This design makes the subsequent multi-scale feature fusion more effective, improving the accuracy and robustness of the final detection results. Therefore, step 1 not only optimizes the feature extraction process, but also significantly enhances the performance of the entire detection system, laying the foundation for efficient and accurate traditional pattern detection.
[0062] Step 2: Input the obtained CSPPC feature map into the deformable attention DAT module for feature enhancement to generate the DAT enhanced feature map;
[0063] In this embodiment, the DAT enhanced feature map is generated based on the following specific logic: first, multiple uniformly distributed reference points are selected on the CSPPC feature map, a dynamic sampling offset based on the query vector is calculated, the reference point positions are adjusted according to the dynamic sampling offset to obtain deformed sampling points, and the deformed point feature values are extracted at the deformed sampling point positions through bilinear interpolation. The deformed point feature values are weightedly fused in combination with a multi-head attention mechanism to output the DAT enhanced feature map;
[0064] The dynamic sampling offset based on the query vector is calculated based on the following logic: The corresponding query vector is obtained by transformation. The offset network takes the query vector of each reference point as input, learns the relationship between the query vector and the offset vector, and obtains the dynamic sampling offset based on the query vector.
[0065] According to the dynamic sampling offset calculated by the offset network, the initial reference point is offset and moved to the new position to obtain the deformed sampling point;
[0066] For each deformed sampling point obtained, bilinear interpolation is performed to obtain a set of eigenvalues corresponding to the deformed sampling point on the feature map. The eigenvalues of the deformed points are weightedly fused in combination with the multi-head attention mechanism to output the DAT enhanced feature map.
[0067] The offset network consists of two convolutional modules, which capture local features through the deep convolution layer, pass the GELU activation function, and then obtain the final dynamic sampling offset through 1×1 convolution. The transformation is implemented through a 1×1 convolution layer, whose convolution kernel weight matrix It is learned during training that for dimensions The input feature map is first converted into The matrix, and then Matrix multiplication is finally reconstructed into The query vector, 、 and are the height of the input feature map, the width of the input feature map, and the number of channels of the input feature map, respectively. is the query vector dimension.
[0068] Step 2 significantly improves the quality of feature representation by feeding the CSPPC feature map into the Deformable Attention (DAT) module for feature enhancement. This module not only generates an initial uniform distribution of reference points but also flexibly adjusts feature point positions by calculating dynamic sampling offsets, enabling refined feature capture of complex patterns. Combined with a multi-head attention mechanism, the deformed point feature values are weightedly fused, ensuring that important information in the feature map is focused on and utilized. This process enables the model to more effectively capture subtle differences and changes in the image, thereby enhancing the accuracy and robustness of object detection.
[0069] Compared to existing techniques, step 2 provides a more adaptable feature enhancement method. Traditional feature extraction often lacks the flexibility to adapt to objects of varying scales and shapes. However, the Deformable Attention (DAT) module dynamically adjusts the sampling positions of feature points to better adapt to a variety of complex pattern forms. This innovation not only improves the detail of feature extraction but also optimizes the spatial information validity of the feature map, thereby enhancing object detection performance. Furthermore, the DAT module utilizes a multi-head attention mechanism to effectively focus on important regions of the pattern, avoiding information redundancy and loss, further improving detection accuracy. In this solution, the implementation of step 2 provides strong feature enhancement support for the entire detection process, ensuring the effectiveness of subsequent multi-scale feature fusion. By inputting the enhanced feature map generated by the DAT module into the improved D-ASFF network, the fused features are enriched and accurate. This enhanced feature representation not only provides a stronger foundation for recognizing complex patterns but also makes subsequent detection results more reliable and comprehensive, thereby improving the overall detection effectiveness and accuracy of the overall solution. Therefore, step 2 plays a core role in the entire process, laying a solid foundation for efficient and accurate detection of objects with traditional patterns.
[0070] Step 3: Based on the ASFF spatial adaptive fusion network, the DWR expandable residual module is introduced to form an improved D-ASFF network. The DAT enhanced feature map is input into the D-ASFF network for multi-scale feature fusion to obtain the D-ASFF fused feature map. The DWR expandable residual module includes a regional residual branch and a semantic residual branch.
[0071] In this embodiment, the DAT enhanced feature map is input into the D-ASFF network for multi-scale feature fusion. The specific logic is as follows: multi-scale context information is extracted through the expandable residual DWR module, and then an adaptive weight map is generated based on the adaptive spatial feature fusion ASFF network, and features at different levels are weightedly fused to output the D-ASFF fused feature map. The DWR expandable residual module includes a regional residual branch and a semantic residual branch. The output of the DWR module is connected to the adaptive weight map generation module of the ASFF network.
[0072] The regional residual branch refers to the use of dilated convolutions with dilation rates of 1, 3, and 5 to extract local fine-grained features in parallel; the semantic residual branch refers to the use of point-by-point convolution to fuse multi-scale features and retain global semantic information.
[0073] The improved D-ASFF network introduced in step 3, combined with the DWR expandable residual module, significantly enhances the ability to fuse multi-scale features. Through the design of regional residual branches and semantic residual branches, the DWR module effectively extracts contextual information at different scales, ensuring that both fine-grained features and global semantic information are better preserved and utilized during feature fusion. This multi-level feature fusion strategy enables the system to more accurately capture important information in images when handling complex traditional pattern detection tasks, thereby improving detection accuracy and robustness.
[0074] Compared with existing feature fusion techniques, the combined approach of the D-ASFF network and the DWR module provides a more flexible and efficient feature fusion mechanism. Traditional feature fusion methods often suffer from information loss or feature redundancy when processing multi-scale information. This step effectively addresses this issue by generating an adaptive weight map and combining regional and semantic information. This allows features at different scales to maintain their information richness during the fusion process while reducing the computational burden, thereby improving the overall performance and adaptability of the detection model. In this solution, the implementation of step 3 provides key support for the performance improvement of the entire detection system. Multi-scale feature fusion via the D-ASFF network effectively combines features from different levels, providing a more comprehensive and accurate foundation for subsequent detection and classification tasks. This integration mechanism not only enhances the model's adaptability and recognition capabilities for traditional patterns, but also makes the final detection results more reliable. Therefore, step 3 plays a crucial role in improving feature expression and detection performance within the overall solution, providing a strong guarantee for efficient and accurate traditional pattern detection.
[0075] Step 4: Collect key parameters in the image processing process and determine the quality evaluation index of the image to be detected based on the key parameters, wherein the key parameters include feature map clarity index, attention focus index and feature consistency index;
[0076] In this embodiment, the feature map clarity index is calculated by performing a Sobel operator gradient calculation on the feature map of each downsampling level in the D-ASFF network, and taking the standard deviation of the gradient amplitude of all pixels as the feature map clarity index. The formula is as follows:
[0077] ;
[0078] Where, is the feature map clarity index, is the total number of pixels in the feature map of each downsampling level in the D-ASFF network, is the index of the pixel, is the mean of all pixel gradient magnitudes, Represents the gradient amplitude of the i-th pixel in the feature map, which is calculated by the Sobel operator:
[0079] ;
[0080] Where, and Represents the gradient map of the feature map in the x and y directions respectively;
[0081] ;
[0082] ;
[0083] Where, is the horizontal gradient kernel, is the vertical gradient kernel, Represents the transpose operation of the matrix;
[0084] The method for calculating the attention focus index is: counting the attention weight distribution entropy of all attention heads in the DAT module, and taking the inverse characteristic number of the entropy value as the attention focus index. The formula is as follows:
[0085] ;
[0086] Where, is an indicator of attention focus. is the total number of attention heads in the DAT module, is the index of the attention head, It is The entropy of the weight distribution of attention heads:
[0087] ;
[0088] Where, It is In the attention head The normalized attention weights of the query vector, Represents the total number of query vectors of the attention head, where the query vector refers to the coordinates of the deformed sampling points generated by dynamic offset in the DAT module;
[0089] The feature consistency index is calculated by calculating the mutual information of two adjacent scale feature maps in the D-ASFF network and taking the average mutual information of all scale pairs as the feature consistency index. The formula is as follows:
[0090] ;
[0091] Where, is the feature consistency index, Indicates the total number of feature map levels in the D-ASFF network, which is used to reflect the depth of multi-scale fusion. The more levels there are, the more it can cover the decorative features from micro to macro. The value is determined by the number of downsampling stages of the backbone network. is the adjacent scale feature map and The mutual information of is the index of the feature map level, The larger the value, the stronger the synergy between shallow details and deep semantics, and the more consistent the cross-scale expression of the pattern:
[0092] ;
[0093] Where, Represents the joint probability distribution, which is used to reflect the statistical dependence of pattern features across scales. The larger the value, the stronger the correlation between details and semantics. It is estimated by the eigenvalue histogram. and They represent the eigenvalues of the feature maps of adjacent levels in the D-ASFF network at the spatially aligned positions, and is the marginal probability distribution.
[0094] The quality evaluation index of the image to be detected is determined based on the feature map clarity index, attention focus index, and feature consistency index. The formula is as follows:
[0095] ;
[0096] Where, is the quality assessment index, is the feature map clarity index, is an indicator of attention focus. is the feature consistency index, 、 and is the preset weight value, , , The detection of cultural relics patterns is highly dependent on the integrity of edges and textures. It directly reflects the resolvability of the image. Experiments show that clarity has the greatest impact on detection accuracy, especially in low-contrast patterns. This measures the model's ability to locate key areas, which is crucial for distinguishing patterns against complex backgrounds. However, its influence is slightly lower than clarity, so it is given a 30% weight. It ensures the coordination between shallow details and deep semantics, but has a weak direct impact on the final detection results and mainly plays an auxiliary optimization role, so the weight is set to be low.
[0097] Quality Assessment Index It is a comprehensive indicator that aims to quantify the overall quality of the image to be detected. The higher the value, the better the quality of the image to be detected, and vice versa. In this model, three key factors are combined: feature map clarity, attention focus, and feature consistency to provide a comprehensive assessment of image quality. Feature map clarity indicator Directly affects the visual quality of the image. The higher the value, the higher the clarity and the richer the image details, which can help the model identify the target more accurately. Reflects the effectiveness of the model in detecting the target. A high degree of focus means that the model can better focus on important features, thereby improving the accuracy of detection. Feature consistency index It is used to measure the consistency between features of different scales. High feature consistency helps to ensure the effect of multi-scale feature fusion and improve the overall detection performance. 、 、 and There is a positive correlation.
[0098] The formal rationality of this formula is reflected in the following aspects: First, the logarithmic function is used in the formula To process the feature map clarity index , which can effectively reduce the high-definition value This approach allows the quality assessment to maintain a reasonable growth even in the case of high clarity. In square root form , which aims to emphasize the positive contribution of focus improvement to image quality without over-exaggerating its impact, so that the performance of the model at different focus levels can be reasonably evaluated. Finally, the feature consistency index Use square form , which is designed to strengthen the influence of feature consistency on quality assessment, ensuring that when feature consistency is high, quality assessment can be significantly improved. This processing method highlights the importance of feature consistency in image quality and can effectively improve the sensitivity of the overall assessment. At the same time, the given weight value , , This clarifies the relative importance of each indicator in quality assessment, reflects a reasonable balance between clarity, focus, and consistency, and ensures a more accurate comprehensive evaluation result based on the combined effects of these factors. In summary, this formula is formally sound and can comprehensively and effectively reflect the quality of the image being tested.
[0099] Table 1: Quality Assessment Index Statistics
[0100] ;
[0101] See also Figure 2-Figure 4 In this data analysis, according to the statistical data in Table 1, the quality assessment index shows a significant positive correlation with the feature image clarity index, attention focus index and feature consistency index, which verifies the rationality of the formula design. The specific analysis is as follows: 、 、 The gradual improvement of The values show a monotonically increasing trend, indicating that the contribution of each indicator to image quality is consistent. When it increases from 0.01 to 0.9, From 0.0911 to 0.875, the increase is 860%. The improvement in clarity is beneficial to The impact is most significant. ), The value increases slowly (such as serial numbers 1-6), indicating that when the image quality is poor, a small improvement in a single indicator will have limited impact on the overall effect; while in the high index range (such as ), The speed increase (such as serial number 16-20) meets the actual demand of “high-quality images are easier to achieve stable detection”. and When synchronously promoted (such as sequence number 11-15), The increase is higher than that of a single indicator, indicating that the coordinated optimization of attention focus and feature consistency can significantly enhance detection reliability. For example, No. 14 ( , , )of The value has exceeded the quality threshold, indicating that the best effect is achieved when the three factors are balanced. The increase in It increases significantly (0.238→0.339), which verifies the key role of multi-scale feature fusion in preserving pattern details. When it is the quality dividing line, it corresponds to 、 、 (No. 14), which is highly consistent with the requirements of “clear edge + high attention focus + cross-scale alignment” in the digital scene of cultural heritage. (Sequence number 1-5), when the image is severely blurred or features are broken, the system triggers a re-acquisition prompt to effectively avoid false detection. This formula accurately quantifies the image quality through nonlinear weighting, and its output is The value is strongly correlated with the actual detection performance and can effectively guide the engineering application of pattern detection.
[0102] The key to step 4 is determining the quality assessment index of the image to be inspected by collecting key parameters from the image processing process. This approach not only focuses on the detection results output by the model but also incorporates a quantitative assessment of image quality to ensure the effectiveness and feasibility of the entire detection process. By calculating the feature map clarity index, attention focus index, and feature consistency index, it is possible to comprehensively assess the image quality during the feature extraction and detection process, providing an important basis for subsequent detection decisions. This quantitative evaluation method ensures the adaptability and accuracy of the detection system.
[0103] Compared to traditional object detection methods, step 4 significantly enhances the detection system's intelligence and accuracy by introducing the concept of image quality assessment. Existing technologies often rely on fixed detection thresholds and standards, lacking dynamic adaptation to input image quality. This step, by calculating key parameters, enables real-time assessment and adjustment of processing strategies for varying image quality. This flexibility ensures the system maintains high detection performance across varying shooting conditions and image quality, reducing detection errors due to poor image quality. In this solution, the implementation of step 4 provides essential quality control mechanisms for the overall detection process. By monitoring image quality in real time, prompts for re-acquisition are issued when issues are identified, preventing the final detection results from being affected by poor input image quality. This mechanism not only improves the reliability of detection results but also optimizes subsequent processing steps, reducing unnecessary errors and wasted resources. Therefore, step 4 plays a crucial role in quality assessment and control within the overall solution, providing crucial support for efficient and accurate detection of traditional pattern objects.
[0104] Step 5: Add a 4x downsampled P2 detection head on the basis of the original detection head, generate a high-resolution feature map through the P2 detection head, and perform cross-scale fusion with the D-ASFF fusion feature map to obtain a cross-scale fused feature map. The cross-scale fusion adopts channel splicing and It is achieved by convolution dimensionality reduction;
[0105] In this embodiment, a high-resolution feature map is generated by the P2 detection head and cross-scale fusion is performed with the D-ASFF fusion feature map. The specific logic is as follows:
[0106] The 4-fold downsampled feature map output by the backbone network is processed by 3×3 convolution and 1×1 convolution in sequence to generate The high-resolution feature map of resolution is obtained by bilinearly upsampling the D-ASFF fusion feature map by 2 times. The high-resolution feature map and the up-sampled feature map are spliced along the channel dimension to form The fusion feature map is finally compressed to the number of channels through 1×1 convolution. , output the cross-scale fused feature map, where is the number of feature map channels. The backbone network is the improved feature extraction backbone in the CDDA-YOLOv8 model, and its Stage 2 outputs a 4x downsampled feature map.
[0107] The key advantage of step 5 lies in the introduction of a 4x downsampled P2 detection head to generate a high-resolution feature map, which is then cross-scale fused with the D-ASFF fused feature map. This approach not only effectively improves the resolution of the feature map but also combines information from different scales, enhancing the model's ability to detect small objects and details. Through cross-scale fusion, the model can better capture the complex features of traditional patterns, achieving more accurate object localization and classification.
[0108] Compared with the existing technology, step 5 significantly improves the detection system's adaptability to small targets by adding a P2 detection head and a cross-scale fusion mechanism. Traditional detection methods usually do not work well when facing small targets or low-resolution feature maps, which can easily lead to false detection or missed detection. Through the design of this step, the model can fully utilize the combination of high-resolution feature maps and multi-scale information, effectively improving detection accuracy and reliability. This innovative fusion strategy enables the system to perform better in a variety of practical application scenarios. In this solution, the implementation of step 5 injects stronger detail capture capabilities into the overall detection system, ensuring the system's efficiency and accuracy when facing complex patterns. By generating high-resolution feature maps and performing cross-scale fusion, this step improves the richness and comprehensiveness of the model's processing information, providing a more solid foundation for subsequent target detection. Overall, the design of step 5 not only enhances the performance of the model, but also optimizes the detection process, improves its practical value in actual applications, and ensures efficient and accurate detection of patterned targets.
[0109] Step 6: Input the cross-scale fused feature map into the detection head module, and output the final pattern target detection result and quality assessment index. The pattern target detection result includes category and location information. When the quality assessment index is lower than the preset quality assessment threshold, a prompt message is output simultaneously to recommend re-capturing the image.
[0110] In this embodiment, the method for determining the quality assessment threshold is: collecting a variety of pattern image data sets as samples, using the same method to calculate the quality assessment index of each image, selecting the 75% quantile of the sample quality assessment index as the dividing threshold between "clear" and "blurred", that is, the quality assessment threshold, and when the quality assessment index of the image to be detected is lower than the preset quality assessment threshold, synchronously outputting a prompt message suggesting re-capturing the image.
[0111] The key advantage of step 6 is that the cross-scale fused feature map is input into the detection head module, which then outputs the final pattern detection results and quality assessment index. This process not only accurately detects target category and location information, but also promptly issues a prompt to recapture the image if the quality assessment index falls below a set threshold. This mechanism effectively ensures the reliability of detection results and enables the system to dynamically adjust to changes in image quality when dealing with complex traditional patterns, thereby improving overall detection performance.
[0112] Compared to existing technologies, step 6 enhances the intelligence of object detection by integrating cross-scale feature map fusion and quality assessment mechanisms. Traditional technologies often focus solely on detection results while ignoring the impact of image quality on detection accuracy, which can easily lead to false detections or missed detections in low-quality images. This step, by introducing a quality assessment index, proactively issues alerts when quality falls short, enhancing the system's automation and intelligence. This innovative design makes the detection system more user-friendly and effectively avoids detection failures due to image quality issues. In this solution, the implementation of step 6 provides a crucial feedback mechanism for the overall detection process, monitoring the quality of detection results in real time to ensure the system's efficiency and stability in practical applications. By outputting category, location information, and a quality assessment index, users can promptly understand the detection results and make adjustments or re-acquire images as necessary. This design not only improves the adaptability and reliability of the detection system but also provides valuable data support for subsequent image processing, enhancing the application effectiveness and practical value of the overall solution in the detection of traditional multi-scale features.
[0113] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0114] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software depends on the specific application and design constraints of the technical solution.
[0115] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0116] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A multi-scale feature fusion detection method for patterns based on light convolution and deformation attention, characterized in that: The specific steps include: Step 1: Input the image to be detected into the improved CDDA-YOLOv8 model, perform lightweight feature extraction through the cross-stage partial convolution CSPPC module, and generate a CSPPC feature map; Step 2: Input the obtained CSPPC feature map into the deformable attention DAT module for feature enhancement to generate the DAT enhanced feature map; Step 3: Based on the ASFF spatial adaptive fusion network, the DWR expandable residual module is introduced to form an improved D-ASFF network. The DAT enhanced feature map is input into the D-ASFF network for multi-scale feature fusion to obtain the D-ASFF fused feature map. The DWR expandable residual module includes a regional residual branch and a semantic residual branch. Step 4: Collect key parameters in the image processing process and determine the quality evaluation index of the image to be detected based on the key parameters, wherein the key parameters include feature map clarity index, attention focus index and feature consistency index; Step 5: Add a P2 detection head with 4x downsampling to the original detection head, generate a high-resolution feature map through the P2 detection head, and perform cross-scale fusion with the D-ASFF fusion feature map to obtain a cross-scale fused feature map. The cross-scale fusion is achieved by channel splicing and 1×1 convolution dimensionality reduction. Step 6: Input the cross-scale fused feature map into the detection head module, and output the final pattern target detection result and quality assessment index. The pattern target detection result includes category and location information. When the quality assessment index is lower than the preset quality assessment threshold, a prompt message is output simultaneously to recommend re-capturing the image. The specific logic for generating the CSPPC feature map is as follows: after adjusting the number of channels of the input feature map through the convolution layer, it is divided into the first and second features along the channel dimension. The first feature is directly transferred through the skip connection, and the second feature enters the container module containing the partial convolution operation for processing. The first feature is concatenated with the processed second feature, and then integrated through the convolution layer to output the CSPPC feature map. The input feature map refers to the first feature representation obtained after the image to be detected is processed by the initial convolution layer; The specific logic for generating the DAT enhanced feature map is as follows: first, multiple evenly distributed reference points are selected on the CSPPC feature map, the dynamic sampling offset based on the query vector is calculated, the reference point position is adjusted according to the dynamic sampling offset to obtain the deformed sampling point, and the deformed point feature value is extracted at the deformed sampling point position through bilinear interpolation. The deformed point feature value is weightedly fused in combination with the multi-head attention mechanism to output the DAT enhanced feature map.
2. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 1 is characterized by: The CDDA-YOLOv8 model is an improved detection model based on the YOLOv8 architecture. Its improvements include: replacing the original C2f module with a cross-stage partial convolution CSPPC module, integrating a deformable attention DAT module into the backbone network, replacing the original feature pyramid network with a dynamic adaptive spatial feature fusion D-ASFF network, and adding a P2 small object detection head.
3. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 1 is characterized by: The dynamic sampling offset based on the query vector is calculated based on the following logic: the CSPPC feature map is transformed by Wq to obtain the corresponding query vector. The offset network takes the query vector at each reference point as input, learns the relationship between the query vector and the offset vector, and obtains the dynamic sampling offset based on the query vector. According to the dynamic sampling offset calculated by the offset network, the initial reference point is offset and moved to the new position to obtain the deformed sampling point; For each deformed sampling point obtained, bilinear interpolation is performed to obtain a set of eigenvalues corresponding to the deformed sampling point on the feature map. The eigenvalues of the deformed points are weightedly fused in combination with the multi-head attention mechanism to output the DAT enhanced feature map. Among them, the offset network consists of two convolution modules, which capture local features through the deep convolution layer, pass the GELU activation function, and then obtain the final dynamic sampling offset through 1×1 convolution; the Wq transformation is implemented through the 1×1 convolution layer, and its convolution kernel weight matrix Wq is learned during training. For the input feature map with a dimension of H×W×C, it is first converted into a (HW)×C matrix through a flattening operation, and then multiplied by the Wq matrix, and finally reconstructed into H×W×d k The query vector, H, W and C are the height, width and number of channels of the input feature map respectively, d k is the query vector dimension.
4. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 1 is characterized by: The DAT enhanced feature map is input into the D-ASFF network for multi-scale feature fusion. The specific logic is as follows: multi-scale context information is extracted through the expandable residual DWR module, and then an adaptive weight map is generated based on the adaptive spatial feature fusion ASFF network. The features of different levels are weightedly fused and the D-ASFF fused feature map is output. The DWR expandable residual module includes a regional residual branch and a semantic residual branch. The output of the DWR module is connected to the adaptive weight map generation module of the ASFF network.
5. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 4 is characterized by: The regional residual branch refers to the use of dilated convolutions with dilation rates of 1, 3, and 5 to extract local fine-grained features in parallel; the semantic residual branch refers to the use of point-by-point convolution to fuse multi-scale features and retain global semantic information.
6. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 1, characterized in that: The feature map clarity index is calculated by performing a Sobel operator gradient calculation on the feature map of each downsampling level in the D-ASFF network, and taking the standard deviation of the gradient amplitude of all pixels as the feature map clarity index. The formula is as follows: Where SMS is the feature map clarity index, N is the total number of pixels in the feature map of each downsampling level in the D-ASFF network, i is the index of the pixel point, μ G is the mean value of the gradient amplitude of all pixels in the feature map, G i Represents the gradient amplitude of the i-th pixel in the feature map, which is calculated by the Sobel operator: Where, I x and I y Represents the gradient map of the feature map in the x and y directions respectively; Where, is the horizontal gradient kernel, is the vertical gradient kernel, T represents the transpose operation of the matrix; The method for calculating the attention focus index is: counting the attention weight distribution entropy of all attention heads in the DAT module, and taking the inverse characteristic number of the entropy value as the attention focus index. The formula is as follows: Where AFM is the attention focus index, M is the total number of attention heads in the DAT module, j is the index of the attention head, and H j is the weight distribution entropy of the j-th attention head: Where p j,k is the normalized attention weight of the kth query vector in the jth attention head, K represents the total number of query vectors in the attention head, where the query vector refers to the coordinates of the deformed sampling point generated by dynamic offset in the DAT module; The feature consistency index is calculated by calculating the mutual information of two adjacent scale feature maps in the D-ASFF network and taking the average mutual information of all scale pairs as the feature consistency index. The formula is as follows: Where FCM is the feature consistency index, L represents the total number of feature map levels in the D-ASFF network, and MI(F l , F l+1 ) is the adjacent scale feature map F l and F l+1 The mutual information of , l is the index of the feature map level: Where p(h,v) represents the joint probability distribution, estimated by the eigenvalue histogram, h and v represent the eigenvalues of the feature maps of adjacent levels in the D-ASFF network at the spatially aligned positions, and p(h) and p(v) are the marginal probability distributions.
7. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 6 is characterized by: The quality evaluation index of the image to be detected is determined based on the feature map clarity index, attention focus index, and feature consistency index. The formula is as follows: Where QI is the quality assessment index, SMS is the feature map clarity index, AFM is the attention focus index, FCM is the feature consistency index, ω1, ω2 and ω3 are preset weight values, ω1>ω2>ω3>0, and ω1+ω2+ω3=1 is satisfied.
8. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 1 is characterized by: The P2 detection head generates a high-resolution feature map and performs cross-scale fusion with the D-ASFF fusion feature map. The specific logic is as follows: The 4x downsampled feature map output by the backbone network is processed with 3×3 convolution and 1×1 convolution in sequence to generate a high-resolution feature map with a resolution of 160×160×C. At the same time, the D-ASFF fusion feature map is bilinearly upsampled by 2 times to obtain an upsampled feature map of 160×160×C. The high-resolution feature map and the upsampled feature map are spliced along the channel dimension to form a fused feature map of 160×160×2C. Finally, the number of channels is compressed to C through 1×1 convolution, and the cross-scale fused feature map is output, where C is the number of feature map channels. The backbone network is the improved feature extraction backbone in the CDDA-YOLOv8 model, and its Stage 2 outputs a 4x downsampled feature map.
9. The method for detecting multi-scale features of patterns based on light convolution and deformation attention according to claim 1, characterized in that: The method for determining the quality assessment threshold is as follows: a diverse dataset of decorative image data is collected as samples, the quality assessment index of each image is calculated using the same method, the 75% quantile of the sample quality assessment index is selected as the demarcation threshold between "clear" and "blurred", i.e., the quality assessment threshold. When the quality assessment index of the image to be detected is lower than the preset quality assessment threshold, a prompt message is synchronously outputted suggesting that the image be recaptured.
Citation Information
Patent Citations
Yao-nationality pattern symbol recognition method based on target detection
CN111931792A
Traditional pattern segmentation method based on multispectral fusion strategy
CN118072006A