Stamp target detection method based on channel-space attention

By introducing the channel-space attention layer module and multi-scale features into seal detection, the problems of poor seal detection effect and low anti-interference in the prior art are solved, and higher detection accuracy and anti-interference are achieved, and efficient needs of modern business are met.

CN119942508APending Publication Date: 2025-05-06GUANGZHOU POWER ELECTRICAL TECH CO LTD

Patent Information

Application Number
CN202510021247.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems in seal detection with poor detection effect, low anti-interference and accuracy, especially when dealing with small targets and diversified seals.

Method used

The seal object detection method based on channel-space attention is adopted. By introducing the attention layer module, the deep block convolution features are extracted based on channel and spatial dimensions, and combined with multi-scale features to enhance detection capabilities.

Benefits of technology

It improves the detection accuracy and anti-interference of small targets, and can focus more accurately on seal details, meeting the efficient needs of modern business for document review and archiving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942508A_ABST
    Figure CN119942508A_ABST
Patent Text Reader

Abstract

The invention discloses a seal target detection method based on channel-space attention, and the method comprises a deployment end and a server end, the deployment end is a deployment reasoning stage, and the server end is a seal detection optimization training stage. According to the method, multi-scale information and attention features are combined, attention information of a channel and a space dimension is extracted by introducing an attention information module based on the channel dimension and the space dimension, deep attention information and shallow detection features are integrated by introducing a cross-scale fusion module, and target detection is achieved; in order to increase the diversity and generalization performance of the model, a further center random rotation and target background random mask strategy is made for a seal target, so that the diversity of a seal sample is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a seal target detection method, in particular to a seal target detection method based on channel-spatial attention, and belongs to the technical field of seal detection. Background Art

[0002] With the rapid development of the economy, the cooperation between state units, enterprises and individuals has become increasingly close. As a legally binding certificate, seals play an irreplaceable role in contract documents in modern society. Therefore, the review of seals in documents is particularly important. However, traditional seal review mostly relies on manual methods, which is time-consuming, costly and inefficient, and cannot meet the current demand for efficient and convenient economic activities.

[0003] In order to improve the efficiency of seal review in the approval of power grid construction archives of power supply bureaus, some researchers have tried to design a series of image processing algorithms (such as color, shape, and edge-based detection operators) based on the color and shape of seals to achieve automatic seal recognition. However, due to the diversity of seals in color and shape, or the partial obstruction of text, these methods usually have problems such as poor detection effect, low anti-interference and low accuracy.

[0004] In order to meet the needs of efficient and convenient business and realize the rapid review and archiving of documents, researchers have designed a series of image processing operators based on color, shape, edge operators, etc. for seal extraction and detection based on the color, shape and other characteristics of the seal. However, due to the influence of many factors such as the color and shape of the seal, the occlusion of the text, etc., this method generally has poor detection effect, low anti-interference ability and low accuracy.

[0005] Traditional manual review has high labor costs, long review time and low work efficiency. The current manual seal detection is no longer sufficient to meet the needs of intelligence and automation. With the advancement of deep learning, feature extraction technology based on deep learning has been widely used in the field of target detection, with good generalization ability and anti-interference. The current target detection algorithms are divided into two categories: two-stage algorithms (such as RCNN series) and one-stage algorithms (such as SSD, YOLO, etc.). The two-stage algorithm has high detection accuracy and is particularly suitable for small target detection, but the detection process is complex and time-consuming; the one-stage algorithm has high detection efficiency. In the prior art, a lightweight seal target detection method based on YOLOv5 is disclosed in the announcement number CN117576373A, including: obtaining electronic bidding documents and establishing a seal image data set of electronic bidding documents; labeling eight shapes of seal image samples using labelImg software, and performing data enhancement and clustering preprocessing on the seal image samples; dividing the seal image data set after data enhancement into a training set and a verification set in an 8:2 ratio; improving the seal recognition model based on YOLOv5; using the labeled seal image data set to debug and optimize the improved YOLOv5 seal recognition model to obtain the optimal detection model. The existing detection method based on deep learning performs well for large targets and low interference, but the accuracy is relatively low when detecting small targets, which poses a challenge to the detection of small targets such as seals. Deep generalization features are obtained through multi-level convolutional neural networks. However, stacked deep features easily ignore details and information about small targets. In order to further improve the detection of targets, this paper proposes a seal target detection method based on channel-spatial attention. Summary of the invention

[0006] The purpose of the present invention is to provide a seal target detection method based on channel-spatial attention in order to solve at least one of the above technical problems.

[0007] The present invention achieves the above-mentioned purpose through the following technical scheme: a seal target detection method based on channel-spatial attention, including a deployment end and a server end, the deployment end is a deployment reasoning stage, the deployment reasoning stage inputs the document and performs seal detection, when encountering an abnormal situation, the abnormal page number and information are fed back to the staff in time; the staff checks and confirms whether it is detected, and if it is detected that there is no seal or there is a missing, it is required to add the seal; if it is confirmed that the seal is not detected, the data is sent to the server end; the server end is a seal detection optimization training stage; in the seal detection optimization training stage, the received data is annotated, and then sent to data processing and augmentation, and then model training and tuning are performed, and finally the model with a good model is evaluated and updated and deployed;

[0008] The seal target detection method includes the following steps:

[0009] S1. Use the normal feature multi-scale model and introduce the attention layer module to extract the attention features of the deep block convolution features based on the channel and spatial dimensions.

[0010] S2. The extracted block convolution features are fused with the detection features and the extracted channel and spatial attention features and then input into the feature detection. The target is detected by combining anchor and non-maximum suppression (NMS).

[0011] As a further solution of the present invention: the attention layer module adopts the self-attention feature method to obtain the attention features of the feature channel and space. The structure of the attention layer module can be divided into five parts: data preprocessing, multi-scale backbone network, attention layer, fusion module and target optimization.

[0012] As a further solution of the present invention: data preprocessing specifically includes:

[0013] In addition to conventional enhancement methods, data enhancement also introduces strategies based on center angle rotation and target background random masking strategies. These strategies are designed to address the situation where seals are not always straight and may be rotated to varying degrees, and seals may be partially damaged due to the influence of ink and stamping force.

[0014] Data sampling: The collected normal equipment operation sound signal is clipped with a fixed time length t and sampled with a fixed frequency M to obtain time domain sound signal data of a fixed length.

[0015] As a further solution of the present invention: the strategy based on the center angle rotation means that during the training process, the target area of ​​the training sample will be rotated by 12 angles, and an image will be obtained every 30° on average to obtain diversified rotation data;

[0016] The target background random masking strategy means that in the process of data processing, some words or background images will be randomly generated to overlap with the seal area, the overlapping area will be calculated, and the alignment will be whitened to obtain a similar incomplete seal; the degree of seal incompleteness is controlled below 20%; if the seal incompleteness is high, then the seal should be judged invalid.

[0017] As a further solution of the present invention: the multi-scale backbone network uses stacked convolutional layers to extract image features, thereby extracting feature information of different scales; the multi-scale pyramid features are beneficial to target retrieval and extraction of abstract features of object targets in target detection tasks, and the backbone network uses VGG19 or ResNet34 trained on Imagenet to extract features.

[0018] As a further solution of the present invention: the attention layer is divided into two parts: spatial attention feature extraction and channel attention feature extraction, where:

[0019] Spatial attention feature extraction specifically includes:

[0020] The features to be detected extracted from the backbone network are processed by global average pooling to obtain channel-based feature information. The information conversion is shown in the following formula:

[0021]

[0022] in represents the channel characteristics, Q (n,w,h) represents the input feature to be detected, n represents the feature channel, w represents the feature width, h represents the feature height; GA represents the global average operation;

[0023] The feature to be detected Q (n,w,h) After integrating the last two dimensions, the dimension becomes Q (n,w×h) , after transposing, we get Q (w ×h,n) Features, dimensionally integrate channel features into Q (w×h,n) and The features are matrix multiplied and then passed through softmax to obtain feature K' s ∈R (wxh,1) , and then K' s The feature is split into dimensions to obtain the spatial attention feature K' s ∈R (1,w,h) ; Align the spatial attention features with the feature information dimension of the previous layer by upsampling K' s ∈R (1,w',h') ; The specific information conversion formula is shown as follows:

[0024]

[0025] Where Q (n,w,h) Represents the feature to be detected; K c (n,1,1) represents channel features; rs represents dimension conversion operation; d represents the number of elements, coefficient factor; up represents upsampling operation, which is used for spatial feature alignment; K' s Represents the spatial attention feature, whose dimension is aligned with the spatial dimension of the feature information of the previous layer;

[0026] Channel attention feature extraction specifically includes:

[0027] The features to be detected extracted from the backbone network are subjected to a 1x1 convolutional layer operation to obtain a feature map with a channel dimension of 1; the information conversion formula of the spatial feature map is as follows:

[0028]

[0029] Among them, Conv1x1 represents the convolution operation with a convolution kernel of 1; Q (n,w,h) It represents the features to be detected. It represents the spatial characteristics;

[0030] The feature to be detected Q (n,w,h) The last two dimensions are integrated, and the dimension becomes Q (n,w×h) ; After transposing, we get Q (w ×h,n) Features, dimensionally integrate channel features into Through K s Q T After matrix multiplication, softmax is used to obtain K' c ∈R (n,1,1) ; Then, the 1x1 convolution kernel is used to align the channel dimension to obtain the feature K' c ∈R (m,1,1) , the specific formula is shown as follows:

[0031]

[0032] in represents spatial features; rs represents dimension integration and transformation operations; Q (n,w,h) Represents the feature to be detected; d represents the number of elements, the coefficient factor; Conv1x1 represents a 1x1 convolution kernel, which is used to align the feature dimension; K' c Represents the channel attention feature, whose dimension is aligned with the channel dimension of the feature information of the previous layer.

[0033] As a further solution of the present invention: the spatial attention features and channel attention features obtained by the fusion module through the attention layer will be used to fuse the features extracted by the backbone network in the previous layer; the features extracted by the backbone network are respectively multiplied element by element with the spatial attention features and the channel attention features to obtain the features to be detected; the specific expression is shown in the following formula:

[0034]

[0035] Where Q (m,w',h') Represents the backbone network characteristics; Represents spatial features; represents channel characteristics; d s ,d cRepresents space, channel characteristic factor coefficient; Q' (m,w',h') Represents the feature to be detected; × represents element-by-element multiplication.

[0036] As a further solution of the present invention: the target optimization uses L1 loss and cross entropy loss to optimize the loss; in the target training optimization stage of the model, the detection result output by the model is compared with the true label, wherein the classification loss is calculated by the cross entropy loss; the positioning loss is calculated by the L1 loss, and the positioning loss includes the center point coordinates and the width and height (x, y, w, h); the overall loss calculation is shown in the following formula:

[0037] L=L cls +λL loc

[0038] Where L represents the total loss; L cls represents the classification loss; L loc represents the positioning loss; λ represents the coefficient factor, which is generally taken as 1.0; the training goal is to make min(L)→0.

[0039] The beneficial effects of the present invention are as follows: the present invention combines multi-scale information and attention features, introduces attention modules of channel and space dimensions on the basis of the traditional SSD detection algorithm to obtain richer spatial and channel features, and then combines multi-scale feature fusion to enhance the detection capability. This method can focus on the seal details more accurately, improve the detection accuracy and anti-interference of small targets, and is expected to meet the efficient needs of modern business for document review and archiving;

[0040] The present invention introduces an attention information module based on channel dimension and space dimension to extract attention information of channel and space dimension, and integrates deep attention information and shallow detection features by introducing a cross-scale fusion module to achieve the integration of multi-scale information. On the other hand, in order to increase the diversity and generalization performance of the model, a further center random rotation and target background random masking strategy are performed for the seal target to enhance the diversity of the seal samples.

[0041] In addition, due to the diversity and expansion of seal samples, the data set cannot cover all seal samples. The present invention designs an online model optimization training deployment scheme to perform real-time training and model optimization for seal samples to further improve the accuracy and robustness of model seal detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic diagram of the process flow of the online model training optimization and deployment reasoning solution of the present invention;

[0043] Figure 2 This is a schematic diagram of the detection process based on the channel-spatial attention network of the present invention;

[0044] Figure 3 Schematic diagram of the attention layer processing structure of the present invention;

[0045] Figure 4 This is a schematic diagram of the fusion module structure of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Embodiment 1

[0048] like Figure 1 As shown, a seal target detection method based on channel-spatial attention includes a deployment end and a server end. The deployment end is a deployment reasoning stage. In the deployment reasoning stage, the document is input and the seal detection is performed. When an abnormal situation is encountered, the abnormal page number and information are promptly fed back to the staff; the staff checks and confirms whether it is detected. If it is detected that there is no seal or there is a missing seal, it is required to add the seal; if it is confirmed that the seal is not detected, the data is sent to the server end; the server end is a seal detection optimization training stage; in the seal detection optimization training stage, the received data is annotated, and then sent to data processing and augmentation, and then model training and tuning are performed, and finally the model with a good model is evaluated and updated and deployed;

[0049] The seal target detection method includes the following steps:

[0050] S1. Use the normal feature multi-scale model and introduce the attention layer module to extract the attention features of the deep block convolution features based on the channel and spatial dimensions.

[0051] S2. The extracted block convolution features are fused with the detection features and the extracted channel and spatial attention features and then input into the feature detection. The target is detected by combining anchor and non-maximum suppression (NMS).

[0052] Embodiment 2

[0053] like Figures 2 to 3 As shown, in addition to all the technical features of the first embodiment, this embodiment also includes:

[0054] The attention layer module uses self-attention features to obtain the attention features of feature channels and space. The structure of the attention layer module can be divided into five parts: data preprocessing, multi-scale backbone network, attention layer, fusion module and target optimization.

[0055] Furthermore, data preprocessing specifically includes:

[0056] Data enhancement: In addition to conventional enhancement methods, we also introduced strategies based on center angle rotation and target background random masking. These strategies are designed to enhance the detection of seal defects, as seals are not always properly stamped and may be rotated to varying degrees, or partially damaged due to the influence of ink and stamping force. This also reduces the cost of manual labeling.

[0057] Data sampling: The collected normal equipment operation sound signal is clipped with a fixed time length t and sampled with a fixed frequency M to obtain time domain sound signal data of a fixed length.

[0058] Furthermore, the strategy based on center angle rotation means that during the training process, the target area of ​​the training sample will be rotated at 12 angles, and an image will be acquired every 30° on average to obtain diversified rotation data;

[0059] The target background random masking strategy means that in the process of data processing, some words or background images will be randomly generated to overlap with the seal area, the overlapping area will be calculated, and the alignment will be whitened to obtain a similar incomplete seal; the degree of seal incompleteness is controlled below 20%; if the seal incompleteness is high, then the seal should be judged invalid.

[0060] Furthermore, the multi-scale backbone network uses stacked convolutional layers to extract image features, thereby extracting feature information of different scales; common backbone networks include VGG, ResNet, GoogLeNet, and even VIT, Swin, etc. For target detection, multi-scale pyramid features are conducive to target retrieval and abstract feature extraction of object targets in target detection tasks. The backbone network uses VGG19 or ResNet34 trained on Imagenet for feature extraction.

[0061] Furthermore, the attention layer is divided into two parts: spatial attention feature extraction and channel attention feature extraction, where:

[0062] Spatial attention feature extraction specifically includes:

[0063] The features to be detected extracted from the backbone network are processed by global average pooling to obtain channel-based feature information. The information conversion is shown in the following formula:

[0064]

[0065] in represents the channel characteristics, Q (n,w,h) represents the input feature to be detected, n represents the feature channel, w represents the feature width, h represents the feature height; GA represents the global average operation;

[0066] The feature to be detected Q (n,w,h) After integrating the last two dimensions, the dimension becomes Q (n,w×h) , after transposing, we get Q (w ×h,n) Features, dimensionally integrate channel features into Q (w×h,n) and The features are matrix multiplied and then passed through softmax to obtain feature K' s ∈R (wxh,1) , and then K' s The feature is split into dimensions to obtain the spatial attention feature K' s ∈R (1,w,h) ; Align the spatial attention features with the feature information dimension of the previous layer by upsampling K' s ∈R (1,w',h') ; The specific information conversion formula is shown as follows:

[0067]

[0068] Where Q (n,w,h) Represents the feature to be detected; K c (n,1,1) represents channel features; rs represents dimension conversion operation; d represents the number of elements, coefficient factor; up represents upsampling operation, which is used for spatial feature alignment; K' s Represents the spatial attention feature, whose dimension is aligned with the spatial dimension of the feature information of the previous layer;

[0069] Channel attention feature extraction specifically includes:

[0070] The features to be detected extracted from the backbone network are subjected to a 1x1 convolutional layer operation to obtain a feature map with a channel dimension of 1; the information conversion formula of the spatial feature map is as follows:

[0071]

[0072] Among them, Conv1x1 represents the convolution operation with a convolution kernel of 1; Q (n,w,h)It represents the features to be detected. It represents the spatial characteristics;

[0073] The feature to be detected Q (n,w,h) The last two dimensions are integrated, and the dimension becomes Q (n,w×h) ; After transposing, we get Q (w ×h,n) Features, dimensionally integrate channel features into Through K s Q T After matrix multiplication, softmax is used to obtain K' c ∈R (n,1,1) ; Then, the 1x1 convolution kernel is used to align the channel dimension to obtain the feature K' c ∈R (m,1,1) , the specific formula is shown as follows:

[0074]

[0075] in represents spatial features; rs represents dimension integration and transformation operations; Q (n,w,h) Represents the feature to be detected; d represents the number of elements, the coefficient factor; Conv1x1 represents a 1x1 convolution kernel, which is used to align the feature dimension; K' c Represents the channel attention feature, whose dimension is aligned with the channel dimension of the feature information of the previous layer.

[0076] Furthermore, the spatial attention features and channel attention features obtained by the fusion module through the attention layer will be used to fuse the features extracted by the previous backbone network; the features extracted by the backbone network are multiplied element by element with the spatial attention features and the channel attention features to obtain the features to be detected; the specific expression is shown in the following formula:

[0077]

[0078] Where Q (m,w',h') Represents the backbone network characteristics; Represents spatial features; represents channel characteristics; d s ,d c Represents space, channel characteristic factor coefficient; Q' (m,w',h') Represents the feature to be detected; × represents element-by-element multiplication.

[0079] Furthermore, the target optimization uses L1 loss and cross entropy loss to optimize the loss; in the target training optimization stage of the model, the detection results output by the model are compared with the true label, where the classification loss is calculated by the cross entropy loss; the positioning loss is calculated by the L1 loss, and the positioning loss includes the center point coordinates and the width and height (x, y, w, h); the overall loss calculation is shown in the following formula:

[0080] L=L cls +λL loc

[0081] Where L represents the total loss; L cls represents the classification loss; L loc represents the positioning loss; λ represents the coefficient factor, which is generally taken as 1.0; the training goal is to make min(L)→0.

[0082] Working principle: It combines multi-scale information and attention features, extracts attention information in channel and spatial dimensions by introducing attention information modules based on channel and spatial dimensions, and integrates deep attention information and shallow detection features by introducing cross-scale fusion modules to achieve target detection. On the other hand, in order to increase the diversity and generalization performance of the model, further center random rotation and target background random masking strategies are performed on the seal target to enhance the diversity of seal samples.

[0083] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0084] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A seal target detection method based on channel-spatial attention, including a deployment end and a server end, characterized in that: The deployment end is the deployment reasoning stage, in which the document is input and the seal is detected. When an abnormal situation is encountered, the abnormal page number and information are fed back to the staff in a timely manner; the staff checks and confirms whether to detect it. If it is detected that there is no seal or there are omissions, it is required to add the seal; If it is confirmed that the seal is not detected, the data will be sent to the server; The server side is the seal detection optimization training stage; in the seal detection optimization training stage, the received data is annotated, and then sent to data processing and augmentation, and then model training and tuning are carried out, and finally the good model is evaluated and updated and deployed; The seal target detection method comprises the following steps: S1. Use the normal feature multi-scale model and introduce the attention layer module to extract the attention features of the deep block convolution features based on the channel and spatial dimensions. S2. The extracted block convolution features are fused with the detection features and the extracted channel and spatial attention features and then input into the feature detection. The target is detected by combining anchor and non-maximum suppression (NMS).

2. The seal target detection method according to claim 1, characterized in that: In S1, the attention layer module adopts the self-attention feature method to obtain the attention features of the feature channel and space. The structure of the attention layer module can be divided into five parts: data preprocessing, multi-scale backbone network, attention layer, fusion module and target optimization.

3. The seal target detection method according to claim 2, characterized in that: The data preprocessing specifically includes: In addition to conventional enhancement methods, data enhancement also introduces strategies based on center angle rotation and target background random masking strategies. These strategies are designed to address the situation where seals are not always straight and may be rotated to varying degrees, and seals may be partially damaged due to the influence of ink and stamping force. Data sampling: The collected normal equipment operation sound signal is clipped with a fixed time length t and sampled with a fixed frequency M to obtain time domain sound signal data of a fixed length.

4. The seal target detection method according to claim 3, characterized in that: The strategy based on center angle rotation means that during the training process, the target area of ​​the training sample will be rotated by 12 angles, and an image will be acquired every 30° on average to obtain diversified rotation data; The target background random mask strategy means that in the process of data processing, some words or background images will be randomly generated to overlap with the seal area, the overlapping area will be calculated, and the alignment will be whitened to obtain a similar incomplete seal; the degree of incompleteness of the seal is controlled below 20%; If the seal is highly damaged, it should be deemed invalid.

5. The seal target detection method according to claim 2, characterized in that: The multi-scale backbone network uses stacked convolutional layers to extract image features, thereby extracting feature information of different scales; the multi-scale pyramid features are conducive to target retrieval and extraction of abstract features of object targets in target detection tasks. The backbone network uses VGG19 or ResNet34 trained on Imagenet to extract features.

6. The seal target detection method according to claim 2, characterized in that: The attention layer is divided into two parts: spatial attention feature extraction and channel attention feature extraction, where: Spatial attention feature extraction specifically includes: The features to be detected extracted from the backbone network are processed by global average pooling to obtain channel-based feature information. The information conversion is shown in the following formula: in, represents the channel characteristics, Q (n,w,h) represents the input feature to be detected, n represents the feature channel, w represents the feature width, h represents the feature height; GA represents the global average operation; The feature to be detected Q (n,w,h) After integrating the last two dimensions, the dimension becomes Q (n,w×h) , after transposing, we get Q (w×h,n) Features, dimensionally integrate channel features into Q (w×h,n) and The features are matrix multiplied and then passed through softmax to obtain feature K' s ∈R (wxh,1) , and then K' s The feature is split into dimensions to obtain the spatial attention feature K' s ∈R (1 ,w,h) ; Align the spatial attention features with the feature information dimension of the previous layer by upsampling K' s ∈R (1,w',h') ; The specific information conversion formula is shown as follows: Where Q (n,w,h) Indicates the feature to be detected; represents channel features; rs represents dimension conversion operation; d represents the number of elements, coefficient factor; up represents upsampling operation, which is used for spatial feature alignment; K' s Represents the spatial attention feature, whose dimension is aligned with the spatial dimension of the feature information of the previous layer; Channel attention feature extraction specifically includes: The features to be detected extracted from the backbone network are subjected to a 1x1 convolutional layer operation to obtain a feature map with a channel dimension of 1; the information conversion formula of the spatial feature map is as follows: Among them, Conv1x1 represents the convolution operation with a convolution kernel of 1; Q (n,w,h) It represents the features to be detected. It represents the spatial characteristics; The feature to be detected Q (n,w,h) The last two dimensions are integrated, and the dimension becomes Q (n,w×h) ; After transposing, we get Q (w×h,n) Features, dimensionally integrate channel features into Through K s Q T After matrix multiplication, softmax is used to obtain K' c ∈R (n,1,1) ; Then, the 1x1 convolution kernel is used to align the channel dimension to obtain the feature K' c ∈R (m,1,1) , the specific formula is shown as follows: in represents spatial features; rs represents dimension integration and transformation operations; Q (n,w,h) Represents the feature to be detected; d represents the number of elements, the coefficient factor; Conv1x1 represents a 1x1 convolution kernel, which is used to align the feature dimension; K' c Represents the channel attention feature, whose dimension is aligned with the channel dimension of the feature information of the previous layer.

7. The seal target detection method according to claim 2, characterized in that: The spatial attention features and channel attention features obtained by the fusion module through the attention layer will be used to fuse the features extracted by the backbone network in the previous layer; the features extracted by the backbone network are respectively multiplied element by element with the spatial attention features and the channel attention features to obtain the features to be detected; The specific expression is shown as follows: Where Q (m,w',h') Represents the backbone network characteristics; K' s (1,w',h') Represents spatial features; Represents channel characteristics; d s ,d c Represents space, channel characteristic factor coefficient; Q' (m,w',h') Indicates the feature to be detected; × means element-wise multiplication.

8. The seal target detection method according to claim 2, characterized in that: The target optimization uses L1 loss and cross entropy loss to optimize the loss; in the target training optimization stage of the model, the detection results output by the model are compared with the true label, where the classification loss is calculated by the cross entropy loss; the positioning loss is calculated by the L1 loss, and the positioning loss includes the center point coordinates and the width and height (x, y, w, h); the overall loss calculation is shown in the following formula: L=L cls +λL loc Where L represents the total loss; L cls represents the classification loss; L loc represents the positioning loss; λ represents the coefficient factor, which is 1.0; the training goal is to make min(L)→0.

Citation Information

Patent Citations

  • Lightweight seal target detection method based on YOLOv5

    CN117576373A

Cited By

  • Intelligent seal identification method and device, electronic equipment and storage medium

    CN121330689A