Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation

By introducing multi-stage dynamic pruning and feature compensation technology into the object detection model, combined with the advantages of Swin Transformer and YOLOv9, the training and inference efficiency of the object detection model in resource-constrained environments is solved, and efficient and accurate object detection is achieved.

CN120164077APending Publication Date: 2025-06-17XIDIAN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510232075.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the field of object detection, it is difficult for the prior art to accelerate model training and inference and reduce model parameters while ensuring detection performance, especially in resource-constrained environments.

Method used

The Swin-YOLOv9 object detection method based on multi-stage dynamic pruning and feature compensation is adopted, and the long-distance dependency modeling is used to model long-distance dependency. Through the multi-stage dynamic token screening mechanism and feature compensation mechanism, meaningful tokens are dynamically screened and feature compensation is performed to reduce redundant feature information, and the model is accelerated and lightweighted.

Benefits of technology

It improves the detection accuracy and robustness of the model in multi-objective and complex contexts, reduces the dependence on large-scale computing resources, realizes deployment and application under a wider range of hardware conditions, and significantly reduces the parameter scale and computing volume during inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164077A_ABST
    Figure CN120164077A_ABST
Patent Text Reader

Abstract

The invention discloses a Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation, and the method is characterized in that the method comprises the following steps: 1, carrying out the multi-stage dynamic pruning and feature compensation; the method comprises the following steps: 1, preprocessing an input target image; 2, designing a target detection network, wherein the target detection network is used for carrying out feature extraction on the preprocessed target image; 3, training the extracted features based on a multi-stage dynamic Token screening mechanism and a feature compensation mechanism of an attention mechanism, dynamically screening meaningful tokens, and performing feature compensation; and step 4, performing token structured pruning based on an attention mechanism on the token after feature compensation to realize target detection. On the premise of keeping or even improving the target detection performance, the pruning strategy is optimized, the redundant parameters and the calculation amount of the model are reduced, and therefore the training efficiency is improved, and the requirement for calculation resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and in particular relates to a Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation. Background Art

[0002] Object detection aims to identify and locate specific object categories from images or videos, and is the basis for many applications such as autonomous driving, intelligent monitoring, and robot navigation. In recent years, the introduction of deep learning models, especially convolutional neural networks (CNNs) and attention mechanisms, has greatly improved the performance of object detection.

[0003] The YOLO series of models (You Only Look Once) have attracted widespread attention for their end-to-end detection methods and real-time detection capabilities. The latest version, YOLOv9, further improves the accuracy and speed of detection. However, as the depth and complexity of the model increase, the computational cost of its training and reasoning also increases significantly. On the other hand, Swin Transformer, as a visual model based on the Transformer architecture, has been successfully applied to a variety of visual tasks by introducing a sliding window mechanism, demonstrating excellent feature extraction capabilities.

[0004] In practical applications, the lightweight and efficient models are becoming increasingly important, especially in resource-constrained environments such as mobile devices and embedded systems.

[0005] Therefore, how to accelerate model training and reasoning and reduce the number of model parameters while ensuring detection performance has become an urgent problem to be solved.

[0006] At present, in the field of target detection, research on model acceleration and lightweighting mainly focuses on the following aspects:

[0007] The first method is to design a lightweight network architecture. For example, MobileNet and ShuffleNet reduce the number of model parameters and computation by introducing technologies such as deep separable convolution and channel shuffling. These models reduce performance to a certain extent, but still cannot meet the needs of high-precision detection.

[0008] The second method is model pruning technology. Pruning reduces the number of model parameters and computational complexity by removing redundant network connections or neurons. Traditional pruning methods are usually based on the size or sparsity of weights, such as L1 norm pruning and L2 norm pruning. However, these methods may lead to a significant decrease in model accuracy and require tedious pruning and fine-tuning processes.

[0009] The third method is knowledge distillation. By enabling a lightweight student model to learn the output or intermediate features of a large teacher model, the purpose of improving performance is achieved. However, the knowledge distillation process is complex, the training time is long, and the effect depends on the quality of the teacher model.

[0010] In addition, although the introduction of Transformer into the field of object detection has improved the feature representation ability, the high computational complexity of Transformer itself also brings challenges. Although there are studies attempting to combine Swin Transformer with YOLO series models, due to the large model size, the training and inference speeds are slow, making it difficult to be deployed in real-time applications.

[0011] In summary, there are still many deficiencies in the existing technical solutions in terms of model lightweighting and acceleration. Traditional pruning methods may result in performance loss, lightweight models are difficult to balance accuracy and efficiency, and the introduction of Swin Transformer increases the computational cost. Summary of the Invention

[0012] In order to overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a Swin-YOLOv9 object detection method based on multi-stage dynamic pruning and feature compensation, which utilizes the characteristics of the Transformer attention mechanism to achieve long-range dependency modeling and improve the detection accuracy and robustness of the model in multi-object and complex background scenarios.

[0013] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0014] A Swin-YOLOv9 object detection method based on multi-stage dynamic pruning and feature compensation, comprising the following steps;

[0015] Step 1: Preprocess the input image to provide a suitable input format for subsequent network processing;

[0016] Step 2: Use the object detection network to extract features from the preprocessed target image;

[0017] Step 3: Adopt a multi-stage dynamic Token screening mechanism and a feature compensation mechanism based on the attention mechanism to train the extracted features, dynamically screen meaningful tokens and perform feature compensation;

[0018] Step 4: Perform attention mechanism-based token structured pruning on the tokens after feature compensation, and use the pruned features for object detection.

[0019] The specific content of Step 1 is:

[0020] (1a) Scale the input target image to the fixed size required by the object detection network to ensure that all image input sizes are consistent. Scale it while maintaining the aspect ratio and then perform padding to achieve this.

[0021] (1b) Normalize the input image to help the model converge faster.

[0022] (1c) Adjust the normalized input image to a fixed size by means of central cropping so that it can be trained with an appropriate batch size.

[0023] (1d) Perform data augmentation on the adjusted image, including operations such as rotation, flipping, scaling, translation, and color change, which can generate more training samples and thus help the model better generalize to new unseen data.

[0024] The said step 2

[0025] The object detection network is divided into a backbone (feature extractor) and a head (detection head), where:

[0026] (i) The backbone (feature extractor) consists of four feature extraction stages of Swin Transformer and a multi-stage dynamic Token screening mechanism and a feature compensation mechanism based on the attention mechanism. The backbone is used to extract multi-level features from the input image for the head to perform object detection.

[0027] (ii) The object detection head head follows the detection head network architecture of YOLOv9. The head is used to perform object detection in the image according to the feature information provided by the object detection backbone.

[0028] The said Swin Transformer is used for feature extraction of the object detection model.

[0029] The said multi-stage Token screening mechanism based on the attention mechanism is used to reduce the problems of slow training speed caused by excessive feature information in the feature extraction stage of the object detection model and poor training effect caused by redundant feature information; the token feature compensation is used to make up for the information loss problem caused by the initialization importance token and the multi-stage dynamic screening token. Using the token feature compensation can significantly obtain better training accuracy.

[0030] The said step 3 is specifically: respectively based on the multi-stage dynamic Token screening module and the feature compensation module of the attention mechanism;

[0031] The said multi-stage dynamic Token screening mechanism performs initialization importance token and multi-stage dynamic screening token;

[0032] The feature compensation mechanism performs token feature compensation.

[0033] Initialize the importance tokens and the multi-stage dynamic screening tokens to cooperate to complete feature information compression; token feature compensation is used to reduce information loss caused by feature information compression.

[0034] (3.1) Initialize the importance tokens:

[0035] Take the output of the first stage of the Swin Transformer as the input of this stage. The input feature information is the token feature after window partitioning as T = {t1, t2, t3, …, t N}. According to the known window attention score matrix in the first stage of the Swin Transformer, in this stage, a part with a higher sum of global attention scores is used as important tokens. Select the initial important tokens as shown in Equation (1) and Equation (2), where A is the attention score matrix within the window, t j ∈ T, S j is the sum of the attention scores in the j-th column of A, β is the set of tokens represented by the top α% of the higher scores in S, and I is the selected importance tokens.

[0036]

[0037] I = {t j ∣ S j ∈ β} (2)

[0038] (3.2) Multi-stage dynamic screening tokens:

[0039] (i) Establish a new window attention score matrix: Use the attention mechanism between the tokens selected by the initialized importance tokens and the unselected tokens to obtain a new attention score matrix as A;

[0040] (ii) Select the top γ% of the tokens with the lowest cumulative attention impact and re-classify these tokens as important tokens;

[0041] The specific steps of (ii) are as follows: Let U be the set of unimportant tokens and I be the set of important tokens; for each unimportant token u i ∈ U, calculate its cumulative attention impact C i on the important tokens, as shown in Equation (3):

[0042]

[0043] (iii) After sorting the unimportant token sets in C from largest to smallest, the tokens in the top γ% are I new The important tokens selected in the first stage are I old The new important tokens are I, as shown in Equation (4):

[0044] I = I new + I old (4)

[0045] Among them, the strategy of multi-stage dynamic token screening is to perform multiple repeated operations to gradually obtain important tokens;

[0046] (3.3) Token feature compensation:

[0047] (i) Based on the updated important token set I′ and unimportant token set U′, calculate the attention score matrix A ∈ R |U′|×|I′| , as shown in Equation (5):

[0048]

[0049] where u j ∈ U′, i k ∈ I′, is the query vector of the attention mechanism, is the value vector of the attention mechanism, and d is the cosine distance between u j and i k .

[0050] (ii) Based on the new attention score matrix A ∈ R |U′|×|I′| , assign each unimportant token u j ∈ U′ to the important token group G k with the highest attention score, as shown in Equation (6):

[0051]

[0052] where i k ∈ I′, u j ∈ U′.

[0053] (iii) Based on the grouping G, perform feature fusion, as shown in Equation (7):

[0054]

[0055] where is the value vector of the unimportant token u j , is the fused feature vector and serves as the fused information passed to the next layer.

[0056] Step 4 is specifically as follows:

[0057] (i) Initialize importance tokens:

[0058] Take the output of the first stage of the Swin Transformer as the input of this stage, uniformly sample the input tokens as the benchmark of the importance tokens, and the input token sequence: T = {t1, t2, t3,..., t N}; set the sampling parameter α ∈ (0, 1], then select an initial importance token every token, and the set of initial importance tokens is shown in Equation (8):

[0059]

[0060] where represents the set of initialized important tokens, N represents the number of original tokens, and a part of the tokens are selected from the input token sequence as the initial importance tokens through uniform sampling;

[0061] (ii) Calculate attention scores:

[0062] Set the window size to M and the basic window size to W, then the divided window size is: For each non-importance Token (t i ∈ T non-imp ) within each window, calculate the maximum attention score between it and the initial importance tokens, as shown in Equation (9):

[0063]

[0064] (iii) Screen important tokens:

[0065] Sort all non-importance Tokens in ascending order of A i as shown in Equation (10):

[0066] {t (1) , t (2) ,..., t (L)}, A (1) ≤ A (2) ≤ … ≤ A (L) (10)

[0067] Screen the top β% of the tokens as the new importance tokens, as shown in Equation (11):

[0068]

[0069] Combine the initial important tokens with the new important tokens as the output of the overall token structured pruning, as shown in Equation (11):

[0070]

[0071] Take the final result T imp As the input information for the second stage of the Swin Transformer within the backbone in the object detection model.

[0072] Advantages of the present invention:

[0073] 1. Compared with the traditional YOLO series, the present invention solves the problem of difficult to efficiently model long-range dependencies in large-scale scenarios or complex target distributions. The Swin Transformer combines the detection advantages of YOLO and utilizes the characteristics of the Transformer attention mechanism to achieve long-range dependency modeling, improving the detection accuracy and robustness of the model in multi-object and complex background scenarios.

[0074] 2. Compared with complex attention structures such as DETR, the present invention solves the problems of excessive computing power requirements and insufficient inference speed on ordinary computer devices. Through the efficient Swin Transformer structure and targeted network design, the present invention introduces a multi-stage dynamic token screening mechanism based on the attention mechanism, compresses token information, reduces the dependence on large-scale computing resources, and realizes deployment and application under a wider range of hardware conditions; introduces a token feature mechanism to solve the problem of target image information loss during token compression in the object detection model, improving the accuracy of the object detection model.

[0075] 3. The present invention introduces an adaptive token pruning strategy, which can analyze and prune unimportant tokens in real time, significantly reducing the network calculation amount and video memory occupancy. Compared with existing pruning methods based on static or single attention criteria, this strategy can better balance performance and speed.

[0076] 4. The present invention achieves a balance among object detection performance, model scale, and inference efficiency, and is applicable to fields with high requirements for accuracy and real-time performance such as security monitoring, unmanned driving, and industrial quality inspection; it also has good feasibility for implementation in resource-constrained application scenarios (such as embedded and mobile). Therefore, the present invention has higher value and competitiveness in commercial applications with high real-time requirements and diverse hardware resources. Brief Description of the Drawings

[0077] Figure 1 It is a schematic diagram of the training stage process.

[0078] Figure 2 It is a schematic diagram of the process in the verification and testing stage.

[0079] Figure 3 It is a schematic diagram of the target detection network architecture.

[0080] Figure 4 It is a schematic diagram of the multi-stage dynamic Token screening mechanism and feature compensation mechanism based on the attention mechanism.

[0081] Figure 5 It is a schematic diagram of token structured pruning based on the attention mechanism.

[0082] Figure 6 It is a schematic diagram of target detection simulation. Specific implementation manners

[0083] The present invention will be further described in detail below with reference to the accompanying drawings.

[0084] As Figure 1 、 Figure 2 shown, a Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation includes the following steps;

[0085] (1) Preprocessing of the input image:

[0086] (1.1) Scale the input image to the fixed size required by the target detection model to ensure that all image input sizes are consistent. Scale it while maintaining the aspect ratio, and then perform padding (filling) to achieve this.

[0087] (1.2) Normalize the input image to help the model converge faster.

[0088] (1.3) Adjust the input image to a fixed size by means of central cropping so that it can be trained with an appropriate batch size.

[0089] (1.4) Perform data augmentation, including operations such as rotation, flipping, scaling, translation, and color change, which can generate more training samples, thereby helping the model better generalize to new unseen data.

[0090] (2) Overall architecture design of the target detection network:

[0091] As Figure 3 shown, target detection is mainly divided into backone (feature extractor) and head (detection head):

[0092] (i) Backbone design: The backbone (feature extractor) mainly consists of four feature extraction stages of Swin Transformer, a multi-stage dynamic Token screening mechanism and a feature compensation mechanism based on the attention mechanism ( Figure 3 TSM in

[0093] ), and the purpose of the backbone is to extract features for the detection head to detect targets.

[0094] (ii) Head design: The object detection head head follows the detection head network architecture of YOLOv9, and the purpose of the head is to perform object detection in the image based on the feature information provided by the object detection backbone.

[0095] Such as Figure 4 shown, the multi-stage dynamic Token screening mechanism and feature compensation mechanism module based on the attention mechanism is divided into initializing important tokens, multi-stage dynamic screening of tokens, and token feature compensation.

[0096] Initializing important tokens and multi-stage dynamic screening of tokens cooperate to achieve the purpose of feature compression, which can accelerate the training speed of the model; the purpose of token feature compensation is to reduce the information loss problem caused by feature information compression, which can improve the model accuracy.

[0097] (3.1) Initializing important tokens:

[0098] Such as Figure 4 shown, the input feature information is the token feature T = {t1, t2, t3,..., t N} after window partitioning. According to the known intra-window attention score matrix in the first stage of Swin Transformer, the initial important tokens are selected as shown in Equation (1) and Equation (2), where A is the intra-window attention score matrix, t j ∈ T, S j is the sum of the attention scores in the j-th column of A, β is the set of tokens represented by the top α% of the higher scores in S, and I is the selected important token.

[0099]

[0100] I = {t j ∣S j ∈ β} (2)

[0101] (3.2) Multi-stage dynamic screening of tokens:

[0102] (i) Establish a new window attention score matrix: Use the attention mechanism between the tokens selected by the initialized importance token and the unselected tokens to obtain a new attention score matrix A.

[0103] (ii) Select the top γ% of the tokens with the lowest cumulative attention impact and reclassify these tokens as important tokens. Specifically: Let U be the set of unimportant tokens and I be the set of important tokens; for each unimportant token u i ∈U, calculate its cumulative attention impact C i on the important tokens, as shown in Equation (3):

[0104]

[0105] (iii) After sorting the set of unimportant tokens in C from largest to smallest, the tokens in the top γ% are I new , the important tokens selected in the first stage are I old , and the new important tokens are I, as shown in Equation (4):

[0106] I = I new + I old (4)

[0107] Among them, the strategy of multi-stage dynamic screening of tokens can be repeated multiple times to gradually obtain important tokens.

[0108] (3.3) Token feature compensation:

[0109] (i) Based on the updated set of important tokens I′ and the set of unimportant tokens U′, calculate the attention score matrix A ∈ R |U′|×|I′| , as shown in Equation (5):

[0110]

[0111] where u j ∈U′, i k ∈I′, Q is the query vector of the attention mechanism, and K is the value vector of the attention mechanism.

[0112] (ii) Based on the new attention score matrix A ∈ R |U′|×|I′| , assign each unimportant token u j ∈U′ to the group G of important tokens with its maximum attention score k , as shown in Equation (6):

[0113]

[0114] where i k ∈ I′, u j ∈ U′.

[0115] (iii) Feature fusion is performed based on the group G, as shown in Equation (7):

[0116]

[0117] where is the value vector of the unimportant token u j , and is the fused feature vector and serves as the fused information passed to the next layer.

[0118] (4) Token structured pruning based on the attention mechanism: This step belongs to the inference stage. It acts after the first stage of the Swin Transformer in the backone stage. By pruning operations, redundant tokens are removed to reduce the computational load, thereby improving the inference efficiency. The pruning in the inference stage is performed on the model optimized in the training stage. This step can improve the inference speed with a little reduction in accuracy, facilitating deployment on embedded devices with insufficient computing power.

[0119] (i) Initialize the importance tokens:

[0120] As Figure 5 shown, uniformly sample the input tokens as the benchmark for the importance tokens. Given the original Token sequence: T = {t1, t2, t3,..., t N}, set the sampling parameter α ∈ (0, 1]. Then, select an initial importance token every tokens. The initial importance token set is as shown in Equation (8):

[0121]

[0122] where represents the initialized importance token set, and N represents the number of original tokens. By uniform sampling, a part of the tokens are selected from the input token sequence as the initial importance tokens.

[0123] (ii) Calculate the attention scores:

[0124] Set the window size as M and the base window size as W. Then, the divided window size is: For each unimportant Token (t i ∈ T non-imp ) within each window, calculate the maximum attention score between it and the initial importance tokens, as shown in Equation (9):

[0125]

[0126] (iii) Screening important tokens:

[0127] Sort all non-important tokens in ascending order of A i as shown in Equation (10):

[0128] {t (1) , t (2) ,..., t (L)}, A (1) ≤ A (2) ≤... ≤ A (L) (10)

[0129] Screen the top β% of the tokens as the new important tokens, as shown in Equation (11):

[0130]

[0131] Merge the initial important tokens and the new important tokens as the output of the overall token structured pruning, as shown in Equation (11):

[0132]

[0133] The present invention proposes an innovative structure that integrates Swin Transformer and YOLOv9 to achieve global context awareness and long-distance target association modeling. Through cross-level feature expression optimization, the modeling ability of long-distance target relationships and context is enhanced;

[0134] The present invention introduces a multi-stage dynamic token screening mechanism and a shallow redundant feature suppression strategy. Through multi-stage design, important token features are gradually screened out, significantly reducing the computational complexity during the training and inference stages of the model, improving the training efficiency and accelerating the model convergence;

[0135] The present invention proposes a feature compensation mechanism to avoid the loss of important semantic information. Aiming at the key semantic loss caused by the possible discarding of low-importance token features during dynamic token screening, a feature compensation module is established to moderately reconstruct or information backflow the ignored low-importance features;

[0136] The present invention proposes an attention mechanism-based token pruning strategy, which significantly reduces the parameter scale and computational complexity during inference while maintaining the target detection performance, improving the deployment efficiency.

[0137] Such as Figure 6As shown in the figure, it is the detection effect diagram of the object detection network based on the present invention on COCO. The objects detected are framed in the figure, and the label of the frame is the category and classification prediction probability of the object inside the frame. For example, in the schematic diagram, person 0.93 means that the object inside the detection frame is of the person category, and the classification prediction probability is 0.93.

Claims

1. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation, characterized in that: The steps include: Step 1: Preprocess the input target image; Step 2: Use the target detection network to extract features from the preprocessed target image; Step 3: Use a multi-stage dynamic token screening mechanism based on the attention mechanism and a feature compensation mechanism based on the attention mechanism to train the extracted features, dynamically screen meaningful tokens and perform feature compensation; Step 4: Perform token structured pruning based on the attention mechanism on the token after feature compensation, and use the pruned features for target detection.

2. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 1, characterized in that: The step 1 is specifically as follows: (1a) Scale the input target image to the fixed size required by the target detection network; (1b) Normalize the input image; (1c) adjusting the normalized input image to a fixed size by center cropping; (1d) Perform data augmentation on the adjusted image.

3. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 1, characterized in that: Step 2 The target detection network is divided into backone and head, where: The backbone consists of four feature extraction stages of Swin Transformer, a multi-stage dynamic Token screening mechanism based on the attention mechanism, and a feature compensation mechanism. The backbone is used to extract multi-level features from the input image and extract features for the head to detect the target. The head uses the detection head network architecture of YOLOv9. The head is used to detect objects in the image based on the multi-level features provided by the backone.

4. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 3, characterized in that: The step 3 specifically includes: a multi-stage dynamic Token screening module and a feature compensation module based on the attention mechanism respectively; The multi-stage dynamic token screening mechanism initializes importance tokens and multi-stage dynamic screening tokens; The feature compensation mechanism performs token feature compensation.

5. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 4, characterized in that: The initialization importance token: The output of the first stage of Swin Transformer is used as the input of this stage. The input feature information is the token feature T after window division = {t1, t2, t3, …, t N }, select the initial important token as shown in formula (1) and formula (2): I={t j ∣S j ∈β} (2) Where A is the attention score matrix within the window, t j ∈T,S j is the sum of the attention scores of the jth column in A, β is the set of tokens represented by the top α% with higher scores in S, and I is the selected importance token.

6. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 5, characterized in that: The multi-stage dynamic screening token: (i) Establish a new window attention score matrix: Use the attention mechanism between the tokens selected by the initial importance token and the unselected tokens to find the new attention score matrix A; (ii) Select the top γ% tokens with the lowest cumulative attention influence and reclassify these tokens as important tokens.

7. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 6, characterized in that: The step (ii) is specifically as follows: let U be a set of unimportant tokens, and I be a set of important tokens; for each unimportant token u i ∈U, calculate its cumulative attention impact C on important tokens i , as shown in formula (3): After selecting the unimportant token set in C and sorting it from large to small, the tokens in the first γ% are I new The important token selected in the first stage is I old , the new important token is I, as shown in formula (4): I=I new +I old (4) The strategy of dynamically screening tokens in multiple stages is repeated multiple times to gradually obtain important tokens.

8. A Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 7, characterized in that: Token feature compensation: (i) Based on the updated important token set I′ and the unimportant token set U′, calculate the attention score matrix A∈R |U′|×|I′| , as shown in formula (5): where u j ∈U′,i k ∈I′, is the query vector of the attention mechanism, is the value vector of the attention mechanism, d is u j with i k The cosine distance of (ii) Based on the new attention score matrix A∈R |U′|×|I′| , for each unimportant tokenu j ∈U′ is assigned to the important token group G with the maximum attention score k , as shown in formula (6): Part i k ∈I′, u j ∈U′; (iii) Feature fusion is performed based on group G, as shown in formula (7): in For unimportant token j The value vector of is the fused feature vector and is used as the fusion information passed to the next layer.

9. The Swin-YOLOv9 target detection method based on multi-stage dynamic pruning and feature compensation according to claim 5, characterized in that: The step 4 is specifically as follows: (i) Initialize importance token: The output of the first stage of Swin Transformer is used as the input of this stage. The input tokens are uniformly sampled as the benchmark of the importance tokens. The input token sequence is: T = {t1, t2, t3, …, t N }, set the sampling parameter α∈(0,1], then every Tokens select an initial importance token, then the initial importance token set is shown in formula (8): in Represents the initialization of the important token set, N represents the number of original tokens, and a part of the tokens are selected from the input token sequence as the initial important tokens through uniform sampling; (ii) Attention score calculation: Set the window size to M and the base window size to W, then the partitioned window size is: For each unimportant Token(t i ∈T non-imp ), calculate the maximum attention score between it and the initial importance token, as shown in formula (9): (iii) Screening important tokens: Press A for all unimportant tokens i The sorting from small to large is shown in formula (10): {t (1) ,t (2) ,...,t (L) },A (1) ≤A (2) ≤…≤A (L) (10) Filter the top β% of tokens as new important tokens, as shown in formula (11): The initial important token and the new important token are combined as the output of the entire token structured pruning, as shown in formula (11): The final result T imp As the input information of the second stage of Swin Transformer in the backbone of the target detection model.

Citation Information

Cited By

  • Visual Transform-based dynamic screening medical image target tracking method and device

    CN120823216A

  • A dynamic screening medical image target tracking method and device based on visual Transformer

    CN120823216B

  • Multi-layer attention model reasoning acceleration method and device based on anchor point compression mechanism

    CN121351903A

  • Image recognition method, electronic equipment, readable storage medium and program product

    CN122454298A