A target detection method for sharing guidance modules

Through the object detection method of the shared guidance module, all information of the positioning branch is used to guide the classification branch, the problems of unsatisfactory detection and reduced speed in the YOLOV8 algorithm are solved, and more efficient object detection is achieved.

CN116844085BActive Publication Date: 2025-07-08GAOZHONG INFORMATION TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310770261.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2025-07-08
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

The unused distance focusing loss in the YOLOV8 algorithm results in unsatisfactory detection effect, and the multi-branch structure leads to a reduced algorithm speed, ignoring the guiding significance of low-probability distances on classification and increased computational consumption.

Method used

The object detection method of the shared guidance module is adopted, and through feature fusion and weight sharing, all information of the positioning branch is used to guide the classification branches, including low probability and medium probability values, reducing memory consumption and improving running speed.

Benefits of technology

The target detection effect is improved, the classification guidance ability is enhanced, and the operation speed is improved without reducing the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844085B_ABST
    Figure CN116844085B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method for a shared guidance module, which is specifically divided into the following steps: S1. Image input: Input the image into the algorithm model; S2. Feature extraction: Extract features from the input image; S3. Feature fusion: Fuse the extracted features to form more than one fused feature; S4. Decoding and prediction: Decode the fused feature and obtain the prediction result; Add a guidance module to the fused feature; More than one guidance module shares weights; The present invention improves the effect of target detection by using localization guidance for classification; The input of the guidance module does not only use Top-k and mean values, but uses all the information output by localization, improving the ability of localization guidance for classification; By sharing weights, the memory consumption is reduced without reducing the target detection effect, and the running speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to an object detection method for a shared guidance module. Background Art

[0002] With the large-scale application of intelligent monitoring, the importance of object detection algorithms has also been greatly improved. For example, for video-based vehicle illegal parking, pedestrian crossing the boundary, etc., object detection algorithms are required to output the position and category of the object. Therefore, it is very important to improve the object detection effect.

[0003] In the YOLOV8 algorithm, the Distribution Focal Loss is used to improve the detection effect of the algorithm. The Distribution Focal Loss can calculate the distance distribution from the four sides of the detection box to the center point. Through this distribution, the uncertainty of the bounding box can be judged, thereby guiding the classification branch. However, this method is not used in the YOLOV8 algorithm, resulting in an unsatisfactory detection effect.

[0004] In the Distribution Focal Loss V2, it is proposed to use Top-k (the k values with the highest distance in the distance distribution from the four sides to the center point) and the mean value (the mean value of the distance distribution from the four sides to the center point) to guide classification, which has been effectively improved in multiple detection algorithms;

[0005] Although the Distribution Focal Loss V2 has been improved in multiple detection algorithms, it has two problems: First, only using Top-k and the mean value to guide classification ignores that the distances with small probabilities also have guiding significance for classification, because the distances with small probabilities may represent occlusion or blur here and cannot be simply discarded; Second, the YOLOV8 algorithm is multi-branched, and different branches need to extract features in the Head to guide classification, resulting in a reduction in the algorithm speed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an object detection method for a shared guidance module to solve the technical problems mentioned in the above background art.

[0007] The object detection method of the shared guidance module of the present invention is realized through the following technical solutions:

[0008] Specifically, it is divided into the following steps:

[0009] S1. Image input: Input the image into the algorithm model;

[0010] S2. Feature extraction: Extract features from the input image;

[0011] S3. Feature Fusion: Fuse the extracted features to form more than one fused feature;

[0012] S4. Decoding and Prediction: Decode the fused feature and obtain the prediction result;

[0013] Add a guidance module to the fused feature; more than one guidance module shares weights; more than one guidance module includes a localization branch and a classification branch; the guidance module is added between the localization branch and the classification branch.

[0014] As a preferred technical solution, the localization branch includes the first convolutional module with three or more layers; the guidance module includes the second convolutional module with three or more layers; the classification branch includes the third convolutional module with three or more layers;

[0015] The guidance module inputs the information on the localization branch; a Sigmoid module is provided at the end of the second convolutional module with three or more layers; the information is input into the Sigmoid module after the features are extracted by the second convolutional module with three or more layers;

[0016] The Sigmoid module normalizes the features as weights and multiplies them with the classification branch to obtain the output of the classification branch.

[0017] As a preferred technical solution, the normalization function of the Sigmoid module:

[0018] As a preferred technical solution, S2. Feature Extraction: Extract features through the Backbone; the Backbone outputs features of three scales, which are respectively used to predict large, medium, and small targets.

[0019] As a preferred technical solution, S3. Feature Fusion: The features of the three scales are fused through the Neck and then respectively input into feature extraction; the feature extraction includes Head1, Head2, and Head3.

[0020] The beneficial effects of the present invention are:

[0021] 1. By using localization to guide classification, the effect of object detection is improved;

[0022] 2. The input of the guidance module does not only use Top-k and the mean, but uses all the information output by localization, which improves the ability of localization to guide classification;

[0023] 3. By sharing weights, the memory consumption is reduced without reducing the object detection effect, and the running speed is improved. Description of the Drawings

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0025] Figure 1 It is the algorithm flow chart of the object detection method of the shared guidance module of the present invention;

[0026] Figure 2 It is the algorithm network structure diagram of the object detection method of the shared guidance module of the present invention. Detailed implementation manners

[0027] All features disclosed in this specification, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.

[0028] Any feature disclosed in this specification (including any additional claims, abstract, and drawings) can be replaced by other equivalent or features with similar purposes, unless specifically stated otherwise. That is, unless specifically stated, each feature is only an example of a series of equivalent or similar features.

[0029] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by terms such as "one end", "the other end", "outer side", "upper", "inner side", "horizontal", "coaxial", "center", "end", "length", "outer end", etc. are based on the orientation or positional relationships shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0030] In addition, in the description of the present invention, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0031] Spatial relative position terms such as "upper", "above", "lower", "below", etc. used in the present invention are for the purpose of facilitating description of the relationship of one unit or feature relative to another unit or feature as shown in the accompanying drawings. The spatial relative position terms may be intended to include different orientations of the device in use or operation other than the orientation shown in the figures. For example, if the device in the figure is flipped, the unit described as being "below" or "beneath" other units or features will be "above" other units or features. Therefore, the exemplary term "below" can encompass both the upper and lower orientations. The device can be oriented in other ways (rotated 90 degrees or other orientations), and the spatially related descriptive terms used herein can be interpreted accordingly.

[0032] In the present invention, unless otherwise clearly specified and defined, terms such as "arranged", "socketed", "connected", "penetrated", "plugged in", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0033] As Figure 1 - Figure 2 shown, a target detection method for a shared guidance module of the present invention specifically includes the following steps:

[0034] S1. Image input: Input the image into the algorithm model; collect detection pictures to form a training set;

[0035] S2. Feature extraction: Extract features from the input image. After feature extraction by the Backbone, the Backbone outputs features of three scales, which are respectively used to predict large, medium, and small targets;

[0036] S3. Feature fusion: Respectively fuse the three-scale features extracted through the Neck to form more than one fusion feature; the fusion features include Head1, Head2, and Head3;

[0037] S4. Decoding and prediction: Decode the fusion features and obtain the prediction results;

[0038] Add a guidance module to the fusion feature; more than one guidance module shares weights; more than one guidance module includes a localization branch and a classification branch; the guidance module is added between the localization branch and the classification branch.

[0039] In this embodiment, the localization branch includes a first convolutional module with three or more layers; the guidance module includes a second convolutional module with three or more layers; the classification branch includes a third convolutional module with three or more layers;

[0040] The guidance module inputs information from the localization branch; a Sigmoid module is provided at the end of the second convolutional module with three or more layers; the information is input into the Sigmoid module after the features are extracted by the second convolutional module with three or more layers; the Sigmoid module normalizes the features as weights and multiplies them with the classification branch to obtain the output of the classification branch;

[0041] The input of the guidance module does not only use Top-k and the mean, but uses all the information of the localization output, including low-probability and medium-probability values, etc., providing richer features for guiding classification; considering that the addition of the guidance module will increase the computational consumption, and the way of localization guiding classification is similar for Head1, Head2, and Head3, so the guidance modules in the three Heads share weights to reduce memory consumption and improve the running speed.

[0042] As Figure 2 shown, in this embodiment, the normalization function of the Sigmoid module: conv represents the convolutional module, and "*" represents multiplication.

[0043] In this embodiment, S2, feature extraction: Features are extracted through the Backbone; the Backbone outputs features of three scales, which are used to predict large, medium, and small targets respectively.

[0044] In this embodiment, S3, feature fusion: Features of the three scales are fused through the Neck and then input into feature extraction respectively; feature extraction includes Head1, Head2, and Head3.

[0045] The working process is as follows: Features of each scale are respectively sent into a decoding and prediction module, and the localization branch and the classification branch process them respectively. The output of the localization branch is sent into the guidance module for feature extraction. To reduce memory consumption and improve the running speed, the guidance modules in the three Heads of the present invention share weights; finally, a value between 0 and 1 is output as a weight through the Sigmoid function, and it is multiplied with the classification probability output by the classification branch to obtain the final output of the classification branch.

[0046] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any change or replacement that can be thought of without creative work should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope defined by the claims.

Claims

1. A target detection method for sharing a guidance module, characterized in that: Specifically, it is divided into the following steps: S1. Image input: Input the image into the algorithm model; S2. Feature extraction: Extract features from the input image; S3. Feature fusion: Fuse the extracted features to form more than one fused feature; S4. Decoding prediction: Decode the fused feature and obtain the prediction result; Add a guidance module to the fused feature; more than one guidance module shares weights; more than one guidance module includes a localization branch and a classification branch; the guidance module is added between the localization branch and the classification branch; The localization branch includes the first convolutional module with three or more layers; the guidance module includes the second convolutional module with three or more layers; the classification branch includes the third convolutional module with three or more layers; The guidance module inputs the information on the localization branch; a Sigmoid module is arranged at the end of the second convolutional module with three or more layers; the information is input into the Sigmoid module after the features are extracted by the second convolutional module with three or more layers; The Sigmoid module normalizes the features as weights and multiplies them with the classification branch to obtain the output of the classification branch; Normalization function of the sigmoid module: ; In the said S2. Feature extraction: Extract features through the Backbone; The Backbone outputs features of three scales, which are respectively used to predict three types of targets, large, medium, and small; In the said S3. Feature fusion: The features of three scales are fused through the Neck and then respectively input into feature extraction; Feature extraction includes Head1, Head2, and Head3.

Citation Information

Patent Citations

  • Target detection method based on rapid neural architecture search

    CN112464960A

  • Target detection method and system based on Transform and fusion attention mechanism

    CN115908772A