An adaptive feature fusion method based on video detection

By inserting the spatial feature fusion module into the video detection model and dynamically adjusting the degree of fusion of feature information, the problem of insufficient target detection accuracy of existing video detection models is solved, and the model performance is improved and training efficiency is optimized.

CN119540695BActive Publication Date: 2025-06-06GUANGDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411601765.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-06-06
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing lightweight video detection model has insufficient accuracy in object detection, and the feature fusion method relies on subjective settings, resulting in bias in training.

Method used

An adaptive feature fusion method based on video detection is proposed. By inserting the spatial feature fusion module, combining multi-scale feature information, and through information processing in the pre-training and formal training stages, the spatial fusion parameter Fi is formed, and the degree of fusion of feature information is dynamically adjusted.

Benefits of technology

It effectively improves the accuracy of the video detection model in object detection, reduces the cost of training time, and ensures the high detection accuracy of the model when processing new data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540695B_ABST
    Figure CN119540695B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of video detection technology, and in particular to an adaptive feature fusion method based on video detection, comprising the following steps: a video detection model is pre-trained according to a training set, and the model performs a certain number of iterations in the pre-training; the video detection model starts formal training using the training set, and processes the information of the video detection model to extract key information in the video detection model training; key procedural information is extracted from the formal training of the video detection model, and a spatial fusion parameter Fi is formed; and the fusion parameter Fi is used to adjust the degree of fusion of feature information in a spatial feature fusion module, so that the video detection model can adaptively adjust the training direction during training. The present invention adaptively fuses information from different channels through a spatial feature fusion module to achieve information compensation, effectively improve the accuracy of the video detection model in target detection, and ultimately achieve the purpose of improving model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video detection, and in particular to an adaptive feature fusion method based on video detection. Background Art

[0002] Video detection is an important research field in artificial intelligence. Currently, video detection technology is widely used in target detection or behavior detection in scenarios such as campuses, factories, and hospitals to provide support for intelligent monitoring. For edge devices, especially portable cameras, due to their limited hardware support, the devices cannot have high computing power support. Therefore, lightweight video detection models are widely used on edge devices. At present, the main optimization idea of ​​lightweight video detection models is to use pruning, quantization and other methods. Although these methods reduce the demand for computing power of video detection models, they also sacrifice the performance of video detection models. In addition to these more traditional methods, feature fusion methods are also used to optimize the structure of video detection models, but these methods still rely more on subjective settings in information fusion, which makes the video detection model have deviations in training. Summary of the invention

[0003] The present invention aims to at least solve the technical problems existing in the prior art. To this end, the present invention proposes an adaptive feature fusion method based on video detection, which effectively improves the accuracy of the model in target detection and ultimately achieves the purpose of improving the performance of the model.

[0004] According to some embodiments of the present invention, an adaptive feature fusion method based on video detection includes the following steps:

[0005] S100, optimizing the network of the video detection model and inserting a spatial feature fusion module;

[0006] S200, the video detection model is pre-trained according to the training set, and the model performs a certain number of iterations in the pre-training, so that the training parameters of the spatial feature fusion module have a certain stability, and the module is initialized;

[0007] S300, the video detection model starts formal training using the training set, and processes the information of the video detection model to extract key information in the video detection model training;

[0008] S400, extracting key process information from the formal training of the video detection model and forming a spatial fusion parameter Fi;

[0009] S500, using the fusion parameter Fi to adjust the fusion degree of feature information in the spatial feature fusion module, highlighting important training information, so that the video detection model can adaptively adjust the training direction during training and conduct targeted learning.

[0010] An adaptive feature fusion method based on video detection according to some embodiments of the present invention has at least the following beneficial effects:

[0011] The present invention introduces a spatial feature fusion module, which can effectively combine multi-scale feature information so that the video detection model can accurately capture the details of the target object in different scenarios. The model improves the accuracy of target recognition. The pre-training stage ensures that the training parameters of the spatial feature fusion module have a certain stability by preliminarily training the model, laying a good foundation for subsequent formal training. This process helps to speed up the overall convergence speed of the model and reduce the time cost required for training. By dynamically adjusting the degree of spatial fusion of feature information, the present invention can enable the video detection model to pay more attention to learning key information that contributes to the task during the training process, and use the fusion parameter Fi formed by key procedural information to adjust the degree of fusion of feature information, ensuring that the model can still maintain a high detection accuracy when processing new data.

[0012] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, step S100 specifically includes: a spatial feature fusion module fuses information of different channels in a frame of video information so that a certain channel has feature information of other channels in horizontal training.

[0013] According to some embodiments of the present invention, an adaptive feature fusion method based on video detection is provided, wherein the video detection model is Yolo v 7 models.

[0014] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, the training set is VisDrone2019.

[0015] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, the video detection model is pre-trained according to a training set, including: initializing parameters of a spatial feature fusion module through pre-training.

[0016] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, the information of the video detection model includes feature information extracted by a spatial feature fusion module and output information of video detection model training.

[0017] According to some embodiments of the present invention, an adaptive feature fusion method based on video detection is provided, wherein the video detection module includes a convolution block and a spatial shift operation, wherein the spatial shift operation is used to increase the receptive field during the convolution operation, and the extracted feature layer, after being processed by the convolution block and the spatial shift operation, transmits the information to the mask layer, and the mask layer randomly masks the information and adjusts the size of the feature layer, and then fuses it with the feature layers of other channels.

[0018] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, the overall expression of the video detection module is as follows:

[0019]

[0020] Among them, X0 is the training information extracted from other channels, MCi is the output of the spatial feature fusion module, C() is the convolution operation, i is the sequence number of the channel, S() is the spatial Shift operation, and W0 is the mask layer, which means element multiplication.

[0021] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, Fi is expressed as follows:

[0022] Fi = sigmoid(Yi)

[0023] Among them, sigmoid() is the activation function, Yi is the output information of the model training, and the activation function is used to normalize the output information to the range of 0 and 1 according to the importance of the output information. The closer to 1, the better the learning effect of the current video detection model for this category, and the closer to 0, the worse.

[0024] According to an adaptive feature fusion method based on video detection in some embodiments of the present invention, the fusion parameter Fi is used to adjust the fusion degree of feature information in the spatial feature fusion module: the expression of the feature layer Po after information fusion is:

[0025]

[0026] Among them, X is the feature information that needs to be fused. If the Fi of the channel is larger, the 1-Fi is smaller, then the feature information of the channel has been well mastered by the model, and the feature information on the channel can be reduced in the feature fusion stage; if the Fi of the channel is smaller, the 1-Fi is larger, indicating that the feature data of the channel is difficult to master for the current model, and it is necessary to enhance this part of the information in the feature fusion stage.

[0027] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0029] Figure 1 The figure is a flowchart of the steps of an embodiment of the present invention.

[0030] Figure 2 This is a flowchart of the spatial feature fusion module processing feature information in an embodiment of the present invention.

[0031] Figure 3 The figure is a flowchart of an embodiment of the present invention. DETAILED DESCRIPTION

[0032] Embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0033] In the description of the present invention, it should be understood that the descriptions involving orientations, such as up, down, left, right, front, back, etc., and the orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the modules or components referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0034] In the description of the present invention, if there is a description of first and second, it is only for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0035] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0036] like Figure 1-Figure 3 As shown, an embodiment of the present invention provides an adaptive feature fusion method based on video detection.

[0037] An adaptive feature fusion method based on video detection comprises the following steps:

[0038] S100, optimizing the network of the video detection model and inserting a spatial feature fusion module;

[0039] S200, the video detection model is pre-trained according to the training set. The model performs a certain number of iterations in the pre-training so that the training parameters of the spatial feature fusion (MCSFF) module have a certain stability and the module is initialized.

[0040] S300, the video detection model starts formal training using the training set, and processes the information of the video detection model to extract key information in the video detection model training;

[0041] S400, extracting key process information from the formal training of the video detection model and forming a spatial fusion parameter Fi;

[0042] S500, using the fusion parameter Fi to adjust the fusion degree of feature information in the spatial feature fusion module, highlighting important training information, so that the video detection model can adaptively adjust the training direction during training and conduct targeted learning.

[0043] The present invention introduces a spatial feature fusion module, which can effectively combine multi-scale feature information, so that the video detection model can accurately capture the details of the target object in different scenarios. The model improves the accuracy of target recognition. The pre-training stage ensures that the training parameters of the spatial feature fusion module have a certain stability by preliminarily training the model, laying a good foundation for subsequent formal training. This process helps to speed up the overall convergence speed of the model and reduce the time cost required for training. By dynamically adjusting the spatial fusion degree of feature information, the present invention can make the video detection model pay more attention to learning the key information that contributes to the task during the training process, and use the fusion parameter Fi formed by key process information to adjust the fusion degree of feature information, ensuring that the model can still maintain a high detection accuracy when processing new data.

[0044] The present embodiment describes an adaptive feature fusion method based on video detection, wherein the video detection model is a Yolov7 model. The Yolov7 model can be divided into four parts according to the function, namely, input, backbone network, neck, and prediction head. After the feature information is processed by the neck module feature fusion, feature information of three sizes, large, medium, and small, is obtained. These three types of feature information are sent to three prediction heads respectively, and the training results are output after detection. The three types of feature information are connected to the three prediction heads to form three channels. The spatial feature fusion module is inserted into the channel between the neck and the prediction head, and the feature information of the three sizes corresponds to the three spatial feature fusion modules.

[0045] The adaptive feature fusion method based on video detection described in this embodiment, the step S100 specifically includes: the spatial feature fusion module fuses the information of different channels in a certain frame of video information, so that a certain channel has the feature information of other channels in horizontal training.

[0046] In the adaptive feature fusion method based on video detection described in this embodiment, in step S200, in order to achieve faster convergence in formal training, the spatial feature fusion module will initialize parameters through model pre-training, and the data set used in the pre-training process is VisDrone2019.

[0047] After a certain amount of iterative training, the parameters of the spatial feature fusion module tend to be stable. At this point, the video detection model can start formal training. The dataset used for formal training is VisDrone2019, which is divided into training set, test set and validation set in a ratio of 6:3:1. In formal training, the model uses the training set.

[0048] In the adaptive feature fusion method based on video detection described in this embodiment, the information of the video detection model includes feature information extracted by the spatial feature fusion module and output information of the video detection model training.

[0049] Two key process information will be extracted during the formal training of the video detection model. The first is the feature information output through the neck, which is divided into three types: large, medium, and small according to different sizes. In order to improve the extraction efficiency, the spatial feature fusion module will not extract feature information directly from the channel, but will extract feature information through an information register. This information register will store the extraction results of the first n times in the three channels. In this project, the value of n is 3, that is, the information register will store the extraction results of the first 3 times in the three channels. The other type of information that needs to be extracted is the detection information of the three channels. This detection information is the confidence of the channel for each category. The overall expression of the video detection module is as follows:

[0050]

[0051] Among them, X0 is the training information extracted from other channels, MCi is the output of the spatial feature fusion module, C() is the convolution operation, i is the sequence number of the channel, S() is the spatial Shift operation, W0 is the mask layer, which means element multiplication. The spatial feature fusion module fuses the information of other channels to supplement the current channel information.

[0052] In the adaptive feature fusion method based on video detection described in this embodiment, the expression Fi is as follows:

[0053] Fi = sigmoid(Yi)

[0054] Among them, sigmoid() is the activation function, Yi is the output information of the model training, and the activation function is used to normalize the output information to the range of 0 and 1 according to the importance of the output information. The closer to 1, the better the learning effect of the current video detection model for this category, and the closer to 0, the worse.

[0055] The present embodiment describes an adaptive feature fusion method based on video detection, wherein the video detection module includes a convolution block and a spatial shift operation, wherein the spatial shift operation is used to increase the receptive field during the convolution operation, and the extracted feature layer passes the information to the mask layer after being processed by the convolution block and the spatial shift operation, and the mask layer randomly masks the feature information and adjusts the size of the feature layer, and then fuses it with the feature layers of other channels.

[0056] The spatial feature fusion module extracts feature information from the information register, performs adaptive fusion, and then fuses it with the feature information of this channel. For small-sized feature information, the spatial feature fusion module first fuses the feature information of the medium-sized and large-sized channels. The degree of fusion of these two feature information is controlled by the fusion parameter Fi. The fused feature layer will undergo spatial shift operation and convolution block processing. In order to make the detection performance of the video detection model more stable, the fusion of the two feature information will be randomly masked and the size of the feature layer will be adjusted so that it can be smoothly input into the prediction head for detection. The feature information fusion of other channels is consistent with the small-sized channel.

[0057] The expression of the feature layer Po after information fusion is:

[0058]

[0059] Among them, X is the feature information that needs to be fused. If the Fi of the channel is larger, the 1-Fi is smaller, then the feature information of the channel has been well mastered by the model, and the feature information on the channel can be reduced in the feature fusion stage; if the Fi of the channel is smaller, the 1-Fi is larger, indicating that the feature data of the channel is difficult to master for the current model, and it is necessary to enhance this part of the information in the feature fusion stage.

[0060] During training and detection, the video detection model can learn too little information due to occlusion of target information, shooting quality and other reasons. By fusing feature information from different channels, the training information can be supplemented from different dimensions, and the degree of information supplementation can be adjusted using the output information of the channel, effectively avoiding overfitting or underfitting of the model during training, so as to improve the performance of the model in video detection.

[0061] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. An adaptive feature fusion method based on video detection, characterized in that: The following steps are involved: S100, optimizing the network of the video detection model and inserting a spatial feature fusion module; S200, the video detection model is pre-trained according to the training set, and the model performs a certain number of iterations in the pre-training, so that the training parameters of the spatial feature fusion module have a certain stability, and the module is initialized; S300, the video detection model starts formal training using the training set, and processes the information of the video detection model to extract key information in the video detection model training; S400, extracting key process information from the formal training of the video detection model and forming a spatial fusion parameter Fi; S500, using the fusion parameter Fi to adjust the fusion degree of the feature information in the spatial feature fusion module, highlighting the important information of the training, so that the video detection model can adaptively adjust the training direction during training and perform targeted learning; Step S100 specifically includes: the spatial feature fusion module fuses the information of different channels in a certain frame of video information, so that a certain channel has the feature information of other channels in horizontal training; The information of the video detection model includes feature information extracted by the spatial feature fusion module and output information of the video detection model training; The spatial feature fusion module includes a convolution block and a spatial shift operation. The spatial shift operation is used to increase the receptive field during the convolution operation. After the extracted feature layer is processed by the convolution block and the spatial shift operation, the information is transmitted to the mask layer. The mask layer randomly masks the information and adjusts the size of the feature layer, and then fuses it with the feature layers of other channels. The expression of Fi is as follows: in, is the activation function, The activation function is used to normalize the output information of the model training to the range of 0 and 1 according to the importance of the output information. The closer to 1, the better the learning effect of the current video detection model for the category, and the closer to 0, the worse.

2. The method for adaptive feature fusion based on video detection according to claim 1, characterized in that: The video detection model is the Yolov7 model.

3. The method for adaptive feature fusion based on video detection according to claim 1, characterized in that: The dataset used in the training set is VisDrone2019.

4. The adaptive feature fusion method based on video detection according to claim 1, characterized in that: The video detection model is pre-trained according to the training set, including: initializing the parameters of the spatial feature fusion module through pre-training.

5. The method for adaptive feature fusion based on video detection according to claim 1, characterized in that: The overall expression of the spatial feature fusion module is as follows: in, is the training information extracted from other channels, is the output of the spatial feature fusion module, is the convolution operation, is the channel number, For the space shift operation, It is a mask layer, which means element-wise multiplication.

6. The adaptive feature fusion method based on video detection according to claim 1, characterized in that: The fusion parameter Fi is used to adjust the fusion degree of feature information in the spatial feature fusion module: The expression is: in, For feature information that needs to be fused, if the Fi of the channel is larger and 1-Fi is smaller, the feature information of the channel has been well mastered by the model, and the feature information on the channel can be reduced in the feature fusion stage; if the Fi of the channel is smaller and 1-Fi is larger, it indicates that the feature data of the channel is difficult to master for the current model, and it is necessary to enhance this part of the information in the feature fusion stage.

Citation Information

Patent Citations

  • Video stream target detection method and device based on modular lightweight network

    CN116403168A

  • Camouflage target detection method based on edge information adaptive feature fusion network

    CN118071998A