Video target detection method, device, equipment and medium based on feature aggregation

By introducing feature aggregation method in video object detection, combining YOLOv5m network and feature aggregation network, the problems of false detection and missed detection and slow detection speed in the prior art are solved, and more efficient video object detection is achieved.

CN114299425BActive Publication Date: 2025-05-09GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111572460.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-05-09
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

The existing video object detection methods have shortcomings in utilizing video time context information, resulting in misdetection and misdetection problems, and the method based on the two-stage detection algorithm is slow.

Method used

Using a video object detection method based on feature aggregation, the object detection network is constructed, including the YOLOv5m network and the feature aggregation network, and the object detection is performed using the time context information of the video. This method obtains the recommended feature set of video frames, and performs feature fusion in the feature aggregation layer to generate the aggregate feature set, and finally outputs the target detection result through the convolution layer.

Benefits of technology

The target detection speed and effect are improved, and the detection speed is faster than the two-stage detection algorithm, while using video time context information to reduce false detection and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299425B_ABST
    Figure CN114299425B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention relate to the field of target detection technology, and disclose a method, device, equipment and medium for video target detection based on feature aggregation. The method includes: constructing a target detection network; acquiring a first frame image and a second frame image of a video, and determining a recommended feature set of a first support frame and a recommended feature set of a second support frame through the target detection network; inputting the t-th frame image into the target detection network, combining the recommended feature set of the first support frame and the recommended feature set of the second support frame, and outputting a target detection result; judging whether to update the first frame image, the second frame image and their corresponding recommended feature sets of the first support frame and the second support frame according to the t-th frame image. Implementing the embodiments of the present invention can improve the target detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a video target detection method, device, equipment and medium based on feature aggregation. Background Art

[0002] YOLO is a one-stage image object detection algorithm. Compared with the two-stage detection algorithm represented by Faster RCNN, it has a higher detection speed. Although YOLO can be directly used for video object detection, it cannot use the temporal context information of the video, which can effectively reduce the false detection and missed detection caused by small object area, motion blur, posture change, etc. in single-frame image detection, and improve the algorithm performance.

[0003] Existing video target detection methods include various implementation schemes such as tracking-based, optical flow-based, and proposed region feature-based. However, most of these schemes are based on improvements on the two-stage target detection algorithm, and the algorithm speed is relatively slow. Summary of the invention

[0004] In view of the above-mentioned defects, the embodiments of the present invention disclose a method, apparatus, device and medium for video target detection based on feature aggregation, which can improve the target detection speed.

[0005] A first aspect of an embodiment of the present invention discloses a video object detection method based on feature aggregation, the method comprising:

[0006] Constructing a target detection network, the target detection network includes a first network and a second network, the output information of the first network is sent to the input end of the second network, the first network is formed by removing the last level of convolutional layer of the YOLOv5m network structure, and the second network is a feature aggregation network;

[0007] Get the first frame image I of the video 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame

[0008] The t-th frame image I t Input the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, t≥3;

[0009] Determine whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0010] As a preferred embodiment, in the first aspect of the embodiment of the present invention, the first frame image I of the video is obtained. 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame include:

[0011] The first frame image I 1 and the second frame image I 2 Input the first network respectively to obtain multi-scale feature maps Where l is the scale, l = 1, 2, 3, S l ×S l is the number of grids in the feature map at scale l, D l is the feature dimension of the feature map grid at scale l;

[0012] The multi-scale feature map The first convolutional layer outputs three confidence scores for each grid feature, each score corresponds to a preset template frame, and the multi-scale feature map is output through the first convolutional layer. The corresponding scores

[0013] The score P l 1 , P l 2 They are respectively input into the feature extraction layer of the second network, and the feature extraction layer first uses the non-maximum suppression method to extract the score P l 1 , P l 2 Process them separately and then extract the N with the highest score l grid features and are combined into the proposed feature set of the first support frame and the proposed feature set for the second support frame

[0014]

[0015]

[0016] Save the first frame image I 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0017] As a preferred embodiment, in the first aspect of the embodiment of the present invention, the t-th frame image is input into the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, including:

[0018] The t-th frame image I t Input the first network to get a multi-scale feature map

[0019] Multi-scale feature map Input the first convolutional layer and feature extraction layer of the second network in turn to obtain the recommended feature set of the test frame

[0020]

[0021] The proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set

[0022]

[0023] The aggregate feature set The second convolutional layer of the second network is input to obtain the target detection result.

[0024] As a preferred embodiment, in the first aspect of the embodiment of the present invention, the proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set include:

[0025] Calculate adaptive weights

[0026]

[0027] Where k = 1, 2, t, 1 ≤ i ≤ Nl

[0028] The adaptive weight Perform normalization:

[0029]

[0030] Compute aggregate feature sets The i-th aggregate feature

[0031]

[0032] Get the aggregate feature set

[0033] As a preferred embodiment, in the first aspect of the embodiment of the present invention, it is determined whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame include:

[0034] Calculate the t-th frame image I t With the second frame image I 2 relevance;

[0035] If the correlation is less than a preset threshold, the second frame image I 2 and the proposed feature set for the second support frame Replace the first frame image I respectively 1 and the proposed feature set for the first support frame The t-th frame image I t and aggregate feature sets Replace the second frame image I 2 and the proposed feature set for the second support frame For the t+1th frame image I t+1 Perform target detection;

[0036] If the correlation is greater than or equal to the preset threshold, the first frame image I is kept 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame No change, continue to process the t+1th frame image I t+1 Perform target detection.

[0037] A second aspect of an embodiment of the present invention discloses a video object detection device based on feature aggregation, which includes:

[0038] A construction unit is used to construct a target detection network, wherein the target detection network includes a first network and a second network, wherein output information of the first network is sent to an input end of the second network, wherein the first network is formed by removing the last convolutional layer of a YOLOv5m network structure, and the second network is a feature aggregation network;

[0039] An acquisition unit, used to acquire a first frame image I of the video 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame

[0040] The target detection unit is used to detect the t-th frame image I t Input the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, t≥3;

[0041] An updating unit, configured to determine whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0042] As a preferred embodiment, in the second aspect of the embodiment of the present invention, the acquisition unit includes:

[0043] The first input subunit is used to input the first frame image I 1 and the second frame image I 2 Input the first network respectively to obtain multi-scale feature maps Where l is the scale, l = 1, 2, 3, S l ×S l is the number of grids in the feature map at scale l, D l is the feature dimension of the feature map grid at scale l;

[0044] The second input subunit is used to convert the multi-scale feature map The first convolutional layer outputs three confidence scores for each grid feature, each score corresponds to a preset template frame, and the multi-scale feature map is output through the first convolutional layer. The corresponding scores

[0045] The third input subunit is used to convert the score P l 1 , P l 2 They are respectively input into the feature extraction layer of the second network, and the feature extraction layer first uses the non-maximum suppression method to extract the score P l 1 , P l 2 Process them separately and then extract the N with the highest score l grid features and are combined into the proposed feature set of the first support frame and the proposed feature set for the second support frame

[0046]

[0047]

[0048] A saving subunit, used to save the first frame image I 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0049] As a preferred embodiment, in the second aspect of the embodiment of the present invention, the target detection unit includes:

[0050] The fourth input subunit is used to input the t-th frame image I t Input the first network to get a multi-scale feature map

[0051] The fifth input subunit is used to convert the multi-scale feature map Input the first convolutional layer and feature extraction layer of the second network in turn to obtain the recommended feature set of the test frame

[0052]

[0053] The sixth input subunit is used to input the proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set

[0054]

[0055] The seventh input subunit is used to convert the aggregate feature set The second convolutional layer of the second network is input to obtain the target detection result.

[0056] A third aspect of an embodiment of the present invention discloses an electronic device, comprising: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute a video target detection method based on feature aggregation disclosed in the first aspect of an embodiment of the present invention.

[0057] A fourth aspect of an embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute a video object detection method based on feature aggregation disclosed in the first aspect of an embodiment of the present invention.

[0058] A fifth aspect of an embodiment of the present invention discloses a computer program product. When the computer program product runs on a computer, the computer executes a video object detection method based on feature aggregation disclosed in the first aspect of an embodiment of the present invention.

[0059] A sixth aspect of an embodiment of the present invention discloses an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer executes a video target detection method based on feature aggregation disclosed in the first aspect of an embodiment of the present invention.

[0060] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0061] The embodiment of the present invention applies the feature aggregation method to the YOLO network, so that the YOLO network can use the temporal context information of the video, improves the target detection effect, and compared with the two-stage detection algorithm, the detection speed is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0063] Figure 1 It is a flow chart of a method for video object detection based on feature aggregation disclosed in an embodiment of the present invention;

[0064] Figure 2 is a result graph of the target detection network disclosed in an embodiment of the present invention;

[0065] Figure 3 It is a structural schematic diagram of a video object detection device based on feature aggregation disclosed in an embodiment of the present invention;

[0066] Figure 4 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0067] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0068] It should be noted that the terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention are used to distinguish different objects rather than to describe a specific order. The terms "including" and "having" in the embodiments of the present invention and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0069] The embodiments of the present invention disclose a method, apparatus, device and medium for video target detection based on feature aggregation, which applies the feature aggregation method to the YOLO network, so that the YOLO network can utilize the temporal context information of the video, improves the target detection effect, and compared with the two-stage detection algorithm, the detection speed is improved, which is described in detail below with reference to the accompanying drawings.

[0070] Embodiment 1

[0071] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for video object detection based on feature aggregation disclosed in an embodiment of the present invention. Figure 1 As shown, the video target detection method based on feature aggregation includes the following steps:

[0072] S110, constructing a target detection network, wherein the target detection network includes a first network and a second network, wherein the output information of the first network is sent to the input end of the second network, wherein the first network is formed by removing the last level of convolutional layer of the YOLOv5m network structure, and the second network is a feature aggregation network.

[0073] Please refer to Figure 2As shown, in the embodiment of the present invention, the network structure after removing the last level of convolution layer of the YOLOv5m network structure is used as the first network, and the feature aggregation network is used as the second network to construct a target detection network, and the output information of the first network is a multi-scale feature map.

[0074] S120, obtaining the first frame image I of the video 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame

[0075] Taking the method of obtaining the recommended feature set of the first supporting frame by using the first frame image as an example for explanation, the method of obtaining the recommended feature set of the second supporting frame by using the second frame image is similar.

[0076] The first frame of the video I 1 Input the first network to get a multi-scale feature map Where l is the scale, S l ×S l is the number of grids in the feature map at scale l, D l is the feature dimension of the feature map grid at scale l. In this embodiment, S1=19, S2=38, S3=76, D1=768, D2=384, and D3=192.

[0077] Multi-scale feature map Input the first convolutional layer of the second network. The first convolutional layer (1×1 convolutional layer) outputs 3 confidence scores for each grid feature. Each score corresponds to a preset template box, so the first convolutional layer outputs a score

[0078] The score P l 1 Input the feature extraction layer of the second network, which first uses the non-maximum suppression method to extract the score P l 1 Process and then extract the N with the highest scores l grid features and combine them into the proposed feature set for the first support frame In this embodiment, N1=25, N2=50, and N3=100.

[0079] Save the first frame image I 1 and the proposed feature set for the first support frame

[0080] S130, the t-th frame image It Input the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, t≥3.

[0081] Similar to the process of obtaining the proposed feature set for the first frame image, first obtain the tth frame image I t The proposed feature set is called the proposed feature set of the test frame. The specific process is:

[0082] A, first take the tth frame image I of the video t Input the first network to get a multi-scale feature map Then Input the first convolutional layer of the second network to get the confidence score P l t , the score P l t Input the feature extraction layer of the second network to obtain the recommended feature set of the test frame

[0083] B, the proposed feature set of the first support frame obtained in step S110 and the proposed feature set for the second support frame And the recommended feature set of the test frame obtained in this step Input the aggregation network layer of the second network to obtain the aggregated feature set

[0084] The specific steps are as follows:

[0085] B1, calculate adaptive weights

[0086]

[0087] Where k = 1, 2, t, 1 ≤ i ≤ N l ;

[0088] B2, the adaptive weight Perform normalization:

[0089]

[0090] B3, calculate the aggregate feature set The i-th aggregate feature

[0091]

[0092] B4, get the aggregated feature set

[0093] C, aggregate feature set The second convolutional layer (1×1 convolutional layer) of the second network is input to obtain the final output, and the content of the output result is consistent with the content of the original YOLOv5m output result.

[0094] S140, determining whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0095] The updating method is determined by judging the correlation between the test frame and the support frame.

[0096] Specifically:

[0097] Calculate the t-th frame image I t With the second frame image I 2 There are many methods for calculating image correlation. For example, a grayscale-based matching algorithm, a feature-based matching algorithm (for example), a relationship-based matching algorithm, etc. can be used.

[0098] Among them, the grayscale-based matching algorithm used may be, for example, the mean absolute difference algorithm (MAD), the absolute error sum algorithm (SAD), the error square sum algorithm (SSD), the mean error square sum algorithm (MSD), the normalized product correlation algorithm (NCC), the sequential similarity detection algorithm (SSDA), the hadamard transform algorithm (SATD), etc.

[0099] The feature-based matching algorithm used may be, for example, point feature matching algorithms such as Harris, Moravec, KLT, SIFT, SURF, BRIEF, SUSAN, FAST, CENSUS, FREAK, BRISK, ORB, optical flow method, A-KAZE, etc.; and edge feature matching algorithms such as LoG operator, Robert operator, Sobel operator, Prewitt operator, Canny operator, etc.

[0100] The relationship-based matching algorithm can be matched in the form of establishing a semantic network.

[0101] If the correlation is less than a preset threshold, the second frame image I 2 and the proposed feature set for the second support frame Replace the first frame image I respectively 1 and the proposed feature set for the first support frame The t-th frame image I t and aggregate feature sets Replace the second frame image I2 and the proposed feature set for the second support frame For the t+1th frame image I t+1 Perform target detection; that is, for the t+1th frame image I t+1 When performing target detection, step B in step S130 uses the updated recommended feature set of the first support frame and the proposed feature set for the second support frame

[0102] If the correlation is greater than or equal to the preset threshold, the first frame image I is kept 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame No change, continue to process the t+1th frame image I t+1 Perform target detection.

[0103] Embodiment 2

[0104] See also Figure 3 , Figure 3 Schematic diagram of the structure of a video object detection device based on feature aggregation disclosed in an embodiment of the present invention. Figure 3 As shown, the video object detection device based on feature aggregation may include:

[0105] A construction unit 210 is used to construct a target detection network, wherein the target detection network includes a first network and a second network, wherein the output information of the first network is sent to the input end of the second network, wherein the first network is formed by removing the last convolutional layer of the YOLOv5m network structure, and the second network is a feature aggregation network;

[0106] The acquisition unit 220 is used to acquire the first frame image I of the video. 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame

[0107] The target detection unit 230 is used to detect the t-th frame image I t Input the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, t≥3;

[0108] The updating unit 240 is used to determine whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0109] Preferably, the acquisition unit 220 may include:

[0110] The first input subunit is used to input the first frame image I 1 and the second frame image I 2 Input the first network respectively to obtain multi-scale feature maps Where l is the scale, l = 1, 2, 3, S l ×S l is the number of grids in the feature map at scale l, D l is the feature dimension of the feature map grid at scale l;

[0111] The second input subunit is used to convert the multi-scale feature map The first convolutional layer outputs three confidence scores for each grid feature, each score corresponds to a preset template frame, and the multi-scale feature map is output through the first convolutional layer. The corresponding scores

[0112] The third input subunit is used to convert the score P l 1 , P l 2 They are respectively input into the feature extraction layer of the second network, and the feature extraction layer first uses the non-maximum suppression method to extract the score P l 1 , P l 2 Process them separately and then extract the N with the highest score l grid features and are combined into the proposed feature set of the first support frame and the proposed feature set for the second support frame

[0113]

[0114]

[0115] A saving subunit, used to save the first frame image I 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame

[0116] Preferably, the target detection unit 230 may include:

[0117] The fourth input subunit is used to input the t-th frame image I t Input the first network to get a multi-scale feature map

[0118] The fifth input subunit is used to convert the multi-scale feature map Input the first convolutional layer and feature extraction layer of the second network in turn to obtain the recommended feature set of the test frame

[0119]

[0120] The sixth input subunit is used to input the proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set

[0121]

[0122] The seventh input subunit is used to convert the aggregate feature set The second convolutional layer of the second network is input to obtain the target detection result.

[0123] Embodiment 3

[0124] See also Figure 4 , Figure 4 Schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. Figure 4 As shown, the electronic device may include:

[0125] A memory 310 storing executable program codes;

[0126] a processor 320 coupled to the memory 310;

[0127] The processor 320 calls the executable program code stored in the memory 310 to execute part or all of the steps in the video object detection method based on feature aggregation in the first embodiment.

[0128] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute part or all of the steps in a video object detection method based on feature aggregation in Embodiment 1.

[0129] The embodiment of the present invention further discloses a computer program product, wherein when the computer program product is run on a computer, the computer is enabled to execute part or all of the steps in a video object detection method based on feature aggregation in the first embodiment.

[0130] An embodiment of the present invention also discloses an application publishing platform, wherein the application publishing platform is used to publish a computer program product, wherein when the computer program product runs on a computer, the computer executes part or all of the steps in a video target detection method based on feature aggregation in embodiment one.

[0131] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0132] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed over multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a memory and includes several requests for a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in a computer device) to perform some or all of the steps of the method described in each embodiment of the present invention.

[0135] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0136] A person of ordinary skill in the art can understand that some or all of the steps in the various methods of the embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0137] The above is a detailed introduction to a method, device, equipment and medium for video target detection based on feature aggregation disclosed in an embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A video object detection method based on feature aggregation, characterized in that: include: Constructing a target detection network, the target detection network includes a first network and a second network, the output information of the first network is sent to the input end of the second network, the first network is formed by removing the last level of convolutional layer of the YOLOv5m network structure, and the second network is a feature aggregation network; Get the first frame of the video I 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame The tth frame image I of the video t Input the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, t≥3; Determine whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame Get the first frame image I of the video 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame include: The first frame image I 1 and the second frame image I 2 Input the first network respectively to obtain multi-scale feature maps Where l is the scale, l = 1, 2, 3, S l ×S l is the number of grids in the feature map at scale l, D l is the feature dimension of the feature map grid at scale l; The multi-scale feature map The first convolutional layer outputs three confidence scores for each grid feature, each score corresponds to a preset template frame, and the multi-scale feature map is output through the first convolutional layer. The corresponding scores The score P l 1 , P l 2 They are respectively input into the feature extraction layer of the second network, and the feature extraction layer first uses the non-maximum suppression method to extract the score P l 1 , P l 2 Process them separately and then extract the N with the highest score l grid features and are combined into the proposed feature set of the first support frame and the proposed feature set for the second support frame Save the first frame image I 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame The tth frame image is input into the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, including: The t-th frame image I t Input the first network to get a multi-scale feature map Multi-scale feature map Input the first convolutional layer and feature extraction layer of the second network in turn to obtain the recommended feature set of the test frame The proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set The aggregate feature set The second convolutional layer of the second network is input to obtain the target detection result.

2. The video target detection method based on feature aggregation according to claim 1 is characterized in that: The proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set include: Calculate adaptive weights Where k = 1, 2, t, 1 ≤ i ≤ N l The adaptive weight Perform normalization: Compute aggregate feature sets The i-th aggregate feature Get the aggregate feature set 3. The video target detection method based on feature aggregation according to claim 1, characterized in that: Determine whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame include: Calculate the t-th frame image I t With the second frame image I 2 relevance; If the correlation is less than a preset threshold, the second frame image I 2 and the proposed feature set for the second support frame Replace the first frame image I respectively 1 and the proposed feature set for the first support frame The t-th frame image I t and aggregate feature sets Replace the second frame image I 2 and the proposed feature set for the second support frame For the t+1th frame image I t+1 Perform target detection; If the correlation is greater than or equal to the preset threshold, the first frame image I is kept 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame No change, continue to process the t+1th frame image I t+1 Perform target detection.

4. A video object detection device based on feature aggregation, characterized in that: It includes: A construction unit is used to construct a target detection network, wherein the target detection network includes a first network and a second network, wherein output information of the first network is sent to an input end of the second network, wherein the first network is formed by removing the last convolutional layer of a YOLOv5m network structure, and the second network is a feature aggregation network; An acquisition unit is used to acquire the first frame image I of the video. 1 and the second frame image I 2 , and determine the first frame image I through the target detection network 1 and the second frame image I 2 The proposed feature set corresponding to the first support frame and the proposed feature set for the second support frame The target detection unit is used to detect the tth frame image I of the video t Input the target detection network, combined with the proposed feature set of the first support frame and the proposed feature set for the second support frame Output target detection results, t≥3; An updating unit, used to determine whether to update the first frame image I according to the tth frame image 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame The acquisition unit comprises: The first input subunit is used to input the first frame image I 1 and the second frame image I 2 Input the first network respectively to obtain multi-scale feature maps Where l is the scale, l = 1, 2, 3, S l ×S l is the number of grids in the feature map at scale l, D l is the feature dimension of the feature map grid at scale l; The second input subunit is used to convert the multi-scale feature map The first convolutional layer outputs three confidence scores for each grid feature, each score corresponds to a preset template frame, and the multi-scale feature map is output through the first convolutional layer. The corresponding scores The third input subunit is used to convert the score P l 1 , P l 2 They are respectively input into the feature extraction layer of the second network, and the feature extraction layer first uses the non-maximum suppression method to extract the score P l 1 , P l 2 Process them separately and then extract the N with the highest score l grid features and are combined into the proposed feature set of the first support frame and the proposed feature set for the second support frame A saving subunit, used to save the first frame image I 1 , the second frame image I 2 and their corresponding proposed feature sets for the first support frame and the proposed feature set for the second support frame The fourth input subunit is used to input the t-th frame image I t Input the first network to get a multi-scale feature map The fifth input subunit is used to convert the multi-scale feature map Input the first convolutional layer and feature extraction layer of the second network in turn to obtain the recommended feature set of the test frame The sixth input subunit is used to input the proposed feature set of the first support frame Proposed feature set for the second support frame And the proposed feature set for the test frame Input the feature aggregation layer of the second network to obtain the aggregated feature set The seventh input subunit is used to convert the aggregate feature set The second convolutional layer of the second network is input to obtain the target detection result.

5. An electronic device, characterized in that: include: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the video target detection method based on feature aggregation as described in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the video object detection method based on feature aggregation as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Intelligent steel SLAG detection method and system based on convolutional neural network

    AU2020102091A4

  • Target detection method and device, storage medium and electronic device

    CN113673604A