Video object detection and domain adaptation method based on motion features and appearance features

By combining appearance and motion characteristics in video object detection and adopting domain adaptation methods, the problems of poor robustness and domain adaptation in the prior art are solved, and more efficient and accurate video object detection is achieved.

CN114863249BActive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210347649.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-05-06
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

The existing video object detection methods are poorly robust in complex and changing video scenarios, making it difficult to effectively extract motion change information in the video, resulting in the detection framework not suitable for tasks such as abnormal behavior detection. At the same time, in the absence of positive sample data, the model performance deteriorates and makes it difficult to achieve domain adaptation.

Method used

Using a video object detection method based on motion characteristics and appearance characteristics, appearance characteristics are extracted through the backbone network, and combined with the motion feature extraction network and appearance feature aggregation network, to integrate appearance and motion information for object detection. In addition, the model is optimized to improve the generalization ability across scenarios through domain adaptation methods of motor space attention and adversarial learning.

Benefits of technology

This method can improve the robustness and accuracy of detection in complex video scenarios, is suitable for detection tasks that require motion information, and significantly improve the performance of the model in the absence of positive sample data, achieving better domain adaptation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863249B_ABST
    Figure CN114863249B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for video target detection and domain adaptation based on motion features and appearance features. The method first extracts motion features and enhanced appearance features based on multiple frames of target frames, then fuses the appearance and motion features to obtain aggregate features and uses them for detecting targets of interest, and thereby automatically captures video frames with targets of interest from the video and determines their locations. The present invention also includes a domain adaptation method for video target detection, which first predicts motion spatial attention with motion features so that aggregate features pay more attention to motion foreground areas that are less relevant to the scene, and then weakens the specific scene information contained in the features by adversarial training of aggregate features, prototyping based on instance features, and aligning features, thereby improving the performance of the video target detection model in scenarios where positive sample training data in the target domain is missing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and pattern recognition, and in particular relates to a video target detection and domain adaptation method based on motion features and appearance features. Background Art

[0002] As multimedia technology becomes more and more common, thanks to the rapid development of computer vision technology and deep learning technology, it is possible to complete some tasks intelligently based on video signals. Intelligent analysis and processing of videos can not only greatly reduce the burden of manpower and save costs, but also achieve more stable and reliable results than manual processing in some tasks.

[0003] At present, some methods for detecting and locating targets of interest based on video input signals first pre-extract areas where target foregrounds may exist based on background difference methods, and then detect targets of interest in a single video frame through subsequent classification. This foreground region extraction method has poor robustness for complex and changeable video scenes, and is prone to low-quality region extraction or missed detection. In addition, most existing methods focus on the extraction of appearance features in videos. These methods do not fully extract the motion change information contained in the video. This problem makes the detection framework unsuitable for tasks such as abnormal behavior detection and automobile exhaust detection that are difficult to effectively complete based on appearance features alone. On the other hand, in some cases, the probability of the target of interest appearing in the video is likely to be relatively low. However, most existing frameworks only use very limited videos (positive sample data) containing targets of interest to train the model, and this approach is likely to cause the model to easily mistakenly detect non-targets of interest in practical applications.

[0004] In addition, in actual detection model application deployment, it is easy to encounter the situation that some video scenes cannot provide videos of the target of interest as positive sample data for detection model training for a period of time. Because different videos usually have large differences in scenes, video quality, etc., the trained model will show serious performance degradation in the scenario where positive sample training data is missing. This problem is similar to the domain adaptation problem in computer vision, and currently there is little attention paid to this problem in video target detection. Summary of the invention

[0005] In response to the problems of video target detection algorithms, the present invention provides a video target detection method based on motion features and appearance features, which can fully extract the appearance and motion information contained in the video and complete the detection and positioning of the target of interest in any video frame.

[0006] To achieve the above purpose, the video target detection method based on motion features and appearance features in the present invention adopts the following technical solutions:

[0007] A first aspect of an embodiment of the present invention provides a video object detection method based on motion features and appearance features, which specifically includes the following steps:

[0008] (1) Convert any input video into a picture set consisting of video frames, detect the target of interest in any target video frame I, extract the target video frame I and its 2p adjacent video frames, totaling 2p+1 video frames, and perform target detection on video frame I;

[0009] (2) Use the backbone network to extract the appearance features of each frame and obtain 2p+1 appearance features;

[0010] (3) Each adjacent frame I n Appearance features A n The appearance feature A of the target video frame I is input into the motion feature extraction network E m To extract the corresponding motion feature M n , while motion feature extraction network E m Output the corresponding pixel-level motion information map f of the predicted motion n ;

[0011] (4) The pixel-level motion information graph f n For each adjacent frame I n Appearance features A n Align to the appearance feature A of the target video frame I to obtain the spatially aligned appearance feature A' n ;

[0012] (5) Using appearance feature aggregation network E aa The appearance features are fused to obtain the appearance feature F a , the appearance feature F a Input appearance feature refinement network R a Perform Hadamard product to obtain the refined appearance feature F' a ;

[0013] (6) Using motion feature aggregation network E am For motion features M n Fusion to obtain motion features F m , the motion feature M n Input motion feature refinement network R m Perform Hadamard product to obtain the refined motion feature F' m

[0014] (7) The refined appearance feature F' obtained in step (5) a The refined motion feature F' obtained in step (6) m Input feature aggregation network Eagg , obtain an aggregate feature F that is consistent with the two input feature sizes agg ;

[0015] (8) Aggregate feature F agg Input the target detection network H to obtain the target bounding box prediction result B and its corresponding classification confidence C;

[0016] (9) Train the video object detection network; test the trained video object detection network. If the maximum value C of the classification confidence C max If it is greater than the preset threshold, it is determined that there is an object of interest in the target video frame I and the target's bounding box prediction result B is output; otherwise, it is determined that there is no object of interest in the frame.

[0017] Furthermore, the backbone network is a ResNet-50, ResNet-101 or VGG-16 network.

[0018] Furthermore, the motion feature extraction network E in step (3) m It can be any current neural network that can achieve the following mapping:

[0019] M n , f n =E m (A,A n )

[0020] Among them, sports information graph n Can be used for the following appearance feature A of a certain adjacent frame n Align to the space of the target frame appearance feature A that needs to be detected:

[0021] A′ n =Alig n (A n , f n )

[0022] The spatial alignment operation Align(·) can be any current mapping that can complete the feature pixel spatial position adjustment operation.

[0023] Furthermore, the process of training the video object detection network is as follows:

[0024] Calculating confidence loss and bounding box regression loss

[0025] The confidence prediction result C is input into the collaborative classification network S to obtain the prediction probability P of whether the target frame I contains the target of interest:

[0026] The label y of the target object of interest is determined based on whether the target frame I actually exists.* And combine the prediction probability P output by the collaborative classification network to calculate the collaborative classification loss L CLS ;

[0027] Using the confidence loss calculated above Bounding Box Regression Loss And the collaborative classification loss L CLS Optimizing video object detection networks.

[0028] Furthermore, the collaborative classification loss L CLS It is a binary classification loss.

[0029] A second aspect of an embodiment of the present invention provides a domain adaptation method for video object detection based on motion features and appearance features, which specifically includes the following steps:

[0030] (1) Refine the motion features into network R m Output motion space attention Att m With the aggregate feature F agg Perform Hadamard product to obtain the optimized aggregation feature F' agg ;

[0031] (2) The aggregated features F in the video object detection network are agg Replaced with the optimized aggregate feature F' agg ; Train the adjusted and optimized video object detection network; and then test the trained video object detection network.

[0032] Preferably, the process of training the adjusted and optimized video object detection network is specifically as follows:

[0033] For the aggregate feature F' agg Perform adversarial domain adaptation and calculate the adversarial learning loss L adv ;

[0034] Using confidence loss Bounding Box Regression Loss Co-classification loss L CLS And adversarial learning loss L adv Training, adjusting and optimizing the video object detection network to obtain a preliminarily trained video object detection network;

[0035] The features used to predict the classification confidence C are completely decomposed into instance-level features in the spatial dimension, and are subdivided into categories including high classification confidence and corresponding to the target of interest tp, high classification confidence but corresponding to the background fp, low classification confidence and corresponding to the background tn, and low classification confidence but corresponding to the target of interest fn according to whether they correspond to the target area of ​​interest and the classification confidence;

[0036] The representative positive prototype features P are constructed by using the instance features with high classification confidence and corresponding to the target of interest tp and the instance features with low classification confidence and corresponding to the background tn. p and negative prototype feature P n ;

[0037] Calculate the loss function L p , this function is currently any function that can be pulled close to P p Distance from instance features in fn and push away P p Function of the distance from the instance feature in fp;

[0038] Calculate the loss function L n , this function is currently any function that can be pulled close to P n Distance from instance features in fp and push away P n Function of the distance from the instance features in fn;

[0039] Based on the preliminarily trained video object detection network, the confidence loss Bounding Box Regression Loss Co-classification loss L CLS , adversarial learning loss L adv , loss function L p And the loss function L n , the model is further fine-tuned and trained to obtain the final video object detection network.

[0040] Preferably, the adversarial domain adaptation is a domain adaptation method based on a gradient reversal layer GRL and a domain classification task.

[0041] A third aspect of an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned video target detection method based on motion features and appearance features and the domain adaptation method of video target detection based on motion features and appearance features.

[0042] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the above-mentioned video target detection method based on motion features and appearance features and the domain adaptation method of video target detection based on motion features and appearance features are implemented.

[0043] The beneficial effects of the domain adaptation method adapted to the aforementioned video target detection disclosed in the present invention are: using motion space attention to make the features extracted by the model pay more attention to the foreground area that is less relevant to the scene, and using adversarial implicit feature alignment and novel explicit instance feature alignment based on positive and negative prototype features to further narrow the cross-scene differences in the features extracted by the model, which can improve the generalization performance of the video target detection network. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flow chart of the video target detection method based on motion features and appearance features in the present invention;

[0045] Figure 2 Schematic diagram of the model structure of the video object detection method based on motion features and appearance features in the present invention;

[0046] Figure 3 It is a flow chart of a domain adaptation method adapted to a video target detection method based on motion features and appearance features in the present invention;

[0047] Figure 4 Schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0048] In order to make the technical solution of the present invention more clearly described, the video object detection and domain adaptation method based on motion features and appearance features in the present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0049] refer to Figure 1 , which is a flow chart of the video target detection method based on motion features and appearance features disclosed in the present invention. Figure 2 , which is a schematic diagram of the model structure of the video target detection method based on motion features and appearance features disclosed in the present invention.

[0050] Input a video (which may not contain the target of interest) and convert it into a set of n video frames consisting of images {I1, I2, ..., I n}, using the video target detection method based on motion features and appearance features disclosed in the present invention for one of the target video frames I i The following steps are required to detect the target of interest:

[0051] Step 1.1: Extract target video frame I i and 2p adjacent video frames, wherein the adjacent 2p video frames in the embodiment of the present invention are the first p adjacent frames and the second p adjacent frames, where p is a self-defined positive integer, and the total is 2p+1 video frames {I i-p , ..., I i-1 , I i , Ii+1 , ..., I i+p}; and perform target detection in video frame I;

[0052] Step 1.2: Input the video frames obtained in step 1.1 into the backbone network E one by one b Extract the appearance features of each frame and obtain 2p+1 appearance features {A i-p , ..., A i-1 , A i , A i+1 , ..., A i+p}, backbone network E in the embodiment of the present invention b You can use networks such as ResNet-50 ResNet-101 or VGG-16 that are commonly used in deep learning;

[0053] Step 1.3: The appearance features A of each adjacent frame j and the target frame appearance feature A i Connect in the channel dimension and input the motion feature extraction network E consisting of convolutional layers and activation layers m , obtain 2p motion features {M i-p , ..., M i-1 , M i+1 , ..., M i+p} and the corresponding 2p pixel-level motion information graphs similar to optical flow {f i-p , ..., f i-1 , f i+1 , ..., f i+p}, where the motion information map is obtained by predicting the corresponding motion features through a single-layer convolution; For target detection in a target video frame I, 2p aligned adjacent frame appearance features and 2p motion features are obtained.

[0054] Preferably, the motion feature extraction network E in step 1.3 m It can be any current neural network that can achieve the following mapping:

[0055] M n , f n =E m (A,A n )

[0056] Among them, sports information graph n Can be used for the following appearance feature A of a certain adjacent frame n Align to the space of the target frame appearance feature A that needs to be detected:

[0057] A n =Align(A n , f n )

[0058] The spatial alignment operation Align(·) can be any current mapping that can complete the feature pixel spatial position adjustment operation.

[0059] Step 1.4: Use each optical flow-like motion information map f j The corresponding adjacent frame appearance features A j To the target frame appearance feature A i Projection is performed to obtain 2p adjacent frame appearance features {A' i-p , ..., A′ i-1 , A' i+1 , ..., A′ i+p};

[0060] Step 1.5: Connect the spatially aligned 2p adjacent frame appearance features obtained in step 1.4 and one target frame appearance feature in the channel dimension and input them into the appearance feature aggregation network E consisting of convolutional layers and activation layers. aa Get the unique appearance feature F of the target frame a , connect all 2p motion features in the channel dimension and input them into the motion feature aggregation network E composed of convolutional layers and activation layers am Get the unique motion feature F of the target frame m ; where E am It can be any neural network that can input 2p features of the same size and output a feature of the same size. aa It can be any neural network that can currently input 2p+1 features of the same size and output a feature of the same size.

[0061] Step 1.6: The unique appearance feature F of the target frame a With motion feature F m Input the appearance feature refinement network R respectively a and motion feature refinement network R m , respectively obtain the refined appearance features F' a With refined movement characteristics F' m The two feature refinement methods are consistent. The specific feature refinement method is to first generate the appearance space attention Att a and motion spatial attention Att m Then, the refined appearance feature F' is obtained by performing Hadamard product with the corresponding features through spatial attention a With the motion feature F' m .

[0062] Preferably, motion space attention Att mThe motion feature F can be input by any current spatial attention module m The appearance space attention Att a F can be input by any current spatial attention module a After prediction is obtained.

[0063] Step 1.7: Refine the appearance feature F' a With refined movement characteristics F' m Connect in the channel dimension and input into the feature aggregation network E consisting of convolutional layers and activation layers agg , obtain the unique aggregate feature F of the target frame and consistent with the two input feature sizes agg ; Feature aggregation network E agg It can be any neural network that can achieve this mapping.

[0064] Step 1.8: Aggregate features F agg Input the target detection network H to obtain the target detection bounding box prediction result B and its corresponding classification confidence C. The target detection network H can be any current target detection network, such as FCOS, RetinaNet and other networks. Figure 1 In the embodiment of the present invention, the target detection network H selected is a one-stage target detection network based on an anchor frame, and its bounding box regression part and classification confidence prediction part are both networks composed of convolutional layers and activation layers.

[0065] Step 1.9 trains the video object detection network; tests the trained video object detection network. If the maximum value C of the classification confidence C is max If it is greater than the preset threshold, it is determined that there is an object of interest in the target video frame I and the target frame prediction result B is output, otherwise it is determined that there is no object of interest in the frame. In the embodiment of the present invention, the preset threshold th=0.75.

[0066] The process of training the video object detection network is as follows:

[0067] Step (a): Denote the target frame as I, and use its interesting target bounding box annotation information and the detection network output to calculate the following confidence loss by referring to the existing target detection method: (Taking a single class of objects of interest as an example) and bounding box regression loss Among them A pos With A neg They represent the index set of positive sample anchor boxes that match the target of interest in the target frame I and the index set of negative sample anchor boxes that do not match the target, respectively. pos =0.999 and w neg =0.001 represents the preset positive and negative sample loss weights, pi With p j They represent the classification confidence of the corresponding positive and negative anchor boxes output by the model, respectively. γ = 3.0 is a parameter that controls the training to pay more attention to samples with poor classification results (the larger γ is, the more attention the training pays to samples with poor classification results). * Is the label of whether the target frame I contains the target of interest, y * 1 means that the target frame contains the target of interest, and the indicator function I(y * = = 1) the output value will be 1, otherwise the indicator function output value is 0. g∈{w, h, x, y} represents the four types of bounding box parameters, w, h, x, y correspond to the width, height, center point horizontal coordinate and center point vertical coordinate respectively. i,g and They respectively represent the predicted value and true label value of the g-type parameter of the positive sample anchor box corresponding to index i.

[0068]

[0069]

[0070]

[0071] Step (b): Input the confidence prediction result C into the collaborative classification network S to obtain the prediction possibility P of whether a single target frame contains the target of interest. The collaborative classification network S can be composed of a convolutional layer, an activation layer and a fully connected layer, and the result P output for a video frame is a scalar.

[0072] Step (c): Determine whether the target frame I actually contains the label y of the target of interest. * Combined with the collaborative classification network output P, ​​the collaborative classification loss L is calculated as follows CLS ;

[0073] L CLS (I) = -y * log(y)-(1-y * )log(1-y)

[0074] Preferably, the collaborative classification loss L CLS It can be any current binary classification loss.

[0075] Step (d): Use the confidence loss calculated above Bounding Box Regression Loss And the collaborative classification loss L CLS Optimizing video object detection networks.

[0076] refer to Figure 3 , which is a schematic diagram of a domain adaptation method adapted to the aforementioned video target detection method in the present invention.

[0077] The domain adaptation method disclosed in the present invention adapted to the aforementioned video object detection can be described in more detail as the following steps:

[0078] Step 2.1: Refine the network R using motion features m The motion spatial attention Att generated in the intermediate step m The aggregated feature F unique to the target frame agg Perform Hadamard product to optimize the aggregate feature F' agg Pay more attention to the motion foreground area that is less relevant to the scene.

[0079] Step 2.2: Aggregate features F in the video object detection network agg Replaced with the optimized aggregate feature F' ag g; training the adjusted and optimized video object detection network; and then testing the trained video object detection network.

[0080] The specific process of training the adjusted and optimized video object detection network is as follows:

[0081] The better aggregated feature F' of the target frame I obtained in step 2.1 agg Perform adversarial cross-scene feature alignment based on the gradient reversal layer GRL. Aggregate feature F' agg The discriminator D composed of the fully connected layer after the GRL inverted gradient predicts the category of the scene to which all feature pixels belong, and calculates the following adversarial learning loss L based on the actual scene category adv . Among them, W is F' agg The height and width of , q represents the number of source domain scenes with no missing training data (source domain scene categories are coded from 1 to g, and scene categories with missing data are coded as 0), Indicates the classification label of the scene to which the target frame belongs (if the target frame belongs to the scene of encoding j, then T (j) =1 and all other values ​​in T are 0);

[0082]

[0083] Using adversarial learning loss L adv And the confidence loss Bounding Box Regression Loss Co-classification loss L CLS Training step 2.1 adjusted video object detection network;

[0084] The goal of this step is to use the training data to further fine-tune the video object detection framework obtained from the preliminary training. The tth round of fine-tuning training consists of the following steps:

[0085] First, the target frame corresponding feature F used to predict the classification confidence C in the target detection network H of the framework c In the spatial dimension, it is completely decomposed into H×W local region instance-level vector features {V k |k∈{1, 2, ..., H×W}};

[0086] According to the detection confidence c corresponding to each instance feature k And whether it corresponds to the true label y′ of the target area of ​​interest k (1 corresponds to the object of interest, 0 corresponds to the background) Determine whether each instance feature is correctly classified as foreground or background. k >0.5, the instance is predicted to contain an object of interest, otherwise the instance is predicted to be a background category;

[0087] The positive and negative prototype features of the tth round are constructed using the instance features corresponding to the correctly classified target of interest and the instance features corresponding to the background area. The construction method can be any currently feasible prototype construction method. The positive and negative prototype features of the tth round can be obtained by sliding average. Specifically, the temporary positive and negative prototypes of the tth round are obtained by averaging the correctly classified positive and negative instance features. and Then the positive and negative prototypes of the tth round are composed of the prototypes of the previous round. With the current round of temporary prototype It is calculated as follows, where α is the adjusted cosine similarity between the previous round prototype and the current round temporary prototype of the same category;

[0088]

[0089]

[0090] By calculating the positive sample prototype loss L as follows p To explicitly reduce the distance between the instance features corresponding to the misclassified target region of interest and the positive prototype features, and to explicitly expand the distance between the instance features corresponding to the misclassified target region of interest and the negative prototype features. Where fp and fn represent the misclassified instance feature index sets, |fp| and |fn| represent the number of two instance features, k is the index of the instance feature, and λ n =0.1 is the weight of the loss function calculated from the instance features of the misclassified background area;

[0091]

[0092] The negative sample prototype loss L is calculated as follows nTo explicitly reduce the distance between the instance features corresponding to the misclassified background area and the negative prototype features, and to explicitly expand the distance between the instance features corresponding to the misclassified background area and the positive prototype features;

[0093]

[0094] Calculate the aforementioned adversarial learning loss L adv And the confidence loss Bounding Box Regression Loss Co-classification loss L CLS , with prototype loss L p With L n Further fine-tuning and training of the video object detection network obtained through preliminary training is implemented to obtain a video object detection framework with improved performance in scenarios where positive sample training data is missing.

[0095] The video target detection method based on motion features and appearance features disclosed in the present invention is applied to a self-built multi-scenario automobile exhaust detection task to obtain experimental test data as shown in Table 1.

[0096] Table 1

[0097]

[0098] The scene 5 in the above experiment is set as the target domain (positive sample training data of the target of interest, automobile exhaust, is missing in the training), and the other four scenes are set as the source domain (with complete training data). As shown in Table 2 below, the target detection index of the video target detection method based on motion features and appearance features disclosed in the present invention in the target domain scene 5 will be severely attenuated, while the domain adaptation method disclosed in the present invention can significantly improve the performance of the video target detection method based on motion features and appearance features in the target domain scene 5.

[0099] Table 2

[0100]

[0101] Corresponding to the above-mentioned embodiments of the video object detection and domain adaptation method based on motion features and appearance features, the present invention also provides embodiments of the video object detection and domain adaptation device based on motion features and appearance features.

[0102] See also Figure 4 An embodiment of the present invention provides a video target detection and domain adaptation device based on motion features and appearance features, including one or more processors for implementing the video target detection and domain adaptation method based on motion features and appearance features in the above embodiment.

[0103] The embodiments of the video object detection and domain adaptation device based on motion features and appearance features of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 4 As shown, it is a hardware structure diagram of any device with data processing capability where the video object detection and domain adaptation device based on motion features and appearance features of the present invention is located. Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0104] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0105] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0106] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the video object detection and domain adaptation method based on motion features and appearance features in the above embodiment is implemented.

[0107] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A video object detection method based on motion features and appearance features, characterized in that: The specific steps include: (1) Convert any input video into a picture set consisting of video frames, detect the target of interest in any target video frame I, extract the target video frame I and its 2p adjacent video frames, totaling 2p+1 video frames, and perform target detection on video frame I; (2) Use the backbone network to extract the appearance features of each frame and obtain 2p+1 appearance features; (3) Each adjacent frame I n Appearance features A n The appearance feature A of the target video frame I is input into the motion feature extraction network E m To extract the corresponding motion feature M n , while motion feature extraction network E m Output the corresponding pixel-level motion information map f of the predicted motion n ; (4) The pixel-level motion information graph f n For each adjacent frame I n Appearance features A n Align to the appearance feature A of the target video frame I to obtain the spatially aligned appearance feature A' n ; (5) Using appearance feature aggregation network E aa The appearance features are fused to obtain the appearance feature F a , the appearance feature F a Input appearance feature refinement network R a Perform Hadamard product to obtain the refined appearance feature F' a ; (6) Using motion feature aggregation network E am For motion features M n Fusion to obtain motion features F m , the motion feature M n Input motion feature refinement network R m Perform Hadamard product to obtain the refined motion feature F' m (7) The refined appearance feature F' obtained in step (5) a The refined motion feature F' obtained in step (6) m Input feature aggregation network E agg , obtain an aggregate feature F that is consistent with the two input feature sizes agg ; (8) Aggregate feature F agg Input the target detection network H to obtain the target bounding box prediction result B and its corresponding classification confidence C; (9) Train the video object detection network; test the trained video object detection network. If the maximum value C of the classification confidence C max If it is greater than the preset threshold, it is determined that there is an object of interest in the target video frame I and the target's border prediction result B is output; otherwise, it is determined that there is no object of interest in the frame.

2. The video target detection method based on motion features and appearance features according to claim 1 is characterized in that: The backbone network is a ResNet-50, ResNet-101 or VGG-16 network.

3. The video target detection method based on motion features and appearance features according to claim 1 is characterized in that: The motion feature extraction network E in step (3) m is any current neural network that can implement the following mapping: M n ,f n =E m (CHALLENGE ACCEPTED n ) Among them, sports information graph n Can be used for the following appearance feature A of a certain adjacent frame n Align to the space of the target frame appearance feature A that needs to be detected: A’ n =Align(A n ,f n ) The spatial alignment operation Align(·) is any current mapping that can complete the feature pixel spatial position adjustment operation.

4. The video target detection method based on motion features and appearance features according to claim 1 is characterized in that: The process of training the video object detection network is as follows: Calculating confidence loss and bounding box regression loss Input the confidence prediction result C into the collaborative classification network S to obtain the prediction probability P of whether the target frame I contains the target of interest; The label y of the target object of interest is determined based on whether the target frame I actually exists. * And combine the prediction probability P output by the collaborative classification network to calculate the collaborative classification loss L CLS ; Using the confidence loss calculated above Bounding Box Regression Loss And the collaborative classification loss L CLS Optimizing video object detection networks.

5. The video target detection method based on motion features and appearance features according to claim 4 is characterized in that: The collaborative classification loss L CLS It is a binary classification loss.

6. A domain adaptation method applicable to the video object detection method based on motion features and appearance features as described in any one of claims 1 to 5, characterized in that: The specific steps include: (1) Refine the motion features into network R m Output motion space attention Att m With the aggregate feature F agg Perform Hadamard product to obtain the optimized aggregation feature F' agg ; (2) The aggregated features F in the video object detection network are agg Replaced with the optimized aggregate feature F' agg ; Train the adjusted and optimized video object detection network; and then test the trained video object detection network.

7. The domain adaptation method according to claim 6, characterized in that: The specific process of training the adjusted and optimized video object detection network is as follows: For the aggregate feature F' agg Perform adversarial domain adaptation and calculate the adversarial learning loss L adv ; Using confidence loss Bounding Box Regression Loss Co-classification loss L CLS And adversarial learning loss L adv Training, adjusting and optimizing the video object detection network to obtain a preliminarily trained video object detection network; The features used to predict the classification confidence C are completely decomposed into instance-level features in the spatial dimension, and are subdivided into categories including high classification confidence and corresponding to the target of interest tp, high classification confidence but corresponding to the background fp, low classification confidence and corresponding to the background tn, and low classification confidence but corresponding to the target of interest fn according to whether they correspond to the target area of ​​interest and the classification confidence; The representative positive prototype features P are constructed by using the instance features with high classification confidence and corresponding to the target of interest tp and the instance features with low classification confidence and corresponding to the background tn. p and negative prototype feature P n ; Calculate the loss function L p , this function is currently any function that can be pulled close to P p Distance from instance features in fn and push away P p Function of the distance from the instance feature in fp; Calculate the loss function L n , this function is currently any function that can be pulled close to P n Distance from instance features in fp and push away P n Function of the distance from the instance features in fn; Based on the preliminarily trained video object detection network, the confidence loss Bounding Box Regression Loss Co-classification loss L CLS , adversarial learning loss L adv , loss function L p And the loss function L n , the model is further fine-tuned and trained to obtain the final video object detection network.

8. The domain adaptation method for video object detection based on motion features and appearance features according to claim 7, characterized in that: The adversarial domain adaptation is a domain adaptation method based on a gradient reversal layer GRL and a domain classification task.

9. An electronic device comprising a memory and a processor, wherein: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the video target detection method based on motion features and appearance features described in any one of claims 1-5 and the domain adaptation method described in any one of claims 6-8.

10. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the video target detection method based on motion features and appearance features as described in any one of claims 1 to 5 and the domain adaptation method as described in any one of claims 6 to 8 are implemented.

Citation Information

Patent Citations

  • Motor vehicle exhaust smoke video identification system

    CN103714363A

  • Road traffic accident detection method based on visual attention mechanism and ConvLSTM network

    CN112084928A