An intelligent scene recognition method and device based on horizontal and vertical pooling

Through the intelligent scene recognition method based on vertical and horizontal pooling, the video frames to be recognized are scaled and convolutional neural network processing, and combined with the typical classification results of the binary classifier, the problems of poor applicability and low accuracy of video frame types in the prior art are solved, and efficient recognition of video frames with different aspect ratios is achieved.

CN114898240BActive Publication Date: 2025-05-30COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110104597.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-26
Publication Date
2025-05-30
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

Existing video frames or image scene recognition methods have poor applicability and low recognition accuracy. Especially when processing video frames with different aspect ratios, methods based on deep learning algorithms can only recognize a single aspect ratio, while methods based on convolutional neural networks have reduced recognition accuracy due to the need to adjust the size.

Method used

Using an intelligent landscape recognition method based on vertical and horizontal pooling, the video frames to be identified are scaled according to their aspect ratio and preset height, and the trained convolutional neural network is used to convolution, vertical pooling and horizontal pooling of the scaled video frames to be convolutional, vertical pooling and horizontal pooling to obtain a six-dimensional vector, and combined with the typical classification results of the binary classifier, the final landscape recognition is performed.

Benefits of technology

It realizes efficient scene-type recognition of video frames with different aspect ratios, improves recognition accuracy and strong applicability, and avoids the defects of low recognition accuracy and unrecognized recognition in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898240B_ABST
    Figure CN114898240B_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent scene recognition method and device based on vertical and horizontal pooling, belonging to the technical field of video processing, and solves the problems that the prior art can only recognize the scene of video frames with a single aspect ratio and has a low recognition accuracy. The method includes: scaling the video frame to be recognized according to the aspect ratio and a preset height; using a trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling on the scaled video frame to be recognized in sequence, obtaining a six-dimensional vector corresponding to the scene type, and obtaining a pre-classification result of the scene of the video frame to be recognized according to the six-dimensional vector; obtaining the distribution characteristics of the six-dimensional vector, and using a trained binary classifier to perform scene typicality classification on the video frame to be recognized based on the distribution characteristics; and obtaining the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre-classification result. This method can recognize the scenes of video frames with different aspect ratios and has a high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to an intelligent scene recognition method and device based on vertical and horizontal pooling. Background Art

[0002] In the era of high-speed information output, the social demand for video products is increasing day by day. Whether in the traditional media or new media fields, the requirements for video editing efficiency are constantly rising, and the efficiency of manual editing has begun to lag behind the pace of video product production. In recent years, artificial intelligence has developed rapidly, and there has gradually emerged a trend of automation and intelligence in video editing.

[0003] During the process of intelligent video editing, the machine needs to learn the methods of how editors identify objects, analyze behaviors, understand pictures, and comprehend video content, as well as the skills of constructing video scenes and organizing audio-visual languages using video elements. Scene is one of the important features of video materials. During the process of organizing audio-visual languages, scene is an important reference factor for shot connection. Efficient and accurate recognition of scenes is one of the challenges faced by intelligent video editing technology.

[0004] Scene refers to the difference in the range size of the subject presented in the picture due to the different distances between the camera and the subject. According to the definition of scene and the industry's habit of classifying scenes, the scenes of videos are usually divided into close-up, medium close-up, medium shot, panoramic shot, and long shot. During the production of converged media, video materials have various disparate aspect ratios, such as standard-definition videos with an aspect ratio of 4:3, high-definition videos with an aspect ratio of 16:9, movie widescreen with an aspect ratio greater than 2:1, and the currently popular 1:2 aspect ratio for mobile live broadcasts and short videos. This requires that the scene recognition algorithm can recognize the scenes of video frames with different aspect ratios. In the prior art, for the method of recognizing the scenes of video frames, one is to recognize the scenes of video frames with a single aspect ratio based on a deep learning algorithm; the other is to adjust video frames with different aspect ratios to a unified size and input them into a convolutional neural network for scene recognition.

[0005] The prior art has at least the following defects: First, the scene recognition method based on a deep learning algorithm can only recognize video frames or pictures with a single aspect ratio, and has poor applicability; second, for the scene recognition method based on a convolutional neural network, it is necessary to adjust video frames with different aspect ratios to a unified size, but due to factors such as the squeezing and deformation of video frames, the accuracy of scene recognition will be reduced. In addition, if video frames or images with different aspect ratios are directly input into the convolutional neural network model, it will cause the number of neurons at the fully connected layer to not match and thus unable to perform recognition. Summary of the Invention

[0006] In view of the above analysis, the present invention aims to provide an intelligent scene recognition method and device based on vertical and horizontal pooling to solve the problems of poor applicability and low recognition accuracy of existing video frame or image scene recognition methods.

[0007] On the one hand, the present invention provides an intelligent scene recognition method based on vertical and horizontal pooling, including the following steps:

[0008] Scale the video frame to be recognized according to the aspect ratio and a preset height, where the aspect ratio is the aspect ratio of the video frame to be recognized;

[0009] Use the trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling on the scaled video frame to be recognized in sequence, obtain a six-dimensional vector corresponding to the scene type, and obtain the preliminary classification result of the scene of the video frame to be recognized according to the six-dimensional vector;

[0010] Obtain the distribution characteristics of the six-dimensional vector, and use the trained binary classifier to perform scene typicality classification on the video frame to be recognized based on the distribution characteristics;

[0011] Obtain the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the preliminary classification result of the scene.

[0012] Further, the step of using the trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling on the scaled video frame to be recognized in sequence to obtain a six-dimensional vector corresponding to the scene type specifically includes:

[0013] Use the trained convolutional neural network to perform convolution on the scaled video frame to be recognized to obtain the initial feature map of the video frame to be recognized;

[0014] Extract the sub-feature maps corresponding to each channel from the initial feature map, and the number of sub-feature maps is the same as the number of channels of the initial feature map;

[0015] Perform max pooling and average pooling on each column vector of the sub-feature map to obtain the first sub-feature map, and traverse each sub-feature map to obtain the first sub-feature map corresponding to each channel;

[0016] Perform global max pooling on each horizontal vector of the first sub-feature map to obtain the second sub-feature map, and traverse each first sub-feature map to obtain the second sub-feature map corresponding to each channel;

[0017] Obtain the feature map of the video frame to be recognized according to the second sub-feature map corresponding to each channel, and perform fully connected processing and normalization processing in sequence to obtain a six-dimensional vector corresponding to the scene type.

[0018] Further, the scene pre-classification result of the video frame to be recognized is obtained according to the six-dimensional vector in the following manner:

[0019] R CNN = argmaxO i , i ∈ {1, 2, 3, 4, 5, 6},

[0020] i

[0021] where R CNN represents the scene pre-classification result of the video frame to be recognized. When R CNN is 1, 2, 3, 4, 5, and 6 respectively, the corresponding scene pre-classification results of the video frame to be recognized are close-up, medium shot, mid-shot, panorama, long shot, and others; O i represents the eigenvalue corresponding to the i-th dimension in the six-dimensional vector.

[0022] Further, the step of performing max pooling and average pooling on each column vector of the sub-feature map to obtain the first sub-feature map includes:

[0023] Performing max pooling of the target feature size on each column vector of the sub-feature map to correspondingly obtain the first column vector of the target feature size, and the dimension of the first column vector is less than the height of the initial feature map;

[0024] Performing average pooling on each column vector of the sub-feature map to obtain the non-position eigenvalue corresponding to each column vector;

[0025] Combining each first column vector and the corresponding non-position eigenvalue to obtain a second column vector, and the dimension of the second column vector is 1 greater than the dimension of the first column vector;

[0026] Obtaining the corresponding first sub-feature map based on each second column vector.

[0027] Further, when performing max pooling on each column vector of the sub-feature map, the corresponding pooling stride and pooling field are:

[0028]

[0029] kernel size = h - (O size - 1) × stride,

[0030] where stride represents the pooling stride, kernel size represents the pooling field, h represents the height of the initial feature map, and O size represents the target feature size of max pooling.

[0031] Further, the distribution characteristics of the six-dimensional vector include a first eigenvalue, a second eigenvalue, and a third feature; the first eigenvalue is the largest eigenvalue in the six-dimensional vector, the second eigenvalue is the ratio of the second-largest eigenvalue to the largest eigenvalue in the six-dimensional vector, and the third eigenvalue is the entropy value of the six-dimensional vector: O i represents the eigenvalue corresponding to the i-th dimension in the six-dimensional vector.

[0032] Further, the scene recognition result of the video frame to be recognized is obtained based on the scene typicality classification result and the scene pre-classification result, specifically including:

[0033] If the scene typicality classification result of the video frame to be recognized is a typical scene, then the scene pre-classification result is used as the scene recognition result of the video frame to be recognized;

[0034] If the scene typicality classification result of the video frame to be recognized is an atypical scene, and the first scene corresponding to the largest eigenvalue in the six-dimensional vector and the second scene corresponding to the second-largest eigenvalue in the six-dimensional vector are adjacent scenes, then the first scene and the second scene are used as the scene recognition result of the video frame to be recognized.

[0035] Further, the trained binary classifier is obtained by training based on a linear logistic regression model.

[0036] On the other hand, the present invention provides an intelligent scene recognition device based on vertical and horizontal pooling, including:

[0037] A preprocessing module that scales the video frame to be recognized according to the aspect ratio and a preset height, and the aspect ratio is the aspect ratio of the video frame to be recognized;

[0038] A scene pre-classification module that uses the trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling processing on the scaled video frame to be recognized in sequence, obtains a six-dimensional vector corresponding to the scene type, and obtains the scene pre-classification result of the video frame to be recognized according to the six-dimensional vector;

[0039] A typicality classification module that obtains the distribution characteristics of the six-dimensional vector, and uses the trained binary classifier to perform scene typicality classification on the video frame to be recognized based on the distribution characteristics;

[0040] A scene recognition module that obtains the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre-classification result.

[0041] Further, the scene pre-classification module is specifically used for:

[0042] Perform convolution processing on the scaled video frame to be recognized using the trained convolutional neural network to obtain the initial feature map of the video frame to be recognized;

[0043] Extract the sub-feature map corresponding to each channel from the initial feature map, and the number of the sub-feature maps is the same as the number of channels of the initial feature map;

[0044] Perform max pooling and average pooling processing on each column vector of the sub-feature map to obtain the first sub-feature map, and traverse each sub-feature map to obtain the first sub-feature map corresponding to each channel;

[0045] Perform global max pooling processing on each horizontal vector of the first sub-feature map to obtain the second sub-feature map, and traverse each first sub-feature map to obtain the second sub-feature map corresponding to each channel;

[0046] Obtain the feature map of the video frame to be recognized according to the second sub-feature map corresponding to each channel, and perform fully connected processing and normalization processing in sequence to obtain a six-dimensional vector corresponding to the scene type.

[0047] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:

[0048] 1. The intelligent scene recognition method and device based on vertical and horizontal pooling proposed by the present invention scale the video frame to be recognized according to the aspect ratio and preset height of the video frame to be recognized, so as to be able to realize the scene recognition of video frames or images with different aspect ratios, with strong applicability, avoiding the limitation that the scene recognition method based on deep learning can only recognize video frames or images with a single aspect ratio, and the defects of low recognition accuracy or inability to recognize of the scene recognition method based on convolutional neural network.

[0049] 2. The intelligent scene recognition method and device based on vertical and horizontal pooling proposed by the present invention adjust the proportions of the vertical position features, horizontal position features and other non-position features on which the scene classification is based to the optimal state by performing vertical pooling and horizontal pooling processing on the feature map of the video frame to be recognized respectively, thereby greatly improving the recognition accuracy of the video frame scene.

[0050] 3. The intelligent scene recognition method and device based on vertical and horizontal pooling proposed by the present invention first perform pre-classification on the video frame to be recognized through the convolutional neural network model of vertical and horizontal pooling to obtain a preliminary classification result, and use a binary classifier based on a linear regression model to classify the scene typicality of the video frame to be recognized, and then obtain the final scene recognition result of the video frame to be recognized according to the preliminary scene classification result and the scene typicality classification result, thereby further improving the accuracy of the scene recognition result.

[0051] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the following specification. Moreover, some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained from the content specifically pointed out in the specification and the accompanying drawings. Description of the Drawings

[0052] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0053] Figure 1 It is a schematic diagram of scaling a video frame to be recognized according to the aspect ratio and a preset height in an embodiment of the present invention;

[0054] Figure 2 It is a flowchart of an intelligent scene classification recognition method based on horizontal and vertical pooling in an embodiment of the present invention;

[0055] Figure 3 It is a schematic diagram of a sub-feature map extracted in an embodiment of the present invention;

[0056] Figure 4 It is a flowchart of vertically pooling the sub-feature map in an embodiment of the present invention;

[0057] Figure 5 It is a schematic diagram of a first sub-feature map obtained in an embodiment of the present invention;

[0058] Figure 6 It is a schematic diagram of a second sub-feature map obtained in an embodiment of the present invention;

[0059] Figure 7 It is a schematic diagram of an intelligent scene classification recognition device based on horizontal and vertical pooling in an embodiment of the present invention.

[0060] Reference Signs:

[0061] 110 - Preprocessing Module; 120 - Scene Pre-classification Module; 130 - Typicality Classification Module; 140 - Scene Classification Recognition Module. Detailed Embodiments

[0062] The following will specifically describe the preferred embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0063] A specific embodiment of the present invention discloses an intelligent scene classification recognition method based on horizontal and vertical pooling.

[0064] When performing the pre-classification of the scene type of the video frame to be recognized, it is first necessary to train the convolutional neural network model for scene type pre-classification based on the video frame data sample set. The convolutional neural network model is a model based on vertical and horizontal pooling, that is, the pooling layer of the convolutional neural network model includes two processing processes of performing vertical pooling and horizontal pooling on the video frame data.

[0065] Specifically, the video frame data sample set includes video frames with different aspect ratios and the scene types of the video frames. The scene types include "close-up", "medium close-up", "mid-shot", "panorama", "long shot", and "other". Preferably, considering that the convolutional neural network model can only process and recognize video frames of one height, therefore, as Figure 1 shown, before training the convolutional neural network model, it is necessary to uniformly scale the video frames with different aspect ratios according to the aspect ratio of each video frame and the preset height, so as to obtain video frames with the same height and unchanged aspect ratio, ensuring that they can all be processed and recognized by the same convolutional neural network, and at the same time, the video frames will not be deformed after scaling, thus causing information loss or errors, etc. Preferably, the preset height is set to the classic empirical value of 224 pixels of the convolutional neural network model.

[0066] Specifically, the video frames with the same height are used as the input of the convolutional neural network model, and the corresponding scene type of the video frame is used as the output, and the parameters such as the convolutional layer, vertical pooling, horizontal pooling, and connection layer in the convolutional neural network model are trained, so as to obtain the trained convolutional neural network.

[0067] Preferably, as Figure 2 shown, when performing scene type classification on the video frame to be recognized, it specifically includes:

[0068] S110. Scale the video frame to be recognized according to the aspect ratio and the preset height. Among them, the aspect ratio is the aspect ratio of the video frame to be recognized, and the preset height is the classic empirical value of 224 pixels.

[0069] S120. Use the trained convolutional neural network to perform processes such as convolution, vertical pooling, and horizontal pooling on the scaled video frame to be recognized in sequence, obtain a six-dimensional vector corresponding to the scene type, and obtain the pre-classification result of the scene type of the video frame to be recognized according to the six-dimensional vector.

[0070] Specifically, within the range where the scene type can be defined, the scene type is divided into six types, and the six eigenvalue in the six-dimensional vector respectively represent the six scene types. Among them, the scene type corresponding to the maximum eigenvalue is the pre-classification result of the scene type of the video frame to be recognized.

[0071] S130. Obtain the distribution characteristics of the six-dimensional vector, and perform scene type typicality classification on the video frame to be recognized based on the distribution characteristics by using the trained binary classifier.

[0072] Specifically, the distribution feature represents the distribution feature of the scene type represented by each eigenvalue in the six-dimensional vector.

[0073] The trained binary classifier is obtained by training based on a linear logistic regression model. Specifically, it includes: selecting a video frame data sample set for training the binary classifier, and this video frame data sample set is a sample set that has not participated in the training of the convolutional neural network model; this sample set includes video frames and the corresponding true scene types of the video frames. Input each video frame in this sample set into the trained convolutional neural network to obtain the predicted result of the scene type of each video frame, and compare the scene type result of the video frame with its true scene type. If they are consistent, label the scene typicality of this video frame as a typical scene; if they are inconsistent, label the scene typicality of this video frame as an atypical scene. Use the distribution feature corresponding to the video frame as the input of the binary classifier, and use the scene typicality as the output of the binary classifier to train the binary classification, so as to use the trained binary classifier to determine whether the video frame to be recognized is a typical scene or an atypical scene.

[0074] S140. Obtain the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre-classification result.

[0075] Among them, a typical scene refers to a scene that can be clearly defined within the definable range of the scene. An atypical scene refers to a scene that cannot be clearly defined. Exemplarily, when the scene of a video frame is in an intermediate state between two adjacent scenes (that is, the ratio of the eigenvalues corresponding to the two scenes is close to 1), it is very difficult to accurately judge which scene it is, so its scene cannot be clearly defined. However, the convolutional neural network will also judge the scene type of this video frame as the scene type corresponding to the slightly larger eigenvalue, resulting in a misjudgment. There are also some video frames with special compositions, and the scene types of these video frames cannot be determined. Therefore, these video frames are classified as atypical scenes. Therefore, combining the scene typicality classification result and the scene pre-classification result to judge the scene classification of the video frame to be recognized can greatly improve the accuracy of scene recognition.

[0076] Preferably, in S120, the steps of using the trained convolutional neural network to perform convolution, vertical pooling, horizontal pooling, etc. on the scaled video frame to be recognized to obtain a six-dimensional vector corresponding to the scene type specifically include:

[0077] S1201. Perform convolution processing on the scaled video frame to be recognized using the trained convolutional neural network to obtain the initial feature map of the video frame to be recognized. Specifically, input the video frame to be recognized into the convolutional layer of the trained convolutional neural network to obtain the initial feature map F, the height of which is h, the width is w, and the number of channels is c. Among them, the number of channels is the color channel number of the initial feature map. Exemplarily, if the color mode of the initial feature map is the RGB mode, the corresponding number of channels is 3.

[0078] S1202. Extract the sub-feature map F corresponding to each channel from the initial feature map x , (x = 1, 2... c). Exemplarily, the extracted sub-feature map is as Figure 3 shown. Among them, the number of sub-feature maps is the same as the number of channels of the initial feature map.

[0079] S1203. Divide the matrix corresponding to the sub-feature map F x into w column vectors a of h×1 dimension i , (i = 1, 2, 3... w), and perform max pooling and average pooling processing on each column vector a x of the sub-feature map F i respectively to obtain the first sub-feature map, and traverse each of the sub-feature maps to obtain the first sub-feature map corresponding to each channel. Specifically, as Figure 4 shown:

[0080] S12031. Perform max pooling processing with the target feature size O x on each column vector a i of the sub-feature map F size respectively, and correspondingly obtain the first column vector a size with the target feature size O V1,i , (O size <h), that is, the dimension of this first column vector a V1,i is O size . Preferably, the value of the target feature size is determined by being able to obtain sufficient vertical direction feature information. Exemplarily, when the whole body of a person appears in the video frame, the target feature size is set to 3 to meet the above conditions.

[0081] Preferably, when performing max pooling processing, the corresponding pooling stride and pooling domain are respectively:

[0082]

[0083] kernel size =h - (O size - 1)×stride,

[0084] where stride represents the pooling stride, kernelsize Denote the pooling domain, h represents the height of the initial feature map, and O size represents the target feature size of the max pooling.

[0085] S12032. For each column vector a x of the sub-feature map F i perform average pooling processing respectively to obtain the non-position feature value a corresponding to each column vector V2,i .

[0086] S12033. Combine each first column vector a V1,i and the corresponding non-position feature value a V2,i to obtain a second column vector a V,i , and the dimension of this second column vector is 1 greater than that of the first column vector.

[0087] Based on each of the second column vectors, obtain a matrix of (O size + 1)×w dimensions, and then obtain the corresponding first sub-feature map F 1,x , as Figure 5 shown.

[0088] S1204. For each horizontal vector b 1,x of the matrix corresponding to the first sub-feature map F j , (j = 1, 2, 3... O size + 1) perform global max pooling processing respectively to obtain the corresponding feature value b H,j , traverse each horizontal vector to obtain (O size + 1) feature values, thus forming a (O size + 1)×1-dimensional column vector, and obtain the corresponding second sub-feature map F 2,x according to this column vector. Exemplarily, the obtained second sub-feature map is as Figure 6 shown; traverse each first sub-feature map to obtain the second sub-feature map corresponding to each channel.

[0089] S1205. Obtain the feature map of the video frame to be recognized according to the second sub-feature map corresponding to each channel. Specifically, combine the column vectors of each second sub-feature map corresponding to each channel to obtain a matrix of (O size + 1)×c, and then obtain the corresponding feature map of the video frame to be recognized; and perform fully connected processing and normalization processing on this feature map in sequence to obtain a six-dimensional vector O corresponding to the scene type. Preferably, use the Softmax activation function to perform normalization processing on the feature map.

[0090] Preferably, obtain the pre-classification result of the scene of the video frame to be recognized through the following method:

[0091]

[0092] Among them, R CNN represents the pre-classification result of the scene type of the video frame to be recognized. When R CNN is 1, 2, 3, 4, 5, or 6 respectively, the corresponding pre-classification results of the scene type of the video frame to be recognized are close-up, medium shot, medium view, panoramic view, long shot, and others; O i represents the eigenvalue corresponding to the i-th dimension in the six-dimensional vector. Exemplarily, if the eigenvalue corresponding to the 2nd dimension in the six-dimensional vector is 0.98 and it is the largest eigenvalue, then R CNN is 2, and the corresponding pre-classification result of the scene type is medium shot.

[0093] Preferably, in S130, the distribution characteristics of the six-dimensional vector include the first eigenvalue, the second eigenvalue, and the third characteristic. Among them, the first eigenvalue is the largest eigenvalue in the six-dimensional vector, that is, the maximum value in the six-dimensional vector; the second eigenvalue is the ratio of the second largest eigenvalue to the largest eigenvalue in the six-dimensional vector, and the second largest eigenvalue is the second largest value in the six-dimensional vector; the third eigenvalue is the entropy value of the six-dimensional vector: O i represents the eigenvalue corresponding to the i-th dimension in the six-dimensional vector.

[0094] Preferably, based on the distribution characteristics, the trained binary classifier is used to perform the scene typicality classification on the video frame to be recognized, specifically including: based on the first eigenvalue, the second eigenvalue, and the third eigenvalue, the trained binary classifier is used to classify the scene typicality of the video frame to be recognized. Considering that when the scene of the video frame is in the intermediate state between two adjacent scene types, it is difficult to accurately determine which scene type it belongs to. Therefore, during the training process of the binary classifier, the scene typicality of this type of video frame is classified as an atypical scene. During subsequent judgment, both of these two adjacent scene types can be used as the scene recognition result of this video frame, thereby further improving the accuracy of scene recognition. In addition, for some video frames with special compositions, it is impossible to determine the scene type of this type of video frame. Then, during the training process of the binary classifier, the scene typicality of this type of video frame is classified as an atypical scene, and the scene recognition result of this video frame is set to be empty.

[0095] Specifically, obtaining the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre-classification result includes:

[0096] If the video frame to be recognized is a typical scene, then the scene pre-classification result is used as the scene recognition result of the video frame to be recognized.

[0097] If the video frame to be recognized is of an atypical scene type, and the first scene type corresponding to the maximum eigenvalue in the six-dimensional vector and the second scene type corresponding to the second largest eigenvalue in the six-dimensional vector are adjacent scene types, then the first scene type and the second scene type are used as the scene recognition result of the video frame to be recognized. Exemplarily, if the first scene type corresponding to the maximum eigenvalue in the six-dimensional vector is a close-up and the second scene type corresponding to the second largest eigenvalue is a medium shot, then the scene recognition result of the video frame to be recognized is a close-up and a medium shot.

[0098] When recognizing the scene types of video frames with aspect ratios of 1:2, 4:3, 16:9, and >2:1 respectively based on the above method embodiment, the accuracy rates can reach 96.57%, 96.72%, 97.93%, and 98.1% respectively, and the recognition speed can reach 50 - 80 frames per second.

[0099] Another embodiment of the present invention discloses an intelligent scene recognition device based on vertical and horizontal pooling.

[0100] Since the device embodiment is based on the same principle as the above method embodiment, the details can be referred to the method embodiment and will not be elaborated here.

[0101] Specifically, as Figure 7 shown, the device includes:

[0102] A preprocessing module 110 that scales the video frame to be recognized according to the aspect ratio and a preset height, where the aspect ratio is the aspect ratio of the video frame to be recognized.

[0103] A scene pre-classification module 120 that uses a trained convolutional neural network to perform convolution, vertical pooling, horizontal pooling, etc. on the scaled video frame to be recognized in sequence, obtains a six-dimensional vector corresponding to the scene type, and obtains the scene pre-classification result of the video frame to be recognized according to the six-dimensional vector.

[0104] A typicality classification module 130 that obtains the distribution characteristics of the six-dimensional vector and uses a trained binary classifier to perform scene typicality classification on the video frame to be recognized based on the distribution characteristics.

[0105] A scene recognition module 140 that obtains the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre-classification result.

[0106] Preferably, the scene pre-classification module 120 is specifically used for:

[0107] Performing convolution processing on the scaled video frame to be recognized using a trained convolutional neural network to obtain the initial feature map of the video frame to be recognized.

[0108] Extract sub-feature maps corresponding to each channel from the initial feature map, where the number of the sub-feature maps is the same as the number of channels of the initial feature map.

[0109] Perform max pooling and average pooling on each column vector of the sub-feature maps respectively to obtain first sub-feature maps, and traverse each sub-feature map to obtain first sub-feature maps corresponding to each channel.

[0110] Perform global max pooling on each horizontal vector of the first sub-feature maps respectively to obtain second sub-feature maps, and traverse each first sub-feature map to obtain second sub-feature maps corresponding to each channel.

[0111] Obtain the feature map of the video frame to be recognized according to the second sub-feature maps corresponding to each channel, and perform fully connected processing and normalization processing in sequence to obtain a six-dimensional vector corresponding to the scene type.

[0112] Compared with the prior art, in the intelligent scene recognition method and device based on vertical and horizontal pooling disclosed in the embodiments of the present invention, firstly, the video frame to be recognized is scaled according to the aspect ratio and preset height of the video frame to be recognized, so that the scene recognition of video frames or images with different aspect ratios can be realized, with strong applicability, avoiding the limitation that the scene recognition method based on deep learning can only recognize video frames or images with a single aspect ratio, and the defects that the scene recognition method based on convolutional neural network has low recognition accuracy or cannot recognize. Secondly, in the present invention, vertical pooling and horizontal pooling are respectively performed on the feature map of the video frame to be recognized, and the proportions of the vertical position feature, horizontal position feature and other non-position features on which scene classification is based are adjusted to the optimal state, thereby greatly improving the accuracy of video frame scene recognition. Finally, for the intelligent scene recognition method and device based on vertical and horizontal pooling proposed in the present invention, the video frame to be recognized is pre-classified by a convolutional neural network model based on vertical and horizontal pooling to obtain a preliminary classification result, and a binary classifier based on a linear regression model is used to classify the scene typicality of the video frame to be recognized, and then the final scene recognition result of the video frame to be recognized is obtained according to the preliminary scene classification result and the scene typicality classification result, thereby further improving the accuracy of the scene recognition result.

[0113] Those skilled in the art can understand that all or part of the processes of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.

[0114] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. An intelligent scene recognition method based on vertical and horizontal pooling, characterized in that, it includes the following steps: Scale the video frame to be recognized according to the aspect ratio and a preset height, where the aspect ratio is the aspect ratio of the video frame to be recognized; Use the trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling on the scaled video frame to be recognized in sequence, obtain a six-dimensional vector corresponding to the scene type, and obtain the preliminary classification result of the scene of the video frame to be recognized according to the six-dimensional vector; Obtain the distribution characteristics of the six-dimensional vector, and use the trained binary classifier to perform scene typicality classification on the video frame to be recognized based on the distribution characteristics; Obtain the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the preliminary classification result of the scene; Among them, the step of using the trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling on the scaled video frame to be recognized in sequence to obtain a six-dimensional vector corresponding to the scene type specifically includes: Use the trained convolutional neural network to perform convolution on the scaled video frame to be recognized to obtain the initial feature map of the video frame to be recognized; Extract the sub-feature maps corresponding to each channel from the initial feature map, and the number of sub-feature maps is the same as the number of channels of the initial feature map; Perform max pooling and average pooling on each column vector of the sub-feature map to obtain the first sub-feature map, and traverse each sub-feature map to obtain the first sub-feature map corresponding to each channel; Perform global max pooling on each horizontal vector of the first sub-feature map to obtain the second sub-feature map, and traverse each first sub-feature map to obtain the second sub-feature map corresponding to each channel; Obtain the feature map of the video frame to be recognized according to the second sub-feature map corresponding to each channel, and perform fully connected processing and normalization processing in sequence to obtain a six-dimensional vector corresponding to the scene type.

2. The intelligent scene recognition method based on vertical and horizontal pooling according to claim 1, characterized in that, The preliminary classification result of the scene of the video frame to be recognized is obtained according to the six-dimensional vector in the following manner: Among them, R CNN represents the pre-classification result of the scene type of the video frame to be recognized. When R CNN is 1, 2, 3, 4, 5, and 6 respectively, the corresponding pre-classification results of the scene type of the video frame to be recognized are close-up, medium shot, medium long shot, long shot, extreme long shot, and others; O i represents the eigenvalue corresponding to the i-th dimension in the six-dimensional vector.

3. The intelligent scene recognition method based on vertical and horizontal pooling according to claim 2, characterized in that, The step of performing max pooling and average pooling on each column vector of the sub-feature map to obtain the first sub-feature map includes: Perform max pooling on each column vector of the sub-feature map with the target feature size, and correspondingly obtain the first column vector with the target feature size, and the dimension of the first column vector is less than the height of the initial feature map; Perform average pooling on each column vector of the sub-feature map to obtain the non-position feature value corresponding to each column vector; Combine each first column vector and the corresponding non-position feature value to obtain a second column vector, and the dimension of the second column vector is 1 greater than the dimension of the first column vector; Obtain the corresponding first sub-feature map based on each second column vector.

4. The intelligent scene recognition method based on vertical and horizontal pooling according to claim 3, characterized in that, When performing max pooling on each column vector of the sub - feature map, the corresponding pooling stride and pooling region are respectively: kernel size = h - (O size - 1) × stride, Among them, stride represents the pooling stride, and kernel size represents the pooling field, h represents the height of the initial feature map, and O size represents the target feature size of the max pooling.

5. The intelligent scene recognition method based on vertical and horizontal pooling according to claim 3, characterized in that, The distribution characteristics of the six-dimensional vector include a first eigenvalue, a second eigenvalue, and a third feature; the first eigenvalue is the maximum eigenvalue in the six-dimensional vector, the second eigenvalue is the ratio of the second largest eigenvalue to the maximum eigenvalue in the six-dimensional vector, and the third eigenvalue is the entropy value of the six-dimensional vector: O i represents the eigenvalue corresponding to the i-th dimension in the six-dimensional vector.

6. The intelligent scene recognition method based on vertical and horizontal pooling according to claim 5, characterized in that, Obtaining the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre - classification result, specifically including: If the scene typicality classification result of the video frame to be recognized is a typical scene, then use the scene pre - classification result as the scene recognition result of the video frame to be recognized; If the scene typicality classification result of the video frame to be recognized is an atypical scene, and the first scene corresponding to the maximum eigenvalue in the six - dimensional vector and the second scene corresponding to the second - largest eigenvalue in the six - dimensional vector are adjacent scenes, then use the first scene and the second scene as the scene recognition result of the video frame to be recognized.

7. The intelligent scene recognition method based on vertical and horizontal pooling according to any one of claims 1 - 6, characterized in that, The trained binary classifier is obtained by training based on a linear logistic regression model.

8. An intelligent scene recognition device based on vertical and horizontal pooling, characterized in that, comprising: A pre - processing module that scales the video frame to be recognized according to the aspect ratio and a preset height, where the aspect ratio is the aspect ratio of the video frame to be recognized; A scene pre - classification module that uses a trained convolutional neural network to perform convolution, vertical pooling, and horizontal pooling processing on the scaled video frame to be recognized in sequence, obtains a six - dimensional vector corresponding to the scene type, and obtains the scene pre - classification result of the video frame to be recognized according to the six - dimensional vector; A typicality classification module that obtains the distribution characteristics of the six - dimensional vector, and uses the trained binary classifier to perform scene typicality classification on the video frame to be recognized based on the distribution characteristics; A scene recognition module that obtains the scene recognition result of the video frame to be recognized based on the scene typicality classification result and the scene pre - classification result; Among them, the scene pre - classification module is specifically used for: Using a trained convolutional neural network to perform convolution processing on the scaled video frame to be recognized to obtain the initial feature map of the video frame to be recognized; Extracting the sub - feature map corresponding to each channel from the initial feature map, and the number of sub - feature maps is the same as the number of channels of the initial feature map; Performing max pooling and average pooling processing on each column vector of the sub - feature map to obtain a first sub - feature map, and traversing each sub - feature map to obtain the first sub - feature map corresponding to each channel; Performing global max pooling processing on each horizontal vector of the first sub - feature map to obtain a second sub - feature map, and traversing each first sub - feature map to obtain the second sub - feature map corresponding to each channel; Obtaining the feature map of the video frame to be recognized according to the second sub - feature map corresponding to each channel, and performing full - connection processing and normalization processing in sequence to obtain a six - dimensional vector corresponding to the scene type.

Citation Information

Patent Citations

  • An implementation method of a toll lane vehicle feature recognition system based on video

    CN109190444A

  • Video classification method and model training method and device thereof, and electronic equipment

    CN110070067A