Large-scale Short Video Retrieval Method, System and Device Based on Deep Learning
The hash code generated by a convolutional neural network based on deep learning is solved for short video retrieval, and the problem of inconsistency between video information and manual annotation is optimized, the space-time complexity of the search is improved, and the search efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202210811333.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-07-11
AI Technical Summary
The existing short video retrieval methods have the problem that video information is inconsistent with manual labeled text, and the machine learning feature extraction is complex, occupying a large amount of storage space and computing resources.
Using a deep learning-based method, the hash code is generated using the improved convolutional neural network model, and short video retrieval is performed through similarity calculation, including keyframe extraction, video standardization processing, feature extraction and hash code mapping.
The space-time complexity of video retrieval is optimized, the semantic correlation of video in the time dimension is captured, and the search efficiency and accuracy are improved.
Smart Images

Figure CN115357754B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video retrieval, and relates to a short video retrieval method, system and device, in particular to a large-scale short video retrieval method, system and device based on deep learning. Background Art
[0002] With the development of Internet technology and intelligent devices, a variety of multimedia data such as text, voice, audio, images and videos have emerged on social networks and other information platforms. In particular, short videos are favored by users. Short videos have emerged due to the explosion of online social media. Short videos are very important information dissemination carriers, and they can convey a large amount of information in a very short time. Due to their fragmented and socialized characteristics, short videos can better meet the personalized needs of users. At the same time, the emergence of various video media has generated various video requirements. How to find videos that are approximately duplicate or of interest to users according to the videos uploaded by users in a large-scale video database and limited device resources has become a meaningful topic in the big data era.
[0003] In recent years, many scholars have used deep learning methods to solve the problem of short video data retrieval. The common approach is to annotate video data with information such as labels manually, and rely on text search to indirectly complete the search task of video data. Although the existing short video retrieval methods have made certain progress, there are still two deficiencies: 1) The short video information and the manually annotated text information are not completely consistent or one-to-one corresponding, and it is difficult to fully elaborate the video information through manual annotation. 2) The video features extracted by machine learning are still very complex, requiring a large amount of storage space for storage and a large amount of computing resources for similarity calculation. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a large-scale short video retrieval method, system and device based on deep learning. By learning the semantic information of short video data, an improved convolutional neural network model is used to generate hash codes, and finally similarity calculation is used to retrieve a given number of short video items.
[0005] The technical solution adopted by the method of the present invention is: a large-scale short video retrieval method based on deep learning, including the following steps:
[0006] Step 1: Extract key frames and perform video normalization processing on the query data and short videos in the retrieval set;
[0007] Step 2: Input the processed data into the short video semantic feature extraction network, and use similarity calculation to obtain T similar videos; where T is a preset value;
[0008] The short video semantic feature extraction network is composed of a feature extraction module and a feature hash code mapping module, and includes a first 3×3×3 convolutional layer, a first 1×2×2 pooling layer, a second 3×3×3 convolutional layer, a second 2×2×2 pooling layer, a third 3×3×3 convolutional layer, a third 2×2×2 pooling layer, a fourth 3×3×3 convolutional layer, a fourth 2×2×2 pooling layer, a fifth 3×3×3 convolutional layer, a fifth 2×2×2 pooling layer, a first 1×4096 fully-connected layer, a second 1×4096 fully-connected layer, and a hash layer;
[0009] The first 3×3×3 convolutional layer, the first 1×2×2 pooling layer, the second 3×3×3 convolutional layer, the second 2×2×2 pooling layer, the third 3×3×3 convolutional layer, the third 2×2×2 pooling layer, the fourth 3×3×3 convolutional layer, the fourth 2×2×2 pooling layer, the fifth 3×3×3 convolutional layer, and the fifth 2×2×2 pooling layer are connected in sequence to jointly form the feature extraction module. Inputting a short video, the original feature vector map of the short video is obtained through the feature extraction module;
[0010] The first 1×4096 fully-connected layer, the second 1×4096 fully-connected layer, and the hash layer are connected in sequence to jointly form the feature hash code mapping module. Inputting the original feature vector map of the short video, the hash code after the short video feature is mapped is obtained through the short video feature hash code mapping module;
[0011] The hash layer is a 1×K fully-connected layer activated by the sigmoid function, where K is the preset hash code length.
[0012] The technical solution adopted by the system of the present invention is: A large-scale short video retrieval system based on deep learning, including a preprocessing module and a retrieval module;
[0013] The preprocessing module is used to perform key frame extraction and video normalization processing on the query data and short videos in the retrieval set;
[0014] The retrieval module is used to input the processed data into the short video semantic feature extraction network and obtain T similar videos by using similarity calculation; where T is a preset value;
[0015] The short video semantic feature extraction network is overall composed of a feature extraction module and a feature hash code mapping module, and includes a first 3×3×3 convolutional layer, a first 1×2×2 pooling layer, a second 3×3×3 convolutional layer, a second 2×2×2 pooling layer, a third 3×3×3 convolutional layer, a third 2×2×2 pooling layer, a fourth 3×3×3 convolutional layer, a fourth 2×2×2 pooling layer, a fifth 3×3×3 convolutional layer, a fifth 2×2×2 pooling layer, a first 1×4096 fully-connected layer, a second 1×4096 fully-connected layer, and a hash layer;
[0016] The first 3×3×3 convolutional layer, the first 1×2×2 pooling layer, the second 3×3×3 convolutional layer, the second 2×2×2 pooling layer, the third 3×3×3 convolutional layer, the third 2×2×2 pooling layer, the fourth 3×3×3 convolutional layer, the fourth 2×2×2 pooling layer, and the fifth 3×3×3 convolutional layer, and the fifth 2×2×2 pooling layer are sequentially connected to jointly form a feature extraction module. Inputting a short video, the original feature vector map of the short video is obtained through the feature extraction module;
[0017] The first 1×4096 fully connected layer, the second 1×4096 fully connected layer, and the hash layer are sequentially connected to jointly form a feature hash code mapping module. Inputting the original feature vector map of the short video, the hashed code after feature mapping of the short video is obtained through the short video feature hash code mapping module;
[0018] The hash layer is a 1×K fully connected layer activated by the sigmoid function, where K is the preset hash code length.
[0019] The technical solution adopted by the device of the present invention is: A large-scale short video retrieval device based on deep learning, including:
[0020] One or more processors;
[0021] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the large-scale short video retrieval method based on deep learning.
[0022] Compared with the prior art, the method proposed by the present invention not only captures the semantic relevance of the video in the time dimension, learns the relative semantic relevance of the deep features, but also greatly optimizes the spatio-temporal complexity of video retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a principle framework diagram of an embodiment of the present invention.
[0024] Figure 2 It is a principle schematic diagram of an embodiment of the present invention.
[0025] Figure 3 It is a structural diagram of a short video semantic feature extraction network according to an embodiment of the present invention.
[0026] Figure 4 It is a schematic diagram of the training process of a short video semantic feature extraction network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0028] Please refer to Figure 1 and Figure 2 , a large-scale short video retrieval method based on deep learning provided by the present invention includes the following steps:
[0029] Step 1: Extract key frames and perform video normalization processing on the query data and short videos in the retrieval set;
[0030] In this embodiment, key frame extraction is based on the characteristics of short video shots, that is, shooting in a one-shot manner, and an equal-interval key frame extraction technology is adopted. Video normalization is based on the size of the video. If the video image is too large, a downsampling operation is required. If the size of the video image is too small, an interpolation operation is required. Finally, the processed images of each frame are spliced to obtain a complete video.
[0031] In this embodiment, for video normalization processing, according to the size of the video, if the video image is larger than the preset value, a downsampling operation is performed. The M×N image is downsampled by s times to obtain an image of (M / s)×(N / s); the specific formula is as follows:
[0032]
[0033] where f is the video frame, b is the different s×s small blocks into which the picture is divided, is the mean pooling operation;
[0034] If the size of the video image is smaller than the preset value, an interpolation operation is performed. Here, to speed up the calculation, the missing pixel value is the mean of the values of the nearest horizontal pixel point and the nearest vertical pixel point; finally, the processed images of each frame are spliced to obtain a complete video.
[0035] If the input video , where is the i th video frame, N is the total number of frames of the video ; then the extracted key frame , where t represents the video length in seconds, FPS represents the frame rate, n is the number of frames selected for the video,
[0036] Step 2: Input the processed data into the short video semantic feature extraction network, and use similarity calculation to obtain T similar videos; where T is a preset value;
[0037] In this embodiment, the hash code of the processed short video is calculated using the short video semantic feature extraction network. The Hamming distances between the hash code of the query data and the hash codes of each sample in the retrieval set are sorted from largest to smallest, and the top T precision values of the ranking list are calculated as the retrieval results.
[0038] Please refer to Figure 3 , the short video semantic feature extraction network of this embodiment is composed of a feature extraction module and a feature hash code mapping module as a whole, including a first 3×3×3 convolutional layer, a first 1×2×2 pooling layer, a second 3×3×3 convolutional layer, a second 2×2×2 pooling layer, a third 3×3×3 convolutional layer, a third 2×2×2 pooling layer, a fourth 3×3×3 convolutional layer, a fourth 2×2×2 pooling layer, a fifth 3×3×3 convolutional layer, a fifth 2×2×2 pooling layer, a first 1×4096 fully connected layer, a second 1×4096 fully connected layer, and a hash layer; the first 3×3×3 convolutional layer, the first 1×2×2 pooling layer, the second 3×3×3 convolutional layer, the second 2×2×2 pooling layer, the third 3×3×3 convolutional layer, the third 2×2×2 pooling layer, the fourth 3×3×3 convolutional layer, the fourth 2×2×2 pooling layer, and the fifth 3×3×3 convolutional layer, and the fifth 2×2×2 pooling layer are sequentially connected to jointly form the feature extraction module. The input short video passes through the feature extraction module to obtain the original feature vector map of the short video; the first 1×4096 fully connected layer, the second 1×4096 fully connected layer, and the hash layer are sequentially connected to jointly form the feature hash code mapping module. The input short video original feature vector map passes through the short video feature hash code mapping module to obtain the hash code after the short video feature is mapped; the hash layer is a 1×K fully connected layer activated by the sigmoid function, and K is the preset hash code length.
[0039] In this embodiment, the deep feature representation of the video , the hash representation layer . The size of the feature vector obtained from the fully connected layer is [N, 4096, 1, 1, 1], that is, the number of nodes num in the fully connected layer is 4096. Assume that the feature vector output from the fully connected layer is FC = {FC1, FC2,..., FC num}}, and the feature vector obtained after passing through the hash layer is , where d is the number of nodes in the hash layer, that is, the dimension of the vector output by the hash layer. To ensure that each class label in the dataset has a unique hash code corresponding to it, at least d min nodes are required, and the calculation formula for the number of nodes is as follows:
[0040]
[0041] Let the maximum value of d be dmax , then d max = num, that is, the number of nodes in the hash layer needs to be lower than the number of nodes in the previous fully connected layer. Since the purpose of the hash layer is to map the fully connected features into a compact binary-like code, it is necessary to reduce the dimension of the features output by the previous layer. Then the value range of d is:
[0042]
[0043] The definition of the hash function is as follows:
[0044]
[0045] Among them, Sign is the sign function, W is the weight of the hash layer, FC is the fully connected layer, and B is the bias.
[0046] In order to map the hash features into binary codes, a threshold function is used to process the output value, and the threshold function is designed as follows:
[0047] .
[0048] Please refer to Figure 4 , the short video semantic feature extraction network of this embodiment is a trained short video semantic feature extraction network; its training process includes the following steps:
[0049] Step 2.1: Obtain a data set and divide it into a training data set and a test data set;
[0050] In this embodiment, the HMDB51 and UCF-101 data sets are used, and 70% of the data set is selected as the training data set , and the remaining 30% is used as the test data set ;
[0051] Step 2.2: For the training data set and the test data set, perform key frame extraction and video normalization processing;
[0052] In this embodiment, key frame extraction is based on the characteristics of short video shots, that is, shooting in a one-shot manner, and an equal-interval key frame extraction technique is adopted. Video normalization is based on the size of the video. If the video image is too large, downsampling operation is required. If the size of the video image is too small, interpolation operation is required. Finally, the processed images of each frame are stitched together to obtain a complete video.
[0053] If the input video , where is the i th video frame, and N is the total number of frames of the video ; then the extracted key frame , where, trepresents the video length in seconds, and FPS represents the frame rate. n is the number of frames selected for the video.
[0054] Step 2.3: Determine the training objective function, optimization algorithm, learning rate, momentum, weight decay, batch size, and the number of network training iterations epoch.
[0055] In this embodiment, the objective function consists of the triplet loss function , the cross-entropy loss function and the smoothed mean average precision loss function ;
[0056] ;
[0057] Among them, are all hyperparameters, and the weight parameters W and bias parameters B of the network are obtained through the backpropagation technology of the training model.
[0058] In this embodiment, the triplet loss function:
[0059] ;
[0060] Among them, is the anchor sample, which is a selected video data of a certain type. is the positive example sample of the same type as the anchor sample. is the negative example sample of a different type from the anchor sample. is the function that maps the sample data to the same vector space. is the minimum distance between the preset negative example sample and the positive example sample; K is the total number of data samples.
[0061] In this embodiment, the cross-entropy loss function for multi-classification ; where K is the total number of data samples and M is the number of categories. is the sign function, taking values 0 or 1. If the true category of the sample i is equal to c, it takes 1. is the predicted probability that the observed sample i belongs to category c.
[0062] In this embodiment, the smoothed mean average precision loss function ; where m is the number of samples. is the smoothed mean estimate value. , is the set of videos in the dataset that belong to the same type as the query video. indicates that the indicator function is replaced by a differentiable Sigmoid function. is a ranking matrix, where the i th row represents the ranking scores of the correlation between the i th sample and the remaining samples; is the smoothing coefficient, and different smoothing coefficients can provide different gradient information;
[0063] Step 2.4: Input the training dataset into the short-video semantic feature extraction network to train the short-video semantic feature extraction network;
[0064] In this embodiment, the Adam algorithm is used for optimization, the learning rate is set to 10 -5 , the momentum is set to 0.9, the weight decay is set to 10 -5 , the batch size is set to 32, and the decay strategy of the learning rate is adjusted according to the non-decrease of the Loss on the validation set. The input frame image size is 3×112×112, 16 frames are convolved each time as a clip, and the initial weights of the convolutional neural network ResNet34_3D are initialized with pre-trained weights. By training the model, the weight parameter W and the bias parameter B of the network are obtained.
[0065] Step 2.5: Use the trained short-video semantic feature extraction network to calculate the hash codes of the samples in the test dataset, sort the Hamming distances between the hash codes of the query sample and each sample in the training dataset from large to small, and calculate the top n precision, obtain the mean average precision (MAP) and the top n retrieval results; If the category of the input video is the same as the category of the retrieved video, it is considered a correct retrieval. When the above loss function value tends to be stable and no longer decreases, a network with good performance is obtained.
[0066] To evaluate the effectiveness of the method of the present invention, the method of the present invention is compared with several state-of-the-art methods in terms of retrieval performance, including FIHTV, VHSL, SVH, DH, DVH, DBNVH, BRVH, DSH, DLBHC, and SRH. In this experiment, hash codes with different numbers of bits (16 bits, 32 bits, 64 bits) are used, and the UCF-101 dataset and the HMDB-51 dataset are adopted. The DVH method uses a convolutional neural network to simultaneously learn the feature representation of the image and the hash function, and then projects their corresponding features into a common representation space. The FIHTV, VHSL, SVH, DH, DBNVH, BRVH, DSH, DLBHC, and SRH methods are executed according to the original text. The environment used in this experiment is Intel® Xeon® Platinum 8260 CPU @ 2.4GHZ, NVIDIA A100 TENSOR CORE GPU, Ubuntu 20.04 system, and Python and the open-source library Pytorch are used for development.
[0067] Table 1
[0068]
[0069] Table 1 shows the comparison experiment results of the present invention and other methods in the video retrieval task on the UCF-101 dataset, where mAP is the mean average precision metric.
[0070] Table 2
[0071]
[0072] Table 2 shows the comparison experiment results of the present invention and other methods in the video retrieval task on the HMDB-51 dataset, where mAP is the mean average precision metric.
[0073] Table 3
[0074]
[0075] Table 3 shows the comparison of the different video retrieval methods of the present invention in terms of speed. The multi-level hashing retrieval means comparing the binary hashing feature codes of 16 bits, 32 bits, and 64 bits in sequence with the video feature code representations in the database to gradually narrow down the search range, and finally using the real-value feature codes for precise comparison in the database. The present invention makes full use of the spatio-temporal semantic information of the video and further improves the retrievability.
[0076] It should be understood that the above description of the preferred embodiment is relatively detailed and should not be considered as a limitation on the protection scope of the patent of the present invention. Those of ordinary skill in the art, under the inspiration of the present invention and without departing from the protection scope defined by the claims of the present invention, can still make substitutions or modifications, which all fall within the protection scope of the present invention. The scope of the present invention claimed shall be subject to the appended claims.
Claims
1. A large-scale short video retrieval method based on deep learning, characterized in that, It includes the following steps: Step 1: Extract key frames and perform video normalization processing on the query data and short videos in the retrieval set; Among them, for the key frame extraction, according to the characteristics of the short video shots, that is, shooting in a one-shot manner, an equal-interval key frame extraction technology is used for key frame extraction; If the input video , where is the i th video frame and N is the total number of frames of the video ; then the extracted key frames , where t represents the video length in seconds, FPS represents the frame rate, n is the number of frames selected for the video, For the video normalization processing, according to the size of the video, if the video image is larger than the preset value, a downsampling operation is performed. The M×N image is downsampled by s times to obtain an image of (M / s)×(N / s); the specific formula is as follows: Among them, f is a video frame, and b is different s×s small blocks into which the picture is segmented. is the mean pooling operation; If the size of the video image is smaller than the preset value, an interpolation operation is performed, and the missing pixel value is the mean of the values of the nearest horizontal pixel point and the nearest vertical pixel point; finally, each processed image frame is stitched to obtain a complete video; Step 2: Input the processed data into the short video semantic feature extraction network, and use similarity calculation to obtain T similar videos; where T is a preset value; The short video semantic feature extraction network is composed of a feature extraction module and a feature hash code mapping module, and includes a first 3×3×3 convolutional layer, a first 1×2×2 pooling layer, a second 3×3×3 convolutional layer, a second 2×2×2 pooling layer, a third 3×3×3 convolutional layer, a third 2×2×2 pooling layer, a fourth 3×3×3 convolutional layer, a fourth 2×2×2 pooling layer, a fifth 3×3×3 convolutional layer, a fifth 2×2×2 pooling layer, a first 1×4096 fully connected layer, a second 1×4096 fully connected layer, and a hash layer; The first 3×3×3 convolutional layer, the first 1×2×2 pooling layer, the second 3×3×3 convolutional layer, the second 2×2×2 pooling layer, the third 3×3×3 convolutional layer, the third 2×2×2 pooling layer, the fourth 3×3×3 convolutional layer, the fourth 2×2×2 pooling layer, the fifth 3×3×3 convolutional layer, and the fifth 2×2×2 pooling layer are connected in sequence to jointly form the feature extraction module. Input the short video, and the short video original feature vector map is obtained through the feature extraction module; The first 1×4096 fully connected layer, the second 1×4096 fully connected layer, and the hash layer are connected in sequence to jointly form the feature hash code mapping module. Input the short video original feature vector map, and the hash code after the short video feature is mapped is obtained through the short video feature hash code mapping module; The hash layer is a 1×K fully connected layer activated by the sigmoid function, and K is the preset hash code length.
2. The large-scale short video retrieval method based on deep learning according to claim 1, wherein: In step 2, use the short video semantic feature extraction network to calculate the hash code of the processed short video, sort the Hamming distances between the query data and the hash codes of each sample in the retrieval set from largest to smallest, and calculate the top T precision as the retrieval result.
3. The large-scale short video retrieval method based on deep learning according to any one of claims 1-2, characterized in that: The short video semantic feature extraction network described in Step 2 is a trained short video semantic feature extraction network; Its training process includes the following steps: Step 2.1: Obtain a data set and divide it into a training data set and a test data set; Step 2.2: Perform key frame extraction and video normalization processing on the training data set and the test data set; Step 2.3: Determine the training objective function, optimization algorithm, learning rate, momentum, weight decay, batch size, and the number of network training iterations epoch; The objective function is composed of a triplet loss function , a cross-entropy loss function and a smoothed average precision loss function ; ; Among them, are all hyperparameters, and the weight parameter W and bias parameter B of the network are obtained by training the model's backpropagation technique; The triplet loss function: ; Among them, is the anchor sample, which is a selected video data of a certain type, is the positive example sample of the same type as the anchor sample, is the negative example sample of a different type from the anchor sample; is the function that maps the sample data to the same vector space; is the preset minimum distance between the negative example sample and the positive example sample; K is the total number of data samples; Cross-entropy loss function for multi-class classification ; where K is the total number of data samples and M is the number of classes; is the sign function, taking values 0 or 1. If the true class of the sample i is equal to c, it takes 1; is the predicted probability that the observed sample i belongs to class c; The smooth average precision loss function ; where m is the number of samples; is the smooth average estimate, , is the set of videos in the dataset that belong to the same type as the query video, indicates that the indicator function is replaced by a differentiable Sigmoid function; is the ranking matrix, and the i -th row represents the ranking scores of the correlation between the i -th sample and the remaining samples; is the smoothing coefficient; Step 2.4: Input the training data set into the short video semantic feature extraction network to train the short video semantic feature extraction network; Step 2.5: Use the trained short video semantic feature extraction network to calculate the hash codes of the samples in the test dataset, sort the Hamming distances between the hash code of the query sample and the hash codes of the samples in the training dataset from large to small, and calculate the precision of the top n ones, obtaining the mean average precision metric MAP and the top n retrieval results; if the category of the input video is the same as the category of the retrieved video, it is considered a correct retrieval. If the value of the above loss function tends to be stable and no longer decreases, a network with good performance is obtained.
4. A large-scale short video retrieval system based on deep learning, characterized in that: It includes a preprocessing module and a retrieval module; The preprocessing module is used to extract key frames and perform video normalization processing on the query data and short videos in the retrieval set; Among them, for the key frame extraction, according to the characteristics of the short video shots, that is, shooting in a one-shot style, the key frame extraction technology of equal-interval extraction is used to extract key frames; If the input video , where is the i th video frame and N is the total number of frames of the video ; then the extracted key frame , where t represents the video length in seconds, FPS represents the frame rate, n is the number of frames selected for the video, For the video normalization processing, according to the size of the video, if the video image is larger than the preset value, a downsampling operation is performed. The M×N image is downsampled by s times to obtain an image of (M / s)×(N / s); the specific formula is as follows: Among them, f is a video frame, and b is different s×s small blocks into which the picture is segmented. is an average pooling operation; If the size of the video image is smaller than the preset value, an interpolation operation is performed, and the missing pixel value is the average value of the horizontal nearest pixel point and the vertical nearest pixel point; finally, the processed images of each frame are spliced to obtain a complete video; The retrieval module is used to input the processed data into the short video semantic feature extraction network and obtain T similar videos by using similarity calculation; where T is a preset value; The short video semantic feature extraction network is composed of a feature extraction module and a feature hash code mapping module, and includes a first 3×3×3 convolutional layer, a first 1×2×2 pooling layer, a second 3×3×3 convolutional layer, a second 2×2×2 pooling layer, a third 3×3×3 convolutional layer, a third 2×2×2 pooling layer, a fourth 3×3×3 convolutional layer, a fourth 2×2×2 pooling layer, a fifth 3×3×3 convolutional layer, a fifth 2×2×2 pooling layer, a first 1×4096 fully connected layer, a second 1×4096 fully connected layer, and a hash layer; The first 3×3×3 convolutional layer, the first 1×2×2 pooling layer, the second 3×3×3 convolutional layer, the second 2×2×2 pooling layer, the third 3×3×3 convolutional layer, the third 2×2×2 pooling layer, the fourth 3×3×3 convolutional layer, the fourth 2×2×2 pooling layer, the fifth 3×3×3 convolutional layer, and the fifth 2×2×2 pooling layer are connected in sequence to jointly form the feature extraction module. Input the short video, and the original feature vector map of the short video is obtained through the feature extraction module; The first 1×4096 fully connected layer, the second 1×4096 fully connected layer, and the hash layer are connected in sequence to jointly form the feature hash code mapping module. Input the original feature vector map of the short video, and the hash code after the short video feature is mapped is obtained through the short video feature hash code mapping module; The hash layer is a 1×K fully connected layer activated by the sigmoid function, and K is the preset hash code length.
5. A large-scale short video retrieval device based on deep learning, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the large-scale short video retrieval method based on deep learning as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Video hash retrieval representation conversion method based on deep learning
CN111563184A
Multi-label video hash retrieval method and equipment based on semantic embedding soft similarity
CN113177141A