Unlabeled Video Hashing Retrieval Method and Device Based on Self-Supervised Learning
By using self-supervised learning and contrast loss functions in the video hash retrieval network, the problem of poor retrieval of unlabeled videos is solved, and high-accuracy and high-performance video retrieval is achieved, which is suitable for large-scale unlabeled data.
Patent Information
- Application Number
- CN202210226862.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-08
AI Technical Summary
The existing video retrieval technology is not effective in the unlabeled scenario, and the traditional method has slow feature extraction speed and large storage volume. The method that relies on manual annotation is costly and has large errors.
The unlabeled video hash retrieval method based on self-supervised learning is used to train the video hash retrieval network using a contrast loss function, intermediate features and hash code features are generated through the feature extraction layer and the hash layer, and network parameters are updated using the stochastic gradient descent method.
It realizes high-accuracy video retrieval without labeling, reduces the quantization error of hash code, improves the search speed and performance, and is suitable for large-scale label-free video data.
Smart Images

Figure CN114722902B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video retrieval, and particularly relates to an unlabeled video hashing retrieval method and device based on self-supervised learning. Background Art
[0002] In recent years, with the rapid development of communication and Internet technologies, the continuous rise of video calls, video software, and video content, videos have become an essential entertainment and social medium for people, and a large amount of video data has been accumulated on the Internet. The current text and image retrieval technologies have been relatively mature, but the video retrieval technology is still very lacking, especially in actual scenarios lacking data annotation. In the vast amount of video data on the Internet, manually annotating videos is an extremely difficult and costly task. Therefore, the video retrieval technology in the unlabeled scenario has become a research hotspot in both academia and industry.
[0003] Video similarity retrieval can be understood as expressing the features of different video materials, and then searching and sorting in the corresponding feature space. There are two ways of feature expression: one is the visual features extracted by traditional methods, such as key point features, color histograms, etc.; the other is to extract low-level basic features or high-level semantic features (deep features) based on deep learning. Traditional methods need to extract visual features in advance before retrieval when facing large-scale data, which not only has slow retrieval speed and poor effect, but also cannot use GPU parallel computing; while the retrieval method based on deep learning is fast and effective, and can be trained on a large scale on the GPU. However, in the real scenario, accurate video annotation is often lacking, resulting in poor retrieval results and low accuracy.
[0004] In the existing video retrieval technologies, Song J et al. adopted a similar self-supervised hashing retrieval method in the literature "Self-Supervised Video Hashing With Hierarchical Binary Auto-Encoder". They used LSTM as the backbone network, input the features of M training video frames into the encoder of the LSTM network to generate corresponding binary hash codes, then used two other LSTM networks to reconstruct the frame features from the forward and backward directions respectively, and finally calculated the reconstruction loss with the features of the original input video frames to achieve video retrieval. In the paper "Unsupervised Deep Video Hashing via Balanced Code for Large-Scale Video Retrieval" published by Wu G et al., TSN was used as the backbone network. Features were extracted from the RGB frames and optical flow frames of the input video through two paths respectively. Then, the features Z output by the 7th fully connected layer FC7 of the RGB path network were clustered to obtain Y, and Y was dimensionally reduced using the CCA method to obtain H. After multiplying by a rotation matrix R and passing through the sign function, the pseudo-hash code B was obtained. Then, the error was calculated between the pseudo-hash code B and the 8th fully connected layer FC8 of the optical flow path network to train the network. Finally, the network parameters of the optical flow path were inherited to the RGB frame path to achieve video retrieval. In the literature "Neighborhood Preserving Hashing for Scalable Video Retrieval", Li S et al. used an LSTM network with an attention mechanism as the backbone network. First, binary hash codes were calculated for the video frame features, and then the video frame features were reconstructed through the LSTM network. A visual content reconstruction loss was calculated between the reconstructed features and the original video frame features, and then the domain similarity loss and the domain information reconstruction loss were calculated to achieve video retrieval. However, the features extracted by the existing retrieval methods are continuous-dimensional features, which require a large amount of storage, have a high time cost, and a slow retrieval speed. The supervised training methods often rely on a large amount of labeled data, but the manual labeling cost is high and the error is large, which easily leads to a low accuracy rate of retrieval and poor results. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the disadvantages and deficiencies of the existing technologies, and provide an unlabeled video hashing retrieval method and device based on self-supervised learning. The method uses a contrast loss function to train the video hashing retrieval network without category annotation information, and updates the network parameters using the stochastic gradient descent method. The obtained retrieval network has a high accuracy rate and accurate and effective results.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] On the one hand, the present invention provides an unlabeled video hashing retrieval method based on self-supervised learning, including the following steps:
[0008] Obtain a video frame dataset and divide it into a training dataset and a test set, perform data augmentation on the training dataset to obtain an augmented dataset;
[0009] Build a video hashing retrieval network, and the video hashing retrieval network includes a feature extraction layer and a hashing layer;
[0010] Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to obtain intermediate features and calculate the contrast loss of the intermediate features;
[0011] Input the intermediate features into the hashing layer to obtain hash code features and calculate the contrast loss of the hash code features;
[0012] Train the video hashing retrieval network, use the stochastic gradient descent method to optimize the loss, update the network parameters until convergence, and obtain a trained video hashing retrieval network;
[0013] Input the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
[0014] As a preferred technical solution, the data augmentation includes random cropping, random color shift, random grayscale change, Gaussian blur, and random horizontal flipping;
[0015] Let the training dataset be denoted as X, then perform the same data augmentation on the training dataset twice to obtain augmented datasets X1 and X2, which are expressed as:
[0016] X1, X2 = augmentation(X)
[0017] where augmentation() represents the data augmentation operation.
[0018] As a preferred technical solution, the feature extraction layer adopts a ResNet network; the hashing layer includes a fully connected layer and an activation function; the activation function is expressed as y = tanh(βx), where β is a parameter.
[0019] As a preferred technical solution, the specific method for obtaining the intermediate features is:
[0020] Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to learn the visual information of the video frames in the dataset, and calculate the intermediate features Z1 and Z2 of X1 and X2 respectively:
[0021] Z1 = F(X1), Z2 = F(X2)
[0022] Among them, F represents the feature extraction layer, Z1 and Z2 are N×C real-valued feature matrices, N is the number of video frames in the training dataset, and C is the number of intermediate channels.
[0023] As a preferred technical solution, the contrastive loss for calculating the intermediate features is specifically:
[0024] Assume that two video frames corresponding to the same video frame in the training dataset in the enhanced datasets Z1 and Z2 are positive sample pairs, and other video frames are negative sample pairs. Use the contrastive loss function to calculate the loss between the intermediate features:
[0025]
[0026] Among them, z i , z j respectively represent the positive sample pair of the i-th video frame in Z1 and the j-th video frame in Z2 corresponding to the same video frame in the training dataset, z i , z k represents the negative sample pair, τ represents the temperature hyperparameter, which is used to adjust the effect of the loss function, represents the i and z j cosine similarity between.
[0027] As a preferred technical solution, the obtaining of the hash code features is specifically:
[0028] Input the intermediate features Z1 and Z2 into the hash layer H to obtain the hash code features B1 and B2:
[0029] B 1 = tanh(βw T Z1)
[0030] B 2 = tanh(βw T Z2)
[0031] Among them, B1 and B2 are N×K hash feature matrices, and the value of each element approaches -1 or 1 to represent binary 0 and 1, and K represents the number of hash code bits.
[0032] As a preferred technical solution, the calculation of the contrastive loss of the hash code features is specifically:
[0033] Assume that the hash code features corresponding to the same video frame in the training dataset in the hash code features B1 and B2 are positive sample pairs, and other video frames are used as negative sample pairs. Use the contrastive loss function to calculate the loss between the hash code features, and the formula is:
[0034]
[0035] Among them, b i , b j represents a positive sample pair in which the i-th hash code feature in B1 and the j-th hash code feature in B2 correspond to the same video frame in the training set, and b i , b k represents a negative sample pair.
[0036] As a preferred technical solution, the updating of the network parameters is specifically as follows:
[0037] The weight parameter of the feature extraction layer is θ, the parameter of the fully connected layer in the hash layer is w, and the parameter of the activation function is β;
[0038] When training the video hash retrieval network, calculate the contrast loss between the intermediate feature and the hash code feature;
[0039] Use the stochastic gradient descent method to update the network parameters, including:
[0040] Update the weight parameter θ of the feature extraction layer, and the update formula is:
[0041]
[0042] where α is the learning rate, and L 1 is the contrast loss function of the intermediate feature;
[0043] Update the parameter w of the fully connected layer and the parameter β of the activation function in the hash layer, and the update formula is:
[0044]
[0045] where, L 2 is the contrast loss function of the hash code feature;
[0046] As the number of training times increases, continuously increase the parameter β of the activation function so that the value output by the hash layer approaches -1 and 1;
[0047] When the network parameters converge, stop training to obtain the trained video hash retrieval network.
[0048] On the other hand, the present invention provides an unlabeled video hash retrieval system based on self-supervised learning, which is applied to the above-mentioned unlabeled video hash retrieval method based on self-supervised learning, and includes a data collection and processing module, a retrieval network establishment module, an intermediate feature extraction module, a hash code feature acquisition module, a retrieval network training module, and a retrieval result output module;
[0049] The data collection and processing module is used to obtain a video frame dataset, divide it into a training dataset and a test set, and perform data augmentation on the training dataset to obtain an augmented dataset;
[0050] The retrieval network establishment module is used to establish a video hashing retrieval network, and the video hashing retrieval network includes a feature extraction layer and a hashing layer;
[0051] The intermediate feature extraction module inputs the augmented dataset into the video hashing retrieval network, uses the feature extraction layer to obtain intermediate features, and calculates the contrast loss of the intermediate features;
[0052] The hash code feature acquisition module inputs the intermediate features into the hashing layer to obtain hash code features and calculates the contrast loss of the hash code features;
[0053] The retrieval network training module is used to train the video hashing retrieval network, optimize the loss using the stochastic gradient descent method, update the network parameters until convergence, and obtain a trained video hashing retrieval network;
[0054] The retrieval result output module inputs the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
[0055] On the other hand, the present invention provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned method for unsupervised video hashing retrieval based on self-supervised learning.
[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0057] 1. The present invention uses a contrast loss function for intermediate features and hash code features, reduces the quantization error of the generated hash codes, trains the video hashing retrieval network without category annotation information, and obtains a network with high retrieval accuracy and good performance;
[0058] 2. During the training process of the present invention, positive and negative sample pairs are constructed using video frame data in the same batch to help the video hashing retrieval network learn more visual representation information and ensure the effectiveness of the retrieval results;
[0059] 3. In traditional methods, since the hashing layer is a binary integer, it is impossible to take the derivative and use the stochastic gradient descent algorithm to update the parameters. However, the present invention uses the activation function y = tanh(βx) in the hashing layer to take the derivative, enabling the entire network model to use the stochastic gradient descent algorithm. Moreover, as the number of training times increases, the activation function β is continuously increased, making the values output by the hashing layer approach -1 and 1 more and more, achieving the effect of hash code output;
[0060] 4. In the existing methods, it is necessary to extract the features of video frames in advance using a feature extraction network before training, while the present invention can be directly trained end-to-end, making the training process more convenient;
[0061] 5. In the case of a very large amount of data, traditional methods have slow training speed and poor training effects, while the present invention can be well used in practical scenarios with a large amount of data and lack of annotation, and has good applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0063] Figure 1 It is a flowchart of a method for unlabeled video hashing retrieval based on self-supervised learning in an embodiment of the present invention;
[0064] Figure 2 It is a structural diagram of a video hashing retrieval network in an embodiment of the present invention;
[0065] Figure 3 It is a structural diagram of a system for unlabeled video hashing retrieval based on self-supervised learning in an embodiment of the present invention;
[0066] Figure 4 It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0068] When referring to "embodiments" in the present application, it means that the specific features, structures, or characteristics described in conjunction with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.
[0069] Such as Figure 1 、 Figure 2As shown, the unsupervised video hashing retrieval method based on self-supervised learning in this embodiment includes the following steps:
[0070] S1. Obtain a video frame dataset and divide it into a training dataset and a test set, and perform data augmentation on the training dataset to obtain an augmented dataset;
[0071] S2. Establish a video hashing retrieval network, where the video hashing retrieval network includes a feature extraction layer and a hashing layer;
[0072] S3. Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to obtain intermediate features and calculate the contrast loss of the intermediate features;
[0073] S4. Input the intermediate features into the hashing layer to obtain hashing code features and calculate the contrast loss of the hashing code features;
[0074] S5. Train the video hashing retrieval network, use the stochastic gradient descent method to optimize the loss, update the network parameters until convergence, and obtain a trained video hashing retrieval network;
[0075] S6. Input the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
[0076] More specifically, in step S1, let the training dataset be denoted as X, and perform the same data augmentation on the training dataset twice, that is, combine methods such as random cropping, random color shift, random grayscale change, Gaussian blur, and random horizontal flipping to perform data augmentation on the training dataset, and obtain augmented datasets X1 and X2, which are expressed as:
[0077] X1, X2 = augmentation(X)
[0078] Among them, augmentation() represents the data augmentation operation.
[0079] More specifically, in step S2, the feature extraction layer of the video hashing retrieval network uses a ResNet network; the hashing layer includes a fully connected layer and an activation function y = tanh(βx), where β is a parameter.
[0080] It should be noted that the feature extraction layer can be constructed using a network with the same function, and is not limited to the ResNet network of this application.
[0081] More specifically, in step S3, obtaining the intermediate features specifically is:
[0082] Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to learn the visual information in the video frames, and calculate the intermediate features Z1 and Z2 of X1 and X2 respectively;
[0083] Z1 = F(X1), Z2 = F(X2)
[0084] Wherein, F represents a feature extraction layer, Z1 and Z2 are N×C real number matrices of features, N is the number of video frames in the training dataset, and C is the number of intermediate channels.
[0085] Next, calculate the contrastive loss of the intermediate features:
[0086] For the N video frame data in the training dataset, 2N enhanced video frame data are obtained after data augmentation; assume that two video frames corresponding to the same video frame in the training dataset in the enhanced datasets Z1 and Z2 are positive sample pairs, and other video frames are negative sample pairs, and use the contrastive loss function to calculate the loss between the intermediate features:
[0087]
[0088] Wherein, z i , z j respectively represent the positive sample pair of the i-th video frame in Z1 and the j-th video frame in Z2 corresponding to the same video frame in the training dataset, z i , z k represents the negative sample pair, τ represents the temperature hyperparameter, which is used to adjust the effect of the loss function, represents the cosine similarity between z i and z j .
[0089] More specifically, the hash code features obtained in step S4 are specifically:
[0090] Since the intermediate feature Z is an N×C real number matrix, and the output of the hash layer should be +1 and -1 representing binary 0 and 1 respectively, it is necessary for the hash layer to convert the real number matrix into an N×K hash feature matrix, where K represents the number of bits of the hash code, usually taking values such as 8, 16, 32, 64, etc.;
[0091] Directly converting real numbers into binary codes is non-differentiable during the gradient backpropagation of the training network. Therefore, a hash layer is designed to make this part differentiable. The hash layer H of this method includes a fully connected layer and an activation function y = tanh(βx). Since y = tanh(βx) is differentiable, the entire training process can proceed normally;
[0092] Therefore, input the intermediate features Z1 and Z2 into the hash layer H to obtain the hash code features B1 and B2:
[0093] B 1 = tanh(βw T Z1)
[0094] B2 = tanh(βw T Z2)
[0095] Among them, B1 and B2 are N×K hash feature matrices, and the value of each element approaches -1 or 1 to represent binary 0 and 1.
[0096] Then calculate the contrast loss of the hash code features:
[0097] Let two hash code features corresponding to the same video frame in the training dataset in the hash code features B1 and B2 be positive sample pairs, and other video frames be negative sample pairs. Use the contrast loss function to calculate the loss between the hash code features. The formula is:
[0098]
[0099] Among them, b i , b j represents the positive sample pair where the i-th hash code feature in B1 corresponds to the j-th hash code feature in B2 for the same video frame in the training set, and b i , b k represents the negative sample pair.
[0100] More specifically, step S5 is specifically as follows:
[0101] The weight parameter of the feature extraction layer is θ, the parameter of the fully connected layer in the hash layer is w, and the parameter of the activation function is β;
[0102] When training the video hash retrieval network, calculate the contrast loss between the intermediate features and the hash code features;
[0103] Use the stochastic gradient descent method to update the network parameters, including:
[0104] Update the weight parameter θ of the feature extraction layer. The update formula is:
[0105]
[0106] Among them, α is the learning rate, and L 1 is the contrast loss function of the intermediate features;
[0107] Update the parameter w of the fully connected layer in the hash layer and the parameter β of the activation function. The update formula is:
[0108]
[0109] Among them, L 2 is the contrast loss function of the hash code features;
[0110] As the number of training times increases, continuously increase the parameter β of the activation function to make the values output by the hash layer approach -1 and 1;
[0111] Stop training when the network parameters converge to obtain a trained video hashing retrieval network.
[0112] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0113] Based on the same idea as the self-supervised learning-based unlabeled video hashing retrieval method in the above embodiments, the present invention also provides a self-supervised learning-based unlabeled video hashing retrieval system, which can be used to execute the above self-supervised learning-based unlabeled video hashing retrieval method. For the sake of convenience of description, in the structural schematic diagram of the self-supervised learning-based unlabeled video hashing retrieval system embodiment, only the parts related to the embodiments of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.
[0114] As Figure 3 shown, another embodiment of the present invention provides a self-supervised learning-based unlabeled video hashing retrieval system, including the following modules:
[0115] The data collection and processing module is used to obtain a video frame dataset and divide it into a training dataset and a test set, and perform data augmentation on the training dataset to obtain an augmented dataset;
[0116] The retrieval network establishment module is used to establish a video hashing retrieval network, and the video hashing retrieval network includes a feature extraction layer and a hashing layer;
[0117] The intermediate feature extraction module inputs the augmented dataset into the video hashing retrieval network, uses the feature extraction layer to obtain intermediate features and calculates the contrast loss of the intermediate features;
[0118] The hash code feature acquisition module inputs the intermediate features into the hashing layer to obtain hash code features and calculates the contrast loss of the hash code features;
[0119] The retrieval network training module is used to train the video hashing retrieval network, optimize the loss using the stochastic gradient descent method, update the network parameters until convergence, and obtain a trained video hashing retrieval network;
[0120] The retrieval result output module inputs the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
[0121] It should be noted that the unsupervised video hashing retrieval system based on self-supervised learning of the present invention corresponds one-to-one with the unsupervised video hashing retrieval method based on self-supervised learning of the present invention. The technical features and their beneficial effects described in the embodiments of the above-mentioned unsupervised video hashing retrieval method based on self-supervised learning are applicable to the embodiments of the unsupervised video hashing retrieval system based on self-supervised learning. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be elaborated here. This is hereby declared.
[0122] In addition, in the implementation manner of the unsupervised video hashing retrieval system based on self-supervised learning in the above embodiments, the logical division of each program module is only for illustration. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the unsupervised video hashing retrieval system based on self-supervised learning is divided into different program modules to complete all or part of the functions described above.
[0123] As Figure 4 shown, in one embodiment, a computer-readable storage medium is provided, storing a program in a memory. When the program is executed by a processor, the unsupervised video hashing retrieval method based on self-supervised learning is implemented, specifically as follows:
[0124] Obtain a video frame dataset and divide it into a training dataset and a test set, and perform data augmentation on the training dataset to obtain an augmented dataset;
[0125] Establish a video hashing retrieval network, where the video hashing retrieval network includes a feature extraction layer and a hashing layer;
[0126] Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to obtain intermediate features and calculate the contrast loss of the intermediate features;
[0127] Input the intermediate features into the hashing layer to obtain hash code features and calculate the contrast loss of the hash code features;
[0128] Train the video hashing retrieval network, use the stochastic gradient descent method to optimize the loss, update the network parameters until convergence, and obtain a trained video hashing retrieval network;
[0129] Input the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
[0130] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories.
[0131] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0132] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention should be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. Unannotated Video Hashing Retrieval Method Based on Self-Supervised Learning, Characterized in that, It includes the following steps: Obtain a video frame dataset and divide it into a training dataset and a test set, perform data augmentation on the training dataset to obtain an augmented dataset; Build a video hashing retrieval network, and the video hashing retrieval network includes a feature extraction layer and a hashing layer; Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to obtain intermediate features and calculate the contrast loss of the intermediate features; Input the intermediate features into the hashing layer to obtain hash code features and calculate the contrast loss of the hash code features; Train the video hashing retrieval network, use the stochastic gradient descent method to optimize the loss, update the network parameters until convergence, and obtain a trained video hashing retrieval network; Input the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
2. The unannotated video hashing retrieval method based on self-supervised learning according to claim 1, Characterized in that, The data augmentation includes random cropping, random color shift, random grayscale change, Gaussian blur and random horizontal flipping; Let the training dataset be represented as X, then perform the same data augmentation on the training dataset twice to obtain augmented datasets X1 and X2, expressed as: X1, X2 = augmentation(X) Where augmentation() represents the data augmentation operation.
3. The unannotated video hashing retrieval method based on self-supervised learning according to claim 2, Characterized in that, The feature extraction layer adopts a ResNet network; the hashing layer includes a fully connected layer and an activation function; the activation function is expressed as y = tanh(βx), where β is a parameter.
4. The unannotated video hashing retrieval method based on self-supervised learning according to claim 3, Characterized in that, The specific process of obtaining the intermediate features is as follows: Input the augmented dataset into the video hashing retrieval network, use the feature extraction layer to learn the visual information of the video frames in the dataset, and calculate the intermediate features Z1 and Z2 of X1 and X2 respectively: Z1 = F(X1), Z2 = F(X2) Where F represents the feature extraction layer, and Z1 and Z2 are N×C feature real number matrices, N is the number of video frames in the training dataset, and C is the number of intermediate channels.
5. The unannotated video hashing retrieval method based on self-supervised learning according to claim 4, Characterized in that, The specific process of calculating the contrast loss of the intermediate features is as follows: Let two video frames corresponding to the same video frame in the training dataset in the augmented datasets Z1 and Z2 be positive sample pairs, and other video frames be negative sample pairs, and use the contrast loss function to calculate the loss between the intermediate features: Among them, z i , z j respectively represent the positive sample pair of the same video frame in the training dataset corresponding to the i-th video frame in Z1 and the j-th video frame in Z2, z i , z k represents the negative sample pair, τ represents the temperature hyperparameter, which is used to adjust the effect of the loss function, represents z i and z j the cosine similarity between them.
6. The unannotated video hashing retrieval method based on self-supervised learning according to claim 5, Characterized in that, The specific process of obtaining the hash code features is as follows: Input the intermediate features Z1 and Z2 into the hashing layer H to obtain hash code features B1 and B2: B 1 = tanh(βw T Z1) B 2 = tanh(βw T Z2) Where B1 and B2 are N×K hash feature matrices, and the value of each element approaches -1 or 1 to represent binary 0 and 1, and K represents the number of hash code bits.
7. The unsupervised video hashing retrieval method based on self-supervised learning according to claim 6, wherein, the contrast loss for calculating the hashing code features is specifically: Let the hashing code features corresponding to the same video frame in the training data set in hashing code features B1 and B2 be positive sample pairs, and other video frames be negative sample pairs. The contrast loss function is used to calculate the loss between the hashing code features. The formula is: Among them, b i , b j represents a positive sample pair where the i-th hash code feature in B1 corresponds to the j-th hash code feature in B2 for the same video frame in the training set, and b i , b k represents a negative sample pair.
8. The unsupervised video hashing retrieval method based on self-supervised learning according to claim 7, wherein, the updating of the network parameters is specifically: The weight parameter of the feature extraction layer is θ, the fully connected layer parameter in the hashing layer is w, and the activation function parameter is β; When training the video hashing retrieval network, calculate the contrast loss between the intermediate features and the hashing code features; Use the stochastic gradient descent method to update the network parameters, including: Update the weight parameter θ of the feature extraction layer. The update formula is: where α is the learning rate, and L 1 is the contrastive loss function of the intermediate features; Update the fully connected layer parameter w and the activation function parameter β of the hashing layer. The update formula is: Among them, L 2 is the contrast loss function of the hash code feature; As the number of training times increases, continuously increase the activation function parameter β to make the values output by the hashing layer approach -1 and 1; Stop training when the network parameters converge to obtain a trained video hashing retrieval network.
9. An unsupervised video hashing retrieval system based on self-supervised learning, wherein, applied to the unsupervised video hashing retrieval method based on self-supervised learning described in any one of claims 1-8, including a data collection and processing module, a retrieval network establishment module, an intermediate feature extraction module, a hashing code feature obtaining module, a retrieval network training module, and a retrieval result output module; The data collection and processing module is used to obtain a video frame data set and divide it into a training data set and a test set, and perform data augmentation on the training data set to obtain an augmented data set; The retrieval network establishment module is used to establish a video hashing retrieval network, and the video hashing retrieval network includes a feature extraction layer and a hashing layer; The intermediate feature extraction module inputs the augmented data set into the video hashing retrieval network, and uses the feature extraction layer to obtain intermediate features and calculate the contrast loss of the intermediate features; The hashing code feature obtaining module inputs the intermediate features into the hashing layer to obtain hashing code features and calculate the contrast loss of the hashing code features; The retrieval network training module is used to train the video hashing retrieval network, optimize the loss using the stochastic gradient descent method, update the network parameters until convergence, and obtain a trained video hashing retrieval network; The retrieval result output module inputs the test set into the trained video hashing retrieval network for video retrieval to obtain retrieval results.
10. A computer-readable storage medium stores a program, wherein, when the program is executed by a processor, it implements the unsupervised video hashing retrieval method based on self-supervised learning described in any one of claims 1-8.
Citation Information
Patent Citations
Face image retrieval method and device based on deep learning and Hash coding
CN110175248A
Video hash retrieval method based on attention mechanism
CN111104555A