An Internet of Things cross-media big data retrieval method based on multi-label deep correlation analysis
By using a deep correlation learning model in the Internet of Things and combining multi-label information for deep correlation analysis of image and text features, the problem of separate processing of feature extraction and learning in the prior art is solved, and more efficient cross-media big data retrieval performance is achieved.
Patent Information
- Application Number
- CN202210142574.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-02-16
AI Technical Summary
The prior art has the problem of feature extraction and feature learning processing separately in the multimedia data processing of IoT, resulting in information loss, and it is difficult for a single tag information to effectively describe the semantic features of multimedia, affecting the retrieval performance.
High-level features based on the existing depth model are adopted, deep correlation learning is combined with multi-label information, low-dimensional representations of images and text are obtained, deep correlation learning models are established, and cross-media retrieval is realized through cosine distance.
It significantly improves the retrieval accuracy of IoT cross-media big data, makes full use of image text embedding with more discriminant capabilities in multi-label information learning, and reduces boundary ambiguity.
Smart Images

Figure CN114491103B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet of Things big data retrieval, and in particular to an Internet of Things cross-media big data retrieval method based on multi-label deep association analysis. Background Art
[0002] With the rapid development of technologies such as the Internet, the Internet of Things, and big data, the continuous emergence of multimedia big data such as images, texts, and videos obtained through the Internet of Things has posed new challenges to existing data retrieval methods. On the one hand, multimedia data provides rich materials for human understanding of the world. On the other hand, the heterogeneity of multimedia data poses challenges to the processing of multimedia data. At present, there are two main methods for processing multimedia data in the Internet of Things. One method is to extract features from multimedia data and then map all features to a low-dimensional subspace to achieve a unified representation of multimedia data. This method separates the steps of feature extraction and feature learning, and some important information will be lost during the implementation process. The second method is to use a deep learning model to directly learn features of multimedia data. This end-to-end learning method has shown superior performance in practical applications because it realizes automatic feature extraction and learning. On the one hand, more compact representation features can be learned through a large amount of labeled data. On the other hand, multimedia data has rich label information, and most existing methods use single label information.
[0003] In view of the above situation, multi-label semantic information has received more and more attention in the cross-media big data retrieval of the Internet of Things. For cross-media big data, multi-label information can better describe the semantic features of multimedia, thereby obtaining better information in retrieval. Summary of the invention
[0004] In response to the problems existing in the prior art, the present invention provides an Internet of Things cross-media big data retrieval method based on multi-label deep association analysis, which is based on the high-level features of the existing deep model and combines multi-label information for deep association learning to obtain low-dimensional representations of images and texts, thereby improving the multi-label cross-media retrieval performance.
[0005] The purpose of the present invention is achieved through the following technical solutions.
[0006] A cross-media big data retrieval method for the Internet of Things based on multi-label deep association analysis includes the following steps:
[0007] 1) Use VGG deep neural network to extract features from the image big data obtained by the Internet of Things; use the bag-of-words model to encode the text big data, and then use the fully connected neural network for feature learning;
[0008] 2) Constructing a semantic similarity matrix using multi-label information;
[0009] 3) Using image, text features and semantic similarity matrix, a deep association learning model is established;
[0010] 4) Use the deep association learning model to output image-text feature vectors, calculate the cosine distance between vectors, and realize cross-media retrieval.
[0011] Furthermore, the image text feature extraction method in step 1) is specifically as follows:
[0012] For the image data in the cross-media big data content of the Internet of Things, the VGG network is used to extract image features, and then a two-layer neural network is used for feature learning; for text data, the bag-of-words model is first used for encoding, and then a two-layer fully connected neural network is input for feature learning.
[0013] Furthermore, the semantic similarity matrix construction formula in step 2 is:
[0014]
[0015] Where S ij Indicates the similarity between the i-th and j-th samples, z i ,z j is the multi-label vector of the i,jth sample, ||z i -z j || represents the 2-norm of the vector, and σ is the weight.
[0016] Furthermore, the specific steps of using deep correlation analysis in step 3) are:
[0017] 3-1) Using the features extracted from the image network and text network, the following deep association analysis model is established:
[0018]
[0019] in
[0020]
[0021]
[0022]
[0023] are the weighted cross-covariance matrix and the weighted variance matrix, L(θ x ,θ y ,w,v) is the objective function, F i ,G j are high-level image and text features, respectively, n x ,n y are the number of images and texts respectively, N = n x ×ny is the total number of images and texts, are the weight vectors calculated by the similarity matrix S, w, v are the projection vectors to be learned, and T represents the transpose;
[0024] 3-2) Iteratively update the variables:
[0025] Using automatic differentiation technology, we can find the model with respect to the parameter θ x ,θ y ,w,v's gradient, the iterative update formula is:
[0026]
[0027]
[0028]
[0029]
[0030] Where τ is the learning rate, is the gradient of the objective function with respect to the network parameters to be learned, is the gradient of the objective function with respect to vector w,v, w k ,v k is the value of the current k-th step.
[0031] Furthermore, the step 4) uses the cosine distance to calculate the similarity between the image and the text, specifically:
[0032]
[0033] Where x i ,y j is the image and text feature vector, <x i ,y j > represents the inner product of two vectors, ||x i || ||y j || represents the 2-norm of a vector.
[0034] Compared with the prior art, the advantages of the present invention are:
[0035] 1. The present invention uses a fully connected network to learn features based on high-level semantic features, and can significantly improve the retrieval accuracy of IoT cross-media big data compared to other methods; experimental results show that this method can achieve good experimental results on data sets of different sizes;
[0036] 2. The IoT cross-media big data retrieval method based on multi-label deep association analysis of the present invention can make full use of multi-label information to learn more discriminative image text embedding. Experimental results show that the boundary fuzziness of this method is significantly reduced compared with other methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flow chart of the present invention.
[0038] Figure 2 It is a schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0039] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0040] A cross-media big data retrieval method for the Internet of Things based on multi-label deep association analysis is divided into four stages, namely, deep feature extraction, construction of semantic similarity matrix, modeling of cross-media features and semantic information using deep association model, and calculation of direct similarity of feature vectors of each media through cosine distance, thereby realizing cross-media big data retrieval for the Internet of Things.
[0041] like Figure 1 and 2 As shown, the specific steps include:
[0042] Step 1: For image data in cross-media content, use the VGG network to extract image features, and then use a 2-layer neural network for feature learning; for text data, first use the bag-of-words model to encode, and then input it into a 2-layer fully connected network for feature learning.
[0043] Step 2: Using the multi-label information of cross-media data, the semantic similarity matrix is calculated using the following formula:
[0044]
[0045] Where S ij Indicates the similarity between the i-th and j-th samples, z i ,z j is the multi-label vector of the i,jth sample, ||·|| represents the 2-norm of the vector, and σ is the weight.
[0046] Step 3: Use deep association analysis to model features and semantic information. The specific steps are as follows:
[0047] Step 3-1: Use the features extracted from the image network and text network to establish the following deep association analysis model:
[0048]
[0049] in
[0050]
[0051]
[0052]
[0053] are the weighted cross-covariance matrix and the weighted variance matrix, L(θ x ,θ y ,w,v) is the objective function, F i ,G j are high-level image and text features, respectively, n x ,n y are the number of images and texts respectively, N = n x ×n y is the total number of images and texts, are the weight vectors calculated by the similarity matrix S, and w and v are the projection vectors to be learned.
[0054] Step 3-2: Iteratively update the variables:
[0055] Using automatic differentiation technology, we can find the model with respect to the parameter θ x ,θ y ,w,v's gradient, the iterative update formula is:
[0056]
[0057]
[0058]
[0059]
[0060] Where τ is the learning rate, is the gradient of the objective function with respect to the network parameters to be learned, is the gradient of the objective function with respect to vector w,v, w k ,v k is the value of the current k-th step.
[0061] Step 4: Use cosine distance to calculate the similarity between the image and the text, specifically:
[0062]
[0063] Where x i ,y j is the image and text feature vector, <x i ,y j> represents the inner product of two vectors, ||x i || ||y j || represents the 2-norm of a vector.
[0064] Example 1
[0065] The method of the present invention was tested using a cross-media big data public dataset, and the results are as follows:
[0066] NUS-WIDE has a total of 269,648 images and corresponding text information. 10,000 images and corresponding text information are selected as the retrieval database, and 1,000 image-text pairs are selected as the test set. After processing, the size of each image becomes 224x224x3, and the corresponding text data is represented as a 1,000-dimensional vector.
[0067] MIRFlickr has a total of 25,000 image-text pairs, of which 2,000 pairs are selected as the test set and the rest as the retrieval database. After processing, the size of each image becomes 224x224x3, and the corresponding text data is represented as a 1,000-dimensional vector.
[0068] The test was conducted on the multimedia datasets NUS-WIDE and MIRFlickr, and compared with unsupervised deep correlation analysis (DCCA) and multi-label cross-modal retrieval (ml-CCA). The following is the comparison results of the method of the present invention (denoted as ml-DSA) and these two methods on two tasks: image to text search (I2T) and text to image search (T2I). The evaluation index uses the mean average precision (mAP, mean average precision).
[0069] Experimental results on NUS-WIDE dataset
[0070]
[0071]
[0072] Experimental results on the MIRFlickr dataset
[0073] method DCCA ml-CCA ml-DSA I2T 0.6558 0.6813 0.7041 T2I 0.6565 0.6826 0.7045
[0074] It can be seen from the above table that the method of the present invention is superior to other methods in cross-media big data retrieval performance.
[0075] The above descriptions are only some embodiments of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principles of the present invention. These improvements should be regarded as within the protection scope of the present invention.
Claims
1. A cross-media big data retrieval method for the Internet of Things based on multi-label deep correlation analysis, It is characterized in that The following steps are involved: 1) Use VGG deep neural network to extract features from the image big data obtained by the Internet of Things; For text big data, use the bag-of-words model to encode and then use a fully connected neural network for feature learning; 2) Constructing a semantic similarity matrix using multi-label information; 3) Using image, text features and semantic similarity matrix, a deep association learning model is established; 4) Use the deep association learning model to output image text feature vectors, calculate the cosine distance between vectors, and realize cross-media retrieval; The specific steps of using deep association analysis in step 3) are: 3-1) Using the features extracted from the image network and text network, the following deep association analysis model is established: in are the weighted cross-covariance matrix and the weighted variance matrix, L(θ x ,θ y ,w,v) is the objective function, F i ,G j are high-level image and text features, respectively, n x ,n y are the number of images and texts respectively, N = n x ×n y is the total number of images and texts, are the weight vectors calculated by the similarity matrix S, w, v are the projection vectors to be learned, and T represents the transpose; 3-2) Iteratively update the variables: Using automatic differentiation techniques, find the gradients of the model with respect to the parameters θ x , θ y , w, v, and the iterative update formula is: Where τ is the learning rate, is the gradient of the objective function with respect to the network parameters to be learned, is the gradient of the objective function with respect to vector w,v, w k ,v k is the value of the current k-th step.
2. According to claim 1, the Internet of Things cross-media big data retrieval method based on multi-label deep association analysis, It is characterized in that The image text feature extraction method in step 1) is specifically as follows: For the image data in the cross-media big data content of the Internet of Things, the VGG network is used to extract image features, and then a two-layer neural network is used for feature learning; for text data, the bag-of-words model is first used for encoding, and then a two-layer fully connected neural network is input for feature learning.
3. According to the Internet of Things cross-media big data retrieval method based on multi-label deep association analysis according to claim 1 or 2, It is characterized in that The semantic similarity matrix construction formula in step 2 is: Where S ij Indicates the similarity between the i-th and j-th samples, z i ,z j is the multi-label vector of the i,jth sample, ||z i -z j || represents the 2-norm of the vector, and σ is the weight.
4. According to claim 1, the Internet of Things cross-media big data retrieval method based on multi-label deep association analysis, It is characterized in that The step 4) uses the cosine distance to calculate the similarity between the image and the text, specifically: Where x i ,y j is the image and text feature vector, <x i ,y j > represents the inner product of two vectors, ||x i ||||y j || represents the 2-norm of a vector.
Citation Information
Patent Citations
Cross-media retrieval method based on deep semantic space
CN108694200A
Modal independent retrieval method and system based on intermediate text semantic enhancement space
CN109376261A