Cross-modal Retrieval Method and System for Marking-enhanced Social Multimedia Data
Through the label enhancement method, a fine-grained similarity matrix is generated and the hash function is optimized, which solves the problem of sample similarity representation and hash code discrimination force in cross-modal hash retrieval, and achieves a more efficient cross-modal retrieval effect.
Patent Information
- Application Number
- CN202210737343.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-06-27
AI Technical Summary
There is semantic correlation between multimedia data of different modes, but due to the heterogeneous and semantic divides, cross-modal hash retrieval faces huge challenges, making it difficult to effectively characterize the similarity between samples and generate hash codes for distinguishing forces.
Using a label enhancement-based method, the label distribution of samples is generated through transfer learning, a fine-grained similarity matrix is constructed, and the hash code is solved through the internal product adaptation objective function, and the hash function is optimized to improve the ability of cross-modal retrieval.
It effectively improves the ability of cross-modal retrieval, and the generated hash code is more distinguishing, which can better characterize the similarity between samples, reduces quantization errors and improves the search speed.
Smart Images

Figure CN115100433B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of cross-modal retrieval, and particularly relates to a cross-modal retrieval method and system for social multimedia data based on tag enhancement. Background Art
[0002] The statements in this part merely provide background technical information related to the present disclosure, and do not necessarily constitute prior art.
[0003] As the most primitive retrieval technology, Nearest Neighbor Search (NN) calculates the distances between query samples and all database samples using a distance function, and finally returns the sample with the minimum distance. However, as the scale of the dataset increases and the dimension of the features rises, the computational cost of this precise search strategy becomes unacceptable. Therefore, Approximate Nearest Neighbor Search (ANN) is used as an alternative. By searching for potentially similar rather than the most similar data samples, it sacrifices a certain degree of accuracy to improve efficiency, thus meeting the retrieval requirements for large-scale data in practical applications.
[0004] In approximate nearest neighbor search methods, the hash learning-based method encodes high-dimensional real-valued features into low-dimensional binary hash codes and tries to maintain the similarity relationship of the data as much as possible. During retrieval, the Hamming distance between the hash code of the query sample and the hash codes of the database samples is calculated to return similar samples as retrieval results. The Hamming distance can be quickly calculated through the exclusive OR (XOR) operation on the central processing unit (CPU). This encoding and indexing mechanism reduces the storage overhead and achieves a faster search speed, making it suitable for the retrieval tasks of large-scale multimedia data, and thus has received extensive attention.
[0005] The inventors found that the content of multimedia data in different modalities may be semantically correlated with each other, and Cross-Modal Retrieval is to return a query result in another modality that is semantically related to a given query sample in one modality. The combination of cross-modal retrieval and hash technology provides an effective solution for realizing cross-modal retrieval of large-scale multimedia data. However, due to the natural problems of heterogeneous gap and semantic gap between different modality data, it brings huge challenges to cross-modal hash retrieval. Summary of the Invention
[0006] To solve the above problems, the present disclosure provides a cross-modal retrieval method and system for social multimedia data based on tag enhancement. The solution can better represent the similarity between samples, generate more discriminative hash codes, and effectively improve the cross-modal retrieval ability.
[0007] According to the first aspect of the embodiments of the present disclosure, a cross-modal retrieval method for social multimedia data based on label enhancement is provided, including:
[0008] Obtain the data samples to be retrieved;
[0009] Extract features from the data samples based on a feature extraction method;
[0010] Based on the features of the data samples to be retrieved, encode them using the pre-trained hash function of the corresponding modality to obtain the hash code representation of the data samples;
[0011] Calculate the similarity value between the hash code representation of the data samples and the hash codes of the samples in the database to be retrieved; obtain the corresponding retrieval results based on the similarity value;
[0012] Among them, for the training of the hash function in different modalities, specifically: based on the sample data categories in the database to be retrieved, obtain the corresponding category words, and convert the category words into category attributes through a word vector model; based on the category attributes and a pre-constructed objective function, solve to obtain a label-enhanced label distribution; construct a fine-grained similarity matrix based on the label distribution, and construct an inner product adaptation objective function based on the similarity matrix. By solving the inner product adaptation objective function, obtain the hash codes of the training samples; use the hash codes of the training samples as supervision information to train the hash functions in different modalities.
[0013] Further, when constructing the inner product adaptation objective function based on the similarity matrix, at the same time, introduce an intermediate variable to replace the hash code in the inner product operation of the inner product adaptation objective function, and introduce a regularization term to minimize the difference between the introduced intermediate variable and the hash code.
[0014] Further, when using the hash codes of the training samples as supervision information to train the hash functions in different modalities, the following objective function is specifically adopted:
[0015]
[0016] where λ is the penalty coefficient of the regularization term to avoid overfitting, W (l) is the mapping matrix of the l-th modality, B is the hash code of the training set, and X (l) is the feature matrix of the training samples in the l-th modality.
[0017] Further, when solving for the label-enhanced label distribution based on the category attributes and the pre-constructed objective function, the following objective function is specifically adopted:
[0018]
[0019] Among them, is the projection matrix, I is the identity matrix, is the rotation matrix, A is the category attribute, L is the logical flag, D is the flag distribution, α is the balance parameter, and θ is the penalty coefficient of the regularization term.
[0020] Furthermore, the data samples are feature-extracted by the feature extraction method, specifically: when the data sample is an image, image feature extraction is performed based on the SIFT or GIST method; when the data sample is text, text feature extraction is performed based on the BoW method.
[0021] Furthermore, the data samples include the image to be retrieved or the text to be retrieved. When the data sample is an image, the retrieved data is the text corresponding to the image; when the data sample is text, the retrieved data is the image corresponding to the text.
[0022] Furthermore, calculating the similarity value between the hash code representation of the data sample and the sample hash codes in the database to be retrieved specifically includes: calculating the Hamming distance between the hash code representation of the data sample and the sample hash codes in the database to be retrieved, sorting the samples in the database in ascending order based on the distance value, and selecting the first k samples as the retrieval results, where k is an integer not less than 1.
[0023] According to the second aspect of the embodiments of the present disclosure, a cross-modal retrieval system for social multimedia data based on label enhancement is provided, including:
[0024] A data acquisition unit for acquiring data samples to be retrieved;
[0025] A feature extraction unit for feature-extracting the data samples based on the feature extraction method;
[0026] An encoding unit for encoding based on the features of the data samples to be retrieved, using the pre-trained hash function of the corresponding modality to obtain the hash code representation of the data samples;
[0027] A retrieval unit for calculating the similarity value between the hash code representation of the data sample and the sample hash codes in the database to be retrieved; and obtaining the corresponding retrieval results based on the similarity value;
[0028] Among them, for the training of the hash function in different modalities, specifically: based on the sample data categories in the database to be retrieved, the corresponding category words are obtained, and the category words are transformed into category attributes through a word vector model; based on the category attributes and a pre-constructed objective function, a labeled distribution with enhanced labels is solved; a fine-grained similarity matrix is constructed based on the labeled distribution, and an inner product adaptation objective function is constructed based on the similarity matrix. By solving the inner product adaptation objective function, the hash code of the training sample is obtained; based on the hash code of the training sample as supervision information, the hash function in different modalities is trained.
[0029] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program running on the memory. When the processor executes the program, the cross-modal retrieval method for social multimedia data based on label enhancement is implemented.
[0030] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the cross-modal retrieval method for social multimedia data based on label enhancement is implemented.
[0031] Compared with the prior art, the beneficial effects of the present disclosure are:
[0032] (1) The present disclosure provides a cross-modal retrieval method and system for social multimedia data based on label enhancement. The solution generates a labeled distribution of samples through label enhancement based on transfer learning, then constructs a more accurate similarity matrix based on the labeled distribution, and constructs an inner product adaptation objective function based on the similarity matrix. By solving the inner product adaptation objective function, the hash code of the training sample is obtained, and then an optimized hash function is solved; the solution can better represent the similarity between samples, generate more discriminative hash codes, and effectively improve the cross-modal retrieval ability.
[0033] (2) The solution provides an efficient discrete optimization algorithm to solve the discrete constraint problem of the hash code, reducing the quantization error; at the same time, the original high-dimensional data is mapped into a low-dimensional Hamming space, greatly improving the retrieval speed and reducing the computational cost.
[0034] The advantages of the additional aspects of the present disclosure will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings forming a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.
[0036] Figure 1 Schematic diagram of the training process of the hash function for the cross-modal retrieval method of social multimedia data based on tag enhancement described in the embodiments of the present disclosure. Detailed implementation manners
[0037] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0038] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0039] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0040] The embodiments in the present disclosure and the features in the embodiments may be combined with each other without conflict.
[0041] Embodiment 1:
[0042] The purpose of this embodiment is to provide a cross-modal retrieval method for social multimedia data based on tag enhancement.
[0043] A cross-modal retrieval method for social multimedia data based on tag enhancement includes:
[0044] Obtain the data samples to be retrieved;
[0045] Extract features from the data samples based on a feature extraction method;
[0046] Based on the features of the data samples to be retrieved, encode using the pre-trained hash function of the corresponding modality to obtain the hash code representation of the data samples;
[0047] Calculate the similarity value between the hash code representation of the data samples and the hash codes of the samples in the database to be retrieved; obtain the corresponding retrieval result based on the similarity value;
[0048] Among them, for the training of the hash function in different modalities, specifically: based on the sample data categories in the database to be retrieved, the corresponding category words are obtained, and the category words are transformed into category attributes through a word vector model; based on the category attributes and a pre-constructed objective function, a labeled enhanced label distribution is solved; a fine-grained similarity matrix is constructed based on the label distribution, and an inner product adaptation objective function is constructed based on the similarity matrix. By solving the inner product adaptation objective function, the hash code of the training sample is obtained; based on the hash code of the training sample as supervision information, the hash function in different modalities is trained.
[0049] Further, when constructing the inner product adaptation objective function based on the similarity matrix, at the same time, an intermediate variable is introduced to replace the hash code in the inner product operation of the inner product adaptation objective function, and a regularization term is introduced to minimize the difference between the introduced intermediate variable and the hash code.
[0050] Further, when training the hash function in different modalities based on the hash code of the training sample as supervision information, the following objective function is specifically adopted:
[0051]
[0052] where λ is the penalty coefficient of the regularization term to avoid overfitting, W (l) is the mapping matrix of the l-th modality, B is the hash code of the training set, and X (l) is the feature matrix of the l-th modality of the training sample.
[0053] Further, when solving the labeled enhanced label distribution based on the category attributes and a pre-constructed objective function, the following objective function is specifically adopted:
[0054]
[0055] where is the projection matrix, I is the identity matrix, is the rotation matrix, A is the category attribute, L is the logical label, D is the label distribution, α is the balance parameter used to balance the relative importance of the first two terms in this objective function (i.e., adjust the relative magnitudes of the loss values generated by the first two terms), and θ is the penalty coefficient of the regularization term to prevent overfitting; the logical label L is actually a label matrix with values of 0 and 1, where 1 indicates that the sample belongs to this class and 0 indicates the opposite; the parameter D is the label distribution, which has the same shape (c×n) as L but takes real values. Actually, it transfers the prior knowledge in the category attribute A to L and then obtains D.
[0056] Further, the feature extraction of the data sample by the feature extraction method is specifically as follows: when the data sample is an image, image feature extraction is performed based on the SIFT or GIST method; when the data sample is text, text feature extraction is performed based on the BoW method.
[0057] Further, the data sample includes an image to be retrieved or text to be retrieved. When the data sample is an image, the retrieved data is the text corresponding to the image; when the data sample is text, the retrieved data is the image corresponding to the text.
[0058] Further, the calculation of the similarity value between the hash code representation of the data sample and the sample hash codes in the database to be retrieved is specifically as follows: calculate the Hamming distance between the hash code representation of the data sample and the sample hash codes in the database to be retrieved, sort the samples in the database in ascending order based on the distance value, and select the top k samples as the retrieval results, where k is an integer not less than 1.
[0059] Specifically, for the sake of easy understanding, the following combines the accompanying drawings and specific examples (i.e., the retrieval of large-scale social multimedia data) to elaborate on the solution described in this embodiment in detail:
[0060] The data used in the method described in this embodiment are multimedia data such as images and texts published by users on the social multimedia platform.
[0061] The training process of the hash retrieval model (i.e., the hash function) is specifically represented as follows:
[0062] Taking the multimedia data such as images and texts in the database as training data, first extract image features such as SIFT / GIST for the image data, extract text features such as BoW for the text data, and then through the training of the hash retrieval algorithm, a unified hash code representation B of the multimedia data such as images and texts, as well as the hash function H1 for the query samples corresponding to the image modality, the hash function H2 for the query samples corresponding to the text modality, etc. can be obtained.
[0063] The retrieval process and results are specifically as follows:
[0064] 1) Image to text. The user inputs an image as a query sample. First, convert the query picture into a hash code representation b through the hash function H1, then calculate the Hamming distance between b and the database hash code B and sort, and return the top-k texts with small distances as the retrieval results.
[0065] 2) Text to image. The user enters a piece of text as a query sample. The query text is first converted into a hash code b through the hash function H2. Then the Hamming distance between b and the database hash code B is calculated and sorted. The top-k images with the smallest distance are returned as the retrieval results.
[0066] Furthermore, the training process of the hash retrieval model is described in detail below. Specifically, the training process includes the following steps:
[0067] Step 1: (Mark enhancement) Input the category words into the word vector model and output them as category attributes; construct an objective function to migrate the category attributes to the mark enhancement proposed in the present invention; solve the objective function using a three-step iterative strategy to obtain the mark distribution, orthogonal rotation matrix and projection matrix generated by the mark enhancement;
[0068] Step 2: (Hash code learning) Use the tag distribution generated in step 1 to construct a finer-grained similarity matrix, called tag distribution similarity, and apply it to the widely used inner product adaptive objective function, and introduce an intermediate variable to achieve efficient optimization; achieve efficient solution through selective iterative update to obtain hash codes and intermediate variables;
[0069] Step 3: (Hash function learning) Use the hash code generated in step 2 as supervision information to train the linear mapping from the features of each modality to the hash code;
[0070] Step 4: (Cross-modal retrieval) Use the linear mapping of the hash code obtained in step 3 to calculate the hash code of the new sample, and calculate its Hamming distance with the public hash code to obtain similar instances across modalities.
[0071] Specifically, the specific process of step 1 is:
[0072] Step 1.1: Substitute the category words o in the dataset into {o1; o2; ....; o c} (where the classifier is the class name of each class, which comes with the dataset), input into the Word2Vec model pre-trained for natural language processing on Wikipedia, and output the class attribute Among them, c is the number of categories of training samples, and k is the dimension of category attributes;
[0073] Step 1.2: The objective function constructed here mainly consists of two parts: 1) Introducing a rotation matrix Make markers distributed With logical label L∈{0,1} c×n Share as much information as possible; 2) Assume that there is a k-dimensional category representation space, which is obtained by projecting the logical label L, that is, P TL. Meanwhile, this category indicates that the space can be decomposed into a semantic category space, i.e., the product of the label distribution D and the semantic category attribute A.
[0074]
[0075] Among them, is the projection matrix, and I is the identity matrix;
[0076] Step 1.3: Use a three-step iterative strategy to solve the objective function proposed in Step 1.2;
[0077] 1) Update R: Fix the hash code B and the projection matrix P unchanged. Equation (1) can be transformed into the following sub-problem regarding R:
[0078]
[0079] Performing singular value decomposition on LD -1 can obtain Thus, the optimal solution of Equation (2) is:
[0080]
[0081] 2) Update P: Let the derivative of the objective function with respect to P be zero, and the optimal solution of P is obtained as:
[0082] P = (αLL T + θI) -1 αLD T A, (4)
[0083] 3) Update D: Take the derivative of Equation (1) with respect to D and set the derivative to zero, and the optimal D can be obtained as:
[0084] D = (αAA T + I) -1 (R T L + αAP T L), (5)
[0085] Iterate the above steps until convergence, and then perform the following soft normalization on D:
[0086] The specific process of Step 2:
[0087] Step 2.1: Apply the label distribution generated in Step 1 to construct a finer-grained similarity matrix, which is called label distribution similarity. Its definition is as follows:
[0088]
[0089] Among them, It is the labeled distribution matrix normalized by column l2 norm. Apply the labeled distribution similarity to the widely used inner product adaptation objective function:
[0090]
[0091] where B is the common hash code of the training set to be learned, and r is the number of bits of the hash code to be learned;
[0092] Step 2.2: Introduce an intermediate variable to replace one hash code B in the symmetric inner product operation. Additionally, to avoid the F offset, apply orthogonal and balance constraints to F, and use a regularization term to minimize the difference between F and B. Finally, formula (7) is transformed into the following form:
[0093]
[0094] Step 2.3: Efficiently solve the formula in Step 2.2 through selective iterative updates
[0095] 1) Update F: Fix B, and formula (8) is transformed into the following sub-problem about F:
[0096]
[0097] Due to the constraint FF T = nI and B ∈ {-1, 1} r×n , the above formula can be equivalently written in the form of the following matrix trace:
[0098]
[0099] To solve problem (10), we define and let Z = rBS + ωB. According to the definition of the similarity matrix in formula (7), we have
[0100]
[0101] Then, perform matrix decomposition on ZJZ T to obtain:
[0102]
[0103] where Σ is a diagonal matrix composed of r' ≤ r positive eigenvalues, V is the matrix composed of the corresponding eigenvectors, is the matrix composed of the eigenvectors corresponding to the remaining zero eigenvalues. Perform Schmidt orthogonalization on to obtain the orthogonal matrix Performing Schmidt orthogonalization on a random matrix yields a random orthogonal matrix Define U = JZ T V∑ -1 / 2, The solution of formula (10) is:
[0104]
[0105] 2) Update B: When F is fixed, problem (8) becomes the following form:
[0106]
[0107] We can easily obtain the optimal solution of the above formula as:
[0108] B = sgn(rFS + ωF), (15)
[0109] According to formula (7), we have
[0110]
[0111] Repeat the above two steps until convergence, and the finally obtained is the hash code of the training set;
[0112] The specific process of step 3:
[0113] Use the hash code generated in step 2 as the supervision information to train a linear mapping from the features of the l-th modality to the hash code, and the objective function is as follows:
[0114]
[0115] Among them, λ is the penalty coefficient of the regularization term to avoid overfitting, and W (l) is the mapping matrix of the l-th modality. By taking the derivative of W (l) and setting the derivative to zero, the optimal solution of formula (17) can be obtained as:
[0116] W (l) = BX (l)T (X (l) X (l)T + λI) -1 , (18)
[0117] Given a query sample q (l) of the l-th modality, its hash code can be obtained through the following hash function:
[0118] H l (q (l) ) = sgn(W (l) q (l) ), (19)
[0119] The specific process of step 4:
[0120] First, use the test set labels and training set labels to find other modal samples that are consistent with the modal to be retrieved; then, given a query sample q of the l-th modality through the hash function obtained in step 3 (l) , calculate its hash code b, calculate the Hamming distance between b and the common hash code B, and sort the Hamming distances; finally, output in order the samples in the training set that are consistent with other modalities different from the modality to be retrieved to obtain the retrieval result.
[0121] Embodiment 2:
[0122] The purpose of this embodiment is to provide a cross-modal retrieval system for social multimedia data based on label enhancement.
[0123] A cross-modal retrieval system for social multimedia data based on label enhancement includes:
[0124] A data acquisition unit for acquiring data samples to be retrieved;
[0125] A feature extraction unit for extracting features from the data samples based on a feature extraction method;
[0126] An encoding unit for encoding based on the features of the data samples to be retrieved by using the hash function of the corresponding modality obtained through pre-training to obtain the hash code representation of the data samples;
[0127] A retrieval unit for calculating the similarity value between the hash code representation of the data samples and the hash codes of the samples in the database to be retrieved; obtaining the corresponding retrieval result based on the similarity value;
[0128] Among them, for the training of the hash function in different modalities, specifically: based on the sample data categories in the database to be retrieved, obtain the corresponding category words, and convert the category words into category attributes through a word vector model; based on the category attributes and a pre-constructed objective function, solve to obtain the label distribution with enhanced labels; construct a fine-grained similarity matrix based on the label distribution, and construct an inner product adaptation objective function based on the similarity matrix. By solving the inner product adaptation objective function, obtain the hash codes of the training samples; use the hash codes of the training samples as supervision information to train the hash functions in different modalities.
[0129] Furthermore, the system in this embodiment corresponds to the method in Embodiment 1, and its technical details have been described in detail in Embodiment 1, so they will not be repeated here.
[0130] In more embodiments, there is also provided:
[0131] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0132] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0133] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0134] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0135] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0136] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0137] The cross-modal retrieval method and system for social multimedia data based on tag enhancement provided in the above embodiment can be implemented and have broad application prospects.
[0138] The above are only the preferred embodiments of this disclosure and are not used to limit this disclosure. For those skilled in the art, this disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A cross-modal retrieval method for social multimedia data based on tag enhancement, characterized in that, Including: Obtain the data samples to be retrieved; Extract features from the data samples based on a feature extraction method; Based on the features of the data samples to be retrieved, encode them using the pre-trained hash function of the corresponding modality to obtain the hash code representation of the data samples; Calculate the similarity value between the hash code representation of the data samples and the sample hash codes in the database to be retrieved; obtain the corresponding retrieval results based on the similarity value; Among them, for the training of the hash function in different modalities, specifically: based on the sample data categories in the database to be retrieved, obtain the corresponding category words, and convert the category words into category attributes through a word vector model; based on the category attributes and a pre-constructed objective function, solve to obtain a label-enhanced label distribution, specifically using the following objective function: Among them, is the projection matrix, I is the identity matrix, is the rotation matrix, A is the class attribute, L is the logical label, D is the label distribution, α is the balance parameter, θ is the penalty coefficient of the regularization term, c is the number of classes of the training samples, and k is the dimension of the class attribute; the logical label L is a label matrix with values of 0 and 1, where 1 indicates that the sample belongs to the class and 0 indicates the opposite; the parameter D is the label distribution, which has the same shape as L; Construct a fine-grained similarity matrix based on the label distribution, and construct an inner product adaptation objective function based on the similarity matrix. By solving the inner product adaptation objective function, obtain the hash code of the training samples; use the hash code of the training samples as supervision information to train the hash functions in different modalities.
2. The cross-modal retrieval method for social multimedia data based on tag enhancement according to claim 1, characterized in that, When constructing the inner product adaptation objective function based on the similarity matrix, at the same time, introduce an intermediate variable to replace the hash code in the inner product operation of the inner product adaptation objective function, and introduce a regularization term to minimize the difference between the introduced intermediate variable and the hash code.
3. The cross-modal retrieval method for social multimedia data based on tag enhancement according to claim 1, characterized in that, When using the hash code of the training samples as supervision information to train the hash functions in different modalities, specifically use the following objective function: in, To avoid overfitting, the penalty coefficient of the regularization term is It is l The mapping matrix of the mode, B is the hash code of the training set, For training sample l The characteristic matrix of the mode.
4. The cross-modal retrieval method for social multimedia data based on tag enhancement according to claim 1, characterized in that, When extracting features from the data samples based on the feature extraction method, specifically: when the data sample is an image, extract image features based on the SIFT or GIST method; when the data sample is text, extract text features based on the BoW method.
5. The cross-modal retrieval method for social multimedia data based on tag enhancement according to claim 1, characterized in that, The data samples include the images or texts to be retrieved. When the data sample is an image, the retrieved data is the text corresponding to the image; when the data sample is text, the retrieved data is the image corresponding to the text.
6. The cross-modal retrieval method for social multimedia data based on tag enhancement according to claim 1, characterized in that, When calculating the similarity value between the hash code representation of the data samples and the sample hash codes in the database to be retrieved, specifically: calculate the Hamming distance between the hash code representation of the data samples and the sample hash codes in the database to be retrieved, sort the samples in the database from smallest to largest based on the distance value, and select the first G samples as the retrieval results, where G is an integer not less than 1.
7. A cross-modal retrieval system for social multimedia data based on tag enhancement, characterized in that, Including: A data acquisition unit for obtaining the data samples to be retrieved; A feature extraction unit for extracting features from the data samples based on a feature extraction method; An encoding unit for encoding the features of the data samples to be retrieved using the pre-trained hash function of the corresponding modality to obtain the hash code representation of the data samples; A retrieval unit for calculating the similarity value between the hash code representation of the data samples and the sample hash codes in the database to be retrieved; obtaining the corresponding retrieval results based on the similarity value; Among them, for the training of the hash function in different modalities, specifically: based on the sample data categories in the database to be retrieved, the corresponding category words are obtained, and the category words are transformed into category attributes through a word vector model; based on the category attributes and a pre-constructed objective function, a labeled distribution with enhanced labels is solved, and the following objective function is specifically adopted: Among them, is the projection matrix, I is the identity matrix, is the rotation matrix, A is the class attribute, L is the logical label, D is the label distribution, α is the balance parameter, θ is the penalty coefficient of the regularization term, c is the number of classes of the training samples, and k is the dimension of the class attribute; the logical label L is a label matrix with values of 0 and 1, where 1 indicates that the sample belongs to the class and 0 indicates the opposite; the parameter D is the label distribution, which has the same shape as L; A fine-grained similarity matrix is constructed based on the labeled distribution, and an inner product adaptation objective function is constructed based on the similarity matrix. By solving the inner product adaptation objective function, the hash codes of the training samples are obtained; based on the hash codes of the training samples as supervision information, the hash functions in different modalities are trained.
8. An electronic device, comprising a memory, a processor and a computer program running on the memory, characterized in that, When the processor executes the program, it implements a cross-modal retrieval method for social multimedia data based on label enhancement as described in any one of claims 1-6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a cross-modal retrieval method for social multimedia data based on label enhancement as described in any one of claims 1-6.