Structured Multimodal Retrieval Method and System for Multiple Multimedia Retrieval Tasks
Through the structured multimodal search method, we learn the linear projection matrix and transformation matrix for multimedia retrieval, and generate structured hash codes, which solves the problem that the multimedia retrieval model cannot be flexiblely expanded in the existing technology, and achieves support for a variety of multimedia retrieval tasks and efficient retrieval effects.
Patent Information
- Application Number
- CN202310001747.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-01-03
AI Technical Summary
Existing hash-based multimedia retrieval models can usually only be designed for specific multimedia retrieval tasks and cannot be flexibly extended to other multimedia retrieval tasks. It is especially difficult to support both cross-modal retrieval and composite multimodal retrieval at the same time.
A structured multimodal retrieval method for multiple multimedia search tasks is proposed. By acquiring the depth features of image modality and text modality, the objective function is constructed to learn linear projection matrix and transformation matrix for cross-modal and multimodal retrieval, and a structured hash code is generated to support a variety of multimedia search tasks.
It realizes the flexibility and accuracy of multimedia search for specific features of each mode in the multimedia search system, supports cross-modal retrieval, and fuses heterogeneous modal features, supports composite multimodal retrieval, which improves the flexibility and accuracy of multimedia search.
Smart Images

Figure CN115952309B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimedia retrieval technology, and in particular to a structured multimodal retrieval method and system for multiple multimedia retrieval tasks. Background Art
[0002] With the development of multimedia technology, how to retrieve semantically relevant multimedia data from large-scale heterogeneous data has become a hot topic for researchers. The hashing method can map high-dimensional data into a low-dimensional Hamming space, and then effectively measure the similarity between samples by simply calculating the Hamming distance between hash codes. With the advantages of small storage space and fast calculation speed, the hashing method has received more and more attention in multimedia retrieval.
[0003] Cross-modal hashing and compound multimodal hashing are the two main hash-based multimedia retrieval techniques. Cross-modal hashing aims to learn a shared hash code to support cross-modal retrieval tasks, and its retrieval tasks are performed across heterogeneous modalities. Compound multimodal hashing uses the complementarity of different multimodal features to learn fused hash codes during the training and retrieval stages, and is mainly designed for compound multimodal retrieval tasks. Previous cross-modal and compound multimodal hashing models have achieved very significant retrieval performance. However, existing hash-based retrieval models are usually designed only for specific multimedia retrieval tasks, that is, these methods are either designed only for cross-modal retrieval tasks or only for compound multimedia retrieval tasks. Due to the different retrieval model settings, it cannot be flexibly extended to other multimedia retrieval tasks. In order to meet the requirements of supporting cross-modal retrieval and compound multimodal retrieval at the same time, it is necessary to configure two hash models in the multimedia retrieval system and store two different types of hash codes at the same time, which is very inconvenient.
[0004] Therefore, how to support multiple multimedia retrieval tasks while flexibly retaining the specific features of each modality in the multimedia data, supporting cross-modal retrieval tasks, integrating heterogeneous modal features, and supporting multimodal retrieval tasks is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical task of the present invention is to provide a structured multimodal retrieval method and system for multiple multimedia retrieval tasks to solve the problem of how to support multiple multimedia retrieval tasks while flexibly retaining the specific features of each modality in the multimedia data, supporting cross-modal retrieval tasks, and integrating heterogeneous modal features to support multimodal retrieval tasks.
[0006] The technical task of the present invention is achieved in the following manner: a structured multimodal retrieval method for multiple multimedia retrieval tasks, the method is specifically as follows:
[0007] Obtain a multimodal retrieval dataset including an image modality and a text modality, and divide the multimodal retrieval dataset into a training dataset, a test dataset, and a database dataset;
[0008] Input the original data of the image modality and the text modality into a deep feature extraction model respectively to extract features from the original data, so as to obtain the deep features of the image modality and the deep features of the text modality;
[0009] Construct the objective function of the structured multimodal hashing method for multiple multimedia retrieval tasks on the training dataset;
[0010] Obtain the linear projection matrix \(R\) of the \(v\)-th modality for cross-modal retrieval through the objective function of the structured multimodal hashing method for multiple multimedia retrieval tasks (v) and the transformation matrix \(W\) for multimodal retrieval;
[0011] During online query, use the objective function of the structured multimodal hashing method for multiple multimedia retrieval tasks, and according to the linear projection matrix \(R\) of the \(v\)-th modality (v) and the transformation matrix \(W\) for multimodal retrieval to obtain the hash codes of the samples in the test dataset and the database dataset, and obtain the Hamming distance between the hash code of each test sample in the test dataset and the hash code of the samples in the database dataset. Sort the samples in the database dataset in ascending order according to the corresponding Hamming distance to obtain the retrieval result.
[0012] Preferably, the specific steps to obtain the deep features of the image modality and the deep features of the text modality are as follows:
[0013] Image modality: Input the image modality into the VGG-16 network model to extract 4096-dimensional image features;
[0014] Text modality: Use the Bag-of-words model to extract 1386-dimensional text features from the labels.
[0015] More preferably, the specific steps to construct the objective function of the structured multimodal hashing method for multiple multimedia retrieval tasks on the training dataset are as follows:
[0016] Through the objective function \(f\) spe Preserve the unique attributes of each modality;
[0017] Through the objective function \(f\) com Preserve the complementarity of multimodal features;
[0018] Through the objective function \(f\) sup Add label guidance to automatically assign semantic information to the shared transformation matrix to bridge the differences between different modalities;
[0019] Construct the final objective function, the formula is:
[0020] f = f spe + f com + f sup ;
[0021] Among them, f spe maps each modality to a separate latent space to learn modality-invariant features, so as to better preserve the specific attributes of each modality; f com sets appropriate weights to make full use of the complementarity of multimodal features; f sup simultaneously uses labels and pair-wise similarity matrix to generate more discriminative hash codes, thus maintaining semantic similarity between the high-dimensional space and the Hamming space.
[0022] Preferably, through the objective function f spe the features of each modality are preserved as follows:
[0023] Construct a non-linear feature embedding The formula is as follows:
[0024]
[0025] Among them, are m anchor points randomly selected from the training dataset; σ is the Gaussian kernel parameter; this process can effectively maintain the correlation of the v-th modality samples; in particular, the storage and computational costs of this process are only O(n), greatly reducing the complexity of optimization;
[0026] After obtaining the non-linear embedding of each modality, construct the objective function f spe : By establishing the relationship between the hash code and the original data, use the modality-specific projection matrix to effectively learn a single hash code; the specific formula is:
[0027]
[0028] s.t. B (v) ∈{-1,1} N×r ;
[0029] Among them, B (v) is the modality-specific hash code of the v-th modality learned, r represents the length of the hash code corresponding to each modality in the structured hash code, V represents the number of modalities; R (v) is the linear projection matrix of the v-th modality; γ (v) is the balance parameter of the v-th modality; N is the total number of samples in the training stage; is the v-th modal feature matrix; d (v) is the feature dimension;
[0030] The complementarity of multi-modal features is preserved through the objective function f com as follows:
[0031] Construct a collaborative non-linear multi-modal feature mapping The formula is as follows:
[0032]
[0033] where consists of V individual feature embeddings, including information from V modalities; in multi-modal hashing learning, by assigning reasonable weights μ (v) , the importance of each different modal feature is effectively measured; each part of the structured hash code contains rich intra-modal information, thus effectively improving the accuracy of cross-modal retrieval; at the same time, by collaborating the structured hash codes to represent the entire multi-modal data, heterogeneous modal features can be effectively fused to achieve composite multi-modal retrieval;
[0034] Construct the objective function f com , and the specific formula is:
[0035]
[0036]
[0037] where [B (1) ; …; B (V) is the defined structured hash code; W is the non-linear projection matrix for multi-modal retrieval tasks; θ is the balance parameter; is the collaborative non-linear multi-modal feature mapping of the original input data;
[0038] Construct the objective function f sup The formula is as follows:
[0039]
[0040] where is the transformation matrix; is the label matrix; [B (1) ; …; B (V) is the structured hash code; is the pairwise similarity matrix; α and β are balance parameters.
[0041] Preferably, the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks is obtained by using an iterative solution algorithm, that is, fixing other variables and solving for one variable to obtain the final optimal solution. The specific optimization process is as follows:
[0042] Update R by fixing other variables (v) , the objective function becomes:
[0043]
[0044] For R (v) Take the partial derivative and set it to zero to get:
[0045]
[0046] Update C by fixing other variables, the objective function becomes:
[0047]
[0048] Take the partial derivative of C and set it to zero to get:
[0049] C = (αYB T B + βI Vr ) -1 (αrB T SY + βB T Y)(Y T Y) -1 ;
[0050] For convenience of representation, let B = [B (1) ; …; B (v) ;
[0051] Update B by fixing other variables: First, solve for the v-th modal hash code B (v) by fixing other variables except B (v) . Remove the terms unrelated to B (v) from the objective function and simplify the function to:
[0052]
[0053] Take the partial derivative of B (v) and set it to zero to get:
[0054]
[0055] Update B by fixing other variables, the objective function becomes:
[0056]
[0057] Rewrite as:
[0058]
[0059] where tr(B T B) is a constant; in order to effectively calculate the hash code and avoid symmetric matrix decomposition, min B∈{-1,1} tr(αBCY T YC T B T - 2αrSYC T B T + βB T B - 2βYC T B t ) is rewritten as:
[0060]
[0061] Adopting an asymmetric hash learning strategy, use the variable to replace one B; at the same time, add D to measure the difference between B and , and simplify the last term in to:
[0062]
[0063] Take the partial derivative of B and set it to zero, obtaining:
[0064]
[0065] Fix other variables and update The objective function becomes:
[0066]
[0067] The update formula of
[0068]
[0069] Fix other variables and update D and η. According to the ALM algorithm, we get:
[0070]
[0071] where the parameter ρ is obtained through tuning to get the optimal parameter;
[0072] Fix other variables and update W, obtaining:
[0073]
[0074] Fix other variables and find the adaptive weight μ, obtaining:
[0075]
[0076] Preferably, cross-modal retrieval is performed using the hash codes generated by the hash function. For the samples to be queried, the prediction formula for their hash codes is as follows:
[0077]
[0078] where, represents the non-linear embedding of N q query samples; the linear projection matrix R (v) is obtained in the training phase and is directly used for online cross-modal retrieval;
[0079] Predict the hash codes of any multi-modal query samples for the composite multi-modal retrieval problem. The formula is as follows:
[0080]
[0081] where, represents the collaborative multi-modal mapping of N q query samples; the linear projection matrix W is obtained in the training phase.
[0082] Preferably, the Hamming distance is as follows:
[0083] Select any test sample in the training dataset, and obtain the Hamming distance between the hash code of the test sample and the hash codes of all samples in the database dataset;
[0084] Sort the samples in the database dataset according to the Hamming distance. Those with a distance less than the set threshold from the test sample are ranked in the front, and those with a distance greater than the set threshold from the test sample are ranked in the back, so as to verify the retrieval accuracy.
[0085] A structured multi-modal retrieval system for multiple multimedia retrieval tasks. The system includes,
[0086] A data preprocessing module for obtaining a multi-modal retrieval dataset. Among them, each sample in the multi-modal retrieval dataset includes paired image-modal and text-modal data, and divides the multi-modal retrieval dataset into a training dataset, a test dataset, and a database dataset;
[0087] A deep feature representation module for respectively inputting the original data of the image modality and the text modality into a deep feature extraction model for feature extraction, and using the extracted features as training features, test features, and database features for objective function construction, hash function learning, and online modality hash retrieval;
[0088] An objective function construction module for constructing the total objective function of the structured multi-modal hash method for multiple multimedia retrieval tasks through the training dataset;
[0089] A hash function learning module, which is used to solve the objective function by using an iterative optimization method;
[0090] An online modal hashing retrieval module, which is used to construct and utilize the objective function of online hashing, obtain the hash codes of samples in the test data set and the database data set, obtain the Hamming distance between the hash code of the test sample of each test data set and the hash code of the sample of the database data set, and sort the samples of the database data set in ascending order according to the corresponding Hamming distance to obtain the retrieval result.
[0091] An electronic device, comprising: a memory and at least one processor;
[0092] Wherein, a computer program is stored on the memory;
[0093] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the structured multi-modal retrieval method for multiple multimedia retrieval tasks as described above.
[0094] A computer-readable storage medium, in which a computer program is stored, and the computer program can be executed by a processor to implement the structured multi-modal retrieval method for multiple multimedia retrieval tasks as described above.
[0095] The structured multi-modal retrieval method and system for multiple multimedia retrieval tasks of the present invention have the following advantages:
[0096] (1) The present invention uses the VGG-16 model and the bag-of-words model (BoW) to extract the deep feature representations of the image and text modalities, and uses them as the input features of the subsequent structured multi-modal hashing model; in order to fully exploit the semantic relevance and complementary information of multi-modal data, a binary hash code is generated independently for each modality, and at the same time, considering the complementarity of multi-modal data, a structured hash code is learned simultaneously. The structured hash code can not only flexibly retain the specific features of each modality in the multimedia data, support cross-modal retrieval, but also fuse heterogeneous modality features to support composite multi-modal retrieval;
[0097] (2) The present invention proposes a unified multi-modal hashing learning framework, which can fuse and strengthen heterogeneous modality data and learn a structured hash code; in particular, the generated hash code can simultaneously process multiple multimedia retrieval tasks;
[0098] (3) The present invention makes full use of available modal features through an independent projection strategy and a collaborative hash code learning strategy, and constructs a shared Hamming space using the shared information and modality-specific information in heterogeneous multimodal data. On this basis, an effective iterative optimization strategy is proposed to directly learn hash codes in an efficient and fast manner. The experimental results on cross-modal and composite multimodal retrieval tasks show that the present invention has good performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The present invention will be further described below with reference to the accompanying drawings.
[0100] Attached Figure 1 is a flowchart of a structured multimodal retrieval method for multiple multimedia retrieval tasks;
[0101] Attached Figure 2 is a schematic diagram showing the differences between the structured multimodal hashing for multiple multimedia retrieval tasks of the present invention and existing cross-modal hashing models and composite multimodal hashing models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0102] The structured multimodal retrieval method and system of the present invention for multiple multimedia retrieval tasks will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0103] Embodiment 1:
[0104] As shown in the attached Figure 1 figure, the structured multimodal retrieval method and system of the present invention for multiple multimedia retrieval tasks are as follows:
[0105] S1. Obtain a multimodal retrieval dataset including an image modality and a text modality, and divide the multimodal retrieval dataset into a training dataset, a test dataset, and a database dataset;
[0106] For the multimodal training set from V modalities, is the v-th modal feature matrix, where d (v) is the feature dimension, and N is the total number of training samples. For ease of description, the present invention takes a multimodal retrieval dataset including two modalities, i.e., images and texts, as an example. Specifically, X (1) and X (2) respectively represent the image and text modal features. The objective in this embodiment is to learn the structured hash code where r represents the length of the hash code corresponding to each modality in the structured hash code, and V represents the number of modalities.
[0107] S2. Input the original data of the image modality and the text modality into a deep feature extraction model respectively to extract features from the original data, so as to obtain the deep features of the image modality and the deep features of the text modality;
[0108] S3. Construct the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks on the training dataset;
[0109] S4. Obtain the linear projection matrix \(R\) of the \(v\)-th modality for cross-modal retrieval through the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks (v) and the transformation matrix \(W\) for multi-modal retrieval;
[0110] S5. During online query, utilize the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks. According to the linear projection matrix \(R\) of the \(v\)-th modality (v) and the transformation matrix \(W\) for multi-modal retrieval, obtain the hash codes of the samples in the test dataset and the database dataset, and obtain the Hamming distance between the hash code of each test sample in the test dataset and the hash code of the samples in the database dataset. Sort the samples in the database dataset in ascending order according to the corresponding Hamming distance to obtain the retrieval result.
[0111] The specific process of obtaining the deep features of the image modality and the deep features of the text modality in step S2 of this embodiment is as follows:
[0112] S201. Image modality: Input the image modality into the VGG-16 network model to extract 4096-dimensional image features;
[0113] S202. Text modality: Use the Bag-of-words model to extract 1386-dimensional text features from the labels.
[0114] The specific process of constructing the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks on the training dataset in step S3 of this embodiment is as follows:
[0115] S301. Through the objective function \(f\) spe Preserve the unique attributes of each modality;
[0116] S302. Through the objective function \(f\) com Preserve the complementarity of multi-modal features;
[0117] S303. Through the objective function \(f\) sup Plus label guidance, automatically assign semantic information to the shared transformation matrix to bridge the differences between different modalities;
[0118] S304. Construct the final objective function, and the formula is:
[0119] \(f = f\) spe +\(f\) com +\(f\) sup ;
[0120] Among them, f spe Map each modality to a separate latent space to learn modality-invariant features, so as to better preserve the unique attributes of each modality; f com Set appropriate weights to make full use of the complementarity of multimodal features; f sup At the same time, utilize labels and pairwise similarity matrices to generate more discriminative hash codes, thereby maintaining semantic similarity between the high-dimensional space and the Hamming space.
[0121] In step S301 of this embodiment, through the objective function f spe The features of each modality are saved as follows:
[0122] S30101. Construct a non-linear feature embedding The formula is as follows:
[0123]
[0124] Among them, are m anchor points randomly selected from the training dataset; σ is the Gaussian kernel parameter; this process can effectively maintain the correlation of the v-th modality samples; in particular, the storage and computational costs of this process are only O(n), greatly reducing the complexity of optimization;
[0125] S30102. After obtaining the non-linear embedding of each modality, construct the objective function f spe : By establishing the relationship between the hash code and the original data, effectively learn a single hash code using the modality-specific projection matrix; the specific formula is:
[0126]
[0127] s.t. B (v) ∈{-1,1} N×r ;
[0128] Among them, B (v) is the modality-specific hash code of the v-th modality learned, r represents the length of the hash code corresponding to each modality in the structured hash code, V represents the number of modalities; R (v) is the linear projection matrix of the v-th modality; γ (v) is the balance parameter of the v-th modality; N is the total number of samples in the training phase; is the v-th modality feature matrix; d (v) is the feature dimension;
[0129] In step S302 of this embodiment, through the objective function f com The complementarity of multimodal features is saved as follows:
[0130] S30201. Construct collaborative non - linear multi - modal feature mapping The formula is as follows:
[0131]
[0132] Wherein, is composed of V individual feature embeddings, including information from V modalities; in multi - modal hashing learning, by assigning reasonable weights μ (v) , effectively measure the importance of each different modal feature; each part of the structured hash code contains rich intra - modal information, thus being able to effectively improve the accuracy of cross - modal retrieval; meanwhile, by combining the structured hash codes to represent the entire multi - modal data, heterogeneous modal features can be effectively fused to achieve composite multi - modal retrieval;
[0133] S30202. Construct the objective function f com , and the specific formula is:
[0134]
[0135]
[0136] Wherein, [B (1) ; …; B (V) is the defined structured hash code; W is a non - linear projection matrix for multi - modal retrieval tasks; θ is a balance parameter; is the collaborative non - linear multi - modal feature mapping of the original input data;
[0137] The objective function f constructed in step S303 of this embodiment sup The formula is specifically as follows:
[0138]
[0139] Wherein, is the transformation matrix; is the label matrix; [B (1) ; …; B (V) is the structured hash code; is the pairwise similarity matrix; α and β are balance parameters.
[0140] In this embodiment, the objective function of the structured multi - modal hashing method for multiple multimedia retrieval tasks in step S4 is obtained by using an iterative solution algorithm, that is, fixing other variables and solving one variable to obtain the final optimal solution. The specific optimization process is as follows:
[0141] S401. Fix other variables and update R (v) , and the objective function becomes:
[0142]
[0143] For R (v) Take the partial derivative with respect to it and set it to zero, we get:
[0144]
[0145] S402. Fix other variables and update C. The objective function becomes:
[0146]
[0147] Take the partial derivative with respect to C and set it to zero, we get:
[0148] C = (αYB T B + βI Vr ) -1 (αrB T SY + βB T Y)(Y T Y) -1 ;
[0149] For convenience of representation, let B = [B (1) ; …; B (v) ;
[0150] S403. Fix other variables and update B: First, solve for the v-th modal hash code B (v) by fixing other variables except B (v) . Remove the terms irrelevant to B (v) from the objective function, and simplify the function to:
[0151]
[0152] Take the partial derivative with respect to B (v) and set it to zero, we get:
[0153]
[0154] Fix other variables and update B. The objective function becomes:
[0155]
[0156] Rewrite as:
[0157]
[0158] where tr(B T B) is a constant; To effectively calculate the hash code and avoid symmetric matrix decomposition, let min B∈{-1,1} tr(αBCYT YC T B T -2αrSYC T B T +βB T B - 2βYC T B t ) Rewrite as:
[0159]
[0160] Adopt an asymmetric hashing learning strategy and use the variable to replace one B; at the same time, add D to measure the difference between B and and simplify the last term in to:
[0161]
[0162] Take the partial derivative of B and set it to zero to get:
[0163]
[0164] S404. Fix other variables and update The objective function becomes:
[0165]
[0166] The update formula of
[0167]
[0168] S405. Fix other variables and update D and η. According to the ALM algorithm, we get:
[0169]
[0170] where the parameter ρ is obtained through parameter tuning to get the optimal parameter;
[0171] S406. Fix other variables and update W to get:
[0172]
[0173] S407. Fix other variables and find the adaptive weight μ to get:
[0174]
[0175] In this embodiment, for cross-modal retrieval using the hash code generated by the hash function in step S5, for the sample to be queried, the prediction formula of its hash code is as follows:
[0176]
[0177] Among them, represents the non - linear embedding of N q query samples; the linear projection matrix R (v) is obtained in the training stage and is directly used for online cross - modal retrieval;
[0178] In step S5 of this embodiment, the hash codes of any multi - modal query samples are predicted for the compound multi - modal retrieval problem, and the formula is as follows:
[0179]
[0180] Among them, represents the collaborative multi - modal mapping of N q query samples; the linear projection matrix W is obtained in the training stage.
[0181] The Hamming distance in step S5 of this embodiment is specifically as follows:
[0182] S501. Select any test sample in the training dataset, and obtain the Hamming distance between the hash code of the test sample and the hash codes of all samples in the database dataset;
[0183] S502. Sort the samples in the database dataset according to the Hamming distance, with the samples whose distance from the test sample is less than the set threshold ranked in the front and the samples whose distance from the test sample is greater than the set threshold ranked in the back, so as to verify the retrieval accuracy.
[0184] This embodiment uses TopK - precision and Mean Average Precision (MAP) as evaluation metrics. The larger the values of TopK - precision and Mean Average Precision (MAP), the better the retrieval performance; the specific definitions are as follows:
[0185] 1). TopK - precision: TopK - precision is used to reflect the change of retrieval precision with the number of retrieved samples; when retrieving the first K samples, TopK - precision represents the proportion of relevant samples among the K samples.
[0186] 2). MAP: Given a retrieval sample set, the average precision (AP) of each retrieval sample is defined as:
[0187]
[0188] Wherein, R is the total number of retrieved samples returned, N is the total number of samples related to the query returned, P(r) represents the precision of the first r retrieved results. If the r-th retrieved sample is related to the query sample, then δ(r) = 1; otherwise, δ(r) = 0. The average value of the AP values of all samples is the MAP.
[0189] Embodiment 2:
[0190] This embodiment provides a structured multi-modal retrieval system for multiple multimedia retrieval tasks. The system includes:
[0191] A data preprocessing module, configured to obtain a multi-modal retrieval data set. Each sample in the multi-modal retrieval data set includes paired image modality and text modality data, and divide the multi-modal retrieval data set into a training data set, a test data set, and a database data set;
[0192] A deep feature representation module, configured to respectively input the original data of the image modality and the text modality into a deep feature extraction model for feature extraction, and use the extracted features as training features, test features, and database features for target function construction, hash function learning, and online modality hash retrieval;
[0193] A target function construction module, configured to construct a total target function of a structured multi-modal hash method for multiple multimedia retrieval tasks through the training data set;
[0194] A hash function learning module, configured to solve the target function by using an iterative optimization method;
[0195] An online modality hash retrieval module, configured to construct and utilize an online hash target function, obtain the hash codes of samples in the test data set and the database data set, obtain the Hamming distance between the hash code of each test sample in the test data set and the hash code of the database data set samples, and sort the database data set samples in ascending order according to the corresponding Hamming distance to obtain the retrieval results.
[0196] Embodiment 3:
[0197] This embodiment also provides an electronic device, including: a memory and a processor;
[0198] Wherein, the memory stores computer execution instructions;
[0199] The processor executes the computer execution instructions stored in the memory, so that the processor executes the structured multi-modal retrieval method for multiple multimedia retrieval tasks in any embodiment of the present invention.
[0200] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0201] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory, the processor realizes various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash memory card, at least one magnetic disk storage period, flash memory device, or other volatile solid-state storage devices.
[0202] Embodiment 4:
[0203] This embodiment also provides a computer-readable storage medium, in which multiple instructions are stored. The instructions are loaded by the processor to make the processor execute the structured multi-modal retrieval method for multiple multimedia retrieval tasks in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program codes for implementing the functions in any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.
[0204] In this case, the program code read from the storage medium itself can implement the functions in any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0205] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0206] In addition, it should be clear that not only can the above-described functions of any one of the embodiments be achieved by executing the program code read by a computer, but also by causing an operating system or the like operating on the computer to perform part or all of the actual operations based on the instructions of the program code.
[0207] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then part and all of the actual operations are performed by causing a CPU or the like installed on the expansion board or the expansion unit to execute based on the instructions of the program code, thereby implementing the functions of any one of the above-described embodiments.
[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A structured multi-modal retrieval method for multiple multimedia retrieval tasks, characterized in that, The method is as follows: Obtain a multi-modal retrieval dataset including image modality and text modality, and divide the multi-modal retrieval dataset into a training dataset, a test dataset, and a database dataset; Input the original data of the image modality and the text modality into a deep feature extraction model respectively to extract features from the original data, so as to obtain the deep features of the image modality and the deep features of the text modality; Construct the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks on the training dataset; Obtain the linear projection matrix \(R\) of the \(v\)-th modality for cross-modal retrieval through the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks (v) and the transformation matrix \(W\) for multi-modal retrieval; During online query, by using the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks, according to the linear projection matrix R of the v-th modality (v) and the transformation matrix W for multi-modal retrieval, the hash codes of the samples in the test data set and the database data set are obtained, and the Hamming distance between the hash codes of the test samples of each test data set and the hash codes of the samples in the database data set is obtained. The samples in the database data set are sorted in ascending order according to the corresponding Hamming distance to obtain the retrieval result; Among them, the construction of the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks on the training dataset is as follows: Through the objective function f spe Save the unique attributes of each modality; Through the objective function f com Preserve the complementarity of multimodal features; Through the objective function f sup With the addition of label guidance, semantic information is automatically assigned to the shared transformation matrix to bridge the differences between different modalities; Construct the final objective function, and the formula is: f = f spe + f com + f sup ; Among them, f spe Maps each modality to a separate latent space to learn modality-invariant features, in order to better preserve the unique attributes of each modality; f com Sets weights to make full use of the complementarity of multimodal features; f sup At the same time, uses labels and pairwise similarity matrices to generate more discriminative hash codes, thereby maintaining semantic similarity between the high-dimensional space and the Hamming space; Use the hash codes generated by the hash function for cross-modal retrieval. For the sample to be queried, the prediction formula of its hash code is as follows: Among them, represents the non-linear embedding of N q query samples; the linear projection matrix R (v) is obtained in the training phase and is directly used for online cross-modal retrieval; Predict the hash codes of any multi-modal query samples for the compound multi-modal retrieval problem, and the formula is as follows: Among them, represents the collaborative multi-modal mapping of N q query samples; the linear projection matrix W is obtained in the training phase.
2. The structured multi-modal retrieval method for multiple multimedia retrieval tasks according to claim 1, characterized in that, The specific method for obtaining the deep features of the image modality and the deep features of the text modality is as follows: Image modality: Input the image modality into the VGG-16 network model to extract 4096-dimensional image features; Text modality: Use the Bag-of-words model to extract 1386-dimensional text features from the labels.
3. The structured multi-modal retrieval method for multiple multimedia retrieval tasks according to claim 1 or 2, characterized in that, Through the objective function f spe The features of each modality are saved as follows: Constructing Nonlinear Feature Embedding The formula is as follows: Among them, are m anchor points randomly selected from the training dataset; σ is the Gaussian kernel parameter; After obtaining the non-linear embeddings of each modality, construct the objective function f spe : By establishing the relationship between the hash code and the original data, effectively learn a single hash code using the modality-specific projection matrix; the specific formula is: s.t.B (v) ∈{-1,1} N×r ; Among them, B (v) is the modality-specific hash code of the v-th learned modality, r represents the length of the hash code corresponding to each modality in the structured hash code, and V represents the number of modalities; R (v) is the linear projection matrix of the v-th modality; γ (v) is the balance parameter of the v-th modality; N is the total number of all samples in the training phase; is the feature matrix of the v-th modality; d (v) is the feature dimension; Through the objective function f com The preservation of the complementarity of multimodal features is specifically as follows: Constructive collaborative non-linear multi-modal feature mapping The formula is as follows: Among them, It consists of V individual feature embeddings, including information from V modalities; in multi-modal hashing learning, by assigning reasonable weights μ (v) , effectively measure the importance of each different modality feature; each part of the structured hash code contains rich intra-modal information, thus being able to effectively improve the accuracy of cross-modal retrieval; at the same time, by combining the structured hash codes to represent the entire multi-modal data, it can effectively fuse heterogeneous modality features and achieve composite multi-modal retrieval; Construct the objective function f com , and the specific formula is: Among them, [B (1) ;...; B (V) is the defined structured hash code; W is a non-linear projection matrix for multi-modal retrieval tasks; θ is a balance parameter; is the collaborative non-linear multi-modal feature mapping of the original input data; Construct the objective function f sup The formula is as follows: Among them, is the transformation matrix; is the label matrix; [B (1) ; …; B (V) is the structured hash code; is the pairwise similarity matrix; α and β are balance parameters.
4. The structured multi-modal retrieval method for multiple multimedia retrieval tasks according to claim 1 or 2, characterized in that, The solution of the objective function of the structured multi-modal hashing method for multiple multimedia retrieval tasks adopts an iterative solution algorithm, that is, fix other variables and solve one variable to obtain the final optimal solution. The specific optimization process is as follows: Fix other variables and update R (v) , the objective function becomes: Take the partial derivative with respect to R (v) Set it to zero and we get: Fix other variables and update C, and the objective function becomes: Take the partial derivative of C and set it to zero, and get: C = (αYB T B + βI Vr ) -1 (αrB T SY + βB T Y)(Y T Y) -1 ; For convenience of representation, let B = [B (1) ; …; B (v) ; Fix other variables and update B: First, solve for the v-th modal hash code B by fixing other variables except B (v) except, and remove the terms irrelevant to B from the objective function (v) , and simplify the function to: (v) irrelevant to B For B (v) Take the partial derivative and set it to zero, we get: Fix other variables and update B, and the objective function becomes: Convert to: where tr(B T B) is a constant; to efficiently compute the hash code and avoid symmetric matrix decomposition, rewrite min B∈{-1,1} tr(αBCY T YC T B T - 2αrSYC T B T + βB T B - 2βYC T B T ) as: Adopt an asymmetric hashing learning strategy and use variable to replace a B; at the same time, add D to measure the difference between B and and simplify the last term in to: Take the partial derivative of B and set it to zero, and get: Fix other variables for update The objective function becomes: The update formula is as follows: Fix other variables and update D and η. According to the ALM algorithm, get: η = ρη; Among them, the parameter ρ obtains the optimal parameter through parameter tuning; Fix other variables and update W, and get: Fix other variables and find the adaptive weight μ, and get:
5. The structured multi-modal retrieval method for multiple multimedia retrieval tasks according to claim 1, characterized in that, The Hamming distance is as follows: Select any test sample in the training dataset, and obtain the Hamming distance between the hash code of the test sample and the hash codes of all samples in the database dataset; Sort the database dataset samples according to the Hamming distance, with those closer to the test sample than the set threshold in the front and those farther from the test sample than the set threshold in the back, so as to verify the retrieval accuracy.
6. A structured multi-modal retrieval system for multiple multimedia retrieval tasks, characterized in that, This system is used to implement the structured multi-modal retrieval method for multiple multimedia retrieval tasks described in any one of claims 1 to 5; this system includes, A data preprocessing module for obtaining a multi-modal retrieval dataset, where each sample of the multi-modal retrieval dataset includes paired data of the image modality and the text modality, and divides the multi-modal retrieval dataset into a training dataset, a test dataset, and a database dataset; A deep feature representation module for inputting the original data of the image modality and the text modality into a deep feature extraction model respectively for feature extraction, and using the extracted features as training features, test features, and database features for objective function construction, hash function learning, and online modality hashing retrieval; A target function construction module for constructing the overall target function of a structured multi-modal hashing method for multiple multimedia retrieval tasks through a training data set; A hash function learning module for solving the target function by using an iterative optimization method; An online modal hashing retrieval module for constructing and utilizing the target function of online hashing, obtaining the hash codes of samples in the test data set and the database data set, obtaining the Hamming distance between the hash code of each test sample in the test data set and the hash code of the samples in the database data set, sorting the samples in the database data set in ascending order according to the corresponding Hamming distance, and obtaining the retrieval result.
7. An electronic device, characterized in that, Comprising: A memory and at least one processor; Wherein, a computer program is stored on the memory; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the structured multi-modal retrieval method for multiple multimedia retrieval tasks according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program can be executed by a processor to implement the structured multi-modal retrieval method for multiple multimedia retrieval tasks according to any one of claims 1 to 5.
Citation Information
Patent Citations
Social image retrieval method and system based on missing multi-modal hash
CN111090765A
Cross-media data association analysis model training and data association analysis method and system
CN111651577A