Multimedia Data Cross-Modal Retrieval Method and System Based on Weighted Hash Codes

By constructing the label distribution matrix and learning the weighted hash code, the problem of insufficient reflection of fine-grained relationships in the existing technology is solved, efficient cross-modal retrieval is achieved, and retrieval performance and scalability are improved.

CN115795065BActive Publication Date: 2025-07-18SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211377750.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-07-18
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

The existing supervised cross-modal hashing method uses logical label coarse granularity, which cannot reflect the fine-grained relationship between instances, resulting in poor retrieval performance, and equal treatment of binary codes limits the ability of hash code to preserve fine semantics.

Method used

By mining the topological structure information of the label, constructing the label distribution matrix, constructing the similarity matrix, learning the weighted hash code to approximate the similarity matrix, using the kernel logistic regression model to learn the hash function, obtain the trained hash search model, and calculate the weighted Heming distance output search result.

Benefits of technology

It improves the ability of cross-modal retrieval, reduces quantization errors, and can complete super-large-scale applications in linear time, improving retrieval performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115795065B_ABST
    Figure CN115795065B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-modal retrieval method and system for multimedia data based on weighted hash codes. The method includes constructing a training sample set; extracting features from the training samples to obtain the features of the training samples in different modalities and training a hash retrieval model. The hash retrieval model includes hash codes and hash functions. Based on the features and semantic labels of the training samples, it learns the hash codes and bit weight matrices of the sample data in the training set, and learns the hash functions of different modalities. Input the data to be retrieved into the trained hash retrieval model. According to the features of the data to be retrieved extracted, combined with the hash function of the corresponding modality of the data to be retrieved, obtain the hash code of the data to be retrieved. According to the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of the sample data in the database, output the retrieval result. The present invention effectively improves the cross-modal retrieval ability by learning the weights of each hash code and emphasizing the unique contributions of different codes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-modal retrieval, and in particular, to a cross-modal retrieval method and system for multimedia data based on weighted hash codes. Background Art

[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the rapid development of information technology, in a social multimedia environment, it is often necessary to retrieve text or other types of data using images. A way of retrieving data in another modality that is semantically related to the data using one modality is called cross-modal retrieval. For cross-modal retrieval, traditional nearest neighbor search (NN) is not feasible because it relies on precise distance calculation. Hashing, as a typical approximation nearest neighbor method (ANN), can achieve a balance between accuracy and efficiency and is a common solution applied to cross-modal retrieval. The hashing method maps high-dimensional data into low-dimensional compact hash codes, which can greatly reduce memory consumption. Moreover, the distance calculation on binary hash codes only needs to be executed by the central processing unit (CPU) through exclusive OR operations, which greatly accelerates the search on large-scale data sets and improves retrieval performance. It is precisely because the hashing method can achieve good retrieval performance while being efficient in terms of time and space that this method is becoming more and more well-known and widely used.

[0004] At the same time, cross-modal hashing can be divided into supervised, unsupervised, and semi-supervised according to whether it uses supervision information. Among them, the supervised cross-modal hashing method uses the semantic information of labels to generate more discriminative hash codes and can obtain better retrieval performance.

[0005] However, existing supervised cross-modal hashing methods use logical labels in the form of {0, 1}, which are usually coarse-grained and cannot reflect the fine-grained relationships between instances. Moreover, there is still a large amount of topological information to be mined in data categories. In addition, treating each bit of the binary code equally will limit the ability of the hash code to preserve more fine-grained semantics, resulting in poor retrieval performance. Summary of the Invention

[0006] To address the deficiencies of the above-mentioned existing technologies, the present invention provides a cross-modal retrieval method and system for multimedia data based on weighted hash codes. By mining the topological structure information of tags to enhance semantic tags, a tag distribution matrix is obtained. Based on this tag distribution matrix, a refined similarity matrix is constructed. According to this tag distribution matrix, a similarity matrix is constructed, and an objective function is built to make the weighted hash code approximate this similarity matrix. Through learning, the hash code and bit weight matrix are obtained, and then the hash functions of different modalities are learned through a kernel logistic regression model to obtain a trained hash retrieval model. The data to be queried is input into this hash retrieval model, and the retrieval result is output. This solution utilizes the topological structure information of semantic tags with logical values, carries more semantic information, and emphasizes the unique contributions of different codes by learning the weights of each hash code, effectively improving the cross-modal retrieval ability.

[0007] In a first aspect, the present disclosure provides a cross-modal retrieval method for multimedia data based on weighted hash codes, including:

[0008] Obtain multimedia data of different modalities and construct a training sample set;

[0009] Extract features from the training samples to obtain the features of the training samples of different modalities and train the hash retrieval model; the hash retrieval model includes a hash code and a hash function. Based on the features and semantic tags of the training samples, learn the hash code and bit weight matrix of the training set sample data; based on the features and hash code of the training samples, learn the hash functions of different modalities;

[0010] Input the data to be retrieved into the trained hash retrieval model. According to the features of the data to be retrieved extracted, combined with the hash function of the corresponding modality of the data to be retrieved, obtain the hash code of the data to be retrieved, calculate the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of the sample data in the database, and output the retrieval result according to this weighted Hamming distance.

[0011] In a further technical solution, the training process of the hash retrieval model specifically includes:

[0012] Use a clustering algorithm to divide the training data in the training set into multiple groups, construct the local category association matrix of each group, and obtain the local tag distribution matrix of each group according to the local category association matrix and semantic tags. Combine them to obtain the global tag distribution matrix of the training samples;

[0013] Use the generated tag distribution matrix to construct a similarity matrix, and build an objective function by approximating the inner product of the weighted hash code to this similarity matrix, and solve to obtain the hash code matrix and bit weight matrix.

[0014] In a further technical solution, the solving process is as follows:

[0015] Initialize the weight matrix as the identity matrix, introduce intermediate variables and balance and irrelevant constraints, and solve to obtain the hash code matrix;

[0016] By taking the derivative of the objective function and setting the derivative to 0, and combining the calculated hash code matrix and similarity matrix, solve to obtain the weight matrix.

[0017] A further technical solution further includes:

[0018] Kernelize the features of the obtained training data of different modalities to obtain a kernelized feature matrix, and based on the learned hash code matrix, learn hash functions of different modalities through a kernel logistic regression model.

[0019] A further technical solution, the calculation process of the weighted Hamming distance specifically includes:

[0020] Obtain the hash codes of the training set sample data and the data to be retrieved, segment the hash codes, and perform an exclusive OR operation on each segment to obtain the byte type value of each segment;

[0021] Use the learned weight matrix to construct a lookup table for each segment;

[0022] Access the corresponding lookup table according to the byte type value of each segment, calculate the floating-point value of each segment, and obtain the weighted Hamming distance between the data to be retrieved and the training set sample data by summing the floating-point values of all segments.

[0023] A further technical solution, outputting the retrieval result according to the weighted Hamming distance specifically includes:

[0024] After calculating the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of all sample data in the training set, sort the samples in the database from small to large according to the weighted Hamming distance, and select the top k samples as the retrieval result to output, where k is an integer not less than 1.

[0025] In a second aspect, the present disclosure provides a cross-modal retrieval system for multimedia data based on weighted hash codes, including:

[0026] A training sample set construction module for obtaining multimedia data of different modalities and constructing a training sample set;

[0027] A hash retrieval model training module for extracting features from the training samples, obtaining the features of the training samples of different modalities, and training the hash retrieval model; the hash retrieval model includes hash codes and hash functions, and based on the features and semantic labels of the training samples, learn the hash codes and weight matrix of the training set sample data; based on the features and hash codes of the training samples, learn hash functions of different modalities;

[0028] A retrieval module, configured to input the data to be retrieved into the trained hash retrieval model, obtain the hash code of the data to be retrieved according to the extracted features of the data to be retrieved and in combination with the hash function of the corresponding modality of the data to be retrieved, calculate the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of the sample data in the database, and output a retrieval result according to the weighted Hamming distance.

[0029] A further technical solution, wherein the training process of the hash retrieval model specifically includes:

[0030] Using a clustering algorithm to divide the training data in the training set into multiple groups, constructing a local class association matrix for each group, obtaining a local label distribution matrix for each group according to the local class association matrix and the semantic labels, and combining to obtain a global label distribution matrix of the training samples;

[0031] Using the generated label distribution matrix to construct a similarity matrix, and constructing an objective function by approximating the similarity matrix with the inner product of weighted hash codes, and solving to obtain a hash code matrix and a bit weight matrix.

[0032] In a third aspect, the present disclosure further provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the method in the first aspect are completed.

[0033] In a fourth aspect, the present disclosure further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the method in the first aspect are completed.

[0034] The above one or more technical solutions have the following beneficial effects:

[0035] 1. The present invention provides a cross-modal retrieval method and system for multimedia data based on weighted hash codes. By mining the topological structure information of labels to enhance semantic labels, a label distribution matrix is obtained. Based on this label distribution matrix, a refined similarity matrix is constructed. According to this label distribution matrix, a similarity matrix is constructed, and an objective function is constructed to make the weighted hash code approximate the similarity matrix. By learning, a hash code and a bit weight matrix are obtained. This solution utilizes the topological structure information of semantic labels with logical values, carries more semantic information, and emphasizes the unique contributions of different codes by learning the weights of each hash code, effectively improving the cross-modal retrieval ability.

[0036] 2. The present invention provides an efficient discrete optimization algorithm to solve the discrete constraint problem of hash codes, reduces quantization errors, and can be completed in linear time, and can be extended to ultra-large-scale application scenarios. Description of the Drawings

[0037] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0038] Figure 1 It is the overall flowchart of the cross-modal retrieval method for multimedia data based on weighted hash codes according to the embodiments of the present invention. Detailed implementation manners

[0039] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0040] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0041] Embodiment 1

[0042] This embodiment provides a cross-modal retrieval method for multimedia data based on weighted hash codes. By mining the topological structure information of data semantic tags to enhance the semantic tags, a tag distribution matrix is obtained. According to this tag distribution matrix, a similarity matrix is constructed, and a hash code and a bit weight matrix that approximate the similarity matrix are learned. Then, a hash function for different modalities is learned through a kernel logistic regression model, and a trained hash retrieval model is obtained. The data to be queried is input into this hash retrieval model, and the retrieval result is output. As Figure 1 shown, the above method of this embodiment specifically includes the following steps:

[0043] Step S1, construct a training set. In this step, multimedia data such as images, texts, audios, and videos published by users on social multimedia platforms are obtained, a database is constructed, and the multimedia data of different modalities in the database are used as training samples to form a training set. In this embodiment, different modality data includes image data, text data, etc.

[0044] Step S2, perform feature extraction on the training samples to obtain the features of the training samples of different modalities, and train the hash retrieval model. This hash retrieval model includes a hash code and a hash function. Based on the features and semantic tags of the training samples, the hash code and the bit weight matrix of the training set sample data are learned; based on the features and hash code of the training samples, hash functions for different modalities are learned.

[0045] In this embodiment, feature extraction is performed on image and text sample data. For image data, image features are extracted using the SIFT or GIST method, and for text data, text features are extracted using the BoW method. The hash retrieval model is trained using the extracted image features and text features of the image-text pair, and the unified hash code representation B of the image and text multimedia data, as well as the hash function H1 corresponding to the image modality and the hash function H2 corresponding to the text modality, are learned.

[0046] The training process of the above hash retrieval model specifically includes the following steps:

[0047] Step S2.1: Use the clustering algorithm to divide the training data in the training set into p groups, construct the local class association matrix for each group, obtain the local label distribution matrix for each group based on the local class association matrix and the semantic labels, and combine them to obtain the global label distribution matrix of the training samples.

[0048] The specific process of step S2.1 is as follows:

[0049] Step S2.1.1: Use the clustering algorithm such as the k-means algorithm to cluster the training sample features, and divide the training data in the training set into p groups. For each group of training data, define the local class association matrix, that is:

[0050]

[0051] where is the semantic label matrix of the training set sample data, c is the number of label categories, n is the number of training set samples, is the row-normalized semantic label matrix of the m-th group. It should be noted that the semantic label matrix of the entire training set refers to a logical label matrix with values of 1 and 0. Specifically, for each element in this matrix, such as L ij = 1 indicates that the j-th sample in the training set belongs to the i-th category, otherwise L ij = 0 indicates that the j-th sample in the training set does not belong to the i-th category; furthermore, for the grouped semantic label matrix L m , [L m ij = 1 indicates that the j-th sample in the m-th group belongs to the i-th category, otherwise [L m ij = 0 indicates that the j-th sample in the m-th group does not belong to the i-th category.

[0052] ​​Step S2.1.2: Define the objective function for learning the label distribution using the category association matrix defined in Step S2.1.1 and the semantic labels of the training set samples. This objective function mainly consists of two parts: 1) The label distribution should be as similar as possible to the semantic labels; 2) The distance between the label distribution vectors of all samples within the same group on any two categories should be consistent with the corresponding semantic category values. The objective function is as follows:

[0053]

[0054] where α is a balance parameter used to adjust the contributions of the two terms before and after. are the label distribution matrix, the semantic label matrix, and the category association matrix respectively. c is the number of label categories, n is the number of training set samples. The matrix subscript m represents the group number after clustering the training data features. The subscript * indicates that the corresponding matrix takes the full column or row. For example, [D m i* represents the entire i-th row of the matrix.

[0055] Step S2.1.3: Take the derivative of the above objective function formula and set the derivative to 0 to obtain the label distribution matrix of the m-th group, that is:

[0056]

[0057] where is the Laplacian matrix of Cm, is a matrix of all 1s. By calculating the local label distribution matrices of all groups and combining them, the global label distribution matrix of the training samples can be obtained. It should be noted that each group can be calculated in parallel to improve efficiency.

[0058] Step S2.2: Use the global label distribution matrix generated in Step S2.1 to construct a similarity matrix containing more semantic information and topological structure information. Construct the objective function by approximating this similarity matrix with the inner product of weighted hash codes; Initialize the bit weight matrix as the identity matrix, introduce intermediate variables to solve the quartic complexity caused by symmetric inner products, and introduce balance and uncorrelated constraints to reduce quantization errors, and solve to obtain the hash code matrix; Take the derivative of this objective function and set the derivative to 0, and combine the calculated hash code matrix and similarity matrix to solve for the bit weight matrix. Through this step, the learning and training of the hash code are completed.

[0059] The specific process of Step S2.2 is as follows:

[0060] Step S2.2.1: Use the label distribution matrix generated in Step S2.1 to construct a similarity matrix containing more semantic information and topological structure information, which is:

[0061] ​

[0062] Among them, is the label distribution matrix after column normalization. To avoid an n×n matrix, when using it, directly calculate using the right side of Equation (4) above.

[0063] Set the hash code matrix to approximate the similarity matrix, and considering the particularity of each hash code, add different weights to each hash code. The matrix composed of the weights added to each hash code in the hash code matrix is the bit weight matrix. Construct the objective function by approximating the similarity matrix with the weighted inner product of the hash codes. The objective function is defined as follows:

[0064]

[0065] where B ∈ {-1, 1} r×n is the hash code matrix of the training set, is the bit weight matrix that emphasizes the importance of the hash codes and is a diagonal matrix, is the similarity matrix defined by the label correlation matrix, and r is the number of bits of the hash code, that is, the length of the hash code.

[0066] Step S2.2.2, initialize the bit weight matrix Λ as the identity matrix, and introduce an intermediate variable Q to replace one of the hash code matrices in the objective function shown in Equation (5) above. To reduce the quantization error, add balance and uncorrelated constraints to the intermediate variable Q, and the objective function becomes:

[0067]

[0068] Based on the above objective function, use a two-step iterative strategy to solve the hash code matrix. The specific solution method is as follows:

[0069] First, randomly initialize the hash code matrix B and the intermediate variable Q with the standard normal distribution;

[0070] Then, fix the hash code matrix B, and transform the objective function of Equation (6) into the form of the matrix trace:

[0071]

[0072] Define Z = rBS and According to the definition of the similarity matrix S, then:

[0073]

[0074] Perform matrix decomposition on ZJZ T to obtain:

[0075]

[0076] Among them, ∑ is a diagonal matrix composed of r′ ≤ r positive eigenvalues, and V is a matrix composed of the corresponding eigenvectors. is a matrix composed of the eigenvectors corresponding to the remaining zero eigenvalues. Perform Schmidt orthogonalization on to obtain an orthogonal matrix Define U = JZ T V∑ -1 / 2 , and the solution of the formula is:

[0077]

[0078] Secondly, similarly, fix the intermediate variable Q and convert formula (6) into the form of matrix trace:

[0079]

[0080] Substitute the definition of the similarity matrix S, and its closed-form solution is:

[0081]

[0082] Among them, sign(·) is the sign function.

[0083] The above solution scheme adopts a two-step iterative strategy. First, initialize the hash code matrix B and the intermediate variable Q, and then iteratively update the hash code matrix B and the intermediate variable Q: when updating the hash code matrix B, fix the intermediate variable Q and solve it using formula (12); when updating the intermediate variable Q, fix the hash code matrix B and solve it using formula (10). The hash code matrix B and the intermediate variable Q start from the initial values and are continuously iteratively updated using formulas (10) and (12) until convergence, that is, the value of the objective function shown in formula (6) no longer decreases. The optimal solution of the hash code matrix is obtained through the above scheme.

[0084] Step S2.2.3, after obtaining the optimal solution of the hash code, let the derivative of the objective function of formula (5) with respect to the weight matrix Λ be zero, and solve to obtain the weight matrix Λ, that is:

[0085] A = diag(r((BB T ))⊙(BB T )) -1 diag(BSB T )) (13)

[0086] Among them, diag(·) represents taking the diagonal elements of the matrix as a new diagonal matrix, and ⊙ is element-wise multiplication.

[0087] Step S2.3: Kernelize the features of the obtained training data of different modalities to obtain a kernelized feature matrix. Based on the hash code matrix learned in step S2.2, learn the hash functions of different modalities through a kernel logistic regression model.

[0088] The specific process of step S2.3 is as follows:

[0089] Step S2.3.1: Kernelize the features of the training data of different modalities, specifically:

[0090]

[0091] Among them, is the feature matrix of the l-th modality of the training set, is the i-th sample, and φ l (·) represents the kernel function, which can specifically use functions such as the radial basis function RBF. n is the number of samples in the training set, d is the feature dimension, and after kernelization, it becomes k-dimensional, that is

[0092] Step S2.3.2: Use the hash code matrix calculated through the two-step iterative strategy in step S2.2.2 and learn the hash function using the kernel logistic regression model. Specifically, the following objective function is adopted:

[0093]

[0094] Among them, ξ is the regularization term coefficient used to avoid overfitting, is the projection matrix of the l-th modality to be learned, B ∈ {-1, 1} r×n is the hash code of the training set, X (l) is the feature of the l-th modality of the training set sample, and φ l (X (l) ) kernelizes the feature matrix of the l-th modality, and r is the number of bits of the hash code.

[0095] Take the derivative of formula (15) and set the derivative to zero to solve for the projection matrix:

[0096] W (l) = Bφ i (X (l) ) T (φ l (X (l) )φ l (X (l) ) T + ξI) -1 (16)

[0097] At this time, for the data to be retrieved in the l-th modality, its hash code can be calculated through the following hash function:

[0098] H l (x (l) ) = sign(W (l) φ l (x (l) )) (17)

[0099] Step S3: Input the data to be retrieved into the trained hash retrieval model. According to the extracted features of the data to be retrieved and combined with the hash function of the corresponding modality of the data to be retrieved, obtain the hash code representation of the data to be retrieved, calculate the weighted Hamming distance between the hash code representation of the data to be retrieved and the hash codes of the sample data in the database, and output the retrieval result according to the weighted Hamming distance.

[0100] The specific process of outputting the retrieval result in the above step S3 is as follows:

[0101] Step S3.1: Calculate the weighted Hamming distance between the hash code of the training set sample data and the hash code of the data to be retrieved. Specifically: Obtain the hash code p of the data to be retrieved and the hash code q of a certain sample data in the training set. Segment them from the back forward every 8 bits, and fill 0 for the insufficient 8 bits. Perform the exclusive OR operation on each segment and represent the result as a byte-type data. For example, for the hash code p of a certain sample data in the training set and the hash code q of the data to be retrieved calculated by formula (12), segment them into s segments from the back forward every 8 bits, and finally fill 0 for the insufficient 8 bits, and finally divide them into [p1; p2;...; p s and [q1; q2;...; q s . Perform the exclusive OR operation on each segment p i and q i , and represent the obtained result as a byte-type value v i , such as v i = 0x01111111 = 127.

[0102] Step S3.2: Use the bit weight matrix learned in the above step S2 to construct the look-up table for each segment, which is specifically represented as:

[0103] w ij = Λ 8(i-1)+j (18)

[0104] where i represents the segment number of the segment, j = {1, 2,..., 8}, representing the number of occurrences of a certain 1 value in the binary form of the byte-type value of the i-th segment obtained in step S3.1.

[0105] Step S3.3: Access the corresponding look-up table according to the byte-type value of each segment, calculate the floating-point value of each segment, and obtain the weighted Hamming distance between the data to be retrieved and the training set sample data by summing the floating-point values of all segments.

[0106] Calculate the floating-point value for each segment based on the byte-type value obtained by performing an exclusive OR operation on each segment. Specifically, based on the byte-type value v of the i-th segment i , calculate the floating-point value u i , where u i is the sum of the values of all bits that are 1 in the corresponding lookup table. For example, if v i = 21 = 0x00010101, the floating-point value u i = w i + w i5 + W i3 . i1 .

[0107] For the hash code p of the training set sample data and the hash code q of the data to be retrieved, their weighted Hamming distance is the sum of the floating-point values u i of each segment, that is:

[0108]

[0109] Step S3.4, calculate the weighted Hamming distance between the hash code q of the data to be retrieved and the hash codes of all sample data in the training set through the above scheme, sort the samples in the database from small to large according to the weighted Hamming distance, and select the top k samples as the retrieval result for output, where k is an integer not less than 1.

[0110] In this embodiment, the process and retrieval result of retrieving the data to be retrieved using the trained hash retrieval model are as follows:

[0111] 1) Image to text. The user inputs an image as a query sample. First, convert the query image into a hash code representation b through the hash function H1, then calculate the Hamming distance between b and the database hash code B and sort them, and return the top-k texts with small distances as the retrieval result.

[0112] 2) Text to image. The user inputs a text as a query sample. First, convert the query text into a hash code representation b through the hash function H2, then calculate the Hamming distance between b and the database hash code B and sort them, and return the top-k images with small distances as the retrieval result.

[0113] Embodiment 2

[0114] This embodiment provides a cross-modal retrieval system for multimedia data based on weighted hash codes, including:

[0115] A training sample set construction module, used to obtain multimedia data of different modalities and construct a training sample set;

[0116] The hash retrieval model training module is used to extract features from training samples, obtain the features of training samples in different modalities, and train the hash retrieval model; the hash retrieval model includes a hash code and a hash function, and based on the features and semantic labels of the training samples, it learns the hash code and bit weight matrix of the training set sample data; based on the features and hash code of the training samples, it learns the hash functions of different modalities.

[0117] The retrieval module is used to input the data to be retrieved into the trained hash retrieval model, obtain the hash code of the data to be retrieved according to the extracted features of the data to be retrieved and in combination with the hash function of the corresponding modality of the data to be retrieved, calculate the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of the sample data in the database, and output the retrieval result according to the weighted Hamming distance.

[0118] Embodiment III

[0119] This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the above-mentioned cross-modal retrieval method of multimedia data based on weighted hash codes are completed.

[0120] Embodiment IV

[0121] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps in the above-mentioned cross-modal retrieval method of multimedia data based on weighted hash codes are completed.

[0122] The steps involved in the above Embodiments II to IV correspond to those in Method Embodiment I. For specific implementation manners, reference may be made to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0123] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0124] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0125] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A cross-modal retrieval method for multimedia data based on weighted hash codes, characterized in that Including: Obtain multimedia data of different modalities and construct a training sample set; Extract features from the training samples to obtain the features of the training samples of different modalities, and train a hash retrieval model; the hash retrieval model includes a hash code and a hash function, and based on the features and semantic labels of the training samples, learn the hash code and bit weight matrix of the training set sample data; Learn hash functions of different modalities based on the features and hash codes of the training samples; The training process of the hash retrieval model includes: Use a clustering algorithm to divide the training data in the training set into multiple groups and construct a local class association matrix for each group; specifically: use a clustering algorithm to cluster the training sample features, divide the training data in the training set into p groups, and for each group of training data, define a local class association matrix, that is: ; wherein, is the semantic label matrix of the training set sample data, c is the number of label categories, and n is the number of training set samples. is the m-th group of row-normalized semantic label matrices; wherein, the semantic label matrix of the entire training set refers to a logical label matrix with values of 1 and 0, that is, for each element in this matrix, = 1 indicates that the j-th sample in the training set belongs to the i-th category, otherwise = 0 indicates that the j-th sample in the training set does not belong to the i-th category; for the grouped semantic label matrix , = 1 indicates that the j-th sample in the m-th group belongs to the i-th category, otherwise = 0 indicates that the j-th sample in the m-th group does not belong to the i-th category. According to the local class association matrix and semantic labels, obtain the local label distribution matrix for each group, and combine them to obtain the global label distribution matrix of the training samples; specifically: use the local class association matrix and the semantic labels of the training set samples to define an objective function for learning the label distribution, which mainly includes two parts: the label distribution is as similar as possible to the semantic labels, and the distance between the label distribution vectors of all samples in the same group on any two classes is consistent with the corresponding semantic class values. This objective function is: ; Among them, is a balance parameter used to adjust the balance between the two contributions before and after, , , are the label distribution matrix, the semantic label matrix, and the category association matrix respectively. c is the number of label categories, n is the number of training set samples. The matrix subscript m represents the group number after clustering the training data features, and the subscript * indicates that the corresponding matrix fills the entire column or row. represents the entire i-th row of the matrix; Take the derivative of the objective function formula and set the derivative to 0 to obtain the label distribution matrix of the mth group, that is: ; Among them, is the Laplacian matrix of a matrix of all 1s. By calculating the local label distribution matrices of all groups and combining them, the global label distribution matrix of the training samples can be obtained; Use the generated label distribution matrix to construct a similarity matrix, and approximate the similarity matrix by the weighted inner product of the weighted hash codes to construct an objective function, and solve to obtain the hash code matrix and the bit weight matrix; Input the data to be retrieved into the trained hash retrieval model, according to the features of the data to be retrieved extracted, combined with the hash function of the corresponding modality of the data to be retrieved, obtain the hash code of the data to be retrieved, and output the retrieval result according to the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of the sample data in the database.

2. The cross-modal retrieval method for multimedia data based on weighted hash codes according to claim 1, wherein The solution process is: Initialize the bit weight matrix as the identity matrix, introduce intermediate variables and balance and irrelevant constraints, and solve to obtain the hash code matrix; Take the derivative of the objective function and set the derivative to 0, and combine the calculated hash code matrix and similarity matrix to solve to obtain the bit weight matrix.

3. The cross-modal retrieval method for multimedia data based on weighted hash codes according to claim 1, wherein It also includes: Kernelize the features of the obtained training data of different modalities to obtain a kernelized feature matrix, and based on the learned hash code matrix, learn hash functions of different modalities through a kernel logistic regression model.

4. The cross-modal retrieval method for multimedia data based on weighted hash codes according to claim 1, wherein The calculation process of the weighted Hamming distance includes: Obtain the hash codes of the training set sample data and the data to be retrieved, segment the hash codes, and perform an exclusive OR operation on each segment to obtain the byte type value of each segment; Use the learned bit weight matrix to construct a lookup table for each segment; Access the corresponding lookup table according to the byte type value of each segment, calculate the floating point value of each segment, and obtain the weighted Hamming distance between the data to be retrieved and the training set sample data by summing the floating point values of all segments.

5. The cross-modal retrieval method for multimedia data based on weighted hash codes according to claim 1, wherein Output the retrieval result according to the weighted Hamming distance, including: After calculating the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of all sample data in the training set, sort the samples in the database from small to large according to the weighted Hamming distance, and select the top k samples as the retrieval results for output, where k is an integer not less than 1.

6. A cross-modal retrieval system for multimedia data based on weighted hash codes, characterized in that, It includes: A training sample set construction module for obtaining multimedia data of different modalities and constructing a training sample set; A hash retrieval model training module for extracting features from training samples, obtaining the features of training samples of different modalities, and training a hash retrieval model; the hash retrieval model includes a hash code and a hash function, and based on the features and semantic labels of the training samples, learns the hash code and bit weight matrix of the training set sample data; Learn hash functions of different modalities based on the features and hash codes of training samples; The training process of the hash retrieval model includes: Using a clustering algorithm to divide the training data in the training set into multiple groups and construct a local class association matrix for each group; specifically: using a clustering algorithm to cluster the training sample features, dividing the training data in the training set into p groups, and for each group of training data, defining a local class association matrix, that is: ; Among them, is the semantic label matrix of the training set sample data, c is the number of label categories, and n is the number of training set samples. is the m-th group of row-normalized semantic label matrices; among them, the semantic label matrix of the entire training set refers to a logical label matrix with values of 1 and 0, that is, for each element in this matrix, = 1 indicates that the j-th sample in the training set belongs to the i-th category, otherwise = 0 indicates that the j-th sample in the training set does not belong to the i-th category; for the grouped semantic label matrix , = 1 indicates that the j-th sample in the m-th group belongs to the i-th category, otherwise = 0 indicates that the j-th sample in the m-th group does not belong to the i-th category. According to the local class association matrix and semantic labels, obtain the local label distribution matrix for each group, and combine them to obtain the global label distribution matrix of the training samples; specifically: use the local class association matrix and the semantic labels of the training set samples to define an objective function for learning the label distribution, and this objective function mainly includes two parts: the label distribution is as similar as possible to the semantic labels, and the distance between the label distribution vectors of all samples in the same group on any two classes is consistent with the corresponding semantic class values. This objective function is: ; Among them, is a balance parameter used to adjust the balance between the two contributions before and after, , , are the label distribution matrix, the semantic label matrix, and the category association matrix respectively. c is the number of label categories, n is the number of training set samples. The matrix subscript m represents the group number after clustering the training data features, and the subscript * indicates that the corresponding matrix fills the entire column or row. represents the entire i-th row of the matrix; Take the derivative of the objective function formula and set the derivative to 0 to obtain the label distribution matrix of the mth group, that is: ; Among them, is 's Laplacian matrix, is a matrix of all 1s. By calculating the local label distribution matrices of all groups and combining them, the global label distribution matrix of the training samples is obtained; Use the generated label distribution matrix to construct a similarity matrix, and construct an objective function by approximating the similarity matrix with the inner product of weighted hash codes, and solve to obtain the hash code matrix and bit weight matrix; A retrieval module for inputting the data to be retrieved into the trained hash retrieval model, obtaining the hash code of the data to be retrieved according to the extracted features of the data to be retrieved and combining the hash function of the corresponding modality of the data to be retrieved, and outputting the retrieval result according to the weighted Hamming distance between the hash code of the data to be retrieved and the hash codes of the sample data in the database.

7. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of a cross-modal retrieval method for multimedia data based on weighted hash codes as described in any one of claims 1-5 are completed.

8. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the steps of a cross-modal retrieval method for multimedia data based on weighted hash codes as described in any one of claims 1-5 are completed.

Citation Information

Patent Citations

  • Cross-modal data discrete hash retrieval method based on similarity maintenance

    CN110059198A

  • Cross-modal retrieval method and system based on robust similarity preservation

    CN115080880A