Dynamic Cross-Modal Hashing Retrieval Method and System Based on Semi-Paired Data
Through a dynamic cross-modal hash retrieval method based on semi-paired data, hash encoding integration is used to use the similarity matrix of paired and unpaired data to solve the problems of data environment changes and unpaired data in cross-modal retrieval, and fast and accurate cross-modal retrieval is achieved.
Patent Information
- Application Number
- CN202411092128.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-08-09
AI Technical Summary
The existing cross-modal hash retrieval method assumes that the data environment is static, cannot adapt to dynamic changes, and cannot effectively process unpaired data, resulting in a degradation in retrieval performance and low training efficiency.
A dynamic cross-modal hash retrieval method based on semi-paired data is adopted. By obtaining a multimodal data set, feature extraction and hash encoding learning are performed, and hash encoding integration is used to construct the overall target hash function to adapt to the dynamic environment.
It improves the accuracy and efficiency of cross-modal retrieval, can maintain retrieval performance in a dynamic data environment, make full use of unpaired data information, and enhances the generalization ability and robustness of hash encoding.
Smart Images

Figure CN119202278B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technologies, and in particular, to a dynamic cross-modal hashing retrieval method and system based on semi-paired data. Background Art
[0002] In the related art, in the field of cross-modal hashing retrieval research, the following problems mainly exist:
[0003] (1) Cross-modal hashing retrieval methods assume that the data environment is static. However, the actual data environment is complex and changeable, and new data will continuously be added. For static cross-modal methods, if the original hashing function remains unchanged, it cannot reflect the dynamic changes of new data compared to old data and cannot adapt to the changes in the new data environment, resulting in a decline in retrieval performance. Moreover, static cross-modal hashing methods often need to retrain the entire hashing model to adapt to the new data environment, but this is undoubtedly very time-consuming. Especially in the case where the data environment changes very frequently, this approach of retraining the entire model is inefficient.
[0004] (2) Cross-modal hashing retrieval methods assume that the training data is one-to-one corresponding. However, in the actual data environment, it cannot be guaranteed that multi-modal data is all one-to-one corresponding, that is, unpaired. Since most cross-modal hashing methods can only be trained when the multi-modal training data is one-to-one corresponding, they cannot handle the unpaired data environment.
[0005] In summary, the technical problems existing in the related art need to be improved. Summary of the Invention
[0006] The embodiments of this application aim to at least solve one of the technical problems in the related art to some extent. For this reason, the main purpose of the embodiments of this application is to propose a dynamic cross-modal hashing retrieval method and system based on semi-paired data, which can improve the accuracy and efficiency of cross-modal hashing retrieval.
[0007] To achieve the above object, on the one hand, the embodiments of this application propose a dynamic cross-modal hashing retrieval method based on semi-paired data, and the method includes the following steps:
[0008] Obtain a multi-modal data set; the multi-modal data set includes a paired data set and an unpaired data set;
[0009] Extract features from the multi-modal data set to obtain multi-modal data feature vectors; the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors;
[0010] Learn the current paired hashing codes of the paired data set based on the paired label information of the paired data set to obtain target paired hashing codes;
[0011] Learning the current unpaired hash code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hash code;
[0012] Integrating the target paired hash code and the target unpaired hash code based on a total objective function to obtain a total target hash code;
[0013] Learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vector to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total target hash code in a database;
[0014] Retrieving a sample to be retrieved based on the target multi-modal hash function and the total target hash code in the database to obtain a target retrieval result.
[0015] In some embodiments, the extracting multi-modal data feature vectors from the multi-modal data set includes:
[0016] Performing a first feature extraction process on the paired data set in the multi-modal data set to obtain the paired feature vector;
[0017] Performing a second feature extraction process on the unpaired data set in the multi-modal data set to obtain the unpaired feature vector.
[0018] In some embodiments, the learning the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code includes:
[0019] Constructing a paired label matrix of the paired data set based on the paired label information of the paired data set;
[0020] Normalizing the paired label matrix to obtain a normalized paired label matrix;
[0021] Calculating the pairwise semantic similarity between label vectors in the normalized paired label matrix to obtain pairwise semantic similarity values;
[0022] Constructing a pairwise similarity matrix based on the pairwise semantic similarity values;
[0023] Learning the current paired hash code of the paired data set based on the pairwise similarity matrix to obtain the target paired hash code.
[0024] In some embodiments, integrating the target paired hash code and the target unpaired hash code based on the total objective function to obtain a total objective hash code includes:
[0025] Merging the target paired hash code and the target unpaired hash code to obtain a target merged hash code;
[0026] Performing optimization processing on the target merged hash code based on the total objective function to obtain the total objective hash code.
[0027] In some embodiments, learning a current multi-modal hash function of the multi-modal data set based on the total objective hash code and the multi-modal data feature vector to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total objective hash code in a database includes:
[0028] Learning the current multi-modal hash function of the multi-modal data set by using a linear regression method based on the total objective hash code and the multi-modal data feature vector to obtain the target multi-modal hash function, and storing the target multi-modal hash function and the total objective hash code in the database; wherein, the target multi-modal hash function is used to represent the target hash function corresponding to various modal data in the multi-modal data set.
[0029] In some embodiments, retrieving a sample to be retrieved based on the target multi-modal hash function and the total objective hash code in the database to obtain a target retrieval result includes:
[0030] Converting a feature vector to be retrieved of the sample to be retrieved into a hash code to be retrieved based on the target multi-modal hash function in the database;
[0031] Calculating a Hamming distance between the hash code to be retrieved and the total objective hash code in the database to obtain a similarity calculation result;
[0032] Retrieving candidate samples in the database according to the similarity calculation result to obtain the target retrieval result.
[0033] In some embodiments, retrieving candidate samples in the database according to the similarity calculation result to obtain the target retrieval result includes:
[0034] Performing similarity ranking on the candidate samples in the database according to the similarity calculation result to obtain a similarity ranking result;
[0035] Taking the candidate samples that meet the target sorting rule and the target quantity in the similarity ranking result as the target retrieval result.
[0036] To achieve the above object, on the other hand, an embodiment of the present application proposes a dynamic cross-modal hashing retrieval system based on semi-paired data, and the system includes the following modules:
[0037] A data acquisition module, configured to acquire a multi-modal data set; the multi-modal data set includes a paired data set and an unpaired data set;
[0038] A feature extraction module, configured to perform feature extraction on the multi-modal data set to obtain multi-modal data feature vectors; the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors;
[0039] A paired hashing code learning module, configured to learn the current paired hashing code of the paired data set based on the paired label information of the paired data set to obtain a target paired hashing code;
[0040] An unpaired hashing code learning module, configured to learn the current unpaired hashing code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hashing code;
[0041] A total target hashing code acquisition module, configured to perform integration processing on the target paired hashing code and the target unpaired hashing code based on a total target function to obtain a total target hashing code;
[0042] A hashing function learning module, configured to learn the current multi-modal hashing function of the multi-modal data set based on the total target hashing code and the multi-modal data feature vectors to obtain a target multi-modal hashing function, and store the target multi-modal hashing function and the total target hashing code in a database;
[0043] A sample retrieval module, configured to retrieve a sample to be retrieved based on the target multi-modal hashing function and the total target hashing code in the database to obtain a target retrieval result.
[0044] To achieve the above object, on the other hand, an embodiment of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described above is implemented.
[0045] To achieve the above object, on the other hand, an embodiment of the present application proposes a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0046] The embodiments of the present application at least include the following beneficial effects: The present application provides a dynamic cross-modal hashing retrieval method and system based on semi-paired data. The solution includes obtaining a multi-modal data set, which contains a paired data set and an unpaired data set; extracting features from the multi-modal data set to obtain multi-modal data feature vectors, which include paired feature vectors and unpaired feature vectors; learning the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code; learning the current unpaired hash code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hash code; integrating the target paired hash code and the target unpaired hash code based on a total objective function to obtain a total target hash code; learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total target hash code in a database; retrieving a sample to be retrieved based on the target multi-modal hash function and the total target hash code in the database to obtain a target retrieval result. In the embodiments of the present application, by using the paired label information of the paired data set to learn the current paired hash code to obtain a target paired hash code, the similarity and semantic consistency of the paired data set in the hash space are maintained, providing a reliable benchmark for subsequent cross-modal retrieval; by introducing the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to learn the target unpaired hash code, not only the intra-modal similarity of the paired data and the unpaired data is utilized, but also the problem of unpaired data is solved by using the semantic information between the paired data and the unpaired data of different modalities, making full use of the information in the unpaired data and compensating for the lack of paired information by constructing a similarity matrix, improving the generalization ability and robustness of the hash code; integrating the target hash codes of the paired data and the unpaired data based on the total objective function to obtain a total target hash code, synthesizing the information of the paired data and the unpaired data, so that the finally obtained hash code can reflect the characteristics and relationships of both types of data simultaneously, improving the accuracy and efficiency of cross-modal retrieval; learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, enabling the hash function to be applicable to data of multiple modalities and capable of capturing the semantic relationships between different modalities, providing data support for cross-modal retrieval; finally, it is possible to dynamically retrieve the sample to be retrieved based on the continuously updated target multi-modal hash function and the total target hash code in the database, realizing fast, accurate, and real-time cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flowchart of the steps of the dynamic cross-modal hashing retrieval method based on semi-paired data provided by an embodiment of the present application;
[0048] Figure 2 It is a schematic flowchart of hash code learning based on semi-paired data provided by an embodiment of the present application;
[0049] Figure 3 It is a schematic structural diagram of the dynamic cross-modal hashing retrieval system based on semi-paired data provided by an embodiment of the present application;
[0050] Figure 4 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application described in detail in the appended claims.
[0052] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0053] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality of includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0055] Before elaborating on the embodiments of this application in detail, some nouns and terms involved in the embodiments of this application are first explained. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0056] (1) Cross-modal retrieval: A retrieval method that retrieves samples of other modalities from a sample of one modality.
[0057] (2) Hashing retrieval: The hashing algorithm maps samples in a dataset from the initial high-dimensional feature space to a low-dimensional binary Hamming space, generates a set of compact hash codes for each sample, and evaluates the similarity between samples based on the Hamming distance between the sample hash codes, thereby achieving fast nearest-neighbor retrieval of query samples. The core of the hashing algorithm is to train a set of hash functions. Each hash function corresponds to a hash hyperplane, which divides the initial feature space into two parts. One side corresponds to hash code 1, and the other side is hash code 0. Therefore, each hash function corresponds to one bit of hash code.
[0058] (3) Dynamic data environment: The dataset or data environment for model training changes over time. For example, as time goes by, new images and texts are continuously added to the original dataset, and old images and texts may also be continuously deleted from the dataset. Such a dynamic data environment is more complex than a static data environment.
[0059] (4) Semi-paired: In the training stage of a cross-modal hashing retrieval model, if the multi-modal data for training the model has each modality corresponding one by one and no data is missing, it can be regarded as paired data. If some modalities of the multi-modal data are missing, it is called semi-paired data.
[0060] (5) SPI H (Semi-paired Hashing for Cross-modal Retrieval in Non-stationary Environment, a semi-paired hashing retrieval method in a dynamic environment), the SPI H method is a cross-modal dynamic hashing retrieval method proposed in the embodiments of this application.
[0061] In the related art, in the field of cross-modal hashing retrieval research, the following problems mainly exist: (1) Cross-modal hashing retrieval methods assume that the data environment is static. However, the actual data environment is complex and changeable, and new data will continuously be added. For static cross-modal methods, if the original hash function remains unchanged, it cannot reflect the dynamic changes of new data compared with old data, and cannot adapt to the changes in the new data environment, resulting in a decline in retrieval performance. Moreover, static cross-modal hashing methods often need to retrain the entire hash model to adapt to the new data environment, but this is undoubtedly very time-consuming. Especially in the case where the data environment changes very frequently, this approach of retraining the entire model is inefficient. (2) Cross-modal hashing retrieval methods assume that the training data is one-to-one. However, in the actual data environment, it cannot be guaranteed that multi-modal data is all one-to-one, that is, unpaired. Since most cross-modal hashing methods can only be trained when the multi-modal training data is one-to-one, they cannot cope with the unpaired data environment.
[0062] In view of this, the embodiments of the present application provide a dynamic cross-modal hashing retrieval method and system based on semi-paired data. The solution includes obtaining a multi-modal data set, which contains a paired data set and an unpaired data set; extracting features from the multi-modal data set to obtain multi-modal data feature vectors, which include paired feature vectors and unpaired feature vectors; learning the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code; learning the current unpaired hash code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hash code; integrating the target paired hash code and the target unpaired hash code based on a total objective function to obtain a total target hash code; learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total target hash code in a database; retrieving a sample to be retrieved based on the target multi-modal hash function and the total target hash code in the database to obtain a target retrieval result. By using the paired label information of the paired data set to learn the current paired hash code to obtain the target paired hash code, the embodiments of the present application maintain the similarity and semantic consistency of the paired data set in the hash space, providing a reliable benchmark for subsequent cross-modal retrieval; by introducing the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to learn the target unpaired hash code, not only the intra-modal similarity of the paired data and the unpaired data is utilized, but also the unpaired data problem is solved by using the semantic information between the paired data and the unpaired data of different modalities, making full use of the information in the unpaired data and compensating for the lack of paired information by constructing a similarity matrix, improving the generalization ability and robustness of the hash code; integrating the target hash codes of the paired data and the unpaired data based on the total objective function to obtain the total target hash code, synthesizing the information of the paired data and the unpaired data, so that the finally obtained hash code can reflect the characteristics and relationships of both types of data at the same time, improving the accuracy and efficiency of cross-modal retrieval; learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain the target multi-modal hash function, making the hash function applicable to data of multiple modalities and capable of capturing the semantic relationships between different modalities, providing data support for cross-modal retrieval; finally, being able to dynamically retrieve the sample to be retrieved based on the continuously updated target multi-modal hash function and the total target hash code in the database, realizing fast, accurate, and real-time cross-modal retrieval.
[0063] The dynamic cross-modal hashing retrieval method based on semi-paired data provided by the embodiments of the present application relates to the technical field of information processing. The dynamic cross-modal hashing retrieval method based on semi-paired data provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the dynamic cross-modal hashing retrieval method based on semi-paired data, etc., but is not limited to the above forms.
[0064] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs (Personal Computers), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0065] Please refer to Figure 1 , Figure 1 which is an optional step flowchart of the dynamic cross-modal hashing retrieval method based on semi-paired data provided by the embodiments of the present application. Figure 1 The method in
[0066] Step S101, obtain a multi-modal data set; the multi-modal data set includes a paired data set and an unpaired data set;
[0067] Among them, a multi-modal dataset refers to a data set that contains multiple different modalities (or types), and the modalities can be text, images, audio, video, etc. In the embodiments of the present application, the multi-modal dataset includes a paired dataset and an unpaired dataset. Among them, the paired dataset contains paired data points, and the paired data points are related or paired in a certain sense. Exemplarily, a certain text data can be paired with a certain image data one by one; the data points contained in the unpaired dataset have no clear pairing relationship, and these data points may come from different sources, different time periods or different objects. Exemplarily, a certain text data cannot be paired with a certain image data and both belong to independent data.
[0068] In a specific application, the dataset or data environment for model training changes over time. Exemplarily, it is assumed that new images and texts are continuously added to the original dataset as time goes by, and it is also possible that old images and texts are continuously deleted from the dataset. Such a dynamic data environment is more complex than a static data environment. In the embodiments of the present application, SPI H (Semi-paired Hashing for Cross-modal Retrieval in Non-stationary Environment, a semi-paired hashing retrieval method in a dynamic environment) is proposed, that is, the dynamic cross-modal hashing retrieval method based on semi-paired data mentioned in the embodiments of the present application. For a dynamic data environment, the retrieval model can adapt to the new data environment, and the hash function and hash code can be updated in response to changes in the data environment, improving the accuracy and efficiency of retrieval. This is the main difference between the dynamic hashing retrieval method and the static hashing retrieval method.
[0069] Step S101 illustrated in the embodiments of the present application provides a diverse data source for subsequent cross-modal retrieval by obtaining a multi-modal dataset that includes a paired dataset and an unpaired dataset.
[0070] Step S102, perform feature extraction on the multi-modal dataset to obtain a multi-modal data feature vector; the multi-modal data feature vector includes a paired feature vector and an unpaired feature vector;
[0071] In some embodiments, step S102 may include: performing a first feature extraction process on the paired dataset in the multi-modal dataset to obtain a paired feature vector; performing a second feature extraction process on the unpaired dataset in the multi-modal dataset to obtain an unpaired feature vector.
[0072] Among them, the first feature extraction process refers to the feature extraction operation performed on the paired data set, aiming to extract useful feature information (i.e., paired feature vectors) for subsequent tasks (such as hash code learning, retrieval, etc.) from the original paired data set. The second feature extraction process is the feature extraction operation performed on the unpaired data set to obtain unpaired feature vectors. Although there is no direct pairing relationship within the unpaired data set, the extracted features can reflect the internal attributes and differences of the data.
[0073] Step S102 illustrated in the embodiment of the present application, by performing feature extraction on the paired data set and the unpaired data set to obtain their respective corresponding feature vectors, provides an effective data representation for subsequent hash code learning, and at the same time can remove redundant information, retain key features, and improve the accuracy and efficiency of learning hash codes.
[0074] Step S103, learning the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain the target paired hash code;
[0075] In some embodiments, step S103 may include: constructing a paired label matrix of the paired data set based on the paired label information of the paired data set; performing a normalization process on the paired label matrix to obtain a normalized paired label matrix; calculating the pairwise semantic similarity between the label vectors in the normalized paired label matrix to obtain pairwise semantic similarity values; constructing a pairwise similarity matrix based on the pairwise semantic similarity values; and learning the current paired hash code of the paired data set based on the pairwise similarity matrix to obtain the target paired hash code.
[0076] Among them, the paired label information is the label information of the paired data set, and the paired label information can represent relationships such as whether two different data perspective points belong to the same category, are similar, or match. Among them, the paired label matrix is a two-dimensional matrix used to represent the paired label information between all data points in the paired data set.
[0077] Among them, the pairwise semantic similarity refers to the similarity degree of two data points at the semantic level. In hash code or retrieval tasks, the pairwise semantic similarity is usually evaluated based on the paired label information. The pairwise semantic similarity value is a quantization index used to represent the pairwise semantic similarity between two data points. The pairwise similarity matrix is a two-dimensional matrix used to represent the pairwise semantic similarity values between all data points in the paired data set, similar to the paired label matrix.
[0078] Among them, the current paired hash code refers to the current hash code of the paired dataset during the learning process. These codes may not be optimal or updated in real time and need to be further optimized or iterated through further learning to obtain the target paired hash code. The target paired hash code refers to the optimal hash code finally obtained through the learning process. The optimal hash code can effectively represent the data points in the paired dataset and show good performance in hash coding or retrieval tasks. The target paired hash code is learned through an optimization algorithm based on paired label information.
[0079] It should be noted that in a dynamic data environment, the target paired hash code is dynamically updated.
[0080] In a specific implementation, the hash code learning process of paired data is as follows:
[0081] To learn the hash code of paired data, first, it is necessary to obtain the label information of the paired data, construct the pairwise similarity matrix of the paired data, and give the corresponding objective function. Among them, the objective function is the loss function. By solving the objective function, the final result of the corresponding matrix can be obtained, that is, the hash code corresponding to different modality paired data can be obtained by minimizing the objective function. In the SPI H method, first, the pairwise similarity matrix of paired data needs to be constructed to guide the learning of hash codes.
[0082] Specifically, the paired label information of the paired dataset is used to construct the paired label matrix, the paired label matrix is used to construct the pairwise similarity matrix, and the hash code corresponding to the paired dataset can be learned using the pairwise similarity matrix and the paired label matrix.
[0083] Among them, two samples sharing similar labels are expected to have similar hash codes with a small Hamming distance. Therefore, the cosine similarity between label vectors is used to evaluate the pairwise semantic similarity between two samples, and the specific calculation is as follows:
[0084]
[0085] Among them, represents the pairwise similarity matrix of paired data; represents the 2-norm normalized label matrix at time step t, where 2-norm represents the second norm (Frobenius norm); 1 represents the all-one matrix; represents the product of an all-one matrix and the transpose of another all-one matrix; represents the transpose.
[0086] Among them, the calculation method of the 2-norm normalized label matrix is as follows:
[0087]
[0088] Among them, represents the i-th column vector of (the total label matrix corresponding to the paired data at time t); represents the i-th column vector of (the 2-norm normalized label matrix with a time step of t).
[0089] Specifically, by minimizing the difference between the hash code similarity and the pairwise semantic similarity, the similarity between paired data can be maintained, and the calculation method is as follows:
[0090]
[0091] Among them, B1 (t) represents the total pairing matrix of the hash code matrix corresponding to the paired data at time t, and r represents the B matrix, that is, the number of bits of the hash code in the hash code matrix (hash code length), represents the pairing similarity matrix of the paired data (also known as the pairwise semantic similarity matrix). Since B1 (t) is a binary matrix. Minimizing the above equation (3) is an NP (Non-deterministic Polynomial time) hard problem. Therefore, in the embodiments of the present application, a real-valued matrix V (t) is introduced to replace one of the hash code matrices, making the optimization process more effective, and the real-valued matrix V (t) retains the latent representation of the hash code B1. Therefore, formula (3) can be transformed as follows:
[0092]
[0093] Among them, is represented as a real-valued matrix, and its value range is from -1 to 1; I represents the identity matrix.
[0094] At the same time, in order to minimize the information loss between B1 (t) and V (t) , the embodiments of the present application add a regularization term of the orthogonal rotation matrix R. The calculation formula for minimizing the information loss is as follows:
[0095]
[0096] Among them, s.t. R (t) R (t)T represents the orthogonality constraint on the orthogonal rotation matrix R.
[0097] In addition, each paired data is associated with a uniquely corresponding tag information, which is beneficial for learning hash codes to preserve more accurate semantic information. Therefore, linear embedding is learned through hash coding to preserve the tag semantic information, and the calculation formula is as follows:
[0098]
[0099] Among them, P (t) represents the linear embedding matrix, which can be learned through iterative optimization and updated online. However, the tag matrix is actually a binary matrix, and the learning efficiency of the linear embedding matrix is low. Directly embedding the latent representation into the tag will result in a high quantization loss. Therefore, the embodiment of the present application reduces the quantization loss and expands the margins of different categories by introducing an adaptive margin factor. (Index matrix) The calculation formula is as follows:
[0100]
[0101] Among them, "+1" represents the positive direction, and "-1" represents the negative direction; j represents the j-th row of the total tag matrix corresponding to the paired data at time t, and k represents the k-th column of the total tag matrix corresponding to the paired data at time t.
[0102] Among them, the margin matrix represents the margin size V of each latent representation (t) , Among them, ε jk represents the value of the j-th row and k-th column of the margin matrix. Therefore, formula (6) can be transformed into the following formula:
[0103]
[0104] Among them, ⊙ represents the Hadamard product operator.
[0105] The step S103 shown in the embodiment of the present application learns the target paired hash code by using the paired tag information of the paired data set, maintains the similarity and semantic consistency of the paired data set in the hash space, and provides a reliable benchmark for subsequent cross-modal retrieval.
[0106] Step S104, based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set, learns the current unpaired hash code of the unpaired data set to obtain the target unpaired hash code;
[0107] Among them, the intra-modal similarity matrix is a measure of the similarity between data points within the same modality (such as all images, all texts, etc.). Specifically, the intra-modal similarity matrix can be a similarity relationship matrix between paired image data and unpaired image data, or a similarity relationship matrix between paired text data and unpaired text data.
[0108] Among them, the inter-modal similarity matrix refers to a measure of the similarity between data points from different modalities. The unpaired dataset may contain data from multiple different modalities (such as images and texts), and the inter-modal similarity matrix is used to capture the associations or similarities between these data points across different modalities. Specifically, the inter-modal similarity matrix can be an inter-modal similarity relationship matrix between unpaired image data and unpaired text data.
[0109] Among them, the current unpaired hash code refers to the current hash code of the unpaired dataset during the learning process. The current unpaired hash code may not be optimal or updated in real time, and further learning is required to optimize or iterate it to obtain the target unpaired hash code. The target unpaired hash code refers to the optimal hash code finally obtained through the learning process. The target unpaired hash code can effectively represent the data points in the unpaired dataset and exhibit good performance in tasks such as hash coding or cross-modal retrieval. The target unpaired hash code is learned based on the intra-modal similarity matrix and the inter-modal similarity matrix.
[0110] In a specific implementation, the specific process of learning the hash code of unpaired data is as follows:
[0111] In the SPIH method, in addition to the similarity information between paired data, the similarity information provided by unpaired data is also of great significance for learning discriminative hash codes. Among them, different from each pair of paired data sharing a unified hash code, unpaired data from different modalities have their own hash codes. For example, unpaired image data in unpaired data has a corresponding hash code, and unpaired text data has a corresponding hash code. Therefore, in the SPIH method provided in the embodiments of the present application, the intra-modal similarity matrix between unpaired data and paired data of each modality, as well as the inter-modal similarity matrix between unpaired image data and unpaired text data, can be provided for hash code learning, and a corresponding objective function is given. By minimizing this objective function, the hash codes corresponding to unpaired data of different modalities can be obtained.
[0112] Among them, the intra-modal similarity matrix between unpaired data and paired data, as well as the inter-modal similarity matrix between unpaired image data and unpaired text data, can be obtained by using the label information of paired data and the label information of unpaired data (such as unpaired image label information and unpaired text label information).
[0113] Specifically, first, a paired label matrix is constructed based on the paired label information of the paired dataset. The paired label matrix is used to construct a pairwise similarity matrix. Then, an unpaired label matrix of the unpaired data is constructed based on the label information of the unpaired data (such as unpaired image label information and unpaired text label information). Based on the paired label matrix and the unpaired label matrix, an intra-modal similarity matrix between the unpaired data and the paired data, and an inter-modal similarity matrix between the unpaired image data and the unpaired text data can be constructed. Finally, based on the paired label matrix, the unpaired label matrix, the intra-modal similarity matrix between the unpaired data and the paired data, and the inter-modal similarity matrix between the unpaired image data and the unpaired text data, the hash code corresponding to the unpaired data can be learned.
[0114] For intra-modal similarity, similar to formula (4), the calculation formula for minimizing the similarity loss between the unpaired data and the paired data of the two modalities is as follows:
[0115]
[0116] Among them, represents the hash code of the unpaired image, represents the hash code of the unpaired text, represents the similarity relationship matrix (semantic similarity matrix) between the paired image data and the unpaired image data, represents the similarity relationship matrix (semantic similarity matrix) between the paired text data and the unpaired text data.
[0117] In order to further explore the semantic information between different modalities of the unpaired data in the embodiments of the present application, the inter-modal similarity can also be learned, specifically as follows:
[0118]
[0119] Among them, represents the similarity relationship matrix (semantic similarity matrix) between the unpaired image data and the unpaired text data. It should be noted that optimizing formula (10) is not an NP problem, and there is an efficient discrete optimization solution.
[0120] Step S104 illustrated in the embodiments of the present application is of great significance for learning the target unpaired hash code by introducing the intra-modal similarity matrix between the paired dataset and the unpaired dataset and the inter-modal similarity matrix of the unpaired dataset. This step not only utilizes the intra-modal similarity of paired and unpaired data, but also solves the unpaired data problem by leveraging the semantic information between paired and unpaired data of different modalities, highlighting their correlation, making full use of the information in unpaired data, and compensating for the lack of paired information by constructing a similarity matrix, thereby improving the generalization ability and robustness of the hash code.
[0121] Step S105, integrating the target paired hash code and the target unpaired hash code based on the total objective function to obtain the total target hash code;
[0122] Regarding the total objective function, combining the above formulas (4), (5), (8), (9) and (10), the total objective function of the SPI H method in the embodiments of the present application is as follows:
[0123]
[0124] Wherein, α1, α2, β1, β2, β3 are all trade-off parameters; assuming that a new data block appears at each time step, and the data block consists of two modalities (i.e., images and texts), among which, some images and texts are paired, and the rest of the images and texts are not paired. Assuming that the number of paired image-text data, unpaired images and unpaired texts remains unchanged within each time step, n1 represents the number of paired image-text data within each time step, n2 represents the number of unpaired image data within each time step, and n3 represents the number of unpaired text data within each time step. By minimizing the above formula (11), the hash codes of paired and unpaired data can be learned.
[0125] In addition, the fine-grained similarity matrix is constructed from the 2-norm normalized label matrix and can be rewritten as follows:
[0126]
[0127] Therefore, the similarity matrix can be rewritten in a block form as follows:
[0128]
[0129] Wherein, is the similarity matrix between old data and old data, is the similarity matrix between old data and new data, is the similarity matrix between new data and old data, is the similarity matrix between new data and new data. Therefore, the total objective function can be further rewritten as follows:
[0130]
[0131] where, represents the label matrix corresponding to the paired data. When is represented as At this time, represents the label matrix corresponding to the unpaired image modality. When is represented as At this time, represents the label matrix corresponding to the unpaired text, where c represents the total number of label categories. represents the hash code corresponding to the paired data, represents the hash code corresponding to the unpaired image data, represents the hash code corresponding to the unpaired text data.
[0132] In some embodiments, step S105 may include: merging the target paired hash code and the target unpaired hash code to obtain a target merged hash code; optimizing the target merged hash code based on the total objective function to obtain a total target hash code.
[0133] Among them, the target merged hash code refers to a preliminary hash code representation, and the target merged hash code is obtained by merging the target paired hash code (that is, the optimal hash code learned from the paired data) and the target unpaired hash code (that is, the optimal hash code learned from the unpaired data).
[0134] Among them, the total target hash code refers to a hash code that can simultaneously reflect the similarity between paired data and the semantic relevance of unpaired data (including between different modalities), aiming to improve the discriminative power, generalization ability, and cross-modal retrieval performance of the hash code by integrating information from different data sources and modalities.
[0135] In a specific implementation, based on minimizing the objective function, the hash codes corresponding to the paired data and unpaired data of different modalities can be obtained, and then the paired hash codes and unpaired hash codes of the same modality can be merged to obtain the hash codes of the same modality. Among them, the overall objective function (total objective function) is the sum of the objective functions involved in step S103 and step S104. In the embodiments of the present application, by solving the overall objective function, the hash codes of different modalities can be learned.
[0136] Step S105 illustrated in the embodiments of the present application integrates the target hash encodings of paired data and unpaired data based on the total objective function to obtain the total objective hash encoding. This process comprehensively considers the information of paired data and unpaired data, enabling the finally obtained hash encoding to simultaneously reflect the characteristics and relationships of both types of data, thereby improving the accuracy and efficiency of cross-modal retrieval.
[0137] Step S106: Learn the current multi-modal hash function of the multi-modal data set based on the total objective hash encoding and the multi-modal data feature vectors to obtain the target multi-modal hash function, and store the target multi-modal hash function and the total objective hash encoding in the database.
[0138] In some embodiments, step S106 may include: learning the current multi-modal hash function of the multi-modal data set using the linear regression method based on the total objective hash encoding and the multi-modal data feature vectors to obtain the target multi-modal hash function, and storing the target multi-modal hash function and the total objective hash encoding in the database; wherein, the target multi-modal hash function is used to represent the target hash functions corresponding to various modal data in the multi-modal data set.
[0139] Among them, the current multi-modal hash function refers to a hash function that already exists or is preliminarily defined before starting the learning of the hash function, and it is necessary to optimize and learn the current multi-modal hash function to obtain the target multi-modal hash function. Among them, the current multi-modal hash function is the current hash function corresponding to various modal data in the multi-modal data set. The target multi-modal hash function is used to represent the target hash functions corresponding to various modal data in the multi-modal data set. The target multi-modal hash function is a hash function obtained through the learning process, which can better map various modal data in the multi-modal data set into the same hash space, and maintain the similarity and semantic relevance between the data in this space. The target multi-modal hash function aims to improve the discriminative power, generalization ability, and cross-modal retrieval performance of the hash encoding.
[0140] Among them, the database is a collection for storing and managing data, allowing users to store, retrieve, modify, and delete data in a structured manner. In the embodiments of the present application, the database is used to store the learned hash functions (such as the target multi-modal hash function) and the corresponding hash encodings (such as the total objective hash encoding) for subsequent data processing, retrieval, and analysis tasks.
[0141] In specific implementation, the specific process of hash function learning is as follows:
[0142] After obtaining the hash codes of all the new data, the hash functions corresponding to each modality at the learning time t can be obtained. For each modality, the hash codes of the paired data and the unpaired data are combined together for learning the hash function. Suppose and represent the hash codes of the image modality and the text modality at time t respectively, then the two can be combined and obtained by the following calculation method:
[0143]
[0144] where, and can represent the total hash code matrix of the image modality and the total hash code matrix of the text modality at time t respectively.
[0145] Given the hash codes of each modality, the hash function corresponding to the modality can be obtained by minimizing the following formula:
[0146]
[0147] where, represents the hash function corresponding to the m-th modality, which is used to map the high-dimensional feature matrix to the low-dimensional Hamming space (hash code). At the same time, can represent the total feature data matrix corresponding to the m-th modality at time t:
[0148]
[0149] where, m = 1, 2, represents the feature matrix of the paired data, represents the unpaired image feature matrix, represents the unpaired text feature matrix, where d1 represents the dimension of the image feature matrix and d2 represents the dimension of the text feature matrix.
[0150] where, at time t, can represent the total feature matrix of the paired data from the beginning to the current time t, which is composed of the total feature matrix before time t and the feature matrix at time t.
[0151] Finally, can be obtained by taking the partial derivatives of formulas (15) and (16) and setting them to 0. The specific calculation formula is as follows:
[0152]
[0153] Step S106 illustrated in the embodiments of the present application learns the current multimodal hash function of the multimodal dataset based on the total target hash code and the multimodal data feature vectors to obtain the target multimodal hash function. This process enables the hash function to be applicable to data of multiple modalities and capture the semantic relationships between different modalities, providing data support for cross-modal retrieval.
[0154] Step S107 retrieves the sample to be retrieved based on the target multimodal hash function and the total target hash code in the database to obtain the target retrieval result.
[0155] In some embodiments, step S107 may include: converting the feature vector to be retrieved of the sample to be retrieved into a hash code to be retrieved based on the target multimodal hash function in the database; calculating the Hamming distance between the hash code to be retrieved and the total target hash code in the database to obtain a similarity calculation result; retrieving the candidate samples in the database according to the similarity calculation result to obtain the target retrieval result.
[0156] In some specific embodiments, retrieving the candidate samples in the database according to the similarity calculation result to obtain the target retrieval result may include: performing similarity ranking on the candidate samples in the database according to the similarity calculation result to obtain a similarity ranking result; using the candidate samples that meet the target ranking rule and the target quantity in the similarity ranking result as the target retrieval result.
[0157] The feature vector to be retrieved is the feature vector extracted according to the feature extraction method corresponding to the sample to be retrieved. Exemplarily, if the sample to be retrieved is image data, an image feature vector is extracted according to the feature extraction method corresponding to the image data; if the sample to be retrieved is text data, a text feature vector is extracted according to the feature extraction method corresponding to the text data. The present application embodiments do not limit the feature extraction method. For example, for images, deep learning models such as convolutional neural networks can be used to extract features; for text, word embedding or pre-trained text encoders can be used to extract features. The hash code to be retrieved is the hash code of the sample to be retrieved obtained based on the hash function corresponding to the modality of the sample to be retrieved.
[0158] The Hamming distance is used to measure the similarity between two hash codes. The smaller the Hamming distance, the more similar the two hash codes are, that is, the samples they correspond to may also be more similar in the original data space.
[0159] For the candidate samples, they are all data samples in the database. During the retrieval process, the similarity between the hash code of the sample to be retrieved and each candidate sample is calculated, and based on the similarity result, it is determined whether a certain candidate sample is of interest to the user.
[0160] Among them, the target sorting rule refers to the rule used to determine the sorting order of candidate samples during the retrieval process; the target quantity can refer to the quantity of target retrieval results that the user hopes to retrieve, or it can refer to the quantity of target retrieval results set by the back-end personnel for retrieval. For the target sorting rule and the target quantity, they can be set according to the actual application situation, and the embodiments of the present application do not limit this.
[0161] In specific implementation, the steps of retrieval can be as follows: First, select the samples to be retrieved, then extract features according to the feature extraction method corresponding to the retrieved samples, obtain the hash code of the retrieved samples through the hash function of the corresponding modality of the retrieved samples, calculate the Hamming distance between the hash code of the retrieved samples and the hash code of the target retrieval modality in the database. A smaller Hamming distance represents a greater similarity to the retrieved samples. After sorting the Hamming distances from small to large, the corresponding samples ranked in the front in the database can be used as retrieval results.
[0162] Step S107 illustrated in the embodiments of the present application realizes fast, accurate, and real-time cross-modal retrieval by storing the target multi-modal hash function and the total target hash code in the database, and dynamically retrieving the samples to be retrieved based on the continuously updated target multi-modal hash function and total target hash code in the database.
[0163] Steps S101 to S107 illustrated in the embodiments of the present application include: obtaining a multi-modal data set, where the multi-modal data set includes a paired data set and an unpaired data set; extracting features from the multi-modal data set to obtain multi-modal data feature vectors, where the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors; learning the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code; learning the current unpaired hash code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hash code; integrating the target paired hash code and the target unpaired hash code based on a total objective function to obtain a total target hash code; learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total target hash code in a database; retrieving a sample to be retrieved based on the target multi-modal hash function and the total target hash code in the database to obtain a target retrieval result. In the embodiments of the present application, by learning the current paired hash code using the paired label information of the paired data set to obtain a target paired hash code, the similarity and semantic consistency of the paired data set in the hash space are maintained, providing a reliable benchmark for subsequent cross-modal retrieval; by introducing the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to learn the target unpaired hash code, not only the intra-modal similarity of the paired data and the unpaired data is utilized, but also the unpaired data problem is solved by using the semantic information between the paired data and the unpaired data of different modalities, making full use of the information in the unpaired data, and compensating for the lack of paired information by constructing a similarity matrix, improving the generalization ability and robustness of the hash code; integrating the target hash codes of the paired data and the unpaired data based on the total objective function to obtain a total target hash code, synthesizing the information of the paired data and the unpaired data, so that the finally obtained hash code can reflect the characteristics and relationships of both types of data at the same time, improving the accuracy and efficiency of cross-modal retrieval; learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, enabling the hash function to be applicable to data of multiple modalities and capable of capturing the semantic relationships between different modalities, providing data support for cross-modal retrieval; finally, being able to dynamically retrieve the sample to be retrieved based on the continuously updated target multi-modal hash function and the total target hash code in the database, realizing fast, accurate, and real-time cross-modal retrieval.
[0164] To explain the principle of the technical solution of the present invention in detail, the following will describe the overall process of the present invention in combination with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation to the present invention.
[0165] The embodiment of the present application proposes a semi-paired cross-modal hashing retrieval method in a dynamic data environment, that is, dynamic cross-modal hashing retrieval based on semi-paired data. This method can perform cross-modal retrieval as close as possible to the real data environment. The SPI H method provided by the embodiment of the present application makes full use of the semantic information of different-modal paired data to learn the hashing codes of the paired data. At the same time, the SPI H method utilizes the semantic connection between unpaired data and paired data, so as to learn the hashing codes of unpaired data in different modalities respectively. After constructing the target hashing code using the above information, through multiple rounds of optimization iteration, the hashing codes corresponding to each modality at time t can be obtained. That is, when new unpaired data arrives at the multi-modal data set, the inter-modal similarity relationship and intra-modal similarity relationship between paired data and unpaired data are respectively constructed, making full use of the semantic labels of the data, and respectively using the adaptive margin factor and introducing an orthogonal real-valued matrix to reduce the quantization loss and accelerate the optimization iteration of the algorithm. The hashing codes of each modality of the new data and other related matrices are obtained through optimization iteration, and then the hashing function corresponding to each modality of the hashing codes is calculated for testing.
[0166] Please refer to Figure 2 , Figure 2 which is a schematic flow chart of hashing code learning based on semi-paired data provided by the embodiment of the present application. As Figure 2 shown, the overall process of hashing code learning based on semi-paired data is as follows: Before the new data arrives, according to the paired data label information of the old data (or called the old data) and the unpaired data label information of the old data, the similarity matrices of the paired data and the unpaired data are respectively constructed. At the same time, an adaptive margin is added to the paired data label information of the old data, and then the total objective function is constructed. Furthermore, by minimizing each objective function in the total objective function (i.e., Figure 2 the semantic embedding step in Figure 2 ), the hashing codes corresponding to the paired data and unpaired data of the old data can be obtained; after the new data arrives, according to the paired data label information of the new data and the unpaired data label information of the new data, the similarity matrices of the paired data and the unpaired data are respectively constructed. Similarly, an adaptive margin is added to the paired data label information of the new data, and then semantic embedding is performed to update the hashing codes of the data set.
[0167] The SPI H method is a hash method with two-step training. In the specific implementation, the training steps can be: In the first step, the SPI H method makes full use of the semantic information of paired data of different modalities to learn the hash code of paired data. At the same time, the SPI H method uses the semantic connection between unpaired data and paired data to learn the hash code of unpaired data in different modalities respectively. After using the above information to construct the target hash code, after multiple rounds of optimization iterations, the hash code corresponding to each modality at time t can be obtained. In the second step, using the hash code obtained in the first step, the SPI H method uses linear regression to learn the projection (hash function) of each modality.
[0168] Specifically, the retrieval step can be: first, prepare a multimodal data set before retrieval, and extract each modal feature vector of the multimodal data set and the label matrix corresponding to the sample and other information through the corresponding feature extraction method, perform hash coding and hash function learning through each modal feature vector and the label matrix corresponding to the sample and other information to obtain the target hash code and target hash function, and then store the target hash code and target hash function in the database and wait for retrieval. Next, select the retrieval sample, extract features according to the feature extraction method corresponding to the retrieval sample, process the extracted features through the hash function of the modality corresponding to the retrieval sample to obtain the hash code of the retrieval sample, calculate the Hamming distance between the hash code of the retrieval sample and the hash code of the target retrieval modality in the database, the smaller the Hamming distance, the greater the similarity with the sample, and after the Hamming distance is sorted from small to large, the corresponding sample in the front can be used as the retrieval result.
[0169] Specifically, the retrieval step can be divided into the following three steps:
[0170] (1) When new unpaired data arrives at the multimodal dataset, the inter-modal similarity relationship and intra-modal similarity relationship between the paired data and the unpaired data are constructed respectively, making full use of the semantic labels of the data. In addition, an adaptive margin factor and an orthogonal real-valued matrix are introduced to reduce the quantization loss and speed up the optimization iteration of the algorithm.
[0171] (2) The hash codes of each mode of the new data and other related matrices are obtained by optimizing the iterations, and then the hash functions corresponding to the hash codes of each mode are calculated for testing.
[0172] (3) In the testing phase, feature extraction technology is first used to extract the feature vector of the test modality. Then, the hash function is used to calculate the corresponding hash code. The Hamming distance between the hash code of the test data and the hash code of the database is compared. The smaller the distance, the more similar it is to the database sample. The samples are sorted from small to large according to the Hamming distance. The first few samples output are the query results corresponding to the test sample.
[0173] It should be noted that this embodiment only briefly illustrates the general process of the dynamic cross-modal hashing retrieval method based on semi-paired data. For the detailed description of each step, reference can be made to the relevant content in the foregoing embodiments, which will not be elaborated here. It can be understood that the present invention is not limited thereto.
[0174] In the embodiments of the present application, a multi-modal data set is obtained; the multi-modal data set includes a paired data set and an unpaired data set; feature extraction is performed on the multi-modal data set to obtain multi-modal data feature vectors; the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors; based on the paired label information of the paired data set, learning is performed on the current paired hash code of the paired data set to obtain a target paired hash code; based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set, learning is performed on the current unpaired hash code of the unpaired data set to obtain a target unpaired hash code; based on the total objective function, integration processing is performed on the target paired hash code and the target unpaired hash code to obtain a total target hash code; based on the total target hash code and the multi-modal data feature vectors, learning is performed on the current multi-modal hash function of the multi-modal data set to obtain a target multi-modal hash function, and the target multi-modal hash function and the total target hash code are stored in a database; based on the target multi-modal hash function and the total target hash code in the database, retrieval is performed on a sample to be retrieved to obtain a target retrieval result. In the embodiments of the present application, by using the paired label information of the paired data set to learn the current paired hash code to obtain a target paired hash code, the similarity and semantic consistency of the paired data set in the hash space are maintained, providing a reliable benchmark for subsequent cross-modal retrieval; by introducing the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to learn the target unpaired hash code, not only the intra-modal similarity of the paired data and the unpaired data is utilized, but also the problem of unpaired data is solved by using the semantic information between the paired data and the unpaired data of different modalities, making full use of the information in the unpaired data, and making up for the lack of paired information by constructing a similarity matrix, improving the generalization ability and robustness of the hash code; based on the total objective function, integration processing is performed on the target hash codes of the paired data and the unpaired data to obtain a total target hash code, integrating the information of the paired data and the unpaired data, so that the finally obtained hash code can reflect the characteristics and relationships of both types of data at the same time, improving the accuracy and efficiency of cross-modal retrieval; based on the total target hash code and the multi-modal data feature vectors, learning is performed on the current multi-modal hash function of the multi-modal data set to obtain a target multi-modal hash function, enabling the hash function to be applicable to data of multiple modalities and capable of capturing the semantic relationships between different modalities, providing data support for cross-modal retrieval; finally, dynamic retrieval can be performed on the sample to be retrieved based on the continuously updated target multi-modal hash function and the total target hash code in the database, realizing fast, accurate, and real-time cross-modal retrieval.
[0175] In summary, the key points of the embodiments of the present application are as follows:
[0176] 1. A method for learning the intra-modal similarity relationship and the inter-modal similarity relationship of paired data and unpaired data respectively in a dynamic and semi-paired data environment.
[0177] 2. An optimized iterative method for solving the final hash function and the final hash code in the embodiments of the present application.
[0178] Please refer to Figure 3 , the embodiments of the present application further provide a dynamic cross-modal hashing retrieval system 300 based on semi-paired data, which can implement the above-mentioned dynamic cross-modal hashing retrieval method based on semi-paired data. The system 300 includes the following modules:
[0179] A data acquisition module 301, configured to acquire a multi-modal data set; the multi-modal data set includes a paired data set and an unpaired data set;
[0180] A feature extraction module 302, configured to extract features from the multi-modal data set to obtain multi-modal data feature vectors; the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors;
[0181] A paired hash code learning module 303, configured to learn the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code;
[0182] An unpaired hash code learning module 304, configured to learn the current unpaired hash code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hash code;
[0183] A total target hash code acquisition module 305, configured to integrate the target paired hash code and the target unpaired hash code based on a total target function to obtain a total target hash code;
[0184] A hash function learning module 306, configured to learn the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, and store the target multi-modal hash function and the total target hash code in a database;
[0185] A sample retrieval module 307, configured to retrieve a sample to be retrieved based on the target multi-modal hash function and the total target hash code in the database to obtain a target retrieval result.
[0186] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0187] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned dynamic cross-modal hashing retrieval method based on semi-paired data. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0188] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0189] Please refer to Figure 4 , Figure 4 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0190] A processor 401, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0191] A memory 402, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402, and the processor 401 is used to call and execute the dynamic cross-modal hashing retrieval method based on semi-paired data in the embodiments of the present application;
[0192] An input / output interface 403, which is used to implement information input and output;
[0193] A communication interface 404, which is used to implement communication and interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0194] A bus 405 transmits information between various components of the device, such as a processor 401, a memory 402, an input / output interface 403, and a communication interface 404.
[0195] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0196] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned dynamic cross-modal hashing retrieval method based on semi-paired data is implemented.
[0197] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0198] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0199] The dynamic cross-modal hashing retrieval method and system based on semi-paired data provided by the embodiments of the present application obtain a multi-modal data set by: the multi-modal data set includes a paired data set and an unpaired data set; extracting features from the multi-modal data set to obtain multi-modal data feature vectors, where the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors; learning the current paired hash codes of the paired data set based on the paired label information of the paired data set to obtain target paired hash codes; learning the current unpaired hash codes of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain target unpaired hash codes; integrating the target paired hash codes and the target unpaired hash codes based on a total objective function to obtain total target hash codes; learning the current multi-modal hash function of the multi-modal data set based on the total target hash codes and the multi-modal data feature vectors to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total target hash codes in a database; retrieving a sample to be retrieved based on the target multi-modal hash function and the total target hash codes in the database to obtain a target retrieval result. By using the paired label information of the paired data set to learn the current paired hash codes to obtain target paired hash codes, the embodiments of the present application maintain the similarity and semantic consistency of the paired data set in the hash space, providing a reliable benchmark for subsequent cross-modal retrieval; by introducing the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to learn the target unpaired hash codes, not only the intra-modal similarity of the paired data and the unpaired data is utilized, but also the problem of unpaired data is solved by using the semantic information between the paired data and the unpaired data of different modalities, fully utilizing the information in the unpaired data and making up for the lack of paired information by constructing a similarity matrix, improving the generalization ability and robustness of the hash codes; integrating the target hash codes of the paired data and the unpaired data based on the total objective function to obtain total target hash codes, synthesizing the information of the paired data and the unpaired data, so that the finally obtained hash codes can reflect the characteristics and relationships of both types of data at the same time, improving the accuracy and efficiency of cross-modal retrieval; learning the current multi-modal hash function of the multi-modal data set based on the total target hash codes and the multi-modal data feature vectors to obtain a target multi-modal hash function, enabling the hash function to be applicable to data of multiple modalities and capable of capturing the semantic relationships between different modalities, providing data support for cross-modal retrieval; finally, being able to dynamically retrieve the sample to be retrieved based on the continuously updated target multi-modal hash function and the total target hash codes in the database, realizing fast, accurate, and real-time cross-modal retrieval.
[0200] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0201] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0202] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware and their appropriate combinations.
[0204] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0205] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0206] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of systems or units can be in electrical, mechanical or other forms.
[0207] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0208] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0209] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs.
[0210] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. A dynamic cross-modal hashing retrieval method based on semi-paired data, characterized in that The method includes the following steps: Obtain a multi-modal data set; the multi-modal data set includes a paired data set and an unpaired data set; Extract features from the multi-modal data set to obtain multi-modal data feature vectors; the multi-modal data feature vectors include paired feature vectors and unpaired feature vectors; Learn the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code; Learn the current unpaired hash code of the unpaired data set based on the intra-modal similarity matrix between the paired data set and the unpaired data set and the inter-modal similarity matrix of the unpaired data set to obtain a target unpaired hash code; Integrate the target paired hash code and the target unpaired hash code based on the total objective function to obtain a total target hash code; Learn the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, and store the target multi-modal hash function and the total target hash code in a database; Retrieve a sample to be retrieved based on the target multi-modal hash function and the total target hash code in the database to obtain a target retrieval result.
2. The method according to claim 1, wherein The extracting features from the multi-modal data set to obtain multi-modal data feature vectors includes: Perform a first feature extraction process on the paired data set in the multi-modal data set to obtain the paired feature vectors; Perform a second feature extraction process on the unpaired data set in the multi-modal data set to obtain the unpaired feature vectors.
3. The method according to claim 1, characterized in that The learning the current paired hash code of the paired data set based on the paired label information of the paired data set to obtain a target paired hash code includes: Construct a paired label matrix of the paired data set based on the paired label information of the paired data set; Normalize the paired label matrix to obtain a normalized paired label matrix; Calculate the pairwise semantic similarity between label vectors in the normalized paired label matrix to obtain pairwise semantic similarity values; Construct a pairwise similarity matrix based on the pairwise semantic similarity values; Learn the current paired hash code of the paired data set based on the pairwise similarity matrix to obtain the target paired hash code.
4. The method according to claim 1, wherein The integrating the target paired hash code and the target unpaired hash code based on the total objective function to obtain a total target hash code includes: Merge the target paired hash code and the target unpaired hash code to obtain a target merged hash code; Optimize the target merged hash code based on the total objective function to obtain the total target hash code.
5. The method according to claim 1, characterized in that, The learning the current multi-modal hash function of the multi-modal data set based on the total target hash code and the multi-modal data feature vectors to obtain a target multi-modal hash function, and storing the target multi-modal hash function and the total target hash code in a database includes: Learn the current multimodal hash function of the multimodal dataset using a linear regression method based on the overall target hash code and the multimodal data feature vectors to obtain the target multimodal hash function, and store the target multimodal hash function and the overall target hash code in the database; wherein, the target multimodal hash function is used to represent the target hash functions corresponding to various modal data in the multimodal dataset.
6. The method according to claim 1, characterized in that Retrieve the sample to be retrieved based on the target multimodal hash function and the overall target hash code in the database to obtain a target retrieval result, including: Convert the feature vector to be retrieved of the sample to be retrieved into a hash code to be retrieved based on the target multimodal hash function in the database; Calculate the Hamming distance between the hash code to be retrieved and the overall target hash code in the database to obtain a similarity calculation result; Retrieve candidate samples in the database according to the similarity calculation result to obtain the target retrieval result.
7. The method according to claim 6, wherein The retrieving the candidate samples in the database according to the similarity calculation result to obtain the target retrieval result includes: Perform similarity ranking on the candidate samples in the database according to the similarity calculation result to obtain a similarity ranking result; Use the candidate samples that meet the target ranking rules and the target quantity in the similarity ranking result as the target retrieval result.
8. A dynamic cross-modal hashing retrieval system based on semi-paired data, characterized in that, The system includes the following modules: A data acquisition module, configured to acquire a multimodal dataset; the multimodal dataset includes a paired dataset and an unpaired dataset; A feature extraction module, configured to extract features from the multimodal dataset to obtain multimodal data feature vectors; the multimodal data feature vectors include paired feature vectors and unpaired feature vectors; A paired hash code learning module, configured to learn the current paired hash code of the paired dataset based on the paired label information of the paired dataset to obtain a target paired hash code; An unpaired hash code learning module, configured to learn the current unpaired hash code of the unpaired dataset based on the intra-modal similarity matrix between the paired dataset and the unpaired dataset and the inter-modal similarity matrix of the unpaired dataset to obtain a target unpaired hash code; An overall target hash code acquisition module, configured to perform integration processing on the target paired hash code and the target unpaired hash code based on an overall target function to obtain an overall target hash code; A hash function learning module, configured to learn the current multimodal hash function of the multimodal dataset based on the overall target hash code and the multimodal data feature vectors to obtain a target multimodal hash function, and store the target multimodal hash function and the overall target hash code in the database; A sample retrieval module, configured to retrieve the sample to be retrieved based on the target multimodal hash function and the overall target hash code in the database to obtain a target retrieval result.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Dynamic cross-modal Hash retrieval method for concept drift
CN118193667A
Unsupervised hashing method for cross-modal video-text retrieval with clip
WO2023004206A1