Unsupervised cross-modal hash retrieval method based on steady state distribution and clustering

By constructing steady-state distribution and clustering methods under unsupervised conditions, the problem of low accuracy in the unsupervised cross-modal hashing method is solved, and the alignment and global optimization of image and text features in a unified hash space is achieved, which improves the accuracy and robustness of cross-modal retrieval.

CN120277205AActive Publication Date: 2025-07-08CENT SOUTH UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510766022.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing unsupervised cross-modal hashing method has low accuracy in cross-modal retrieval, making it difficult to effectively cross the semantic gap between modes. The built graph structure is prone to noisy connections, lacks global semantic structure modeling capabilities, and the clustering process lacks cross-modal consistency.

Method used

Using a method based on steady-state distribution and clustering, the image and text features are mapped to the unified hash space through nonlinear transformation, and the alignment loss function, cluster-level comparison loss function and steady-state loss function are constructed. Combined with the quantized loss function, the total loss function is constructed to train the unsupervised cross-modal hash retrieval model.

Benefits of technology

It improves the accuracy of cross-modal retrieval, enhances semantic consistency and robustness, ensures that features are effectively matched in the common semantic space, and improves the reliability of hash representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277205A_ABST
    Figure CN120277205A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised cross-modal Hash retrieval method based on steady-state distribution and clustering, and the method comprises the steps: constructing an alignment loss function according to a similarity matrix, a first coding Hash code and a second coding Hash code; obtaining a first soft assignment value corresponding to the first coding hash code and a second soft assignment value corresponding to the second coding hash code through a pseudo classifier, and constructing a cluster-level comparison loss function according to the first soft assignment value and the second soft assignment value; fusing the first coded Hash code and the second coded Hash code to obtain a fused coded Hash code; according to the first coding Hash code, the second coding Hash code and the fusion coding Hash code, constructing a steady-state loss function, and constructing a quantization loss function; combining the alignment loss function, the cluster-level comparison loss function, the steady-state loss function and the quantization loss function to construct a total loss function; and determining a trained unsupervised cross-modal hash retrieval model according to the total loss function convergence. According to the method and the device, the accuracy of cross-modal retrieval can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cross-modal retrieval, and in particular, to an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering. Background Art

[0002] With the rapid growth of multimedia data, cross-modal retrieval technology has become a research hotspot in the field of information retrieval. This technology aims to query relevant content of another modality (such as an image) through one modality (such as text), so as to achieve cross-modal information alignment and efficient retrieval. In recent years, hashing methods have been widely used in cross-modal retrieval tasks due to their high encoding efficiency, low storage cost, and fast retrieval speed. However, unsupervised cross-modal hashing methods still face significant challenges in feature alignment and semantic preservation due to the lack of explicit semantic annotation information.

[0003] To solve the problem of supervision dependence, unsupervised cross-modal hashing methods have emerged. These methods usually learn a shared semantic space by constructing an inter-modal similarity graph or introducing a pseudo-label mechanism. However, current unsupervised methods face some core challenges: (1) It is difficult to bridge the semantic gap between modalities. There are natural differences in the perceptual structure and feature distribution between images and texts. Traditional unsupervised methods lack effective alignment mechanisms and are difficult to learn consistent cross-modal representations. (2) Noise connections are likely to be generated in the construction of the graph structure. When constructing a sample similarity graph under unsupervised conditions, it often relies on low-level features or initial similarity metrics, which may lead to misconnection of semantically irrelevant samples, thereby affecting the accuracy of semantic structure modeling. (3) There is a lack of the ability to model the global semantic structure. Existing methods mostly focus on local sample relationships and are difficult to comprehensively express the global distribution and aggregation features of data in the latent semantic space. In addition, some methods introduce a clustering mechanism to assist in modal representation learning. Although it alleviates the problem of label loss to a certain extent, the clustering process often lacks cross-modal consistency modeling, resulting in the inability to accurately match the generated hash codes between different modalities.

[0004] In summary, the existing unsupervised cross-modal hashing retrieval methods have relatively low accuracy for cross-modal retrieval. Summary of the Invention

[0005] This application aims to propose an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering, which can improve the accuracy of cross-modal retrieval.

[0006] In a first aspect, an embodiment of this application provides an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering, and the method includes: Obtain multimodal data including image modal data and text modal data, and respectively extract the features of the first modal data and the second modal data to obtain the first extracted feature and the second extracted feature, where the first modal data and the second modal data are any one of the modal data in the multimodal data, and the first modal data and the second modal data are different modal data; Perform non-linear transformation on the first extracted feature and the second extracted feature respectively to obtain a first encoded hash code and a second encoded hash code; Calculate a similarity matrix based on the first extracted feature and the second extracted feature, and construct an alignment loss function according to the similarity matrix, the first encoded hash code and the second encoded hash code; Obtain a first soft assignment corresponding to the first encoded hash code and a second soft assignment corresponding to the second encoded hash code through a pseudo-classifier, and construct a cluster-level contrast loss function according to the first soft assignment and the second soft assignment; Fuse the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code; Construct a steady-state loss function according to the first encoded hash code, the second encoded hash code and the fused encoded hash code; Construct a quantization loss function according to the first encoded hash code and the second encoded hash code; Combine the alignment loss function, the cluster-level contrast loss function, the steady-state loss function and the quantization loss function to construct a total loss function; Determine a trained unsupervised cross-modal hashing retrieval model according to the convergence of the total loss function, so as to perform cross-modal hashing retrieval according to the trained unsupervised cross-modal hashing retrieval model.

[0007] Compared with the prior art, the first aspect of this application has the following beneficial effects: This method obtains multimodal data including image modal data and text modal data, and extracts the features of the first modal data and the second modal data respectively to obtain the first extracted feature and the second extracted feature. Among them, the first modal data and the second modal data are any one of the modal data in the multimodal data, and the first modal data and the second modal data are different modal data. The first extracted feature and the second extracted feature are respectively subjected to a non-linear transformation to obtain a first encoded hash code and a second encoded hash code. Through the non-linear transformation, the first extracted feature and the second extracted feature can be mapped to a unified hash space, realizing the alignment of the first modal data and the second modal data in the hash space. Then, a similarity matrix is calculated based on the first extracted feature and the second extracted feature, and an alignment loss function is constructed according to the similarity matrix, the first encoded hash code and the second encoded hash code, which can ensure that the features of the first modal data and the second modal data can be effectively matched in a common semantic space. Furthermore, a first soft assignment corresponding to the first encoded hash code and a second soft assignment corresponding to the second encoded hash code are obtained through a pseudo-classifier, and a cluster-level contrast loss function is constructed according to the first soft assignment and the second soft assignment, which can enhance the cross-modal semantic consistency. The first encoded hash code and the second encoded hash code are fused to obtain a fused encoded hash code. A steady-state loss function is constructed according to the first encoded hash code, the second encoded hash code and the fused encoded hash code, and a quantization loss function is constructed according to the first encoded hash code and the second encoded hash code. From the perspective of global optimization, a more robust feature expression method can be constructed, which has the characteristics of strong stability and insensitivity to outliers, thus improving the reliability of the hash representation. Finally, by combining the alignment loss function, the cluster-level contrast loss function, the steady-state loss function and the quantization loss function, a total loss function is constructed. According to the convergence of the total loss function, a trained unsupervised cross-modal hash retrieval model is determined, so as to perform cross-modal hash retrieval according to the trained unsupervised cross-modal hash retrieval model. By comprehensively considering and combining the alignment loss function, the cluster-level contrast loss function, the steady-state loss function and the quantization loss function to construct the total loss function, the accuracy of cross-modal retrieval can be improved according to the trained unsupervised cross-modal hash retrieval model.

[0008] In some embodiments, the step of respectively subjecting the first extracted feature and the second extracted feature to a non-linear transformation to obtain a first encoded hash code and a second encoded hash code includes: ; ; Wherein, represents the first encoded hash code, represents the hyperbolic tangent function, represents a scaling factor that adaptively adjusts with the number of training rounds, and Represents the weight matrix of the mapping layer, Represents the regularization layer, Represents the non-linear activation function, And Represents the weight matrix of the fully connected layer, Represents the first extracted feature, And Represents the bias matrix of the fully connected layer, And Represents the bias matrix of the mapping layer, Represents the second encoded hash code, Represents the second extracted feature.

[0009] In some embodiments, calculating a similarity matrix based on the first extracted feature and the second extracted feature, and constructing an alignment loss function according to the similarity matrix, the first encoded hash code, and the second encoded hash code, includes: Calculating a first similarity matrix based on the first extracted feature; Calculating a second similarity matrix based on the second extracted feature; Fusing the first similarity matrix and the second similarity matrix to obtain a fused similarity matrix; Calculating a first loss according to the fused similarity matrix and the first encoded hash code; Calculating a second loss according to the fused similarity matrix and the second encoded hash code; Calculating a third loss according to the fused similarity matrix, the first encoded hash code, and the second encoded hash code; Combining the first loss, the second loss, and the third loss to construct an alignment loss function.

[0010] In some embodiments, constructing a cluster-level contrast loss function according to the first soft assignment and the second soft assignment, includes: ; ; ; Wherein, Represents the contrast loss of the th cluster assignment statistical vector corresponding to the first modality data, Represents the contrast loss of the th cluster assignment statistical vector corresponding to the second modality data, Represents the cluster-level contrast loss function, Represents the first soft assignment of the th cluster corresponding to the first modality data, Denote the second soft assignment of the cluster corresponding to the second modal data, denote the cluster-level temperature parameter, denote the number of clusters, denote the function for calculating the cosine similarity, denote the first soft assignment of the cluster corresponding to the first modal data, denote the second soft assignment of the cluster corresponding to the second modal data, denote the information entropy corresponding to the first modal data, denote the information entropy corresponding to the second modal data, denote the exponential function.

[0011] In some embodiments, constructing the steady-state loss function according to the first encoded hash code, the second encoded hash code, and the fused encoded hash code includes: Normalize the first encoded hash code to obtain a first state vector; Normalize the second encoded hash code to obtain a second state vector; Normalize the fused encoded hash code to obtain a third state vector; Calculate the cosine similarity according to the fused encoded hash code, and construct a state transition probability according to the cosine similarity; Iterate the third state vector multiple times through the state transition probability to determine the steady-state distribution of the state vector; Construct a steady-state loss function according to the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector.

[0012] In some embodiments, constructing the steady-state loss function according to the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector includes: ; where denotes the steady-state loss function, denotes the number of samples corresponding to each modality, denotes the third state vector, denotes the steady-state distribution of the state vector, denotes the first state vector, denotes the second state vector, denotes the logarithmic function.

[0013] In some embodiments, constructing a quantization loss function according to the first encoded hash code and the second encoded hash code includes: ; Wherein, represents the quantization loss function, represents the first encoded hash code, represents the sign function, represents the second encoded hash code.

[0014] In some embodiments, fusing the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code includes: Concatenating the first encoded hash code and the second encoded hash code to obtain a concatenated encoded hash code; Non-linearly fusing the concatenated encoded hash code to obtain a fused encoded hash code.

[0015] In a second aspect, an embodiment of the present application further provides an unsupervised cross-modal hashing retrieval system based on steady-state distribution and clustering. The system includes: A data acquisition unit, configured to acquire multi-modal data including image modal data and text modal data, and respectively extract features of the first modal data and the second modal data to obtain a first extracted feature and a second extracted feature, where the first modal data and the second modal data are any one of the modal data in the multi-modal data, and the first modal data and the second modal data are different modal data; A non-linear transformation unit, configured to respectively perform non-linear transformation on the first extracted feature and the second extracted feature to obtain a first encoded hash code and a second encoded hash code; A first construction unit, configured to calculate a similarity matrix based on the first extracted feature and the second extracted feature, and construct an alignment loss function according to the similarity matrix, the first encoded hash code, and the second encoded hash code; A second construction unit, configured to obtain a first soft assignment corresponding to the first encoded hash code and a second soft assignment corresponding to the second encoded hash code through a pseudo-classifier, and construct a cluster-level contrast loss function according to the first soft assignment and the second soft assignment; A data fusion unit, configured to fuse the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code; A third construction unit, configured to construct a steady-state loss function according to the first encoded hash code, the second encoded hash code, and the fused encoded hash code; A fourth construction unit, configured to construct a quantization loss function according to the first encoded hash code and the second encoded hash code; A fifth construction unit, configured to construct a total loss function by combining the alignment loss function, the cluster-level contrast loss function, the steady-state loss function, and the quantization loss function; A cross-modal retrieval unit, configured to determine a trained unsupervised cross-modal hashing retrieval model according to the convergence of the total loss function, so as to perform cross-modal hashing retrieval according to the trained unsupervised cross-modal hashing retrieval model.

[0016] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering as described above.

[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering as described above.

[0018] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art, and reference can be made to the relevant descriptions in the above first aspect, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of an embodiment of an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering provided by the present application; Figure 2 is a schematic diagram of the overall architecture in the best embodiment of an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering provided by the present application; Figure 3 is a schematic structural diagram of an embodiment of an unsupervised cross-modal hashing retrieval system based on steady-state distribution and clustering provided by the present application; Figure 4 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0021] In the description of the present application, if the first, second, etc. are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0022] In the description of the present application, it should be understood that for the orientation description, such as up, down, etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present application.

[0023] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0024] First, several nouns involved in the present application are parsed: Deep semantic hashing coding: Using a deep neural network to extract high-level semantic features of data and compress them into discriminative binary hash codes for efficient retrieval.

[0025] Semantic alignment: Refers to the semantic correspondence of different modalities (such as images and texts) in the feature space, enabling similar contents to have similar representations.

[0026] Pseudo-classifier: Simulating a real classifier in unsupervised learning, used to generate pseudo-labels or similarity metrics to guide the model to learn discriminative features.

[0027] Soft assignment: When a sample is assigned to multiple classes or cluster centers, the assignment weights are continuous values instead of only selecting a hard label for one class.

[0028] Cluster assignment statistical vector (ASV): Represents the probability statistical information of the assignment of a sample on the cluster centers, used to characterize its positional relationship relative to the overall data distribution.

[0029] Steady-state distribution: Refers to the state distribution that finally tends to be stable in the iteration of a Markov chain, reflecting the average state probability after the system runs for a long time.

[0030] Transition probability matrix: Describes the probability of a Markov process transitioning from one state to another, and is used to model the evolutionary relationship between states.

[0031] To address the problem of supervision dependence, unsupervised cross-modal hashing methods have emerged. These methods usually learn a shared semantic space by constructing an inter-modal similarity graph or introducing a pseudo-label mechanism. However, current unsupervised methods face some core challenges: (1) It is difficult to bridge the semantic gap between modalities. There are natural differences in the perceptual structure and feature distribution between images and texts. Traditional unsupervised methods lack an effective alignment mechanism and are difficult to learn consistent cross-modal representations. (2) Noise connections are likely to occur in the construction of the graph structure. When constructing a sample similarity graph under unsupervised conditions, it often relies on low-level features or initial similarity metrics, which may lead to incorrect connections of semantically irrelevant samples, thereby affecting the accuracy of semantic structure modeling. (3) Lack of the ability to model the global semantic structure. Existing methods mostly focus on local sample relationships and are difficult to comprehensively represent the global distribution and aggregation features of data in the latent semantic space. In addition, some methods introduce a clustering mechanism to assist in modal representation learning. Although it alleviates the problem of label absence to a certain extent, the clustering process often lacks cross-modal consistency modeling, resulting in the final generated hash codes being unable to accurately match between different modalities.

[0032] To address the problem that existing unsupervised cross-modal hashing retrieval methods have relatively low accuracy for cross-modal retrieval, this application proposes an unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering.

[0033] Refer to Figure 1 , the flowchart of the unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering provided by the embodiments of this application. This unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering is applied to an electronic device, which can be a server or a mobile terminal, etc. As Figure 1 shown, this unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering may include the following steps: Step S100: Obtain multi-modal data including image modal data and text modal data, and respectively extract the features of the first modal data and the second modal data to obtain a first extracted feature and a second extracted feature, where the first modal data and the second modal data are any one of the modal data in the multi-modal data, and the first modal data and the second modal data are different modal data; Step S200: Respectively perform non-linear transformation on the first extracted feature and the second extracted feature to obtain a first encoded hash code and a second encoded hash code; Step S300: Calculate a similarity matrix based on the first extracted feature and the second extracted feature, and construct an alignment loss function according to the similarity matrix, the first encoded hash code, and the second encoded hash code; Step S400: Obtain the first soft assignment corresponding to the first encoded hash code and the second soft assignment corresponding to the second encoded hash code through a pseudo-classifier, and construct a contrastive loss function at the cluster level based on the first soft assignment and the second soft assignment; Step S500: Fuse the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code; Step S600: Construct a steady-state loss function based on the first encoded hash code, the second encoded hash code, and the fused encoded hash code; Step S700: Construct a quantization loss function based on the first encoded hash code and the second encoded hash code; Step S800: Combine the alignment loss function, the contrastive loss function at the cluster level, the steady-state loss function, and the quantization loss function to construct a total loss function; Step S900: Determine a trained unsupervised cross-modal hashing retrieval model according to the convergence of the total loss function, so as to perform cross-modal hashing retrieval according to the trained unsupervised cross-modal hashing retrieval model.

[0034] In this embodiment, by obtaining multimodal data including image modal data and text modal data, and respectively extracting the features of the first modal data and the second modal data, the first extracted feature and the second extracted feature are obtained, where the first modal data and the second modal data are any one of the modal data in the multimodal data, and the first modal data and the second modal data are different modal data. The first extracted feature and the second extracted feature are respectively subjected to a non-linear transformation to obtain a first encoded hash code and a second encoded hash code. Through the non-linear transformation, the first extracted feature and the second extracted feature can be mapped to a unified hash space, realizing the alignment of the first modal data and the second modal data in the hash space. Then, a similarity matrix is calculated based on the first extracted feature and the second extracted feature, and an alignment loss function is constructed according to the similarity matrix, the first encoded hash code, and the second encoded hash code, which can ensure that the features of the first modal data and the second modal data can be effectively matched in a common semantic space. Furthermore, a first soft assignment corresponding to the first encoded hash code and a second soft assignment corresponding to the second encoded hash code are obtained through a pseudo-classifier, and a cluster-level contrast loss function is constructed according to the first soft assignment and the second soft assignment, which can enhance cross-modal semantic consistency. The first encoded hash code and the second encoded hash code are fused to obtain a fused encoded hash code. A steady-state loss function is constructed according to the first encoded hash code, the second encoded hash code, and the fused encoded hash code, and a quantization loss function is constructed according to the first encoded hash code and the second encoded hash code, which can construct a more robust feature representation method from a global optimization perspective and has the characteristics of strong stability and insensitivity to outliers, thereby improving the reliability of the hash representation. Finally, by combining the alignment loss function, the cluster-level contrast loss function, the steady-state loss function, and the quantization loss function, a total loss function is constructed. According to the convergence of the total loss function, a trained unsupervised cross-modal hashing retrieval model is determined, so as to perform cross-modal hashing retrieval according to the trained unsupervised cross-modal hashing retrieval model. By comprehensively considering and combining the alignment loss function, the cluster-level contrast loss function, the steady-state loss function, and the quantization loss function to construct the total loss function, the accuracy of cross-modal retrieval can be improved according to the trained unsupervised cross-modal hashing retrieval model.

[0035] The above multimodal data may include image modal data, text modal data, audio modal data, video modal data, etc., and this embodiment does not make specific limitations. Images, texts, audios, and videos are data of different modalities, and images, texts, audios, and videos may be from actual scenarios such as cross-modal search, content recommendation, and multimodal intelligent analysis. For example, cross-modal search may be to find the corresponding image through known text information.

[0036] The above-mentioned obtaining the first soft assignment corresponding to the first encoded hash code and the second soft assignment corresponding to the second encoded hash code through the pseudo-classifier may be to perform clustering on the first encoded hash code and the second encoded hash code through the pseudo-classifier to obtain the soft assignment of the cluster (i.e., including the first soft assignment and the second soft assignment).

[0037] In some embodiments, subjecting the first extracted feature and the second extracted feature to non-linear transformation respectively to obtain the first encoded hash code and the second encoded hash code includes: ; ; Wherein, represents the first encoded hash code, represents the hyperbolic tangent function, represents the scaling factor that adaptively adjusts with the number of training rounds, and represent the weight matrix of the mapping layer, represents the regularization layer, represents the non-linear activation function, and represent the weight matrix of the fully connected layer, represents the first extracted feature, and represent the bias matrix of the fully connected layer, and represent the bias matrix of the mapping layer, represents the second encoded hash code, represents the second extracted feature.

[0038] In this embodiment, by subjecting the first extracted feature and the second extracted feature to non-linear transformation respectively, the first extracted feature and the second extracted feature can be mapped to a unified hash space, realizing the alignment of the first-modal data and the second-modal data in the hash space.

[0039] The above formula is first processed by the non-linear activation function, then passed through the regularization layer, then the scaling factor is adjusted, and finally passed through the hyperbolic tangent function to obtain the first encoded hash code and the second encoded hash code.

[0040] In some embodiments, calculating a similarity matrix based on the first extracted feature and the second extracted feature, and constructing an alignment loss function according to the similarity matrix, the first encoded hash code and the second encoded hash code, includes: Calculating a first similarity matrix based on the first extracted feature; Calculating a second similarity matrix based on the second extracted feature; Fusing the first similarity matrix and the second similarity matrix to obtain a fused similarity matrix; Calculate a first loss according to the fusion similarity matrix and the first encoded hash code; Calculate a second loss according to the fusion similarity matrix and the second encoded hash code; Calculate a third loss according to the fusion similarity matrix, the first encoded hash code, and the second encoded hash code; Construct an alignment loss function by combining the first loss, the second loss, and the third loss.

[0041] In this embodiment, by calculating a first similarity matrix based on the first extracted features, calculating a second similarity matrix based on the second extracted features, fusing the first similarity matrix and the second similarity matrix to obtain a fusion similarity matrix, calculating a first loss according to the fusion similarity matrix and the first encoded hash code, calculating a second loss according to the fusion similarity matrix and the second encoded hash code, calculating a third loss according to the fusion similarity matrix, the first encoded hash code, and the second encoded hash code, and constructing an alignment loss function by combining the first loss, the second loss, and the third loss. In this way, by using cosine similarity to measure the similarity between different modal data and fusing the cross-modal similarity matrices, it can be ensured that the first extracted features and the second extracted features can be effectively matched in a common semantic space.

[0042] In some embodiments, construct a cluster-level contrast loss function according to the first soft assignment and the second soft assignment, including: ; ; ; wherein, represents the contrast loss of the -th cluster assignment statistical vector corresponding to the first modal data, represents the contrast loss of the -th cluster assignment statistical vector corresponding to the second modal data, represents the cluster-level contrast loss function, represents the first soft assignment of the -th cluster corresponding to the first modal data, represents the second soft assignment of the -th cluster corresponding to the second modal data, represents the cluster-level temperature parameter, represents the number of clusters, represents the function for calculating cosine similarity, represents the first soft assignment of the -th cluster corresponding to the first modal data, represents the second soft assignment of the The second soft assignment of clusters, represents the information entropy corresponding to the first modal data, represents the information entropy corresponding to the second modal data, Represents an exponential function.

[0043] Above and The formula constructs a cluster-level contrast loss function based on the idea of ​​contrastive learning. Traditional contrastive learning learns discriminative feature representations by shortening the representation distance of positive sample pairs and pushing the representation distance of negative sample pairs away. The loss function is as follows: ; in, and are two augmented views of the same data sample (positive sample pair), is the representation of other samples (negative samples), Represents the temperature parameter.

[0044] Using this idea, the representation of the same cluster in different modes and As positive sample pairs, the representations of different clusters and as well as and As negative sample pairs, a cluster-level contrast loss function is constructed. By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the goal of pulling samples from the same cluster together and pushing samples from different clusters away is achieved.

[0045] Above In order to avoid all samples being assigned to the same cluster, information entropy is introduced and , by maximizing entropy, we force the samples to be evenly distributed to clusters and avoid degenerate solutions.

[0046] In some implementations, constructing a steady-state loss function based on the first encoding hash code, the second encoding hash code, and the fused encoding hash code includes: Normalizing the first encoded hash code to obtain a first state vector; Normalizing the second encoded hash code to obtain a second state vector; Normalize the fused encoding hash code to obtain a third state vector; According to the fusion encoding hash code, the cosine similarity is calculated, and the state transition probability is constructed according to the cosine similarity; Iterate the third state vector multiple times through the state transition probability to determine the steady-state distribution of the state vector; Construct a steady-state loss function based on the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector.

[0047] In this embodiment, the first state vector is obtained by normalizing the first encoded hash code; the second state vector is obtained by normalizing the second encoded hash code; the third state vector is obtained by normalizing the fused encoded hash code; the cosine similarity is calculated based on the fused encoded hash code, and the state transition probability is constructed based on the cosine similarity; the steady-state distribution of the state vector is determined through multiple iterations using the state transition probability; a steady-state loss function is constructed based on the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector. In this way, in the context of deep learning, the steady-state distribution can be regarded as an ideal form of global feature distribution for capturing the overall structure of the data. By determining the steady-state distribution of the state vector through multiple iterations using the state transition probability, a more robust feature representation can be constructed from a global optimization perspective, which can have the characteristics of strong stability and insensitivity to outliers, thereby improving the reliability of the hash representation.

[0048] In some embodiments, constructing a steady-state loss function based on the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector includes: ; where, represents the steady-state loss function, represents the number of samples corresponding to each modality, represents the third state vector, represents the steady-state distribution of the state vector, represents the first state vector, represents the second state vector, represents the logarithmic function.

[0049] The above formula is derived as follows: Traditional hash retrieval methods are usually based on local optimization. This embodiment optimizes from a global perspective and uses the core concept of the steady-state distribution of the Markov chain to improve the robustness to noise and outliers. Under the action of the same transition matrix, the state change will tend to be stable. According to the Perron-Frobenius theorem, an irreducible and aperiodic Markov chain has a unique stationary distribution , and for any initial distribution , the t-th iteration gives , and through iterative update, it will converge to . That is to say, ; ; According to this principle, in this embodiment, the state vector is first initialized , , and then the constructed transition probability matrix is: ; ; where, is the th row of the fused feature , is the th feature and the th feature's cosine similarity, is the transition probability from the th feature to the th feature, is the Euclidean norm. Since and , this proves that this Markov chain satisfies irreducibility and aperiodicity, so there must be a unique stationary distribution .

[0050] After reaching the steady state, the loss function is further designed to guide the model to learn the feature representation close to this steady state structure. To constrain the consistency between the current model output distribution and the ideal steady state distribution, the steady state distribution loss function is introduced: ; where, represents the Kullback-Leibler divergence, which is used to measure the difference between distributions, and is specifically expanded as: .

[0051] In some embodiments, according to the first encoded hash code and the second encoded hash code, a quantization loss function is constructed, including: ; where, represents the quantization loss function, represents the first encoded hash code, represents the sign function, represents the second encoded hash code.

[0052] The above quantization loss function is constructed based on the idea of minimizing the mean square error between the network output and its corresponding binary hash code and making the continuous features gradually approach space.

[0053] In some embodiments, the first encoded hash code and the second encoded hash code are fused to obtain a fused encoded hash code, including: Concatenate the first encoded hash code and the second encoded hash code to obtain the concatenated encoded hash code; Perform non - linear fusion on the concatenated encoded hash code to obtain the fused encoded hash code.

[0054] In this embodiment, by fusing the first encoded hash code and the second encoded hash code, it is possible to achieve the abstract extraction of shared semantics between modalities.

[0055] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments: With the rapid growth of multimedia data, cross - modal retrieval technology has become a research hotspot in the field of information retrieval. To solve the problem of supervision dependence, unsupervised cross - modal hashing methods have emerged as the times require. These methods usually learn the shared semantic space by constructing a similarity graph between modalities or introducing a pseudo - label mechanism. However, the current unsupervised methods face the following core challenges: (1) It is difficult to bridge the semantic gap between modalities. There are natural differences in the perceptual structure and feature distribution between images and texts. Traditional unsupervised methods lack an effective alignment mechanism and are difficult to learn consistent cross - modal representations.

[0056] (2) Noise connections are likely to occur in the construction of the graph structure. When constructing a sample similarity graph under unsupervised conditions, it is often based on low - level features or initial similarity metrics, which may lead to the misconnection of semantically irrelevant samples, thus affecting the accuracy of semantic structure modeling.

[0057] (3) Lack the ability to model the global semantic structure. Existing methods mostly focus on local sample relationships and are difficult to comprehensively express the global distribution and aggregation features of data in the latent semantic space.

[0058] In addition, some methods introduce a clustering mechanism to assist in the learning of modal representations. Although it alleviates the problem of label loss to a certain extent, the clustering process often lacks cross - modal consistency modeling, resulting in the final generated hash codes being unable to accurately match between different modalities.

[0059] In summary, how to effectively model cross - modal global semantic relationships in an unsupervised environment and improve the clustering consistency and semantic alignment effect between different modalities has become a key problem that needs to be broken through urgently in the current cross - modal hashing research.

[0060] Aiming at the problems of weak semantic modeling ability, inaccurate modal alignment, and large quantization error in existing unsupervised cross - modal hashing methods, this embodiment proposes an unsupervised cross - modal hashing retrieval method based on steady - state optimization and semantic clustering. This method constructs a steady - state probability transfer mechanism and a cross - modal contrast clustering framework, jointly optimizes the semantic structure perception and modal alignment process, so as to obtain consistent cross - modal hash codes without label supervision. The method of this embodiment mainly includes the following content: (1) Considering the characteristic differences between the image and text modalities, a convolutional neural network (VGG) and a bag-of-words model (BoW) are respectively used to extract multi-level feature representations, ensuring the expression ability between modalities from the source. Subsequently, a deep non-linear hashing coding network with dual branches of image and text is constructed. Through ReLU activation, Dropout regularization, and Tanh transformation, the original modality features are mapped to a consistent hashing space, and continuous hash codes are output.

[0061] (2) To further enhance the semantic alignment between different modalities, this embodiment designs a cross-modal alignment loss. Through feature normalization and cosine similarity constraint, it ensures that the representations of the image and text modalities in the embedding space are highly consistent. In addition, to solve the semantic modeling problem caused by the lack of labels in the unsupervised environment, a multi-modal clustering mechanism based on a pseudo-classifier is proposed. This mechanism uses soft assignment to generate a cluster assignment statistical vector (ASV), and constructs a cross-modal contrast loss and an entropy regularization term to generate a cluster-level contrast loss, promoting the alignment and distribution diversity of the clustering results between modalities, thereby enhancing the model's ability to capture the multi-modal semantic structure.

[0062] (3) In terms of feature stability, this embodiment introduces the steady-state distribution theory in the Markov chain, probabilistically models the fused features to generate a state vector, and constructs a transition probability matrix to guide the model output to approach a stable feature distribution state. By minimizing the Kullback-Leibler divergence between the current output distribution and the ideal steady-state distribution, the robustness and stability of the feature representation against noise, outliers, and complex semantics are effectively improved.

[0063] (4) In addition, considering that hashing retrieval ultimately relies on a compact binary code representation, this embodiment introduces a quantization loss function to constrain the continuous hash codes to approach their binary forms. This loss term, together with the modal alignment loss, the cluster-level contrast loss, and the steady-state loss, constitutes the total loss function, and is jointly optimized through an end-to-end training mechanism.

[0064] Finally, this embodiment can achieve efficient cross-modal hashing retrieval under unsupervised conditions. Users can optionally select an image or text as a query item, and through fast Hamming distance calculation, match and rank the potentially relevant content in the other modality. Benefiting from the comprehensive optimization of this embodiment in semantic modeling, structural stability, and binary quantizability, the proposed method demonstrates superior performance and broad application potential in scenarios such as image-text matching, image description generation, and cross-modal recommendation.

[0065] To further elaborate on the specific implementation of this embodiment, the following will detail the implementation process and technical details of this embodiment through the fusion and hashing retrieval tasks of two typical modalities, namely images and texts. Although this embodiment mainly takes images and texts as examples in the description, its implementation principle is applicable to the processing of other multi-modal data and can be flexibly applied to retrieval tasks in different fields. Referring to Figure 2 , the specific implementation includes the following implementation steps: Step S1, multi-modal feature extraction.

[0066] To extract rich feature information from multi-modal data to capture the semantic differences and correlations between different modalities, a preliminary multi-modal feature representation is first constructed. For image data (i.e., the first modal data), a deep learning model based on the convolutional neural network (VGG) is used. This model effectively extracts the hierarchical features of the image through multiple convolutional operations, captures image information at different scales through pooling operations, and further enhances the non-linear mapping ability of the network by combining the ReLU activation function. On this basis, a feature vector of an image (i.e., the first extracted feature) is output through a fully connected layer , and this feature vector represents the high-level semantic information of the image. For text data (i.e., the second modal data), the classical bag-of-words model (BoW) is used to model the features of the text data. Through this model, the word frequencies and their co-occurrence relationships in the text are effectively captured, forming a sparse feature (i.e., the second extracted feature) representation .

[0067] Step S2, deep semantic hashing encoding of multi-modal data.

[0068] To alleviate the problem of feature scale differences between different modalities, this embodiment designs a dual-branch non-linear projection network to achieve cross-modal feature alignment and unified encoding. The image branch and the text branch are each composed of a layer of fully connected network, ReLU activation function, and Dropout regularization layer. This network structure respectively passes the extracted feature vectors layer by layer and , ensuring that after the features of each modality undergo non-linear transformation, they can be mapped to a unified hash space. Subsequently, they are mapped to the target hash code length through linear transformation, and then continuous hash codes are output through the Tanh function. The final outputs of the network are respectively: ; ; where , is the number of samples, is the target hash code length; and are the weight matrices of the fully connected layer, and is the bias matrix of the fully connected layer; and is the weight matrix of the mapping layer, and is the bias matrix of the mapping layer; is the scaling factor adaptively adjusted with the number of training epochs.

[0069] This process not only realizes the alignment of the image modality and the text modality in the hash space, but also promotes the model to learn the potential semantic relationships between different modalities during the optimization process through the end-to-end training method. With this design, this embodiment can effectively process heterogeneous data and ensure the consistency of multimodal data in the same semantic space.

[0070] Step S3, cross-modal semantic alignment.

[0071] In multimodal hashing retrieval, in order to ensure that the features of the image and text modalities can be effectively matched in a common semantic space, this embodiment introduces a cross-modal semantic alignment mechanism. For the extracted image features and text features , first normalize them so that they fluctuate within a unified range. The normalized features are used to calculate the cross-modal similarity matrix, which uses cosine similarity to measure the similarity between data of different modalities, and fuse the cross-modal similarity matrix, where is the trade-off parameter for adjusting the importance of information from different modalities: ; ; ; For the deep semantic hash code (i.e., the first encoded hash code) and the deep semantic hash code (i.e., the second encoded hash code), after normalization, construct the following alignment loss: ; ; ; ; Among them, and are the trade-off parameters, represents the first similarity matrix, represents the second similarity matrix, represents the fused similarity matrix, represents the first loss, represents the second loss, Indicates the third loss, represents the alignment loss function.

[0072] This mechanism optimizes the effect of modality alignment by minimizing the distribution difference between the hash codes of the image modality and the text modality, effectively ensuring the representational consistency of different modality features in the hash space, and laying a foundation for the subsequent retrieval process.

[0073] Step S4, Cluster optimization and cross-modal semantic consistency enhancement.

[0074] To further optimize the cross-modal representation, this embodiment proposes a clustering method based on a pseudo-classifier, enabling different modalities to share clustering weights.

[0075] Specifically, instances can be passed through the pseudo-classifier to obtain soft assignments to the clusters, i.e., , where is the number of samples, is the number of clusters, represents two modalities of image and text, denoted as the th column of is: ; Among them, collects the probability values (i.e., soft assignments) of all instances assigned to the th cluster, is the first soft assignment when , is the second soft assignment when , which this embodiment calls the cluster assignment statistical vector (ASV).

[0076] The ASVs from different clusters should be mutually exclusive and, ideally, orthogonal, while the ASVs of different modalities from the same cluster should be consistent to enhance cross-modal semantic consistency. To achieve this goal, this embodiment extends contrastive clustering to multi-modal data, defines a cluster-level contrastive loss, and introduces entropy regularization to prevent all samples from being assigned to the same cluster, thereby improving retrieval robustness. The contrastive loss of the th ASV in the modality is calculated as follows: ; ; Among them, is the cluster-level temperature parameter, is the number of clusters, represents the cosine similarity function calculation. Then, for the entire data, the cluster-level contrastive loss is expressed as: ; Among them, is the information entropy, , by maximizing the entropy, all samples are prevented from being assigned to the same cluster, where is the number of samples, is the number of clusters.

[0077] Step S5, Steady-state distribution optimization and global feature stability.

[0078] In unsupervised cross-modal hashing retrieval, most traditional hashing methods rely on local similarity or reconstruction error, often ignoring the structural information of data at the overall distribution level, which leads to unstable performance of the generated hash codes when facing noise, abnormal samples or complex semantic structures.

[0079] To solve this problem, this embodiment introduces the concept of steady-state distribution in Markov chains, suppresses the influence of outliers through iterative diffusion, and constructs a more robust feature representation from a global optimization perspective. The steady-state distribution can be regarded as an ideal form of global feature distribution in the context of deep learning to capture the overall structure of data. In this embodiment, the fused feature matrix is used to generate a state vector, and multiple iterations are performed through the transition probability matrix until a steady state is reached, making it have the characteristics of strong stability and insensitivity to outliers, thereby improving the reliability of hash representation.

[0080] Step S5.1, Cross-modal deep semantic hash code fusion.

[0081] To solve the problem of semantic fragmentation and insufficient fusion caused by independent modeling of hash representations in the cross-modal retrieval task for the image modality and the text modality, this embodiment fuses the cross-modal deep semantic hash codes for subsequent operations. Specifically, first, the deep semantic hash codes of the two modalities and are concatenated in the channel dimension through feature splicing. To compress the concatenated vector back to a unified hash code dimension and simultaneously achieve the abstract extraction of shared semantics between modalities, the module introduces a fully connected layer (i.e., a linear transformation layer) to map the fused features to an output vector of the target hash code length, and maps the vector to a continuous hash code on (i.e., the fused encoded hash code).

[0082] Step S5.2, Steady-state distribution constraint optimization.

[0083] This embodiment uses the fused hash code for subsequent steady-state distribution optimization. First, the fused features ( is the number of samples, is the target hash code length) are normalized to a probability distribution (i.e., the third state vector), which is referred to as the state vector in this embodiment, is used to express the relative importance of samples in the feature space and eliminate the scale influence of the original features. Among them represents the th row of: ; At the same time, the independent distributions corresponding to the image modality and the text modality are respectively constructed to characterize the relative importance of the hash features of each modality in the feature space, which are respectively expressed as: ; ; Among them, and respectively represent and the th row of, represents the first state vector, represents the second state vector.

[0084] Next, a transition probability matrix between samples is constructed, where is the number of samples. This matrix is normalized to transition probabilities through the Softmax function based on the cosine similarity between features, ensuring a high state transition probability between similar samples: ; ; On this basis, through , where is the state vector after iterations, the distribution state is continuously updated through the iterative method until it approximately converges to the steady-state distribution that satisfies . This process can be regarded as an explicit global optimization process for capturing the overall structure of the data. After reaching the steady state, this embodiment further designs a loss function to guide the model to learn the feature representation close to this steady-state structure. In order to constrain the consistency between the current model output distribution and the ideal steady-state distribution, a steady-state loss function is introduced as: ; Among them, represents the Kullback-Leibler divergence, which is used to measure the difference between distributions and is specifically expanded as: ; Through the above mechanism, this embodiment can explicitly model and optimize the global steady-state distribution of features, making the fused hash representation more stable and having the ability to perceive the global structure, thereby improving the accuracy and robustness of multimodal hash retrieval.

[0085] Step S6, generate the total loss space.

[0086] Although a continuous hash representation can be learned through a deep network, since hash retrieval essentially depends on binary hash codes, if the difference between the continuous output and the ideal binary code is not constrained, the model may learn floating-point features that deviate significantly from the binary code, thus weakening the discriminative power and consistency in the retrieval stage. To solve this problem, this paper introduces a quantization loss term to constrain the continuous hash representation. This loss term makes the continuous features gradually approach the space by minimizing the mean square error between the network output and its corresponding binarized hash code. Formally, the quantization loss is defined as: ; On this basis, combining the foregoing modal alignment loss, cluster-level contrast loss, steady-state loss and other optimization objectives, a joint total loss function is further constructed, and are parameters for adjusting the loss weights: ; Step S7, repeat Step S2 to Step S6 until the model converges.

[0087] This embodiment jointly optimizes the foregoing constructed multiple loss functions and continuously monitors the change trend of the total loss during the training process. When the loss tends to be stable in multiple training rounds or the change amplitude is lower than the set threshold, it is determined that the entire model has reached the convergence state and the training process is terminated.

[0088] Refer to Figure 2 , the entire model is an unsupervised cross-modal hash retrieval model, including a convolutional neural network (VGG), a bag-of-words model (BoW), a two-branch non-linear projection network, a cross-modal semantic alignment mechanism, a soft clustering layer, and a steady-state distribution module, etc.

[0089] In addition, to prevent the model from falling into local optima or overfitting, weight decay, Dropout, and regularization means are introduced during the training process, and at the same time, the quantization degree and semantic preservation effect of the hash output are evaluated stage by stage. Through this training mechanism of repeated iteration and steady optimization, the model can effectively learn and align semantic relationships in the multimodal feature space, thereby stably outputting a hash representation with strong discriminability and good quantizability.

[0090] Step S8, efficient cross-modal retrieval and result ranking.

[0091] This embodiment is based on an optimized multi-modal hashing encoding mechanism. By mapping the image modality and the text modality to a unified hash space respectively, heterogeneous modalities have consistent semantic expressions in this space. Specifically, given any modality as a query item (such as an image), the hash code can efficiently match potential semantically related items in another modality (such as text) through fast Hamming distance calculation.

[0092] In the actual retrieval process, the system first calculates the hash code of the query sample, and then sorts the hash codes of all samples in the target database by Hamming distance. Since the hash code is in a compact binary form, this process can greatly improve the efficiency through bit operations, supporting sub-second response times. At the same time, thanks to the semantic preservation, modality alignment, and steady-state optimization mechanisms introduced in the previous training, the sorting results have stronger consistency and discriminability semantically.

[0093] Therefore, this embodiment can achieve cross-modal retrieval with high accuracy, high response speed, and high robustness, showing significant performance advantages in practical applications such as multi-modal image-text retrieval, image description matching, and visual question answering.

[0094] Compared with the prior art, the technical solution of this embodiment has the following advantages: This embodiment proposes an unsupervised cross-modal hashing retrieval method that integrates semantic alignment, clustering optimization, and steady-state modeling, effectively alleviating the core bottlenecks of the prior art in modality heterogeneity, feature instability, and hash code quantifiability. This method realizes the generation of continuous hash codes through a dual-branch deep encoding network for images and texts, combines cross-modal alignment loss and a pseudo-label-guided multi-modal clustering mechanism to capture the shared semantic structure under unsupervised conditions. At the same time, it globally models the features with the help of the steady-state distribution theory of Markov chains to further improve the stability and robustness of the representation. Finally, by jointly optimizing multiple loss functions, discriminative and highly compressible binary hash codes are effectively obtained to achieve efficient and accurate retrieval matching between image and text modalities. This embodiment has the characteristics of unified structure, efficient training, and strong adaptability, and is widely applicable to practical scenarios such as cross-modal search, content recommendation, and multi-modal intelligent analysis.

[0095] Refer to Figure 3 , this application embodiment also provides an unsupervised cross-modal hashing retrieval system based on steady-state distribution and clustering. The system includes a data acquisition unit 100, a non-linear transformation unit 200, a first construction unit 300, a second construction unit 400, a data fusion unit 500, a third construction unit 600, a fourth construction unit 700, a fifth construction unit 800, and a cross-modal retrieval unit 900, where: A data acquisition unit 100 is configured to acquire multi-modal data including image modal data and text modal data, and respectively extract features of the first modal data and the second modal data to obtain a first extracted feature and a second extracted feature, where the first modal data and the second modal data are any one of the modal data in the multi-modal data, and the first modal data and the second modal data are different modal data; A non-linear transformation unit 200 is configured to respectively perform non-linear transformation on the first extracted feature and the second extracted feature to obtain a first encoded hash code and a second encoded hash code; A first construction unit 300 is configured to calculate a similarity matrix based on the first extracted feature and the second extracted feature, and construct an alignment loss function according to the similarity matrix, the first encoded hash code and the second encoded hash code; A second construction unit 400 is configured to obtain a first soft assignment corresponding to the first encoded hash code and a second soft assignment corresponding to the second encoded hash code through a pseudo-classifier, and construct a cluster-level contrast loss function according to the first soft assignment and the second soft assignment; A data fusion unit 500 is configured to fuse the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code; A third construction unit 600 is configured to construct a steady-state loss function according to the first encoded hash code, the second encoded hash code and the fused encoded hash code; A fourth construction unit 700 is configured to construct a quantization loss function according to the first encoded hash code and the second encoded hash code; A fifth construction unit 800 is configured to construct a total loss function by combining the alignment loss function, the cluster-level contrast loss function, the steady-state loss function and the quantization loss function; A cross-modal retrieval unit 900 is configured to determine a trained unsupervised cross-modal hash retrieval model according to the convergence of the total loss function, so as to perform cross-modal hash retrieval according to the trained unsupervised cross-modal hash retrieval model.

[0096] It should be noted that since an unsupervised cross-modal hash retrieval system based on steady-state distribution and clustering in this embodiment is based on the same inventive concept as the above-mentioned unsupervised cross-modal hash retrieval method based on steady-state distribution and clustering, the corresponding content in the method embodiment is equally applicable to this system embodiment and will not be elaborated here.

[0097] Refer to Figure 4 , this application embodiment also provides an electronic device, and this electronic device includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering as described above in the present disclosure.

[0098] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0099] The electronic device according to the embodiments of the present application will be introduced in detail below.

[0100] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700, and the processor 1600 is called to execute the unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering in the embodiments of the present disclosure.

[0101] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices, and can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0102] An embodiment of the present disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering.

[0103] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0104] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0105] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0108] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0109] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0110] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0111] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs. The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made without departing from the gist of the present application within the knowledge scope of those of ordinary skill in the art.

[0114] The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made without departing from the gist of the present application within the knowledge scope of those of ordinary skill in the art.

Claims

1. An unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering, characterized in that The method includes: Obtain multimodal data including image modal data and text modal data, and respectively extract the features of the first modal data and the second modal data to obtain a first extracted feature and a second extracted feature, where the first modal data and the second modal data are any one of the modal data in the multimodal data, and the first modal data and the second modal data are different modal data; Respectively perform non-linear transformation on the first extracted feature and the second extracted feature to obtain a first encoded hash code and a second encoded hash code; Calculate a similarity matrix based on the first extracted feature and the second extracted feature, and construct an alignment loss function according to the similarity matrix, the first encoded hash code, and the second encoded hash code; Obtain a first soft assignment corresponding to the first encoded hash code and a second soft assignment corresponding to the second encoded hash code through a pseudo-classifier, and construct a cluster-level contrast loss function according to the first soft assignment and the second soft assignment; Fuse the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code; Construct a steady-state loss function according to the first encoded hash code, the second encoded hash code, and the fused encoded hash code; Construct a quantization loss function according to the first encoded hash code and the second encoded hash code; Combine the alignment loss function, the cluster-level contrast loss function, the steady-state loss function, and the quantization loss function to construct a total loss function; Determine a trained unsupervised cross-modal hashing retrieval model according to the convergence of the total loss function, so as to perform cross-modal hashing retrieval according to the trained unsupervised cross-modal hashing retrieval model.

2. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 1, wherein The step of respectively performing non-linear transformation on the first extracted feature and the second extracted feature to obtain a first encoded hash code and a second encoded hash code includes: ; ; Among them, represents the first encoded hash code, represents the hyperbolic tangent function, represents a scaling factor that adaptively adjusts with the number of training rounds, and represents the weight matrix of the mapping layer, represents the regularization layer, represents the non-linear activation function, and represents the weight matrix of the fully connected layer, represents the first extracted feature, and represents the bias matrix of the fully connected layer, and represents the bias matrix of the mapping layer, represents the second encoded hash code, represents the second extracted feature.

3. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 1, characterized in that, The step of calculating a similarity matrix based on the first extracted feature and the second extracted feature, and constructing an alignment loss function according to the similarity matrix, the first encoded hash code, and the second encoded hash code includes: Calculate a first similarity matrix based on the first extracted feature; Calculate a second similarity matrix based on the second extracted feature; Fuse the first similarity matrix and the second similarity matrix to obtain a fused similarity matrix; Calculate a first loss according to the fused similarity matrix and the first encoded hash code; Calculate a second loss according to the fused similarity matrix and the second encoded hash code; Calculate a third loss according to the fused similarity matrix, the first encoded hash code, and the second encoded hash code; Combine the first loss, the second loss, and the third loss to construct an alignment loss function.

4. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 1, wherein The step of constructing a cluster-level contrast loss function according to the first soft assignment and the second soft assignment includes: ; ; ; Among them, represents the contrast loss of the th cluster assignment statistical vector corresponding to the first modal data, represents the contrast loss of the th cluster assignment statistical vector corresponding to the second modal data, represents the contrast loss function at the cluster level, represents the first soft assignment of the th cluster corresponding to the first modal data, represents the second soft assignment of the th cluster corresponding to the second modal data, represents the cluster-level temperature parameter, represents the number of clusters, represents the function for calculating the cosine similarity, represents the first soft assignment of the th cluster corresponding to the first modal data, represents the second soft assignment of the th cluster corresponding to the second modal data, represents the information entropy corresponding to the first modal data, represents the information entropy corresponding to the second modal data, represents the exponential function.

5. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 1, characterized in that, The step of constructing a steady-state loss function according to the first encoded hash code, the second encoded hash code, and the fused encoded hash code includes: Normalize the first encoded hash code to obtain a first state vector; Normalize the second encoded hash code to obtain a second state vector; Normalize the fused encoded hash code to obtain a third state vector; Calculate the cosine similarity based on the fused encoded hash code, and construct a state transition probability based on the cosine similarity; Iterate the third state vector multiple times through the state transition probability to determine the steady-state distribution of the state vector; Construct a steady-state loss function based on the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector.

6. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 5, wherein The constructing a steady-state loss function based on the steady-state distribution of the state vector, the first state vector, the second state vector, and the third state vector includes: ; Among them, represents the steady-state loss function, represents the number of samples corresponding to each modality, represents the third state vector, represents the steady-state distribution of the state vector, represents the first state vector, represents the second state vector, represents the logarithmic function.

7. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 1, characterized in that, The constructing a quantization loss function based on the first encoded hash code and the second encoded hash code includes: ; Among them, represents the quantization loss function, represents the first encoded hash code, represents the sign function, represents the second encoded hash code.

8. The unsupervised cross-modal hashing retrieval method based on steady-state distribution and clustering according to claim 1, characterized in that The fusing the first encoded hash code and the second encoded hash code to obtain a fused encoded hash code includes: Concatenate the first encoded hash code and the second encoded hash code to obtain a concatenated encoded hash code; Non-linearly fuse the concatenated encoded hash code to obtain a fused encoded hash code.

Citation Information

Patent Citations

  • Depth unsupervised cross-modal retrieval method for reconstructing Hash based on modal fusion

    CN115687571A

  • Data retrieval method based on unsupervised cross-modal hash algorithm

    CN117540039A

  • Cross-modal hash retrieval method based on boundary mutual information

    CN117591623A

  • Cross-modal hash retrieval method based on graph convolutional network fusion

    CN119293271A

  • Cross-modal hash retrieval method based on attention mechanism

    CN119830222A