Cross-modal retrieval method and system based on hierarchical information aggregation hash
By using hierarchical information aggregation hashing technology in cross-modal retrieval, a cross-modal retrieval model is constructed, and the problem of missing fine-grained semantic information and semantic associations in multimodal data is solved, and the search efficiency and performance are improved.
Patent Information
- Application Number
- CN202510285765.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively mine the fine-grained semantic information and semantic associations of data in multimodal data retrieval, resulting in low retrieval efficiency and difficulty in obtaining information.
A cross-modal search method based on hierarchical information aggregation hash is adopted to construct a cross-modal search model through feature extraction module, student module and teacher module. A hierarchical information aggregation network and hash function are used to project different modal data features into the same dimension space and public Hamming space to train semantic supervision knowledge guidance.
Effectively considering the semantic information of multimodal data, improving the modeling ability of correlation relationships between samples, and improving the performance and efficiency of multimodal retrieval.
Smart Images

Figure CN120162472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-modal retrieval method and system based on hierarchical information aggregation hashing, and belongs to the technical field of data processing. Background Art
[0002] With the development of production technology and the continuous update and iteration of intelligent devices, the data involved in power intelligent scenarios is showing a diversified trend, covering various forms of data such as text, images, and videos. However, there is a large amount of scattered multi-modal data in power specialties such as equipment, safety supervision, and infrastructure construction, and the data association and information fusion are insufficient, resulting in low cross-modal retrieval efficiency and great difficulty in obtaining information.
[0003] The invention patent with the publication number of "CN117033308A" discloses a multi-modal retrieval method and device based on a specific range. The method includes: obtaining a specific scenario layout file; the specific scenario layout file includes a standard layout file, an image file, and an audio-video file; using a parsing tool to process the standard layout file to obtain the text content in the standard layout file; processing the image file to obtain the text content in the image file; processing the audio-video file to obtain the text content in the audio-video file; processing the text content in the standard layout file, the text content in the image file, and the text content in the audio-video file to establish a retrieval library; using a keyword retrieval method to search the retrieval library to obtain corresponding multi-modal content. The method of the present invention improves the file search efficiency and at the same time improves the file search accuracy.
[0004] However, the above scheme is only a shallow multi-modal retrieval model design structure, which is difficult to consider the rich semantic information in multi-modal data, difficult to mine the fine-grained semantic information of the data itself, and the modeling of the association relationship between samples is not sufficient, and the explicit semantic association between label categories is not fully considered, which affects the multi-modal retrieval performance. Summary of the Invention
[0005] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a cross-modal retrieval method and system based on hierarchical information aggregation hashing.
[0006] The technical solution of the present invention is as follows:
[0007] On the one hand, the present invention provides a cross-modal retrieval method based on hierarchical information aggregation hashing, including the following steps:
[0008] Collect data of different modalities and preprocess them respectively to construct a training data set;
[0009] Construct a cross-modal retrieval model, and the cross-modal retrieval model includes a feature extraction module, a student module, and a teacher module;
[0010] Input the training data set into the feature extraction module to extract the features of different modality data respectively;
[0011] Input the features of different modality data into the teacher module. The teacher module projects the features of different modality data into the same dimensional space through the feature encoding network, and hierarchically aggregates the features of different modality data in the same dimensional space through the hierarchical information aggregation network to obtain the aggregated features;
[0012] Construct the objective function of the teacher module, train the teacher module based on the objective function of the teacher module, obtain the trained teacher module, and extract and construct the teacher module parameters as semantic supervision knowledge;
[0013] Input the features of different modality data into the student module, and then project the features of different modality data into the same common Hamming space through the hash function. At the same time, during the training process of the hash function, conduct guided training through semantic supervision knowledge, and construct the features of different modality data in the same common Hamming space output by the hash function as joint features;
[0014] Construct the objective function of the student module, train the student module based on the objective function of the student module, and obtain the trained student module;
[0015] Construct a multi-modal database based on the cross-modal retrieval model. Input the retrieval content into the cross-modal retrieval model. The student module of the cross-modal retrieval model extracts the features of the retrieval content and the database content and projects them into the common Hamming space, and outputs the retrieval result according to the Hamming distance between the retrieval content and the database content.
[0016] As a preferred embodiment of the present invention, the training data set includes text data and image data.
[0017] As a preferred embodiment of the present invention, the feature encoding network includes a text feature encoding network and an image feature encoding network. The feature encoding network encodes the features in the form of hash codes, as shown in the following formula:
[0018] F x =ImageNet(X;θ x )
[0019] F y =TextNet(X;θ y )
[0020] Where: F x represents the output result of the image feature encoding network; X represents the input data; θ x represents the parameters of the image feature encoding network; F y represents the output result of the text feature encoding network; θy Represents the network parameters of the text feature encoding.
[0021] As a preferred embodiment of the present invention, the hierarchical information aggregation network includes an intra-modal information aggregation network and an inter-modal information aggregation network;
[0022] The intra-modal information aggregation network is used to aggregate the features of data in the same modality;
[0023] The inter-modal information aggregation network is used to aggregate the features of data in different modalities;
[0024] The aggregation results of the intra-modal information aggregation network and the inter-modal information aggregation network are aggregated again as the output of the hierarchical information aggregation network.
[0025] As a preferred embodiment of the present invention, the objective function of the teacher module is specifically shown as the following formula:
[0026]
[0027] Where: Represents the objective function of the teacher module; θ tea Represents the parameters of the teacher module; α and β respectively represent the balance coefficients of the teacher module; Represents the cosine similarity error between the output value of the teacher module and the actual value; Represents the quantization error of the training dataset samples.
[0028] As a preferred embodiment of the present invention, the specific steps of constructing the features of different modality data into joint features are as follows:
[0029] Construct a hash function for the text data and the image data respectively, and project the features of the text data and the image data into the same common Hamming space through the corresponding hash function;
[0030] Construct a bit-by-bit knowledge distillation module, and distill the features of the text data and the image data in the same common Hamming space bit by bit and construct them into joint features, specifically as shown in the following formula:
[0031]
[0032] Where: Represents the joint feature; Represents the feature of the image data in the same common Hamming space; Represents the feature of the text data in the same common Hamming space; W f Represents the fusion weight vector; ⊙ represents the Hadamard product.
[0033] As a preferred embodiment of the present invention, the objective function of the student module is specifically shown as the following formula:
[0034]
[0035] Wherein: represents the objective function of the student module; θ stu represents the parameters of the student module; μ represents the balance coefficient of the student module; represents the reconstruction loss of the intra-modal correlation relationship; represents the difference between the output of the student module and the output of the teacher module.
[0036] On the other hand, the present invention also provides a cross-modal retrieval system based on hierarchical information aggregation hashing, including a data acquisition module, a cross-modal retrieval model training module, and a database cross-modal retrieval module;
[0037] The data acquisition module is used to collect data of different modalities, respectively preprocess them, and construct a training data set;
[0038] The cross-modal retrieval model training module is used to construct a cross-modal retrieval model and train it through the training data set;
[0039] The database cross-modal retrieval module is used to construct a multi-modal database based on the cross-modal retrieval model, input the retrieval content into the cross-modal retrieval model, the student module of the cross-modal retrieval model extracts the features of the retrieval content and the database content and projects them into the common Hamming space, and outputs the retrieval result according to the Hamming distance between the retrieval content and the database content.
[0040] On yet another aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in any embodiment of the present invention.
[0041] On yet another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any embodiment of the present invention.
[0042] The present invention has the following beneficial effects:
[0043] 1. The present invention considers the rich semantic information in multi-modal data, can simultaneously consider the complementarity between different samples and different modalities of power data, and effectively solves the problems of fine-grained semantic information and semantic association loss in multi-modal data. Description of the Drawings
[0044] Figure 1 is a schematic structural diagram of the cross-modal retrieval model of the present invention;
[0045] Figure 2 is a schematic diagram of the hierarchical information aggregation network of the present invention. Detailed Embodiments
[0046] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0047] It should be understood that the step numbers used in the text are only for convenient description and do not limit the order of execution of the steps.
[0048] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0049] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0050] The term " / and / " refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0051] Embodiment 1:
[0052] See Figure 1 , a cross-modal retrieval method based on hierarchical information aggregation hashing, including the following steps:
[0053] Collect data of different modalities, preprocess them separately, and construct them into a training data set;
[0054] Construct a cross-modal retrieval model, which includes a feature extraction module, a student module, and a teacher module; in this embodiment, the cross-modal retrieval model further includes a discriminator (dimension: k→64→32→1), and each layer includes an activation function Relu behind it.
[0055] Input the training data set into the feature extraction module to extract the features of different modality data respectively. For image modality data, extract features through the CNN-F model, and for text modality data, extract features through the BoW model;
[0056] Input the features of different-modal data into the teacher module. The teacher module projects the features of different-modal data into the same-dimensional space through a feature encoding network, and hierarchically aggregates the features of different-modal data in the same-dimensional space through a hierarchical information aggregation network to obtain aggregated features;
[0057] Construct the objective function of the teacher module, train the teacher module based on the objective function of the teacher module, obtain the trained teacher module, and extract and construct the teacher module parameters as semantic supervision knowledge;
[0058] Input the features of different-modal data into the student module, and then project the features of different-modal data into the same common Hamming space through a hash function. At the same time, during the training process of the hash function, conduct guided training through semantic supervision knowledge, and construct the features of different-modal data in the same common Hamming space output by the hash function as joint features;
[0059] Construct the objective function of the student module, train the student module based on the objective function of the student module, and obtain the trained student module;
[0060] Construct a multi-modal database based on the cross-modal retrieval model. Input the retrieval content into the cross-modal retrieval model. The student module of the cross-modal retrieval model extracts the features of the retrieval content and the database content and projects them into the common Hamming space, and outputs the database content with the closest Hamming distance to the retrieval content as the retrieval result.
[0061] As a preferred implementation manner of this embodiment, the training data set includes text data and image data. In this embodiment, the training data set is defined as including text-modal data and image-modal data, denoted as where N is the number of training samples, d * represents the feature dimension of different modalities, *={x,y}; for the i-th multi-modal sample O i ={x i ,y i ,l i}, l i represents the annotation information. If O i belongs to the j-th class, then l ij =1, otherwise l ij =0.
[0062] As a preferred implementation manner of this embodiment, the feature encoding network includes a text feature encoding network and an image feature encoding network. The feature encoding network encodes the features in the form of hash codes, specifically as shown in the following formula:
[0063] F x =ImageNet(X;θ x)
[0064] F y = TextNet(X; θ y )
[0065] Where: F x represents the output result of the image feature encoding network; X represents the input data; θ x represents the parameters of the image feature encoding network; F y represents the output result of the text feature encoding network; θ y represents the parameters of the text feature encoding network;
[0066] In this embodiment, both the text feature encoding network and the image feature encoding network are composed of three fully connected layers (dimension: d x →2048→1024→512, d y →300→512→512), and each fully connected layer is followed by a batch normalization layer and an activation function Tanh.
[0067] As a preferred implementation manner of this embodiment, referring to Figure 2 , in order to avoid using a mandatory modality alignment constraint to learn a shared cross-modal space, which will inevitably lead to the problem of different degrees of information loss in each modality, a hierarchical information aggregation network is designed to construct a multi-modal complementary space to align heterogeneous modalities and model potential multi-modal association relationships. The hierarchical information aggregation network includes an intra-modal information aggregation network and a cross-modal information aggregation network;
[0068] The intra-modal information aggregation network is used to aggregate the features of the same modality data;
[0069] The cross-modal information aggregation network is used to aggregate the features of different modality data;
[0070] The aggregation results of the intra-modal information aggregation network and the cross-modal information aggregation network are aggregated again as the output of the hierarchical information aggregation network;
[0071] In this embodiment, the intra-modal information aggregation network (IM-MAN) and the cross-modal information aggregation network (IM-MAN) open-sourced by tensorflow are adopted. Since more aggregation layers will lead to the over-smoothing problem, both are set to a single aggregation layer. In addition, the input and output dimensions of IM-MAN and CM-MAN are both set to 512.
[0072] As a preferred implementation manner of this embodiment, the objective function of the teacher module is specifically shown as the following formula:
[0073]
[0074] Wherein: represents the objective function of the teacher module; θ tea represents the parameters of the teacher module; α = 100 and β = 0.001 respectively represent the balance coefficients of the teacher module; represents the cosine similarity error between the output value of the teacher module and the actual value; represents the quantization error of the training dataset samples.
[0075] As a preferred implementation manner of this embodiment, the specific steps of constructing the features of different modality data into joint features are as follows:
[0076] Construct a hash function for the text data and the image data respectively. The hash function is composed of three fully connected layers (dimensions: d x →2048→512→k, d y →300→512→k), and each layer is followed by a batch normalization layer and an activation function Tanh;
[0077] Construct a bit-by-bit knowledge distillation module. The features of the text data and the image data in the same common Hamming space are distilled bit by bit and constructed into joint features, as shown in the following formula:
[0078]
[0079] Wherein: represents the joint feature; represents the feature of the image data in the same common Hamming space; represents the feature of the text data in the same common Hamming space; W f represents the fusion weight vector, enabling different samples to share the same distillation weight vector, while different hash bits have different distillation ratios; ⊙ represents the Hadamard product.
[0080] As a preferred implementation manner of this embodiment, the objective function of the student module is specifically shown in the following formula:
[0081]
[0082] Wherein: represents the objective function of the student module; θ stu represents the parameters of the student module; μ = 1 represents the balance coefficient of the student module; represents the loss of reconstructing the intra-modal correlation relationship; represents the difference between the output of the student module and the output of the teacher module, that is, the difference between the hash codes.
[0083] In this embodiment, all parameters of the student module and the teacher module are trained based on the corresponding objective function through the standard backpropagation algorithm, and the learning rates of both the teacher module and the student module are set to 0.0001.
[0084] Embodiment 2:
[0085] A cross-modal retrieval system based on hierarchical information aggregation hashing, comprising a data acquisition module, a cross-modal retrieval model training module, and a database cross-modal retrieval module;
[0086] The data acquisition module is used to collect data of different modalities, preprocess them respectively, and construct a training data set;
[0087] The cross-modal retrieval model training module is used to construct a cross-modal retrieval model and train it through the training data set;
[0088] The database cross-modal retrieval module is used to construct a multi-modal database based on the cross-modal retrieval model, input the retrieval content into the cross-modal retrieval model, the student module of the cross-modal retrieval model extracts the features of the retrieval content and the database content and projects them into the common Hamming space, and outputs the retrieval result according to the Hamming distance between the retrieval content and the database content.
[0089] This system is used to implement the method in Embodiment 1, which will not be elaborated here.
[0090] Embodiment 3:
[0091] This embodiment provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in any embodiment of the present invention.
[0092] Embodiment 4:
[0093] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any embodiment of the present invention.
[0094] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0095] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0096] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0097] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (hereinafter referred to as ROM), random access memories (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.
[0098] The above are only the embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A cross-modal retrieval method based on hierarchical information aggregation hashing, characterized in that: The following steps are involved: Collect data of different modalities and construct training data sets after preprocessing them respectively; Constructing a cross-modal retrieval model, wherein the cross-modal retrieval model includes a feature extraction module, a student module, and a teacher module; Input the training data set into the feature extraction module to extract the features of different modal data respectively; The features of different modal data are input into the teacher module. The teacher module projects the features of different modal data into the same dimensional space through the feature encoding network, and hierarchically aggregates the features of different modal data in the same dimensional space through the hierarchical information aggregation network to obtain aggregated features. Constructing a teacher module objective function, training the teacher module based on the teacher module objective function, obtaining a trained teacher module and extracting teacher module parameters to construct semantic supervision knowledge; The features of different modal data are input into the student module, and then the features of different modal data are projected into the same common Hamming space through the hash function. At the same time, the training is guided by semantic supervision knowledge during the hash function training process, and the features of different modal data in the same common Hamming space output by the hash function are constructed as joint features. Constructing a student module objective function, training the student module based on the student module objective function, and obtaining a trained student module; A multimodal database is constructed based on the cross-modal retrieval model. The retrieval content is input into the cross-modal retrieval model. The student module of the cross-modal retrieval model extracts the features of the retrieval content and the database content and projects them into the common Hamming space. The retrieval results are output according to the Hamming distance between the retrieval content and the database content.
2. According to claim 1, a cross-modal retrieval method based on hierarchical information aggregation hashing is characterized in that: The training data set includes text data and image data.
3. According to claim 2, a cross-modal retrieval method based on hierarchical information aggregation hashing is characterized in that: The feature encoding network includes a text feature encoding network and an image feature encoding network. The feature encoding network encodes features in the form of hash codes, as shown in the following formula: F x NImageNet(X6θ x ) F y =TexNet(X;θ y ) Among them: F x represents the output result of the image feature encoding network; X represents the input data; θ x represents the image feature encoding network parameters; F y Represents the output result of the text feature encoding network; θ y Represents the parameters of the text feature encoding network.
4. According to claim 1, a cross-modal retrieval method based on hierarchical information aggregation hashing is characterized in that: The hierarchical information aggregation network includes an intra-modal information aggregation network and a cross-modal information aggregation network; The intra-modality information aggregation network is used to aggregate features of data of the same modality; The cross-modal information aggregation network is used to aggregate features of data of different modalities; The aggregation results of the intra-modal information aggregation network and the cross-modal information aggregation network are aggregated again as the output of the hierarchical information aggregation network.
5. According to claim 1, a cross-modal retrieval method based on hierarchical information aggregation hashing is characterized in that: The teacher module objective function is specifically shown in the following formula: in: represents the objective function of the teacher module; θ tea represents the teacher module parameter; α and β represent the balance coefficient of the teacher module respectively; Represents the cosine similarity error between the teacher module output value and the actual value; Represents the quantization error of the training dataset samples.
6. A cross-modal retrieval method based on hierarchical information aggregation hashing according to claim 2, characterized in that: The specific steps to construct the features of different modal data into joint features are: Construct a hash function for the text data and the image data respectively, and project the features of the text data and the image data into the same common Hamming space through the corresponding hash function; Construct a bit-by-bit knowledge distillation module, and distill the text data and image data features of the same common Hamming space bit by bit and construct them into joint features, as shown in the following formula: in: Represents joint features; Represent image data features in the same common Hamming space; Represents the text data features of the same public Hamming space; W f represents the fusion weight vector; ⊙ represents the Hadamard product.
7. The cross-modal retrieval method based on hierarchical information aggregation hashing according to claim 1 is characterized in that: The student module objective function is specifically shown in the following formula: in: represents the objective function of the student module; θ stu represents the student module parameter; μ represents the student module balance coefficient; Represents the intra-modal association reconstruction loss; Represents the difference between the output of the student module and the output of the teacher module.
8. A cross-modal retrieval system based on hierarchical information aggregation hashing, characterized in that: It includes data collection module, cross-modal retrieval model training module and database cross-modal retrieval module; The data acquisition module is used to collect data of different modes and construct them into training data sets after preprocessing respectively; The cross-modal retrieval model training module is used to construct a cross-modal retrieval model and train it through a training data set; The database cross-modal retrieval module is used to build a multimodal database based on a cross-modal retrieval model, input the retrieval content into the cross-modal retrieval model, and the student module of the cross-modal retrieval model extracts the retrieval content and database content features and projects them into a common Hamming space, and outputs the retrieval results according to the Hamming distance between the retrieval content and the database content.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal retrieval method and device based on specific range
CN117033308A