A deep adaptive multimodal hash retrieval method and related equipment

By adaptively fusing data information from multiple modes in multimodal hash learning, using deep learning neural networks for hash learning, and introducing semantic supervision, the problem of low hash learning efficiency in the existing technology is solved, and efficient and discriminant hash code generation is achieved.

CN114691897BActive Publication Date: 2025-05-16HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210284064.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-05-16
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

The existing multimodal hash learning methods cannot effectively adaptively integrate data information from multiple modalities, affecting the efficiency of hash learning.

Method used

By selecting multiple target training samples to form the target training batch, the feature extraction network is used to obtain the initial features of each modal, and the weight extraction network is used to obtain the weights of each modal and fusion to obtain the fusion features. Then, the fused features are input to the first hash network to obtain the sample hash code, and the semantic tag is input to the second hash network to obtain the semantic hash code, and the training loss is calculated based on both and the network parameters are updated.

Benefits of technology

The complementary fusion of adaptive weight update and multimodal content is realized, which improves the efficiency of hash learning, and makes the final hash code discriminant and effective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691897B_ABST
    Figure CN114691897B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep adaptive multimodal hash retrieval method and related equipment. The method provided by the present invention first designs a feature learning network for each modal data according to the physical characteristics and properties of each modal data in a hash learning process for multimodal data, determines a learnable weight for each modal feature according to the contribution of each modality in the training sample invested in learning each time to the performance of the final common feature, fuses the features of each modality according to the weight, and realizes information fusion of adaptive weights according to the characteristics of the training sample itself; minimizes the difference between the fused common feature and the hash code, adds scalable semantic features extracted from preset labels in this process, automatically updates the parameters of the hash function, realizes the alignment of the feature space and the hash space, uses label semantic information to supervise parameter updating, and can improve the adaptive fusion capability of multimodal features and the discriminative representation capability of hash learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of hash retrieval technology, and in particular to a deep adaptive multi-modal hash retrieval method and related equipment. Background Art

[0002] With the rapid development of information technology, multimedia data has become more and more diverse, including images, text, audio, etc. Multimodal hash retrieval is to encode data with multiple modalities into compact binary codes for retrieval. Before hash retrieval, hash learning is required. The existing multimodal hash learning method cannot effectively and adaptively integrate data information of multiple modalities, which affects the efficiency of hash learning.

[0003] Therefore, the prior art still needs to be improved and enhanced. Summary of the invention

[0004] In view of the above-mentioned defects of the prior art, the present invention provides a deep adaptive multimodal hash retrieval method and related equipment, aiming to solve the problem in the prior art that multimodal hash retrieval cannot adaptively fuse data information of multiple modalities.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0006] A first aspect of the present invention provides a deep adaptive multimodal hash retrieval method, the method comprising:

[0007] Selecting multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, determining a first feature extraction network set corresponding to the target training sample according to the modality of the data in the target training sample, the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and obtaining initial features of each modality in the target training sample through each of the first feature extraction networks;

[0008] Inputting the initial features of each modality in the target training sample into a weight extraction network, obtaining the weights of each modality in the target training sample output by the weight extraction network, and fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality to obtain a fused feature corresponding to the target training sample;

[0009] Input the fusion feature of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network;

[0010] Obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and updating the parameters of the first hash network according to the training loss of the target training batch;

[0011] The step of selecting multiple target training samples from multiple training samples to form a target training batch is re-executed until the parameters of the first hash network converge, and the hash code corresponding to the sample to be retrieved is obtained by using the first hash network after the parameters converge.

[0012] The weight extraction network includes a feature extraction layer and a weight output layer, and the initial features of each modality in the target training sample are input into the weight extraction network, and the weights of each modality in the target training sample output by the weight extraction network are obtained, including:

[0013] Inputting the initial features of each modality in the target training sample into the feature extraction layer to obtain potential consistent features corresponding to each modality in the target training sample;

[0014] Inputting the potential consistent features corresponding to each modality in the target training sample into the weight output layer to obtain the weight of each modality in the target training sample;

[0015] The fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality includes:

[0016] The potential consistencies of each modality in the target training sample are fused according to the weight corresponding to each modality to obtain a fusion feature corresponding to the target training sample.

[0017] The deep adaptive multimodal hash retrieval method, wherein the step of obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each target training sample, comprises:

[0018] Obtaining a first loss of the target training batch according to a difference between the sample hash code and the fusion feature of each of the target training samples;

[0019] Obtain a second loss of the target training batch according to a difference between the sample hash code and the semantic hash code of each of the target training samples;

[0020] Obtaining a third loss of the target training batch according to the sample hash code of each of the target training samples and the semantic similarity between each of the target training samples;

[0021] The training loss of the target training batch is obtained according to the first loss, the second loss, and the third loss.

[0022] The deep adaptive multimodal hash retrieval method, wherein, before obtaining the third loss of the target training batch according to the semantic hash code of each of the target training samples and the semantic similarity between each of the target training samples, the method further includes:

[0023] The semantic similarity between the target training samples is obtained according to the semantic label corresponding to each target training sample.

[0024] The deep adaptive multimodal hash retrieval method, wherein the updating of the parameters of the first hash network according to the training loss of the target training batch comprises:

[0025] Update parameters of the first hash network, the second hash network, and the weight extraction network according to the loss of the target training sample.

[0026] The deep adaptive multimodal hash retrieval method, wherein the step of selecting a plurality of target training samples from the plurality of training samples to form a target training batch, comprises:

[0027] Obtaining the fusion feature loss of each of the training samples according to the fusion features respectively corresponding to the multiple training samples under the current network parameters and the initial features respectively corresponding to the multiple training samples;

[0028] The plurality of training samples are sorted according to the fusion feature losses respectively corresponding to the plurality of training samples, and the plurality of target training samples are selected according to the sorting result.

[0029] The deep adaptive multimodal hash retrieval method, wherein the step of using the first hash network after parameter convergence to obtain the hash code corresponding to the sample to be retrieved includes:

[0030] The fused features corresponding to the samples to be retrieved are obtained, the fused features of the samples to be retrieved are input into the first hash network after parameter convergence, and the hash codes corresponding to the samples to be retrieved output by the first hash network are obtained.

[0031] A second aspect of the present invention provides a deep adaptive multimodal hash retrieval device, comprising:

[0032] An initial feature extraction module, wherein the feature extraction module is used to select multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, and a first feature extraction network set corresponding to the target training sample is determined according to the modality of the data in the target training sample, wherein the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and initial features of each modality in the target training sample are obtained through each of the first feature extraction networks;

[0033] A fusion feature extraction module, which is used to obtain the weight of each modality in the target training sample according to the initial features of each modality in the target training sample, and fuse the initial features of each modality in the target training sample according to the weight corresponding to each modality to obtain the fusion feature corresponding to the target training sample;

[0034] A hash module, wherein the hash module is used to input the fusion feature of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network;

[0035] A parameter updating module, the parameter updating module is used to obtain the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and update the parameters of the first hash network according to the training loss of the target training batch;

[0036] An iteration module, the iteration module is used to re-execute the step of selecting multiple target training samples from multiple training samples to form a target training batch until the parameters of the first hash network converge;

[0037] A retrieval module is used to obtain a hash code corresponding to a sample to be retrieved by using the first hash network after parameter convergence.

[0038] According to a third aspect of the present invention, a terminal is provided, comprising a processor and a computer-readable storage medium communicatively connected to the processor, wherein the computer-readable storage medium is suitable for storing a plurality of instructions, and the processor is suitable for calling the instructions in the computer-readable storage medium to execute the steps of implementing any of the above-described deep adaptive multimodal hash retrieval methods.

[0039] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the deep adaptive multimodal hash retrieval method described in any one of the above items.

[0040] Compared with the prior art, the present invention provides a deep adaptive multimodal hash retrieval method and related equipment. In the process of hash learning, the initial features of different modalities are extracted according to the neural networks suitable for different modalities, and then the weights of each modality are outputted according to the neural network based on the initial features of each modality, and then fused to obtain the fusion features of the data of each modality, thereby realizing adaptive weight update and complementary fusion of multimodal content. After obtaining the fused features, the fused features are input into the first hash network to obtain feature hash codes, and the semantic labels are also converted into binary codes (semantic hash codes), and the training loss is obtained according to the feature hash codes and the semantic hash codes. The method adopts an end-to-end deep learning neural network for hash learning, introduces semantic supervision in the process of hash learning, improves the efficiency of hash learning, and makes the hash code finally obtained have discriminativeness and effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A flowchart of an embodiment of the deep adaptive multimodal hash retrieval method provided by the present invention;

[0042] Figure 2 A schematic diagram of a learning framework in an embodiment of the deep adaptive multimodal hash retrieval method provided by the present invention;

[0043] Figure 3 A structural principle diagram of an embodiment of a deep adaptive multi-modal hash retrieval device provided by the present invention;

[0044] Figure 4 A schematic diagram of the principles of an embodiment of a terminal provided by the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] The deep adaptive multimodal hash retrieval method provided by the present invention can be applied to a terminal with computing capabilities. The terminal can execute the deep adaptive multimodal hash retrieval method provided by the present invention to determine the multimodal hash retrieval parameters. The terminal can be but is not limited to various computers, mobile terminals, smart home appliances, wearable devices, etc.

[0047] Embodiment 1

[0048] like Figure 1 As shown, in one embodiment of the deep adaptive multimodal hash retrieval method, the steps include:

[0049] S100. Select multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, determine a first feature extraction network set corresponding to the target training sample according to the modality of the data in the target training sample, the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and obtain initial features of each modality in the target training sample through each of the first feature extraction networks.

[0050] In this embodiment, one training batch is used to complete one training, that is, the network parameters are updated once according to the operation results of all training samples in one training batch. Each training sample is multimodal data, that is, each of the training samples includes data of multiple modes. Specifically, each information source can be called a mode, and the multimodal data includes data from multiple information sources, such as video, audio, text, image, etc. When processing multimodal data, feature extraction must be performed first. For data of different modes, different feature extraction networks can be used to extract features to obtain the initial features corresponding to the data of each mode. For example, for image data, CNN network can be used to extract features, for video data, C3D network can be used to extract features, and for text data, BERT network can be used to extract initial features. Of course, it can be understood that there are a variety of optional feature extraction networks for data of different modes in this field, and this field is not limited to the above examples.

[0051] The modalities included in different training samples may be different. For each target training sample, the first feature extraction networks corresponding to the target training sample are determined according to the modality of the data in the target training sample. Then, the data of the corresponding modality in the target training sample is subjected to feature extraction according to the first feature extraction networks corresponding to the target training sample to obtain the initial features of each modality in the target training sample.

[0052] S200. Input the initial features of each modality in the target training sample into a weight extraction network, obtain the weights of each modality in the target training sample output by the weight extraction network, fuse the initial features of each modality in the target training sample according to the weight corresponding to each modality, and obtain the fused features corresponding to the target training sample.

[0053] The initial features of each modality in the target training sample are further processed in the weight extraction network. Specifically, the weight extraction network may include a feature extraction layer and a weight output layer, such as Figure 2As shown, the initial features of each modality in the target training sample are first further extracted through the feature extraction layer. This step can be regarded as projecting the initial features of each modality in the target training sample to the latent consistent space to obtain the latent consistent features. The latent consistent features corresponding to each modality reflect the feature expression of each modality in the latent consistent space. After the latent consistent features are input into the weight output layer, the weight output layer outputs the weights corresponding to each modality. The parameters of the weight output layer are determined based on the performance contribution of the fused features obtained after the fusion of the initial features of each modality to the final generated hash code, that is, during the training process, the parameters of the weight output layer are updated according to the training loss to make the training loss as small as possible. The initial features of each modality in the target training sample are input into the weight extraction network to obtain the weights of each modality in the target training sample output by the weight extraction network, including:

[0054] Inputting the initial features of each modality in the target training sample into the feature extraction layer to obtain potential consistent features corresponding to each modality in the target training sample;

[0055] The potential consistent features corresponding to each modality in the target training sample are respectively input into the weight output layer to obtain the weight of each modality in the target training sample.

[0056] The fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality includes:

[0057] The potential consistencies of each modality in the target training sample are fused according to the weight corresponding to each modality to obtain a fusion feature corresponding to the target training sample.

[0058] It is not difficult to see from the previous description that the weight corresponding to each modality is adaptively adjusted according to the characteristics of the data of each modality, so that the complementary fusion of features can be achieved according to the actual situation of each modality data in the sample.

[0059] After obtaining the weights corresponding to the various modes output by the weight output layer, the initial features of the various modes in the target training sample are fused according to the weights corresponding to each mode, including:

[0060] The potential consistencies of each modality in the target training sample are fused according to the weight corresponding to each modality to obtain a fusion feature corresponding to the target training sample.

[0061] S300, input the fusion features of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network.

[0062] In order to ensure that the generated hash code is discriminative and effective, semantic supervision is introduced in this embodiment. Specifically, for each of the training samples, there is a corresponding semantic label, and the semantic label reflects the semantic category of the training sample. The semantic label can be a vector composed of 0 and 1, and each value in the vector reflects whether the training sample belongs to a certain semantic category.

[0063] After obtaining the fused features of the target training sample, the fused features of the target training sample are input into the first hash network, and the fused features of the target training sample are processed by the hash function in the first hash network to obtain the sample hash code output by the first hash network. The semantic label corresponding to the target training sample is flexibly transformed, specifically, it is input into the second hash network to obtain the semantic hash code output by the second hash network, that is, the semantic label corresponding to the target training sample is converted into a binary code, and the length of the binary code can be arbitrary.

[0064] The obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each target training sample includes:

[0065] Obtaining a first loss of the target training batch according to a difference between the sample hash code and the fusion feature of each of the target training samples;

[0066] Obtain a second loss of the target training batch according to a difference between the sample hash code and the semantic hash code of each of the target training samples;

[0067] Obtaining a third loss of the target training batch according to the sample hash code of each of the target training samples and the semantic similarity between each of the target training samples;

[0068] The training loss of the target training batch is obtained according to the first loss, the second loss, and the third loss.

[0069] In order to make the hash code of the sample effective and discriminative, the difference between the sample hash code and the fusion feature of the same sample calculated through the above steps should be as small as possible, so as to retain more original information. At the same time, the difference between the sample hash code and the semantic hash code of the same sample should be as small as possible, so that the hash code can accurately reflect the semantics. The more similar the sample hash codes of samples with similar semantic labels should be, and the more dissimilar the sample hash codes of samples with dissimilar semantic labels should be, so as to make the hash code discriminative.

[0070] That is, before obtaining the third loss of the target training batch according to the semantic hash code of each target training sample and the semantic similarity between each target training sample, the method further includes the following steps:

[0071] The semantic similarity between the target training samples is obtained according to the semantic label corresponding to each target training sample.

[0072] Specifically, the third loss of the target training batch is obtained according to the sample hash code of each target training sample and the semantic similarity between each target training sample. For each pair of target training samples in the target training batch, the difference between the sample hash codes and the semantic similarity between the two are calculated, and the distance between the difference in the sample hash codes and the semantic similarity is measured to obtain the sample pair loss, and all the sample pair losses in the target training batch are summed to obtain the third loss.

[0073] When obtaining the training loss of the target training batch according to the first loss, the second loss and the third loss, the first loss, the second loss and the third loss may be assigned corresponding weights and then weighted summed to obtain the training loss of the target training batch.

[0074] Please refer again Figure 1 , the method provided in this embodiment further includes the steps of:

[0075] S400. Obtain the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and update the parameters of the first hash network according to the training loss of the target training batch.

[0076] The deep adaptive multimodal hash retrieval method provided in this embodiment adopts an end-to-end training method to update the network parameters according to the training loss of the target training batch. In a possible implementation, only the parameters of the first hash network can be updated, and the parameters of other networks can be fixed. In order to speed up the training process and improve the training effect, the parameters of the weight extraction network and the second hash network can also be learnable, that is, the parameters of the first hash network are updated according to the training loss of the target training batch, including:

[0077] Update parameters of the first hash network, the second hash network, and the weight extraction network according to the loss of the target training sample.

[0078] Furthermore, the parameters of the first feature extraction network may also be learnable, that is, they are updated together with the parameters of other networks during the training process.

[0079] Please refer again Figure 1 , the method provided in this embodiment further includes the steps of:

[0080] S500, re-execute the step of selecting multiple target training samples from the multiple training samples to form a target training batch until the parameters of the first hash network converge, and use the first hash network after the parameters converge to obtain the hash code corresponding to the sample to be retrieved.

[0081] After updating the network parameters once, reselect the training batch for the next training, that is, select the new target training batch, and then update the network parameters according to the above steps. Repeat this process until the network parameters converge, and the training is completed.

[0082] In this embodiment, in order to improve the training efficiency, when selecting the target training sample from the multiple training samples, it is not randomly selected, but selected according to the information entropy corresponding to each training sample. Specifically, the selecting multiple target training samples from the multiple training samples to form a target training batch includes:

[0083] Obtaining the fusion feature loss of each of the training samples according to the fusion features respectively corresponding to the multiple training samples under the current network parameters and the initial features respectively corresponding to the multiple training samples;

[0084] The plurality of training samples are sorted according to the fusion feature losses respectively corresponding to the plurality of training samples, and the plurality of target training samples are selected according to the sorting result.

[0085] The fusion feature loss corresponding to the target training sample reflects the difference between the fusion feature obtained after the initial features corresponding to each modality of the target training sample are fused and the initial feature. The fusion feature loss can be calculated by finding the difference, distance, regularization loss, etc. between the initial feature and the fusion feature. At the end of each training, the parameters of the neural network will be updated. For the training samples that are not selected as the target training batch in this training, the corresponding initial features and fusion features are obtained according to the processing methods in steps S100-S200, and then the fusion feature loss corresponding to each training sample is obtained according to the initial features and fusion features corresponding to each training sample. Then, the fusion losses corresponding to each training sample are sorted, and multiple target training samples in a new training are selected as the target training batch in a new training according to the sorting results.

[0086] The fusion feature loss corresponding to the training sample reflects the learning difficulty of the training sample. The greater the fusion feature loss, the greater the learning difficulty and the greater the information entropy. Conversely, the smaller the information entropy. The fusion feature losses corresponding to the training samples that are not selected as the target training samples in the multiple training samples under the current network parameters are sorted, and the first n training samples with the smallest current fusion feature loss can be used as the target training samples in the next training. The effect of sample learning in order from easy to difficult can be achieved, effectively improving the efficiency of learning.

[0087] The acquiring of a hash code corresponding to a sample to be retrieved by using the first hash network after parameter convergence includes:

[0088] Obtain the fused features corresponding to the samples to be retrieved, input the fused features of the samples to be retrieved into the first hash network after parameter convergence, and obtain the hash code corresponding to the samples to be retrieved output by the first hash network. After the network parameters converge, the training is completed, and the hash code corresponding to the samples to be retrieved is obtained based on the network parameters after the training is completed to realize hash retrieval. Specifically, after the training is completed, the second hash network is not used in the process of obtaining the hash code corresponding to the samples to be retrieved. The specific process is: according to the modality of the data in the samples to be retrieved, the corresponding first feature extraction network is used to extract the initial features, and then the weights of each modality are obtained through the weight extraction network, and the features are fused based on the weights to obtain the fused features corresponding to the samples to be retrieved, and the fused features corresponding to the retrieved samples are input into the first hash network with converged parameters to obtain the hash code of the samples to be retrieved.

[0089] In summary, this embodiment provides a deep adaptive multimodal hash retrieval method. In the process of hash learning, the initial features of different modalities are extracted according to the neural networks applicable to different modalities, and then the neural network is used to output the weights of each modality according to the initial features of each modality, and then fusion is performed to obtain the fusion features of the data of each modality, thereby realizing adaptive weight update and complementary fusion of multimodal content. After obtaining the fusion features, the fusion features are input into the first hash network to obtain the feature hash code, and the semantic label is also converted into a binary code (semantic hash code), and the training loss is obtained according to the feature hash code and the semantic hash code. This method uses an end-to-end deep learning neural network for hash learning, introduces semantic supervision in the process of hash learning, improves the efficiency of hash learning, and makes the final hash code discriminative and effective.

[0090] It should be understood that, although the various steps in the flowcharts given in the accompanying drawings of the present specification are displayed in sequence according to the indications of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0091] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0092] Embodiment 2

[0093] Based on the above embodiments, the present invention also provides a deep adaptive multi-modal hash retrieval device. Figure 3 As shown, the deep adaptive multimodal hash retrieval device includes:

[0094] An initial feature extraction module, wherein the feature extraction module is used to select multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, and a first feature extraction network set corresponding to the target training sample is determined according to the modality of the data in the target training sample, wherein the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and initial features of each modality in the target training sample are obtained through each of the first feature extraction networks, as described in Embodiment 1;

[0095] A fusion feature extraction module, which is used to obtain the weight of each modality in the target training sample according to the initial features of each modality in the target training sample, and fuse the initial features of each modality in the target training sample according to the weight corresponding to each modality to obtain the fusion feature corresponding to the target training sample, as described in Example 1;

[0096] A hash module, wherein the hash module is used to input the fusion feature of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network, as described in Embodiment 1;

[0097] A parameter updating module, the parameter updating module is used to obtain the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and update the parameters of the first hash network according to the training loss of the target training batch, as described in the first embodiment;

[0098] An iteration module, the iteration module is used to re-execute the step of selecting multiple target training samples from multiple training samples to form a target training batch until the parameters of the first hash network converge, as described in the first embodiment;

[0099] A retrieval module, the retrieval module is used to obtain the hash code corresponding to the sample to be retrieved using the first hash network after parameter convergence, as described in the first embodiment.

[0100] Embodiment 3

[0101] Based on the above embodiments, the present invention also provides a terminal, such as Figure 4 As shown, the terminal includes a processor 10 and a memory 20. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0102] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as a hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a deep adaptive multimodal hash retrieval program 30 is stored on the memory 20, and the deep adaptive multimodal hash retrieval program 30 can be executed by the processor 10, thereby realizing the deep adaptive multimodal hash retrieval method in the present application.

[0103] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other chip, used to run the program code or process data stored in the memory 20, such as executing the deep adaptive multimodal hash retrieval method.

[0104] In one embodiment, when the processor 10 executes the deep adaptive multimodal hash retrieval program 30 in the memory 20, the following steps are implemented:

[0105] Selecting multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, determining a first feature extraction network set corresponding to the target training sample according to the modality of the data in the target training sample, the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and obtaining initial features of each modality in the target training sample through each of the first feature extraction networks;

[0106] Inputting the initial features of each modality in the target training sample into a weight extraction network, obtaining the weights of each modality in the target training sample output by the weight extraction network, and fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality to obtain a fused feature corresponding to the target training sample;

[0107] Input the fusion feature of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network;

[0108] Obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and updating the parameters of the first hash network according to the training loss of the target training batch;

[0109] The step of selecting multiple target training samples from multiple training samples to form a target training batch is re-executed until the parameters of the first hash network converge, and the hash code corresponding to the sample to be retrieved is obtained by using the first hash network after the parameters converge.

[0110] The weight extraction network includes a feature extraction layer and a weight output layer, and the initial features of each modality in the target training sample are input into the weight extraction network, and the weights of each modality in the target training sample output by the weight extraction network are obtained, including:

[0111] Inputting the initial features of each modality in the target training sample into the feature extraction layer to obtain potential consistent features corresponding to each modality in the target training sample;

[0112] Inputting the potential consistent features corresponding to each modality in the target training sample into the weight output layer to obtain the weight of each modality in the target training sample;

[0113] The fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality includes:

[0114] The potential consistencies of each modality in the target training sample are fused according to the weight corresponding to each modality to obtain a fusion feature corresponding to the target training sample.

[0115] Wherein, obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each target training sample includes:

[0116] Obtaining a first loss of the target training batch according to a difference between the sample hash code and the fusion feature of each of the target training samples;

[0117] Obtain a second loss of the target training batch according to a difference between the sample hash code and the semantic hash code of each of the target training samples;

[0118] Obtaining a third loss of the target training batch according to the sample hash code of each of the target training samples and the semantic similarity between each of the target training samples;

[0119] The training loss of the target training batch is obtained according to the first loss, the second loss, and the third loss.

[0120] Before obtaining the third loss of the target training batch according to the semantic hash code of each of the target training samples and the semantic similarity between each of the target training samples, the method further includes:

[0121] The semantic similarity between the target training samples is obtained according to the semantic label corresponding to each target training sample.

[0122] Wherein, updating the parameters of the first hash network according to the training loss of the target training batch includes:

[0123] Update parameters of the first hash network, the second hash network, and the weight extraction network according to the loss of the target training sample.

[0124] The step of selecting a plurality of target training samples from the plurality of training samples to form a target training batch includes:

[0125] Obtaining the fusion feature loss of each of the training samples according to the fusion features respectively corresponding to the multiple training samples under the current network parameters and the initial features respectively corresponding to the multiple training samples;

[0126] The plurality of training samples are sorted according to the fusion feature losses respectively corresponding to the plurality of training samples, and the plurality of target training samples are selected according to the sorting result.

[0127] The step of obtaining a hash code corresponding to a sample to be retrieved by using the first hash network after parameter convergence includes:

[0128] The fused features corresponding to the samples to be retrieved are obtained, the fused features of the samples to be retrieved are input into the first hash network after parameter convergence, and the hash codes corresponding to the samples to be retrieved output by the first hash network are obtained.

[0129] Embodiment 4

[0130] The present invention also provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the deep adaptive multimodal hash retrieval method as described above.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A deep adaptive multimodal hash retrieval method, characterized in that: The method comprises: Selecting multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, determining a first feature extraction network set corresponding to the target training sample according to the modality of the data in the target training sample, the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and obtaining initial features of each modality in the target training sample through each of the first feature extraction networks, wherein each of the training samples is multimodal data, and the multimodal data includes video, audio, text, and image; Inputting the initial features of each modality in the target training sample into a weight extraction network, obtaining the weights of each modality in the target training sample output by the weight extraction network, and fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality to obtain a fused feature corresponding to the target training sample; Input the fusion feature of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network; Obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and updating the parameters of the first hash network according to the training loss of the target training batch; Re-execute the step of selecting multiple target training samples from multiple training samples to form a target training batch until the parameters of the first hash network converge, use the first hash network after parameter convergence to obtain the hash code corresponding to the sample to be retrieved, and use the hash code corresponding to the sample to be retrieved to retrieve the sample to be retrieved to obtain a retrieval result.

2. The deep adaptive multimodal hash retrieval method according to claim 1, characterized in that: The weight extraction network includes a feature extraction layer and a weight output layer, and the initial features of each modality in the target training sample are input into the weight extraction network, and the weights of each modality in the target training sample output by the weight extraction network are obtained, including: Inputting the initial features of each modality in the target training sample into the feature extraction layer to obtain potential consistent features corresponding to each modality in the target training sample; Inputting the potential consistent features corresponding to each modality in the target training sample into the weight output layer to obtain the weight of each modality in the target training sample; The fusing the initial features of each modality in the target training sample according to the weight corresponding to each modality includes: The potential consistencies of each modality in the target training sample are fused according to the weight corresponding to each modality to obtain a fusion feature corresponding to the target training sample.

3. The deep adaptive multimodal hash retrieval method according to claim 1, characterized in that: The obtaining the training loss of the target training batch according to the sample hash code and the semantic hash code of each target training sample includes: Obtaining a first loss of the target training batch according to a difference between the sample hash code and the fusion feature of each of the target training samples; Obtain a second loss of the target training batch according to a difference between the sample hash code and the semantic hash code of each of the target training samples; Obtaining a third loss of the target training batch according to the sample hash code of each of the target training samples and the semantic similarity between each of the target training samples; The training loss of the target training batch is obtained according to the first loss, the second loss, and the third loss.

4. The deep adaptive multimodal hash retrieval method according to claim 1, characterized in that: Before obtaining the third loss of the target training batch according to the semantic hash code of each of the target training samples and the semantic similarity between each of the target training samples, the method further includes: The semantic similarity between the target training samples is obtained according to the semantic label corresponding to each target training sample.

5. The deep adaptive multimodal hash retrieval method according to claim 1, characterized in that: The updating of the parameters of the first hash network according to the training loss of the target training batch includes: Update parameters of the first hash network, the second hash network, and the weight extraction network according to the loss of the target training sample.

6. The deep adaptive multimodal hash retrieval method according to claim 1, characterized in that: The selecting a plurality of target training samples from the plurality of training samples to form a target training batch comprises: Obtaining the fusion feature loss of each of the training samples according to the fusion features respectively corresponding to the multiple training samples under the current network parameters and the initial features respectively corresponding to the multiple training samples; The plurality of training samples are sorted according to the fusion feature losses respectively corresponding to the plurality of training samples, and the plurality of target training samples are selected according to the sorting result.

7. The deep adaptive multimodal hash retrieval method according to claim 1, characterized in that: The acquiring of a hash code corresponding to a sample to be retrieved by using the first hash network after parameter convergence includes: The fused features corresponding to the samples to be retrieved are obtained, the fused features of the samples to be retrieved are input into the first hash network after parameter convergence, and the hash codes corresponding to the samples to be retrieved output by the first hash network are obtained.

8. A deep adaptive multimodal hash retrieval device, characterized in that: include: An initial feature extraction module, wherein the feature extraction module is used to select multiple target training samples from multiple training samples to form a target training batch, each of the training samples includes data of at least one modality, and a first feature extraction network set corresponding to the target training sample is determined according to the modality of the data in the target training sample, wherein the first feature extraction network set includes first feature extraction networks corresponding to each modality in the target training sample, and initial features of each modality in the target training sample are obtained through each of the first feature extraction networks, wherein each of the training samples is multimodal data, and the multimodal data includes video, audio, text, and image; A fusion feature extraction module, which is used to obtain the weight of each modality in the target training sample according to the initial features of each modality in the target training sample, and fuse the initial features of each modality in the target training sample according to the weight corresponding to each modality to obtain the fusion feature corresponding to the target training sample; A hash module, wherein the hash module is used to input the fusion feature of the target training sample into a first hash network, obtain the sample hash code output by the first hash network, input the semantic label corresponding to the target training sample into a second hash network, and obtain the semantic hash code output by the second hash network; A parameter updating module, the parameter updating module is used to obtain the training loss of the target training batch according to the sample hash code and the semantic hash code of each of the target training samples, and update the parameters of the first hash network according to the training loss of the target training batch; An iteration module, the iteration module is used to re-execute the step of selecting multiple target training samples from multiple training samples to form a target training batch until the parameters of the first hash network converge; A retrieval module is used to use the first hash network after parameter convergence to obtain the hash code corresponding to the sample to be retrieved, and use the hash code corresponding to the sample to be retrieved to retrieve the sample to be retrieved to obtain a retrieval result.

9. A terminal, characterized in that: The terminal includes: a processor, a computer-readable storage medium communicatively connected to the processor, the computer-readable storage medium being suitable for storing a plurality of instructions, and the processor being suitable for calling the instructions in the computer-readable storage medium to execute the steps of implementing the deep adaptive multimodal hash retrieval method described in any one of claims 1 to 7 above.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the deep adaptive multimodal hash retrieval method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Hash learning method for short text integrated with implicit semantic features

    CN104657350A

  • Big data cross-modal retrieval method and system based on deep integration Hash

    CN107871014A