An image segment hashing sorting method based on deep learning

By designing a deep learning-based image segmented hash sorting method in image retrieval technology, the problem of insufficient similarity in the multivariate similarity sorting problem is solved, and the effect of expanding the similarity expression ability and reducing calculation overhead without increasing the hash code length is achieved.

CN114155403BActive Publication Date: 2025-06-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111217840.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2025-06-13
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

When existing image retrieval technology deals with the problem of multiple similarity sorting, traditional Hamming distance method is difficult to describe sufficient similarity, resulting in increased computational overhead.

Method used

A deep learning-based image segmented hash sorting method is designed. By establishing a deep learning network model G, the hash code is calculated using a new segmented hash metric method, and the model is trained and tested using a dense triple loss function to expand the expressive similarity degree of the hash code.

Benefits of technology

Without increasing the length of the hash code, the expressive similarity degree is expanded, so that the hash code has richer semantic information and reduces the calculation loss in the sorting problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155403B_ABST
    Figure CN114155403B_ABST
Patent Text Reader

Abstract

The present invention provides an image segment hashing sorting method based on deep learning. This method extracts hash codes that can be used to solve the sorting problem of multivariate similarity through the G network, and through a carefully designed multi-segment hash distance metric and a dense triplet loss function, expands the number of expressible similarities without increasing the length of the hash code, so that the hash code has richer semantic information and also maintains less computational loss in the sorting problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technology and computer vision, and more specifically, to an image segment hashing and sorting method based on deep learning. Background Art

[0002] In recent years, with the rapid development of the Internet, the network has become the main way for people to entertain and obtain information. In this process, a large amount of image data has been accumulated on the Internet. Currently, quite mature text retrieval technology can help people obtain information, but there are still deficiencies in using images for retrieval. Image retrieval technology can help people find other images related to a certain image. For example, when shopping online and wanting to find clothes similar to a certain piece of clothing but it is difficult to describe it in words, the online shopping platform usually provides an image search by image interface. The user only needs to provide the image of the product they want, and the system will automatically sort in the database to find similar products for the user to choose. Therefore, image retrieval technology has shown great attraction both in the academic community and the industrial community.

[0003] Currently, common image retrieval technologies are all based on category information for retrieval, that is, only the images similar to the current retrieved image are considered to be images of the same category as this image, while ignoring the semantic distance between some labels. For example, both cats and dogs are animals, so the image of a dog should be ranked before the image of a car relative to a car. This patent is based on a deep learning method to extract image features, train under the supervision information based on semantic distance, and then perform sorting. The application of deep learning models in the field of pictures is relatively mature. Currently, common image feature extraction networks include VGG, ResNet, etc.

[0004] For some of the above problems, the ResNet network was adopted after research. There are many depths of this model, such as the common 18 layers, 34 layers, 50 layers, 101 layers, 152 layers, etc. Generally speaking, the deeper the depth, the more detailed features of the image can be extracted. However, the deeper the depth, the higher the computational cost and the higher the requirements for hardware. After considering various factors, ResNet with 50 layers was adopted for image feature extraction. After testing, it was found that ResNet with 50 layers can already achieve a good effect. By converting the real-valued continuous features of the image into binary hash codes, the retrieval cost can be greatly reduced. However, since the distance metric of traditional binary hash codes uses the Hamming distance, the Hamming distance is sufficient to solve the retrieval problem based on classification labels in the past because only the distinction between like or unlike is needed, and only two similarities are required. However, when the problem is extended to a sorting problem that requires more similarities, the Hamming distance becomes somewhat insufficient. For example, to solve the sorting of 100 categories, 100 similarities are needed, and at least 99-bit hash codes are required to calculate 100 Hamming distances. However, longer hash codes bring greater computational costs. To solve the problem of insufficient similarities, a segmented hash distance metric function was designed to increase the number of expressible similarities without increasing the length of the hash code. Summary of the Invention

[0005] The present invention aims to overcome at least one of the above-mentioned defects (deficiencies) of the prior art and provides a + attributive + name.

[0006] To achieve the above technical effects, the technical solution of the present invention is as follows:

[0007] An image segmented hash sorting method based on deep learning, comprising the following steps:

[0008] S1: Establish a deep learning network model G for image hash code extraction;

[0009] S2: Calculate the distance of the hash code using a new segmented hash metric method;

[0010] S3: Train and test the model using a new sorting loss function;

[0011] S4: Establish a process for providing a background interface, provide a sorting entry, and return the sorting result.

[0012] Further, the specific process of step S1 is as follows:

[0013] S11: Establish the first module of the G network, represent each preprocessed image as a low-dimensional real-valued vector, and pre-train the model ResNet-50 on a large-scale labeled photo. After passing through the ResNet-50 model, a set of real-valued feature vectors with a set length is extracted ;

[0014] S12: Establish the second module of the G network. Use a fully connected layer to map a real - valued feature vector with a set length into an n - bit hash code, where n is an even number. The essence of the hash code is a string of n - bit binary code. The output result of the fully connected layer is still a real number. Take the sign of each real number as the result of the final hash code at that bit, that is, 1 represents a positive number and 0 represents a negative number. Since the sign - taking operation is a non - differentiable operation, during the training process, the tanh function is used as the activation function on the fully connected layer as an approximation of the hash code.

[0015] Furthermore, the specific process of step S2 is as follows:

[0016] S21: Divide the n - bit hash code output by the G network into two segments, each with a length of n / 2;

[0017] S22: Calculate the distance by the newly designed segmented hash metric. For any two n - bit hash codes, calculate the Hamming distance for the upper n / 2 - bit and lower n / 2 - bit hash codes respectively. The Hamming distance calculated for the upper - bit hash code is denoted as d1, and the Hamming distance calculated for the lower - bit hash code is denoted as d2. The final distance:

[0018] .

[0019] Furthermore, the specific process of step S3 is as follows:

[0020] S31: Divide the data set into training data and test data;

[0021] S32: The overall model needs to be trained. The training steps of the G network are as follows: Extract the image hash code from the G network, calculate the distance by the newly designed segmented hash metric, then calculate the ranking loss function and minimize the loss to train the G network model and optimize the parameters of the G network;

[0022] S33: The test steps of the model are: Divide the test data set into a query set and a retrieval set, and use the images in the query set to sort the images in the retrieval set; First, input the data in the retrieval set into the G network, then the G network generates a hash code, and store the hash code result in the database DB; Then input each image in the query set into the G network, calculate the distance between the obtained hash code and the data in DB, and then perform nDCG calculation. The specific calculation method is: For the hash code of the q - th query, calculate the distance from all the data in DB, then sort the distances from small to large, and substitute the similarity between each sample in the order and the query into the nDCG formula for calculation. The calculation formula of nDCG is: , where represents the similarity between the sample ranked at the i-th position and the query, represents the calculation result of the similarity sorted in the ideal order from large to small, that is, a normalization factor. Finally, the nDCG of each query is averaged as the final result.

[0023] Further, the specific process of step S4 is as follows:

[0024] S41: Save the trained ResNet-50 model;

[0025] S42: Create a background service process and reserve an interface for image input;

[0026] S43: Through accessing the interface created in S42, input the image. After that, the background service process in S42 will first preprocess the video into the input format required by the ResNet-50 model in S41; then call the ResNet-50 model saved in S41, input the processed image into the model, and obtain an n-bit hash code; then call the image hash codes stored in the database to calculate the distance by segmented hash metric, and sort them from small to large and return the top k images, that is, the top k most similar images are the retrieval results.

[0027] Further, in step S12, the feature extraction process is as follows:

[0028] First, pre-train the ResNet-50 model with the ImageNet image dataset, and then fine-tune it on the Cifar-100 dataset; after each image passes through the pre-trained ResNet-50 model, a set of continuous feature vectors with a size of 2048 dimensions will be generated, and then converted into an n-dimensional continuous feature vector through a fully connected layer, and then converted into a custom n-length code through a hash layer.

[0029] Further, in step S32, during the training of the G network, the dense triplet loss is used as the loss function, where the distance metric function in the triplet loss is the new segmented hash distance metric, and the definition of the dense triplet loss function and the traditional triplet definition are:

[0030]

[0031] where is the anchor point, is the positive sample, It is a negative sample. Positive and negative samples are no longer just samples of the same and different categories as the anchor point. Instead, as long as a positive sample is semantically closer to the anchor point than a negative sample, a triplet can be formed. The semantic distance between categories is represented by calculating the word embedding distance of the labeled words.

[0032] Furthermore, in step S32, during the training process of the G network, the dense triplet loss is used as the loss function. Among them, the distance metric function in the triplet loss is the new piecewise hash distance metric. The definition of the dense triplet loss function and the traditional triplet is as follows:

[0033] Among them is the anchor point, is the positive sample, is the negative sample. As long as the positive sample is semantically closer to the anchor point than the negative sample, a triplet can be formed. This semantic information is characterized by the currently advanced word embedding technology, and the semantic distance between categories is represented by calculating the word embedding distance of the labeled words. The designed is the piecewise hash distance metric, which solves the problem that the traditional Hamming distance method can describe too few similarities, and solves the problem of insufficient similarity without increasing the number of bits and parameters.

[0034] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0035] The present invention can extract hash codes that can be used to solve the sorting problem of multiple similarities through the G network. And through the carefully designed multi-segment hash distance metric and dense triplet loss function, the number of describable similarities is expanded without increasing the length of the hash code, so that the hash code has richer semantic information and also maintains less computational loss in the sorting problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is the complete graph of the algorithm model of the present invention;

[0037] Figure 2 is the schematic diagram of the piecewise hash distance of the present invention. DETAILED DESCRIPTION

[0038] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0039] To better illustrate this embodiment, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0040] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0041] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0042] As Figure 1-2 shown, an image segmentation hashing sorting method based on deep learning includes the following steps:

[0043] S1: Establish a deep learning network model G for image hash code extraction;

[0044] S2: Calculate the distance of the hash code using a new segmented hashing metric;

[0045] S3: Train and test the model using a new sorting loss function;

[0046] S4: Establish a process for providing a background interface, provide a sorting entry, and return the sorting result.

[0047] Furthermore, the specific process of step S1 is as follows:

[0048] S11: Establish the first module of the G network, represent each preprocessed image as a low-dimensional real vector, and use the pre-trained model ResNet-50 on a large-scale labeled photo. After passing through the ResNet-50 model, a set of real-valued feature vectors with a set length is extracted ;

[0049] S12: Establish the second module of the G network, use a fully connected layer to map the real-valued feature vectors with a set length into an n-bit hash code, where n is an even number. The essence of the hash code is a string of n-bit binary codes. The output result of the fully connected layer is still a real number. Take the sign of each real number as the result of the final hash code at that bit, that is, 1 represents a positive number, and 0 represents a negative number. Since the sign-taking operation is a non-differentiable operation, the tanh function is used as the activation function on the fully connected layer during the training process as an approximation of the hash code.

[0050] The specific process of step S2 is as follows:

[0051] S21: Divide the n-bit hash code output by the G network into two segments, the length of each segment is n / 2;

[0052] S22: Calculate the distance using the newly designed segmented hashing metric. For any two n-bit hash codes, calculate the Hamming distance for the high n / 2-bit and low n / 2-bit hash codes of the two respectively. The Hamming distance calculated for the high-bit hash code is denoted as d1, and the Hamming distance calculated for the low-bit hash code is denoted as d2. The final distance:

[0053] .

[0054] The specific process of step S3 is as follows:

[0055] S31: Divide the data set into training data and test data;

[0056] S32: The overall model needs to be trained. The training steps of the G network are as follows: Extract the image hash code by the G network, calculate the distance by the newly designed segmented hash metric, and then calculate the ranking loss function and minimize the loss to train the G network model and optimize the parameters of the G network;

[0057] S33: The test steps of the model are as follows: Divide the test data set into a query set and a retrieval set, and use the images in the query set to rank the images in the retrieval set; First, input the data in the retrieval set into the G network, and then the G network generates a hash code, and store the hash code result in the database DB; Then, input each image in the query set into the G network, calculate the distance between the obtained hash code and the data in DB, and then perform nDCG calculation. The specific calculation method is: For the hash code of the q-th query, calculate the distance from all the data in DB, and then sort the distances from small to large. Substitute the similarity between each sample in the order and the query into the nDCG formula for calculation. The calculation formula of nDCG is: , where represents the similarity between the sample ranked at the i-th position and the query, represents the calculation result of the similarity sorted in the ideal order from large to small, that is, a normalization factor. Finally, average the nDCG of each query as the final result.

[0058] The specific process of step S4 is as follows:

[0059] S41: Save the trained ResNet-50 model;

[0060] S42: Create a background service process and reserve an interface for image input;

[0061] S43: By accessing the interface created in S42, input the image. Then, the background service process of S42 will first preprocess the video into the input format required by the ResNet-50 model in S41; Next, retrieve the ResNet-50 model saved in S41, input the processed image into the model, and obtain an n-bit hash code; Then, retrieve the image hash codes stored in the database for segmented hash metric calculation of the distance, and sort them from small to large and return the top k images, that is, the top k most similar images are the retrieval results.

[0062] In step S12, the feature extraction process is as follows:

[0063] First, pre-train the ResNet-50 model on the ImageNet image dataset, and then fine-tune it on the Cifar-100 dataset; after each image passes through the pre-trained ResNet-50 model, a set of continuous feature vectors with a size of 2048 dimensions will be generated, and then converted into n-dimensional continuous feature vectors through a fully connected layer, and then converted into a custom n-length code through a hash layer.

[0064] In step S32, during the training process of the G network, the dense triplet loss is used as the loss function, where the distance metric function in the triplet loss is the new piecewise hash distance metric. The definition of the dense triplet loss function and the traditional triplet definition are as follows:

[0065]

[0066] where is the anchor point, is the positive sample, is the negative sample. The positive and negative samples are no longer just samples of the same and different categories as the anchor point, but as long as the positive sample is semantically closer to the anchor point than the negative sample, a triplet can be formed. The semantic distance between categories is represented by calculating the word embedding distance of the label words.

[0067] In step S32, during the training process of the G network, the dense triplet loss is used as the loss function. Among them, the distance metric function in the triplet loss is the new piecewise hash distance metric. The definition of the dense triplet loss function and the traditional triplet definition are as follows:

[0068] where is the anchor point, is the positive sample, is the negative sample. As long as the positive sample is semantically closer to the anchor point than the negative sample, a triplet can be formed. This semantic information is characterized by the currently more advanced word embedding technology, and the semantic distance between categories is represented by calculating the word embedding distance of the label words. The designed is the piecewise hash distance metric, which solves the problem that the traditional Hamming distance method can describe too few similarities, and solves the problem of insufficient similarity without increasing the number of bits and parameters.

[0069] The same or similar reference numerals correspond to the same or similar components;

[0070] The terms used to describe the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0071] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for segment - hashing - based sorting of images using deep learning, characterized in that, it includes the following steps: S1: Establish a deep - learning network model G for image hash - code extraction; the specific process of step S1 is: S11: Establish the first module of the G network. Represent each preprocessed image as a low-dimensional real vector. Use the pre-trained model ResNet-50 on a large-scale labeled photo dataset. After passing through the ResNet-50 model, a set of real-valued feature vectors with a set length is extracted. ; S12: Establish the second module of the G network, and use a fully connected layer to map a real-valued feature vector of a set length into an n-bit hash code, where n is an even number. The essence of the hash code is a string of n-bit binary codes. The output result of the fully connected layer is still a real number. Take the sign of each real number as the result of the final hash code at that bit, that is, 1 represents a positive number and 0 represents a negative number. Since the sign-taking operation is a non-differentiable operation, the tanh function is used as the activation function on the fully connected layer during the training process as an approximation of the hash code; In step S12, the feature extraction process is as follows: First, pre - train the ResNet - 50 model using the ImageNet image dataset, and then fine - tune it on the Cifar - 100 dataset; after each image passes through the pre - trained ResNet - 50 model, a set of continuous feature vectors of size 2048 - dimensional will be generated, then converted into an n - dimensional continuous feature vector through a fully - connected layer, and then converted into a custom - defined n - length code through a hash layer; S2: Calculate the distance of the hash - code using a new segment - hashing metric; the specific process of step S2 is: S21: Divide the n - bit hash - code output by the G network into two segments, the length of each segment is n / 2; S22: Calculate the distance using the newly designed segment - hashing metric. For any two n - bit hash - codes, calculate the Hamming distance for their upper n / 2 - bit and lower n / 2 - bit hash - codes respectively. The Hamming distance calculated for the upper - bit hash - code is denoted as d1, and the Hamming distance calculated for the lower - bit hash - code is denoted as d2. The final distance: ; S3: Train and test the model using a new sorting loss function; the specific process of step S3 is: S31: Divide the dataset into training data and test data; S32: The overall model needs to be trained. The training steps of the G network are as follows: Extract the image hash - code by the G network, calculate the distance using the newly designed segment - hashing metric, and then calculate the sorting loss function and minimize the loss to train the G network model and optimize the parameters of the G network; In step S32, during the training process of the G network, use the dense triplet loss as the loss function, where the distance metric function in the triplet loss is the new segment - hashing distance metric. The definition of the dense triplet loss function and the traditional triplet definition are: Among them is the anchor point, is the positive sample, is the negative sample. The positive and negative samples are no longer just samples of categories that are the same as and different from the anchor point. Instead, as long as the positive sample is semantically closer to the anchor point than the negative sample, a triplet can be formed, and the semantic distance between categories is represented by calculating the word embedding distance of the labeled words; Designed For the segmented hash distance metric, the adopted segmented hash distance metric. Since there are 33 types of high-bit distances and 33 types of low-bit distances after segmentation, and the weights of the high-bit and low-bit distances are different, 1089 distances can be generated, that is, there are 1089 similarities, thus solving the problem of insufficient similarities without increasing the number of bits and the number of parameters. S33: The testing steps of the model are as follows: The test dataset is divided into a query set and a retrieval set, and the images in the retrieval set are sorted using the images in the query set; First, the data in the retrieval set is input into the G network, and then the G network generates hash codes, and the hash code results are stored in the database DB; Then, each image in the query set is input into the G network, the calculated hash codes are used to calculate the distances from the data in DB, and then nDCG calculations are performed. The specific calculation method is: For the hash code of the q-th query, calculate the distances from all the data in DB, then sort the distances from small to large, substitute the similarities between each sample in the order and the query into the nDCG formula for calculation. The calculation formula of nDCG is: , where represents the similarity between the sample ranked at the i-th position and the query, represents the calculation result of the similarities sorted in the ideal order from large to small, which is a normalization factor. Finally, the nDCG of each query is averaged as the final result; S4: Establish a process for providing a background interface, provide a sorting entry and return the sorting result.

2. The method for segment - hashing - based sorting of images using deep learning according to claim 1, characterized in that, the specific process of step S4 is: S41: Save the trained ResNet - 50 model; S42: Create a background service process and reserve an interface for image input; S43: Through accessing the interface in S42, input the image. After that, the background service process in S42 will first pre - process the video into the input format required by the ResNet - 50 model in S41; then retrieve the ResNet - 50 model saved in S41, input the processed image into the ResNet - 50 model, and obtain an n - bit hash - code; then retrieve the image hash - codes stored in the database for segment - hashing metric calculation of the distance, and sort them from small to large and return the top k images, that is, the top k most similar images are the retrieval results.

Citation Information

Patent Citations

  • Image retrieval method utilizing deep semantic to rank hash codes

    CN104834748A

  • Traffic image retrieval method based on depth learning

    CN106407352A