A fabric image retrieval method and device based on deep learning
By constructing a deep learning fabric image feature extraction network model based on DenseNet, using the structure of two hash layers, the existing fabric image retrieval methods are solved, and the fast and accurate retrieval of silk fabric images is achieved.
Patent Information
- Application Number
- CN202210219750.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Existing fabric image retrieval methods are slower or have low accuracy, especially when facing complex scenarios, low-level features-based methods are not effective.
Using a fabric image retrieval method based on deep learning, a feature extraction network model including a DenseNet backbone network, an average pooling layer, a first hash layer and a second hash layer is built, and the tag information is fully utilized through the two hash layers to reduce the hash code length to improve the retrieval speed.
The fast and accurate retrieval of silk fabric images in real scenes is achieved, and the problems of slow speed and low accuracy in traditional methods are solved.
Smart Images

Figure CN114579788B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image retrieval technology, and in particular, relates to a fabric image retrieval method and device based on deep learning. Background Art
[0002] China is a major textile production country. Silk fabrics have been exported to foreign countries since ancient times. With continuous development, the textile industry has made great progress in production techniques, but there are still problems to be solved in the retrieval of silk fabrics. When factories obtain samples from consumers for imitation, they need to manually analyze the samples and search for the same or similar existing fabrics in the warehouse, and then obtain guidance for production. This method is very time-consuming, laborious, and prone to errors, with low efficiency and low precision.
[0003] Currently, fabric retrieval is implemented using text-based image retrieval (TBIR), which requires the use of manually annotated text keywords. This method is very time-consuming, tedious and subjective. There are also content-based image retrieval (CBIR) methods. CBIR methods are mainly based on texture, color and shape, which overcome the shortcomings of TBIR to a certain extent. CBIR methods usually involve two key components: (1) designing a feature extraction algorithm for representing images; (2) selecting a suitable similarity calculation method. Traditional feature extraction methods usually use hand-crafted image descriptions, such as using SIFT and GIST to extract features and retrieve lace and embroidery fabrics. This method has achieved certain success, but when faced with complex scenes, these low-level feature-based methods often do not work well due to insufficient extracted features. Summary of the invention
[0004] The purpose of this application is to provide a fabric image retrieval method and device based on deep learning, so as to improve the retrieval speed and accuracy of existing fabric image retrieval, such as slow speed or low accuracy.
[0005] In order to achieve the above purpose, the technical solution of this application is as follows:
[0006] A fabric image retrieval method based on deep learning, comprising:
[0007] Constructing a fabric image feature extraction network model, wherein the fabric image feature extraction network model includes a DenseNet backbone network layer with a classifier layer removed, an average pooling layer, a first hash layer, and a second hash layer;
[0008] preparing training samples, training the fabric image feature extraction network model, and obtaining a trained fabric image feature extraction network model;
[0009] When retrieving fabric images, the retrieved fabric images are input into the trained fabric image feature extraction network model to obtain the image features corresponding to the retrieved fabric images, and the similarity matching of the Hamming distance with the image features of the fabric images in the retrieval database is performed to obtain the most similar fabric image in the retrieval database and display it as the retrieval result.
[0010] Furthermore, the loss function of the fabric image feature extraction network model is:
[0011] L f =L+αQ+βL 0
[0012] Among them, L f represents the loss function of the fabric image feature extraction network model, α and β are the weight parameters of the loss, and L 0 is the loss of the first hash layer, L and Q are the losses of the second hash layer;
[0013] in:
[0014]
[0015]
[0016]
[0017] Where N is the batch size of training samples, c i represents the output feature of the first hashing layer, m i Belong to [M K ,-M K ],M K is the Hadamard matrix, m i Indicates c i Corresponding to [M K ,-M K ] is the label information in , K is the dimension of the output feature; w ij is the weight coefficient, s ij Represents the image pair x i 、x j similarity, i and j belong to N, γ is the Hamming distance parameter, h i 、h j For the image pair x i 、x j The hash code output by the second hash layer, d(h i ,h j ) means h i and h j The Hamming distance between i |,1) represents h i The Hamming distance between the absolute value of and 1.
[0018] Furthermore, the weight coefficient w ij , satisfying the following formula:
[0019]
[0020] Among them, S is the similarity set, S 1 ={s ij ∈S:s ij =1}, S 0 ={s ij ∈S:s ij =0}.
[0021] Furthermore, the d(h i ,h j ) is calculated as follows:
[0022]
[0023] Furthermore, the image pair x i 、x j An image pair consisting of two training samples is randomly selected from a batch of training samples.
[0024] The present application also proposes a deep learning-based fabric image retrieval device, comprising a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, implement the steps of the deep learning-based fabric image retrieval method.
[0025] The present application proposes a deep learning-based fabric image retrieval method and device, which fully utilizes the label information by constructing two hash layers, and makes the classification center not limited by the length of the hash code. At the same time, by using a shorter hash code, the problem of slow speed in the fabric retrieval process is solved, thereby realizing time-saving, labor-saving and accurate effective retrieval of silk fabric images in real scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a flow chart of the fabric image retrieval method based on deep learning in this application;
[0027] Figure 2 This is a structural representation diagram of the fabric image feature extraction network model according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0029] In one embodiment, Figure 1 As shown in the figure, a fabric image retrieval method based on deep learning is proposed, including:
[0030] Step S1, constructing a fabric image feature extraction network model, wherein the fabric image feature extraction network model includes a DenseNet backbone network layer with a classifier layer removed, an average pooling layer, a first hash layer, and a second hash layer.
[0031] This embodiment adopts an improved DenseNet network, uses the network structure of DenseNet121, removes the classifier layer on the original basis, and forms a backbone network layer for extracting features.
[0032] like Figure 2 As shown, the fabric image feature extraction network model of the present application adds an average pooling layer (AdaptiveAvgPool2d), a first hash layer (Hash1) and a second hash layer (Hash2) after the backbone network layer (Features), and finally uses the hash code output by the Hash2 layer for retrieval. By constructing two hash layers, the label information is fully utilized, and the classification center is not limited by the length of the hash code. At the same time, by using a shorter hash code, the problem of slow speed in the fabric retrieval process is solved.
[0033] Step S2: prepare training samples, train the fabric image feature extraction network model, and obtain a trained fabric image feature extraction network model.
[0034] In this embodiment, a camera is used to shoot fabric images in a real scene, and then the fabric images are annotated to generate training samples.
[0035] Whether it is a training sample, a fabric image that needs to be retrieved, or a fabric image in a retrieval database, a normalization operation is required to set the image size to 256*256 so that the value is distributed between 0 and 1 to facilitate training and retrieval.
[0036] The training of a general network model is to input a batch of training samples into the network to obtain the predicted results, then compare them with the annotations (true values) to calculate the loss, and then perform back propagation to update the network parameters to complete the training. The training of network models is a relatively mature technology in this field and will not be described here.
[0037] An excellent network model depends on its network structure and the designed loss function. In the previous step, the network results of the fabric image feature extraction network model of this embodiment have been described. This step will explain the loss function corresponding to the network model.
[0038] In a specific embodiment, the loss function of the fabric image feature extraction network model of the present application is expressed as:
[0039] L f =L+αQ+βL 0
[0040] Among them, L f L represents the loss function of the fabric image feature extraction network model of this application, and α and β are the weight parameters of the loss. 0 is the loss of the first hash layer, L and Q are the losses of the second hash layer.
[0041] This embodiment proposes a new loss function, which is specifically reflected in the Hash1 layer and the Hash2 layer. As an improvement, the loss function of the Hash1 layer uses the Hadamard matrix, so that the output features are closer to a certain center and the distance between different categories is expanded. The form of Hadamard is shown in the following formula, where A and K represent the order, the value of A is half of K, and the value of A is a power of 2, which is also the characteristic of the Hadamard matrix:
[0042]
[0043]
[0044]
[0045] Cross entropy loss is a commonly used loss function for probabilities in machine learning, and it has shown certain performance in various models, as shown below:
[0046]
[0047] In formula 4, p i,1 , p i,2 They represent the label information and output features of the training samples respectively, and are used to represent the classification loss. The changes made are:
[0048] The softmax layer is removed, and the N samples obtained from the Hash1 layer output feature vector c = [c 1 ,…,c N ], at this time each c i The dimension is K. In this case, K must be a power of 2, such as (512, 1024, 2048), etc., where each c i The label of the corresponding sample is y i , represented by onehot encoding.
[0049] Then construct a Hadamard matrix with dimension K, and get a [MK ,-M K ]=[m 1 ,…,m 2K ], and for each different y i Select an m i The specific method is that the index number of the position where the number 1 is located in the onehot encoding is i in the matrix, so that the m corresponding to the same category i Same, different categories corresponding to m i Different, so we get m with K as the number of categories i , m i As label information, c i As the predicted feature is replaced, the loss function L 0 As shown in the following formula:
[0050]
[0051] Where N is the batch size of training samples, c i represents the output feature of the first hashing layer, m i Belong to [M K ,-M K ],M K is the Hadamard matrix, m i Indicates c i Corresponding to [M K ,-M K ] is the label information in , and K is the dimension of the feature.
[0052] The loss function of the Hash2 layer mainly uses the similarity between data to generate a suitable hash code to represent the similarity loss. First, given a set of training image pairs, use {(x i ,x j ,s ij ):s ij ∈S} to represent, x i ,x j ,s ij Respectively represent the image pair x i ,x j , and their similarities ij , S is the set of all similarities, and the image pairs are obtained in the following way: During the model training period, a batch of images is read and two of them are randomly selected to form an image pair x i ,x j , i and j belong to N.
[0053] Then the hash code H generated for N samples in the image is H = [h 1 ,h 2 ,…,h NThe maximum a posteriori estimate of ] can be expressed as logP(H|S). Using the Bayesian learning framework, we can get that logP(H|S) is proportional to logP(S|H)P(H), where the likelihood function P(S|H) can be expressed as Formula 6, and from this we can get Formula 7:
[0054]
[0055]
[0056] where w ij is the weight coefficient used to balance the data between similar pairs and dissimilar pairs, expressed by Formula 8:
[0057]
[0058] S 1 ={s ij ∈S:s ij =1} represents the set of similar pairs, S 0 ={s ij ∈S:s ij =0} represents the set of dissimilar pairs, P(S ij |h i ,h j ) represents the similarity label s ij Given a pair of hash codes (h i ,h j ), is expressed by Formula 9:
[0059]
[0060] σ is a well-defined probability function. The more common method is to use the sigmoid function. In this application, the modified Cauchy distribution probability density function is used to define σ. The original Cauchy distribution is represented by formula 10, where x 0 is the position parameter that defines the peak position of the distribution, δ is the scale parameter of the half width at half the maximum value, and the modified representation is expressed by formula 11, where γ is the Hamming distance parameter, which is used to control different Hamming distance radii, and formula 12 is obtained:
[0061]
[0062]
[0063]
[0064] The value of the Hamming distance d(h i ,h j ) is expressed as:
[0065]
[0066] This application expects to maximize the likelihood estimation, that is, to minimize the negative log-likelihood function. Then, in the hash2 layer, by substituting the above formula into formula (7), the loss function that is finally expected to be minimized is expressed as L+Q, where:
[0067]
[0068]
[0069] Then the final loss function L 1 As shown in Formula 16, α and β are parameters used to weigh these losses:
[0070] L f =L+αQ+βL 0 (16).
[0071] w ij is the weight coefficient, s ij Represents the image pair x i 、x j similarity, i and j belong to N, γ is the Hamming distance parameter, h i 、h j For the image pair x i 、x j The hash code output by the second hash layer, d(h i ,h j ) means h i and h j The Hamming distance between i |,1) represents h i The Hamming distance between the absolute value of and 1.
[0072] In actual operation, it is necessary to minimize the above loss function. Using the pytorch deep learning framework, as long as the above network model and loss function are defined, it can automatically perform derivatives, forward propagation and back propagation, and train the network model.
[0073] Step S3, when retrieving fabric images, the retrieved fabric images are input into the trained fabric image feature extraction network model to obtain the image features corresponding to the retrieved fabric images, and the similarity matching of the Hamming distance with the image features of the fabric images in the retrieval database is performed to obtain the most similar fabric image in the retrieval database and display it as the retrieval result.
[0074] After the training is completed, the trained fabric image feature extraction network model can be used for image retrieval. That is, the fabric image to be retrieved is input into the trained fabric image feature extraction network model to obtain the image features corresponding to the fabric image to be retrieved, which is the output of the second hash layer. At the same time, the fabric images in the database are retrieved and the image features are obtained in the same way. By performing similarity matching of the Hamming distance, the Hamming distance between the fabric image to be retrieved and the fabric image image features in the retrieval database can be obtained. The closer the Hamming distance is to 0, the more similar it is. Then sorting is performed to find the fabric image in the retrieval database with the smallest Hamming distance to the fabric image to be retrieved as the retrieval result, and output it for display.
[0075] In another embodiment, the present application also provides a deep learning-based fabric image retrieval device, comprising a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, implement the steps of the deep learning-based fabric image retrieval method.
[0076] The specific definition of the fabric image retrieval device based on deep learning can be found in the definition of the fabric image retrieval method based on deep learning above, which will not be repeated here. The above-mentioned fabric image retrieval device based on deep learning can be implemented in whole or in part by software, hardware and a combination thereof. It can be embedded in or independent of the processor in the computer device in the form of hardware, or it can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the above corresponding operations.
[0077] The memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The memory stores a computer program that can be run on the processor, and the processor implements the network topology layout method in the embodiment of the present invention by running the computer program stored in the memory.
[0078] The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0079] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0080] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A fabric image retrieval method based on deep learning, characterized in that: The fabric image retrieval method based on deep learning comprises: Constructing a fabric image feature extraction network model, wherein the fabric image feature extraction network model includes a DenseNet backbone network layer with a classifier layer removed, an average pooling layer, a first hash layer, and a second hash layer; preparing training samples, training the fabric image feature extraction network model, and obtaining a trained fabric image feature extraction network model; When retrieving fabric images, the retrieved fabric images are input into the trained fabric image feature extraction network model to obtain the image features corresponding to the retrieved fabric images, and the similarity matching of the Hamming distance with the image features of the fabric images in the retrieval database is performed to obtain the most similar fabric image in the retrieval database and display it as the retrieval result; Among them, the loss function of the fabric image feature extraction network model is: L f =L+αQ+βL0 Among them, L f represents the loss function of the fabric image feature extraction network model, α and β are the weight parameters of the loss, L0 is the loss of the first hash layer, and L and Q are the losses of the second hash layer; in: Where N is the batch size of training samples, c i represents the output feature of the first hashing layer, m i Belong to [M K ,-M K ],M K is the Hadamard matrix, m i Indicates c i Corresponding to [M K ,-M K ] is the label information in , K is the dimension of the output feature; w ij is the weight coefficient, s ij Represents the image pair x i 、x j similarity, i and j belong to N, γ is the Hamming distance parameter, h i 、h j For the image pair x i 、x j The hash code output by the second hash layer, d(h i ,h j ) means h i and h j The Hamming distance between i |,1) represents h i The absolute value of and the Hamming distance of 1, S is the set of all similarities.
2. The fabric image retrieval method based on deep learning according to claim 1, characterized in that: The weight coefficient w ij , satisfying the following formula: Among them, S is the similarity set, S1={s ij ∈S:s ij =1}, S0={s ij ∈S:s ij =0}.
3. The fabric image retrieval method based on deep learning according to claim 1, characterized in that: The d(h i ,h j ) is calculated as follows:
4. The fabric image retrieval method based on deep learning according to claim 1, characterized in that: The image pair x i 、x j An image pair consisting of two training samples is randomly selected from a batch of training samples.
5. A fabric image retrieval device based on deep learning, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
A method of generating a hash code for image retrieval using a deep convolutional network
CN109800314A
Large-scale image retrieval method based on hierarchical deep Hashing
CN111177432A