An accurate retrieval system that focuses on the main image under complex layout

Through the iterative network of rotation hierarchical iteration and sequencing adaptive compression algorithm, combined with the body core dimensionality reduction and significance extraction, the problems of image retrieval accuracy and speed under complex layouts are solved, efficient and accurate image retrieval is achieved, and user experience is improved.

CN114817597BActive Publication Date: 2025-08-19WUHAN QISHI MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210429814.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-08-19
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

The existing image retrieval system has low search accuracy under complex layouts, cannot accurately understand the user's search topic, and it is difficult to efficiently complete high-precision retrieval when computing resources are limited, and the user experience is poor.

Method used

Image features are extracted using a rotary hierarchical iterative network, combined with sequencing adaptive compression and subject core dimensionality reduction algorithm, through significant extraction and subject background separation, the design hierarchical module runs independently to achieve efficient and accurate retrieval.

Benefits of technology

Improve search accuracy and speed under complex layouts, reduce computing time, enhance user experience, adapt to the big data environment, and achieve fast and accurate image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114817597B_ABST
    Figure CN114817597B_ABST
Patent Text Reader

Abstract

In response to the needs of e-commerce, this application extracts image features through a convolution hierarchical iterative network, and uses sequenced adaptive compression for efficient, precise and accurate retrieval, to realize an e-commerce image retrieval application software based on a convolution hierarchical iterative network and sequenced adaptive compression for large-scale image libraries with complex layouts. By designing a series of algorithms for saliency extraction and subject kernel dimensionality reduction, the software can adapt to conditions with complex backgrounds, which not only improves the retrieval accuracy, but also adapts to more usage environments. It can accurately find the results that users need from massive image libraries within the time range tolerated by users. The image hashing method achieves fast retrieval by mapping high-dimensional data into binary codes, and efficiently expresses image features through a convolution hierarchical iterative network. The two are integrated and absorb their respective advantages to adapt to massive retrieval in the big data era, and can well complete the retrieval task of large-scale complex images. It has great practical significance and a wide range of applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a precise retrieval system focusing on an image subject, and in particular to a precise retrieval system focusing on an image subject under a complex layout, belonging to the technical field of precise retrieval of image content. Background Art

[0002] With the rapid adoption of smart devices and the widespread use of mobile networks, the e-commerce industry is booming, and requirements for shopping apps' functionality and performance are increasing. For example, these apps must be able to accurately search for product information and ensure smooth operation. Often, users may not be able to accurately describe a product's keywords. For example, when browsing the web, a user may see an exquisite cup, a small piece of art recommended by a netizen, or clothing from a street photo. They may simply want the same product, not something similar. These items often cannot be accurately retrieved with simple verbal descriptions. In such cases, if shopping apps feature image search, users only need to upload an image of the product, and the system will return accurate results. This not only provides a good user experience but also significantly saves user time.

[0003] Previously, commercial image retrieval systems were mostly text-based, with their underlying image databases containing manually labeled images. When a user enters a keyword, the system compares it with their image tags and outputs images with the same or similar tags. As devices and environments evolve, the demand for image-based search has become increasingly demanding. However, the slow development of image-based search systems has been hampered by the weak representational power of extracted image features, resulting in search results that fail to meet commercial requirements.

[0004] Subsequently, image search technology has made significant progress, but even now, the retrieval performance of these systems still has many shortcomings. Three main factors limit retrieval accuracy: First, the amount of image data itself is very large, and users have very high requirements for system retrieval accuracy and time. Given limited computer resources, balancing limited resources with high accuracy and speed places high demands on both algorithm and system design. Second, image similarity: everyone has a different definition of similarity. For example, when searching for the same item of clothing, one person may prefer to find the same style but different colors, while another may prioritize the same color. This subjective nature requires understanding the specific dimension of similarity that the user is seeking. Third, with the advancement of image acquisition equipment, images are becoming larger and more complex. With these complex layouts, there are more and more interfering factors, making the user's search topic increasingly unclear. This makes it difficult for the retrieval system to accurately understand the search topic, resulting in a serious decline in search quality and even completely incorrect results.

[0005] In recent years, content-based image retrieval methods have developed rapidly under the research boom of artificial intelligence. First of all, how to extract more representative image features has always been one of the hot topics in computer vision research and development. The image "understood" by the computer is the low-level information expressed at the pixel level, which is very different from the multi-dimensional high-level information content of the image understood by humans. Therefore, it is necessary to design an algorithm that can extract a feature that expresses the image level information as richly as possible, thereby unifying the computer pixel-level data and the human high-dimensional abstract concept. The existing technology algorithms are not detailed enough for the expression of images. No matter which optimization algorithms are used subsequently, their fitting results are in an underfitting state compared with the actual results and the accuracy is not high enough.

[0006] Humans are able to quickly extract desired information while filtering out other information in complex scenes. Simulating this process with computers is a method for extracting objects of interest, known as saliency image detection. Algorithms based on contrast for saliency extraction have become a major research direction in recent years. Contrast refers to the degree of difference in grayscale, color, texture, and other features between different regions of an image. Contrast-based methods can be divided into two categories: local contrast-based methods utilize the relative sparseness of image regions relative to local regions for saliency segmentation. The resulting image has distinct contours at the edges, where contrast is most pronounced. However, in some images, these methods only reveal local features and fail to uniformly detect saliency across the entire image. Global contrast-based methods directly compare the saliency of local regions with that of the entire image. Compared to local saliency methods, these methods achieve more complete detection results without missing objects. However, in complex image layouts, these algorithms can easily misclassify background features as salient objects, resulting in suboptimal segmentation.

[0007] In summary, the existing image content retrieval technology still has several problems and defects. The difficulties and problems to be solved in this application are mainly concentrated in the following aspects:

[0008] First, the existing text-based image retrieval system stores images with manually labeled images in the image database behind the system. The labeling workload is large and the requirements are high. Only the labeled related content can be searched, which is very limited and can no longer meet user needs. The current image retrieval technology implemented by the image search method also has obvious defects. The slow development of the image search system is hindered by the fact that the extracted image features are not strong in representation, and the retrieval results cannot meet commercial requirements. Moreover, with the advent of the big data era, the ever-increasing background database has put forward very high requirements on the performance of the image search algorithm. The current retrieval software takes a long time to accurately find the results required by users from the massive image library, which is difficult for users to tolerate. The existing technology cannot coordinate the contradiction between limited resources and high precision and high speed under limited computer resources, and cannot well complete the retrieval task of large-scale complex images.

[0009] Second, when it comes to the similarity of images retrieved for shopping software products, everyone has a different definition of similarity. For example, when searching for the same piece of clothing, some people may hope to find the same style but in different colors, while others may feel that the same color is more important. In other words, there is obvious subjectivity, and it is necessary to know which dimension of similarity the user wants. Existing technologies cannot extract more representative image features. There is a large gap between the images "understood" by computers and the multi-dimensional, high-level information content of images understood by humans. Existing technologies lack algorithms that can extract features that express image hierarchical information as richly as possible, thereby unifying computer pixel-level data with human high-dimensional abstract concepts. Existing algorithms are not detailed enough for image expression, especially when used in the field of e-commerce search. They lack an understanding of shoppers' search intent and topics, and the retrieval accuracy is not high enough.

[0010] Third, with the development of image acquisition equipment, images are becoming larger and more complex. With these complex layouts, there are more and more interference factors, making the topics and products that users are interested in searching less and less clear. Retrieval systems are unable to accurately understand the search subject. Results are unsatisfactory when faced with image deformation, lighting, or when only partial objects are captured. Retrieval precision is poor. The lack of a method that combines saliency extraction with a hierarchical iterative network has severely degraded search quality in complex environments, and can even produce completely erroneous results. Furthermore, due to inaccurate understanding of the image subject and user search intent, the search scope and volume are unnecessarily expanded, resulting in not only significant errors in search results but also a significant slowdown in search speed, seriously affecting the user experience.

[0011] Fourth, the current e-commerce platform's image search system is not very targeted. There is a lack of e-commerce image retrieval application software based on convolution hierarchical iterative network and sequence adaptive compression for large-scale image libraries with complex layouts. There is also a lack of a series of algorithms designed for saliency extraction and subject kernel dimensionality reduction, which cannot adapt to conditions with complex backgrounds. In terms of feature extraction, there is a lack of a convolution hierarchical iterative network structure, which makes it impossible to extract the output data of the last full-link layer as input for sequence adaptive compression, and it is impossible to use the subject kernel dimensionality reduction algorithm to reduce the dimensionality of the data, which requires a lot of calculations. In terms of image retrieval, it is impossible to reduce the time required for retrieval through sequence adaptive compression. There is a lack of a subject kernel dimensionality reduction hashing method that performs subject kernel dimensionality reduction on the feature matrix of the data, and there is a lack of an optimal gradient rotation method to construct a hash function to reduce the error caused by quantization. There is a lack of a saliency extraction algorithm to separate the subject and background, which cannot make up for the low retrieval accuracy of the algorithm under conditions of complex backgrounds. In particular, when the image is complex, the database is large, and the image subject is not obvious, it is difficult to promote it to practical applications. Summary of the Invention

[0012] The software designed in this application has high precision and fast response time, especially when facing massive backgrounds with large and complex layouts, it has great advantages; first, it better alleviates the contradiction between the high precision and high temporal complexity of the hierarchical iterative algorithm. For high-dimensional data, sequenced adaptive compression is a similar nearest neighbor algorithm that can effectively reduce the temporal and spatial complexity of the algorithm; second, in response to the situation where the retrieval effect of objects in complex layouts is poor, a subject extraction module is added to improve the retrieval effect. The current retrieval software mostly guides users to place the subject to be photographed into the image frame pre-defined by the software to prevent background interference. The disadvantage of this method is that the user needs to constantly adjust the distance between the camera and the subject to be photographed to achieve the optimal retrieval effect. This application improves user experience through the technology of automatic subject-background separation; third, in terms of software design, the algorithms of each module are independent of each other, and the algorithm update and maintenance of a single module will not affect other modules. More effective hierarchical iterative networks and sequenced adaptive compression, subject extraction algorithms, etc. can be selected, which has excellent maintainability.

[0013] To achieve the above technical effects, the technical solutions adopted in this application are as follows:

[0014] The precise retrieval system for focusing on the main body of an image under complex layouts includes: first, a feature theme extraction module, specifically including: inverse diffusion and disordered hierarchical descent, super-factor layering setting, key point acquisition weight initialization, hierarchical fusion naturalization, local network activation, and feature theme extraction; second, a key point retrieval module, specifically including: main body kernel dimensionality reduction hashing, main body kernel dimensionality reduction and projection matrix calculation, input data generation, hash function generation and retrieval; third, a main body extraction module; fourth, a retrieval preprocessing module; and fifth, a precise matching module.

[0015] This application adopts a hierarchical iterative network forward evaluation and sequenced adaptive compression calculation method. First, the high-level data expression of the hierarchical iterative network is extracted. Then, this data is used to perform similar nearest neighbor retrieval using sequenced adaptive compression to obtain a rough retrieval result. Finally, the rough retrieval result is further processed for image precise matching. If the background is complex and the retrieval result is poor, a subject extraction module is added to improve the retrieval accuracy. Through the hierarchical design system, each module of the system is made independent of each other.

[0016] By extracting image features through a convolutional hierarchical iterative network and using sequence-adaptive compression for efficient, precise, and accurate retrieval, we developed an e-commerce image retrieval application software for large-scale, complex image libraries. Furthermore, by designing a series of algorithms for saliency extraction and subject kernel dimensionality reduction, the software is adaptable to complex backgrounds.

[0017] In terms of feature extraction, the network structure is iterated based on the convolution hierarchy, and the output data of the last full-link layer is extracted as the input of the sequenced adaptive compression;

[0018] In image retrieval, we use sequenced adaptive compression to reduce retrieval time. We use a core-kernel dimensionality reduction hashing algorithm to reduce the dimensionality of the data's feature matrix and construct a hash function using an optimal gradient rotation method to reduce quantization-induced errors. During retrieval, we use a feature similarity discrimination algorithm to compare the Hamming distance between hash codes to obtain similar images, and then perform further precise matching to obtain the final result.

[0019] In addition, a saliency extraction algorithm is used to separate the subject and background to improve the retrieval accuracy.

[0020] Preferably, inverse diffusion and disordered hierarchical descent: inverse diffusion calculates the factor hierarchy by the chain rule, that is, given a function f(x), calculates the hierarchy of the function with respect to x, that is, f(x), where x is a multidimensional vector representing all weight factors in the hierarchical iterative network, as well as the input data. Inverse diffusion uses a composite function to find partial derivatives, function f(x, y, z) = (x + y) * z, let the intermediate variable q = x + y, set the initial input value of the function to x = l, y = 2, z = 4, first perform forward calculation, and get q = 3, f = q * z, which is 12; when performing inverse diffusion, first transfer to f = q * z, and Further as well as Get the partial derivative of f with respect to its independent variable, and then find the level;

[0021] For a level point, the function f is the polynomial obtained during training, x, y are input data, z is the weight w, and when it gets the input, it first calculates the output value f and the local level of the output value with respect to the input, that is, and and Then, in the inverse diffusion, the level of the final output pair f of the entire network is obtained. The returned level is multiplied by the obtained local level to obtain the level of each input. Finally, the level value after the inverse diffusion is obtained. Then, the disordered level descent method is used to continuously approach the extreme value, and finally the network converges to the set accuracy.

[0022] Super-factor layered setting: Super-factor is the framework factor of the hierarchical iterative network. In order to obtain the features of multiple attributes of the image, the number of essential collectors is increased, that is, the depth super-factor, and multiple different essential collectors are designed. Each essential collector has its own set of weights. Each set of weights is an iterative layer, which outputs a certain image with specific features after passing the original image through an essential collector. There are n iterative layers to generate n images of interest in different features. These images are regarded as different channels of an image. When doing a positive evaluation, each movement of the iterative layer is only sensitive to and calculates local data, and traverses the width and height. The window size of the essential collector determines the amount of data obtained by the window at one time. The window is an n×n square. In order to obtain the size of the output data body of a certain size, the depth super-factor is set.

[0023] Preferably, the key is to initialize the weights: initialize the weight vector w of the level point by the following formula:

[0024] U(w)=2 / (n in +n out )

[0025] where n in and n outis the number of input and output data, n is the total amount of data, and the inner product between the weight vector w and the input x is assumed to be Check the variance of s:

[0026]

[0027] Assuming that the mean of input and weight is 0, then E[x i ]=E[w i ]=0, the remaining third term U(w i )U(x i ), assuming that all w i , x i All obey the same distribution. The weight w is initialized based on the code segment w = np.random.randn(n) / sqrt(n) to keep the variance of the input data and output data consistent, avoiding the situation where the high-level data tends to 0 and cannot continue training.

[0028] To solve the problem of disappearing levels, the synonym optimization formula is:

[0029]

[0030] n l For the amount of data per layer, layer disappearance is well controlled.

[0031] Preferably, hierarchical fusion naturalization: before training the hierarchical iterative network, the data is naturalized and the hierarchical factors are constrained:

[0032]

[0033] Among them, x i is the i-th input data value, μ represents the mean of all inputs in a level, σ 2 Indicates variance, ε is a value close to 0 to prevent the denominator from being 0. After calculation, we get However, after applying Gaussian constraints, the data’s ability to express the input of the previous layer is weakened. Adding a linear transformation to restore the original data information is as follows:

[0034]

[0035] The two factors β and γ are learned through a hierarchical iterative network to restore the information expression of the data.

[0036] Optimally, local network activation: When training a hierarchical iterative network, only a local network is used in each calculation, and the rest is in an inactive state. By training many different small networks, the total number of factors remains unchanged, and other units randomly selected under a certain probability criterion form a new network. In the next training, the hierarchical point and other hierarchical points form another type of network, improving the generalization ability of the network, eliminating the weakening of the joint adaptability between the hierarchical point nodes, and enhancing the generalization ability;

[0037] Each time a forward calculation is performed, the hierarchical points will be closed with a certain probability, that is, disordered inactivation. In the forward evaluation, if the local network is not used, the formula is:

[0038]

[0039] That is a typical linear calculation, l represents the lth layer of the hierarchical iterative network, i represents a specific data point, and b is the corresponding parameter. If a local network is added, the formula becomes:

[0040]

[0041]

[0042]

[0043] In the disordered deactivation, the level points are activated or set to 0 with a probability of a preset super factor p, and r is the Bernoulli disorder number, ranging from 0 to 1;

[0044] The local network super factor is set to 0.5, and the local network is used to generate a variety of different network expressions to obtain better generalization ability when processing the same data.

[0045] Preferably, feature theme extraction: based on the high representativeness of hierarchical iteration for image information, the trained network is used to extract the deep expression of each image in the hierarchical iterative network as its feature. Specifically, in the last full-link layer, this layer outputs 4096 values. This value is used as a vector to obtain a 1×4096-dimensional data, which can represent the original image. This process is similar to the forward calculation during training. The difference is that it does not perform hierarchical backpropagation, but takes out its data as image features in the full-link layer; the designed hierarchical iterative network structure calls the above function, and the output of the previous function is used as the input of the next function. Finally, the corresponding 1×4096-dimensional data can be obtained in the full-link layer. This data serves as the final feature description of the image.

[0046] Preferably, the gist retrieval module:

[0047] (1) Main core dimensionality reduction hashing

[0048] First, we use the subject kernel dimensionality reduction hashing to generate hash codes for all the extracted features in the image library. Then, we use the subject kernel dimensionality reduction hashing to generate similar hash codes for the images to be retrieved, and use the projection matrix generated when processing the image library. Finally, we perform similarity matching based on the hash codes.

[0049] (2) Dimensionality reduction of the main kernel and calculation of its projection matrix

[0050] First, the covariance matrix of the sample's feature matrix is calculated, and the eigenvalues and corresponding eigenvectors of the covariance matrix are obtained. The larger the eigenvalue, the higher the discrimination of the corresponding dimension. The one with the largest eigenvalue is the principal component. After sorting the remaining eigenvalues, the corresponding feature matrix obtained is the projection matrix. The projection matrix can be used to analyze new input data with the same initial dimension.

[0051] Specifically define X={x1,x2,…x N} is a sample matrix composed of N sample features, and the dimension of each sample is m. First, the mean of the entire data is shown as follows:

[0052]

[0053] The purpose of finding the mean is to centralize the data and then solve the following formula:

[0054]

[0055] The elements on the diagonal of the matrix C are the autovariance of each sample, and the other positions are the covariance between samples. The covariance and autovariance are represented by a matrix respectively. In order to make the samples uncorrelated, their covariance needs to be 0, that is, for the matrix C, let the positions outside its diagonal be 0, and diagonalize C: let the diagonalized matrix be D, and the following formula is obtained:

[0056]

[0057] The matrix D is the covariance matrix of Y. If D is a diagonal matrix, the one with the largest value is the main kernel, and P is the projection matrix to be solved. P is obtained by similarly diagonalizing C. The main kernel dimensionality reduction is represented by the following algorithm:

[0058] Input: Sample data features that need to be reduced in dimension d is a parameter;

[0059] Output: data after dimensionality reduction

[0060] Step 1: Calculate the mean of the sample and subtract the mean from all samples;

[0061] Step 2: Solve the covariance matrix C of the sample;

[0062] Step 3: Calculate the eigenvalues of the covariance matrix and obtain the eigenvectors from the eigenvalues;

[0063] Step 4: Arrange the eigenvalues in order, select the eigenvectors corresponding to the corresponding eigenvalues, and compose them into a projection matrix;

[0064] Step 5: Multiply the new input data by the projection matrix to obtain the reduced dimensionality data;

[0065] (3) Generating input data

[0066] First, process the dimensionality-reduced data, assuming that the input features have been zero-mean, and let H l ={h l,1 ,…,h l,K} to represent the first hash table, where h l,K It represents the Kth hash function in the hash table. For the input data of the lth hash table, it is defined as:

[0067] X l =X-XV l V l T

[0068] Among them, X is all the data after the main kernel dimension reduction, X l is the input data required by the first hash table, V is the projection matrix, V l = {vl,…,x} is the set of the 1st to 1+K-1th eigenvectors on V, v is the eigenvector corresponding to a certain eigenvalue, and the input data of l hash tables are obtained;

[0069] (4) Generation and retrieval of hash functions

[0070] An optimal gradient rotation method is used to construct the hash function. Multi-table hashing requires the construction of multiple sets of hash functions. The construction process of each table is similar. To construct a single hash table, first define the hash map:

[0071] h k (X) = sgn(Xw k +b k )

[0072] b k As a parameter, the data is zero-meaned and input into the hash function h k , then:

[0073] h k (X) = sgn(Xw k )

[0074] w kis a parameter, adding this factor reduces the direct quantization h k The quantization error caused by (X) = sgn(X) can be further expressed as:

[0075]

[0076] Will w k Use W to simply express it, and h k (X) is represented by the letter Q.

[0077] The symbolic function is not differentiable, and the quadratic equation cannot be solved using conventional optimization methods. Therefore, in the optimization process, the F-norm standard measurement is added to solve the above problem. Given two variables, in order to solve the minimum value, one variable is fixed, and the objective function is minimized under this condition. By first fixing W and then fixing Q, the objective function is continuously optimized, and finally a local minimum is obtained.

[0078] (1) When W is fixed, the original formula can be written as:

[0079]

[0080] Where Q is an n×t matrix, a is the dimension of the feature matrix after projection, which is:

[0081]

[0082] The value range of Q is {-1,1}. P and W are both fixed values. To get the maximum value, Q must be positive when P is positive and negative when P is negative, so Q = P.

[0083] (2) When Q is fixed, the original formula can be written as:

[0084]

[0085] Because Q T X is a×a matrix, for Q T X performs singular value decomposition, S, R and U are all orthogonal matrices, then S T WU is also orthogonal, and its maximum value is 1. When S T The WU value is 1, and the corresponding matrix W is obtained:

[0086]

[0087] The following is a method for iteratively optimizing the extreme value of the main kernel dimensionality reduction hash by fixing W and then fixing Q:

[0088] Input: Initialize the unordered rotation matrix The feature matrix X after the main kernel dimensionality reduction;

[0089] Loop: About W: fix W and set Q = sgn(XW);

[0090] About Q: Fixed Q, Q T X performs singular value decomposition, Q T X=UΛS T , w=SU T ;

[0091] Until: The condition for stopping the loop is met. The condition for stopping the loop is the number of loops or a certain quantization error value;

[0092] Output: orthogonal rotation matrix W;

[0093] Obtain the rotation matrix W, that is, the hash function, by solving the formula sgn(Xw k ), get the quantized hash binary code, further hash the samples in the entire image library to get the hash code corresponding to each image, and at the same time, put the images with equal hash codes into the same hash bucket to facilitate retrieval. Finally, calculate the hash code of the image to be retrieved and compare it with the hash bucket in the library. By comparing the Hamming distance between the hash code of the image to be retrieved and the hash bucket, find the hash bucket with the closest distance, and take out the image in the bucket, which is the similar image under hash retrieval.

[0094] Preferably, the subject extraction module: designs a fast detection method for salient subjects to distinguish the target from its background while occupying only a small amount of system resources;

[0095] First, perform image annotation under the disordered rule model and obtain the corresponding energy function:

[0096]

[0097] Among them, i,j represents pixel position, ψ i and ψ ij are the contrast between the target itself and its local area, and the preliminary calculated saliency map is obtained. For ψ ij The mixed Gaussian model is used to replace the histogram model for calculation, and a constant term is added to the covariance matrix to avoid non-convergence. The color model uses the GMM model. When the pixel labels are used as disordered variables to form the Markov field and global observations can be obtained, these labels are modeled; ψ is eliminated. ij (x i ,x j ) in the complex GMM global model prediction, and finally obtain a more accurate saliency map, and get the final result. The second picture in the first row is the result after the initial extraction, the first picture in the second row is the saliency map, and the fourth picture in the second row is the final saliency extraction result;

[0098] This application uses a dense disordered regular model. If a sparse disordered regular model with an associated background and subject color model is highly correlated with the dense disordered regular model, a more efficient saliency detection can be performed using the dense disordered regular model.

[0099] Preferably, the retrieval preprocessing module: when performing a forward evaluation of the hierarchical iterative network, the input image to be read must first be preprocessed, including demeaning compression and adaptive adjustment of the image data matrix; demeaning compression uses the trained hierarchical iterative network to remember the mean characteristics of the original training set, so that when retrieving an image, even if the input image does not come from the original training library, it is necessary to subtract the mean of the image training library;

[0100] The data matrix of the image is adaptively adjusted. The initial input image size is 224×224×3. The resolution adaptive adjustment method is: Assume that the value of the unknown function f at the point (x, y) needs to be obtained, and the condition G is known. 11 =(x1,y1),G 12 =(x1,y2),G 21 =(x2,y1),G 22 =(x2,y2) The f value of the four points is calculated by linear interpolation in the x and y directions. First, interpolation is performed in the x-axis direction:

[0101]

[0102] Similarly, we can get f(x,y2) and interpolate in the y-axis direction. Finally, we can calculate f(x,y):

[0103]

[0104] Convert the image to any resolution so that it can meet the input requirements of the hierarchical iterative network algorithm.

[0105] Preferably, the precise matching module: after obtaining the results of the hash rough search, finds the image with a higher score in the hash search result, and further compares it with the original search image, that is, precise matching.

[0106] For two n-dimensional vectors, we have the following formula:

[0107]

[0108] The actual encoding is calculated using the vector mode, namely:

[0109]

[0110] All the distances obtained by comparison are sorted by size. The smaller the d value, the higher the similarity between the two images. The final image retrieval result can be obtained by inputting images with relatively small distances.

[0111] Compared with the existing technology, the innovation and advantages of this application are:

[0112] First, in response to the needs of e-commerce and internet users for image retrieval capabilities, this application extracts image features through a convolution hierarchical iterative network and performs efficient, precise, and accurate retrieval through sequenced adaptive compression. This results in an e-commerce image retrieval application software based on a convolution hierarchical iterative network and sequenced adaptive compression for large-scale, complex image libraries. Furthermore, by designing a series of algorithms for saliency extraction and subject kernel dimensionality reduction, the software is able to adapt to complex background conditions, improving retrieval accuracy and adapting to a wider range of usage environments. This addresses the problem of users often being unable to accurately describe the desired item using only text when searching for the same product through existing images. The retrieval software can accurately find the results a user needs from a massive image library within a user-tolerable timeframe. The image hashing method achieves rapid retrieval by mapping high-dimensional data into binary codes and efficiently expresses image features through a convolution hierarchical iterative network. By integrating the two and absorbing their respective advantages, the software adapts to massive retrieval in the big data era and can effectively complete the retrieval task of large-scale, complex images. This software has great practical significance and a wide range of applications.

[0113] Second, in terms of feature extraction, this application is based on a convolution hierarchical iterative network structure, and extracts the output data of the last full-link layer as the input of the sequenced adaptive compression. At the same time, in order to reduce the calculation time, the main kernel dimensionality reduction algorithm is used to reduce the dimension of the data. The experimental results show that the dimensionality reduction algorithm has no obvious effect on the retrieval accuracy. The above method extracts more representative image features, so that the image "understood" by the computer is closer to the multi-dimensional high-level information content of the image understood by humans. It can extract a feature that expresses the image hierarchical information as richly as possible, thereby unifying the computer pixel-level data with the human high-dimensional abstract concept. The algorithm has detailed expression of the image, especially in the field of e-commerce search, and has a more accurate understanding of the shopper's search intentions and topics, and greatly improves the retrieval accuracy.

[0114] Third, in terms of image retrieval, the time required for retrieval is reduced through sequenced adaptive compression. A subject kernel dimensionality reduction hashing is adopted. The subject kernel dimensionality reduction is performed on the feature matrix of the data, and a hash function is constructed through an optimal gradient rotation method to reduce the error caused by quantization. A feature similarity discrimination algorithm is used during retrieval to obtain similar images by comparing the Hamming distance between hash codes. The final result is then obtained through further precise matching. This solves the problem that the current image size is getting larger and the layout is getting more and more complex. Under the complex layout, there are more and more interference factors, and the user's search topics and products are becoming less and less obvious. The retrieval system cannot accurately understand the search subject. In the face of image deformation, lighting, and only partial shooting of objects, the retrieval results are ideal and the retrieval accuracy is high. Based on the combination of saliency extraction and hierarchical iterative network, the retrieval quality in complex environments has a strong competitive advantage. Moreover, due to the accurate grasp of the image subject and the user's search intention, the search scope and retrieval volume are accurately narrowed, resulting in small retrieval result errors, and the retrieval speed is greatly accelerated, providing a good user experience.

[0115] Fourth, this application fully utilizes the advantages of the convolution hierarchical iterative network in feature extraction, the sequenced adaptive compression in fast retrieval, and the saliency extraction algorithm in improving retrieval accuracy to meet the user's needs for physical object retrieval. System tests show that the software designed in this application has high accuracy and fast response time, especially in the face of massive backgrounds with large and complex layouts, it has great advantages; first, it better alleviates the contradiction between the high accuracy and high temporal complexity of the hierarchical iterative algorithm. For high-dimensional data, the current computer speed cannot complete the linear search of the data, so the traditional approach is to use the similar nearest neighbor algorithm to replace the original nearest neighbor algorithm, and the sequenced adaptive compression of this application is a similar nearest neighbor algorithm, which can effectively reduce the algorithm's temporal and spatial complexity; second, in response to the situation where the retrieval effect of objects is poor under complex layouts, a subject extraction module is added to improve the retrieval effect. The current retrieval software mostly guides users to place the subject to be photographed into the image frame pre-defined by the software to prevent background interference in this way. The disadvantage of this method is that the user needs to constantly adjust the distance between the camera and the subject to achieve the best retrieval effect. This application improves the user experience through the technology of automatic subject-background separation; thirdly, in software design, the algorithms of each module are independent of each other, and the algorithm update and maintenance of a single module will not affect other modules. More effective hierarchical iterative networks and sequenced adaptive compression, subject extraction algorithms, etc. can be selected, which has excellent maintainability. BRIEF DESCRIPTION OF THE DRAWINGS

[0116] Figure 1 This is a framework diagram of a precise retrieval system that focuses on the main image under complex layout.

[0117] Figure 2It is a schematic diagram of the forward calculation and inverse calculation process.

[0118] Figure 3 This is a flowchart of the image retrieval process under the subject kernel dimensionality reduction hashing.

[0119] Figure 4 is the saliency map obtained by the subject extraction module.

[0120] Figure 5 It is a schematic diagram of the process of adaptively adjusting the data matrix of an image.

[0121] Figure 6 This is a schematic diagram of the retrieval effect without subject extraction in a complex environment.

[0122] Figure 7 This is the wallet retrieval effect diagram after subject extraction in a complex environment.

[0123] Figure 8 This is the fan retrieval effect after subject extraction in a complex environment.

[0124] Figure 9 This is the first set of comparison charts between the search performance of this application and Pailitao.

[0125] Figure 10 This is the second set of comparison charts between the search performance of this application and Pailitao. Specific implementation methods

[0126] The following further describes the technical solution of the precise retrieval system for focusing on the image subject under a complex layout provided by the present application in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present application and implement it.

[0127] With the rapid adoption of smart devices like mobile phones, online shopping is becoming increasingly common, placing increasing demands on the hardware and software performance of these devices. In the field of information retrieval, traditional text-based retrieval methods are no longer able to meet user needs in some cases. For example, when users need to find a product from existing images, text alone often fails to accurately describe the desired item. Therefore, image retrieval technology, which uses image-based search, has become an urgent need. This has led to rapid development of image retrieval technology in recent years. However, with the advent of the big data era, the ever-expanding backend databases have placed extremely high demands on the performance of retrieval algorithms. Retrieval software must accurately find the desired results from massive image libraries within a timeframe that users can tolerate. Image hashing methods achieve fast retrieval by mapping high-dimensional data into binary codes and efficiently represent image features through a convolutional hierarchical iterative network. Combining these two approaches and leveraging their respective strengths, they are well-suited for retrieval tasks involving large-scale, complex images.

[0128] In response to the needs of e-commerce and Internet users for image retrieval functions, this application extracts image features through a convolution hierarchical iterative network and performs efficient, precise and accurate retrieval through sequenced adaptive compression, thereby realizing an e-commerce image retrieval application software for large-scale and complex image libraries based on a convolution hierarchical iterative network and sequenced adaptive compression. In addition, by designing a series of algorithms for saliency extraction and subject kernel dimensionality reduction, the software can adapt to conditions with complex backgrounds, which not only improves the retrieval accuracy, but also enables the software to adapt to more usage environments.

[0129] In terms of feature extraction, this application uses a convolutional hierarchical iterative network structure and extracts the output data of the final full-link layer as the input for sequenced adaptive compression. Furthermore, to reduce computational time, a principal kernel dimensionality reduction algorithm is used to reduce the dimensionality of the data. Experimental results show that the dimensionality reduction algorithm has little impact on retrieval accuracy.

[0130] In terms of image retrieval, the time required for retrieval is reduced through sequenced adaptive compression, a main kernel dimensionality reduction hash is adopted, the main kernel dimensionality reduction is performed on the feature matrix of the data, and a hash function is constructed through an optimal gradient rotation method to reduce the error caused by quantization; during retrieval, a feature similarity discrimination algorithm is used to obtain similar images by comparing the Hamming distance between hash codes, and then the final result is obtained through further precise matching.

[0131] In addition, in order to make up for the low retrieval accuracy of the algorithm when the background of the object is complex, this application adopts a saliency extraction algorithm to separate the subject and the background. After experimental verification, this method can improve the retrieval accuracy.

[0132] 1. Retrieval System Architecture

[0133] The precise retrieval system that focuses on the main body of the image under complex layout includes a feature theme extraction module, a key retrieval module and a main body extraction module. It adopts the hierarchical iterative network forward evaluation and sequence adaptive compression calculation method. It first extracts the high-level data expression of the hierarchical iterative network, and then uses the sequence adaptive compression to perform similar nearest neighbor retrieval on these data to obtain a rough retrieval result. Finally, the rough retrieval result is further processed for image precise matching. If the background is complex and the retrieval result is poor, the main body extraction module is added to improve the retrieval accuracy. Through the hierarchical design system, each module of the system is independent of each other, which is easy to develop and maintain. The whole system operation flow chart is as follows: Figure 1 shown.

[0134] Image input: Input an image to be retrieved into the system, and preprocess the image after input, including de-averaging and image bilinear interpolation to adjust the image resolution.

[0135] Subject extraction: If the background of the input image to be retrieved is complex and affects the final retrieval results, the subject extraction module is used to remove the background interference.

[0136] Feature extraction: Regardless of whether subject extraction is performed, the image to be retrieved must input the image data into an adaptive hierarchical iterative network for forward calculation to obtain image features.

[0137] Sequencing adaptive compression retrieval: Retrieve similar images in the gallery and effectively speed up the retrieval through sequencing adaptive compression.

[0138] Finally, the retrieval results of the sequenced adaptive compression are accurately matched with the image to be retrieved. The matching adopts the Euclidean superposition cosine distance. After the accurate matching is completed, several closest retrieved images are output.

[0139] 2. Feature Theme Extraction Module

[0140] (1) Inverse diffusion and disordered hierarchy descent

[0141] Inverse diffusion uses the chain rule to calculate the factor level, that is, given a function f(x), calculate the level of the function with respect to x, that is, f(x), where x is a multidimensional vector representing all weight factors in the hierarchical iterative network and the input data. Inverse diffusion uses a composite function to find the partial derivative, function f(x,y,z)=(x+y)*z, let the intermediate variable q=x+y, set the initial input value of the function to x=l, y=2, z=4, first perform forward calculation, get q=3, f=q*z, which is 12; when performing inverse diffusion, first transfer to f=q*z, and Further as well as Get the partial derivative of f with respect to its independent variable, and then find the level. Figure 2 In the calculation process shown, the forward calculation is a frameless font, and the inverse calculation is a framed font.

[0142] For a level point, the function f is the polynomial obtained during training, x, y are input data, z is the weight w, and when it gets the input, it first calculates the output value f and the local level of the output value with respect to the input, that is, and and Then, in the inverse diffusion, the level of the final output pair f of the entire network is obtained, and the returned level and the obtained local level are multiplied to obtain the level of each input. Finally, the level value after the entire inverse diffusion is obtained. Then, the disordered level descent method is used to continuously approach the extreme value, and finally the network converges to the set accuracy.

[0143] (2) Super-factor stratification setting

[0144] Super factors are the framework factors of hierarchical iterative networks. They are of a different type from the factors learned during training and generally need to be set manually.

[0145] In order to obtain the features of multiple attributes of the image, the number of essential collectors is increased, that is, the depth super factor, and multiple different essential collectors are designed. Each essential collector has its own set of weights. Each set of weights is an iterative layer, which outputs a certain image with specific features after passing the original image through an essential collector. There are n iterative layers to generate n images of interest in different features. These images are regarded as different channels of an image. When performing positive evaluation, each movement of the iterative layer is only sensitive to and calculates local data, and traverses the width and height. The window size of the essential collector determines the amount of data obtained by the window at one time. The window is an n×n square. In order to obtain the size of the output data body of a certain size, the depth super factor is set.

[0146] (3) Key acquisition weight initialization

[0147] Initialize the weight vector w of the level point using the following formula:

[0148] U(w)=2 / (n in +n out )

[0149] where n in and n out is the number of input and output data, n is the total amount of data, and the inner product between the weight vector w and the input x is assumed to be Check the variance of s:

[0150]

[0151] Assuming that the mean of input and weight is 0, then E[x i ]=E[w i ]=0, the remaining third term U(w i )U(x i ), assuming that all w i , x i All obey the same distribution. The weight w is initialized based on the code segment w = np.random.randn(n) / sqrt(n) to keep the variance of the input data and output data consistent, avoiding the situation where the high-level data tends to 0 and cannot continue training.

[0152] To solve the problem of disappearing levels, the synonym optimization formula is:

[0153]

[0154] n lFor the amount of data per layer, layer disappearance is well controlled.

[0155] (4) Hierarchical integration and naturalization

[0156] Before training the hierarchical iterative network, the data is normalized and the hierarchical factors are constrained:

[0157]

[0158] Among them, x i is the i-th input data value, μ represents the mean of all inputs in a level, σ 2 Indicates variance, ε is a value close to 0 to prevent the denominator from being 0. After calculation, we get However, after applying Gaussian constraints, the data’s ability to express the input of the previous layer is weakened. Adding a linear transformation to restore the original data information is as follows:

[0159]

[0160] The two factors β and γ are learned through a hierarchical iterative network to restore the information expression of the data.

[0161] (5) Local network activation

[0162] When training a hierarchical iterative network, only a local network is used in each calculation, and the rest is in an inactivated state. By training many different small networks, the total number of factors remains unchanged, and other units randomly selected under a certain probability criterion form a new network. During the next training, the hierarchical points and other hierarchical points form other types of networks, thereby improving the generalization ability of the network, eliminating the weakening of the joint adaptability between the hierarchical point nodes, and enhancing the generalization ability.

[0163] Each time a forward calculation is performed, the hierarchical points will be closed with a certain probability, that is, disordered inactivation. In the forward evaluation, if the local network is not used, the formula is:

[0164]

[0165] That is a typical linear calculation, l represents the lth layer of the hierarchical iterative network, i represents a specific data point, and b is the corresponding parameter. If a local network is added, the formula becomes:

[0166]

[0167]

[0168]

[0169] In random deactivation, a level point is activated or set to 0 with a probability of a preset super factor p, where r is the Bernoulli random number, ranging from 0 to 1.

[0170] The local network super factor is set to 0.5, and the local network is used to generate a variety of different network expressions to obtain better generalization ability when processing the same data.

[0171] (6) Feature theme extraction

[0172] Based on the high representation of image information by hierarchical iteration, the trained network is used to extract the deep expression of each image in the hierarchical iterative network as its feature. Specifically, in the last full-link layer, this layer outputs 4096 values. This value is used as a vector to obtain a 1×4096-dimensional data, which can represent the original image. This process is similar to the forward calculation during training. The difference is that it does not perform hierarchical backpropagation, but takes out its data as image features in the full-link layer; the above function is called with the designed hierarchical iterative network structure, and the output of the previous function is used as the input of the next function. Finally, the corresponding 1×4096-dimensional data can be obtained in the full-link layer. This data serves as the final feature description of the image.

[0173] 3. Keyword Search Module

[0174] (1) Main core dimensionality reduction hashing

[0175] Figure 3 This is a flowchart for image retrieval using subject-kernel dimensionality reduction hashing. First, subject-kernel dimensionality reduction hashing is used to generate hash codes for all extracted features in the image library. Then, subject-kernel dimensionality reduction hashing is also used for the images to be retrieved, and similar hash codes are generated using the projection matrix generated when processing the image library. Finally, similarity matching is performed based on the hash codes. The following details the subject-kernel dimensionality reduction hashing process.

[0176] (2) Dimensionality reduction of the main kernel and calculation of its projection matrix

[0177] First, the covariance matrix of the sample's feature matrix is calculated, and the eigenvalues and corresponding eigenvectors of the covariance matrix are obtained. The larger the eigenvalue, the higher the discrimination of the corresponding dimension. The one with the largest eigenvalue is the principal component. After sorting the remaining eigenvalues, the corresponding feature matrix obtained is the projection matrix. The projection matrix can be used to analyze new input data with the same initial dimension.

[0178] Specifically define X={x1,x2,…x N} is a sample matrix composed of N sample features, and the dimension of each sample is m. First, the mean of the entire data is shown as follows:

[0179]

[0180] The purpose of finding the mean is to centralize the data and then solve the following formula:

[0181]

[0182] The elements on the diagonal of the matrix C are the autovariance of each sample, and the other positions are the covariance between samples. The covariance and autovariance are represented by a matrix respectively. In order to make the samples uncorrelated, their covariance needs to be 0, that is, for the matrix C, let the positions outside its diagonal be 0, and diagonalize C: let the diagonalized matrix be D, and the following formula is obtained:

[0183]

[0184] The matrix D is the covariance matrix of Y. If D is a diagonal matrix, the one with the largest value is the main kernel, and P is the projection matrix to be solved. P is obtained by similarly diagonalizing C. The main kernel dimensionality reduction is represented by the following algorithm:

[0185] Input: Sample data features that need to be reduced in dimension d is a parameter;

[0186] Output: data after dimensionality reduction

[0187] Step 1: Calculate the mean of the sample and subtract the mean from all samples;

[0188] Step 2: Solve the covariance matrix C of the sample;

[0189] Step 3: Calculate the eigenvalues of the covariance matrix and obtain the eigenvectors from the eigenvalues;

[0190] Step 4: Arrange the eigenvalues in order, select the eigenvectors corresponding to the corresponding eigenvalues, and compose them into a projection matrix;

[0191] Step 5: Multiply the new input data by the projection matrix to obtain the reduced dimensionality data.

[0192] (3) Generating input data

[0193] To improve recall, we designed and built multiple hash tables to enhance retrieval precision. Projection revealed that while data energy is primarily concentrated in the principal component, projecting data toward that principal component with a low critical value can easily lead to missed samples that are far from the target. Projecting data toward other components can also lead to missed samples that are closer. Therefore, we combined the principal kernel and multiple edge kernels to form a multi-table hash to mitigate this.

[0194] First, process the dimensionality-reduced data, assuming that the input features have been zero-mean, and let H l ={h l,1 ,…,h l,K} to represent the first hash table, where h l,K It represents the Kth hash function in the hash table. For the input data of the lth hash table, it is defined as:

[0195] X l =X-XV l V l T

[0196] Among them, X is all the data after the main kernel dimension reduction, X l is the input data required by the first hash table, V is the projection matrix, V l ={vl,…,x} is the set of the 1st to 1+K-1th eigenvectors on V, v is the eigenvector corresponding to a certain eigenvalue, and the input data of l hash tables are obtained.

[0197] (4) Generation and retrieval of hash functions

[0198] Although the main kernel dimensionality reduction algorithm greatly reduces the time required for the algorithm, if the features after the main kernel dimensionality reduction are directly hashed and binarized, the local characteristics of the original image will often be lost due to the quantization error caused by strong binarization. This application constructs a hash function through an optimal gradient rotation method to solve the above problem.

[0199] Multi-table hashing requires building multiple sets of hash functions. The construction process for each table is similar. To build a single hash table, first define the hash map:

[0200] h k (X) = sgn(Xw k +b k )

[0201] b k As a parameter, the data is zero-meaned and input into the hash function h k , then:

[0202] h k (X) = sgn(Xw k )

[0203] w k is a parameter, adding this factor reduces the direct quantization h k The quantization error caused by (X) = sgn(X) can be further expressed as:

[0204]

[0205] Will w k Use W to simply express it, and h k (X) is represented by the letter Q.

[0206] The symbolic function is not differentiable, and the quadratic equation cannot be solved using conventional optimization methods. Therefore, during the optimization process, the F-norm standard measurement is added to solve the above problem. Given two variables, in order to solve the minimum value, a certain variable is fixed, and the objective function is minimized under this condition. By first fixing W and then fixing Q, the objective function is continuously optimized, and finally a local minimum is obtained.

[0207] (1) When W is fixed, the original formula can be written as:

[0208]

[0209] Where Q is an n×t matrix, a is the dimension of the feature matrix after projection, which is:

[0210]

[0211] The value range of Q is {-1,1}, and P and W are both fixed values. To obtain the maximum value, Q must be positive when P is positive and negative when P is negative, so Q=P.

[0212] (2) When Q is fixed, the original formula can be written as:

[0213]

[0214] Because Q T X is an a×a matrix, for Q T X performs singular value decomposition, S, R and U are all orthogonal matrices, then S T WU is also orthogonal, and its maximum value is 1. When S T The WU value is 1, and the corresponding matrix W is obtained:

[0215]

[0216] The following is a method for iteratively optimizing the extreme value of the main kernel dimensionality reduction hash by fixing W and then fixing Q:

[0217] Input: Initialize the unordered rotation matrix The feature matrix X after dimensionality reduction by the main kernel.

[0218] Loop: About W: fix W and set Q = sgn(XW);

[0219] About Q: Fixed Q, Q T X performs singular value decomposition, Q T X=UΛS T, w=SU T ;

[0220] Until: The condition for stopping the loop is met. The condition for stopping the loop is the number of loops or a certain quantization error value;

[0221] Output: Orthogonal rotation matrix W.

[0222] Obtain the rotation matrix W, that is, the hash function, by solving the formula sgn(Xw k ), get the quantized hash binary code, further hash the samples in the entire image library to get the hash code corresponding to each image, and at the same time, put the images with equal hash codes into the same hash bucket to facilitate retrieval. Finally, calculate the hash code of the image to be retrieved and compare it with the hash bucket in the library. By comparing the Hamming distance between the hash code of the image to be retrieved and the hash bucket, find the hash bucket with the closest distance, and take out the image in the bucket, which is the similar image under hash retrieval.

[0223] 4. Subject Extraction Module

[0224] After designing the hierarchical iteration module and hashing module, the system has basic retrieval capabilities. However, sometimes the subject captured by the user may only occupy a quarter of the entire image, or even less. In these cases, the subject captured by the user often has a complex background. This useless background information can significantly interfere with feature extraction and affect retrieval results. Therefore, this application designs a method for rapid salient subject detection that distinguishes the target from its background while using only a small amount of system resources.

[0225] First, perform image annotation under the disordered rule model and obtain the corresponding energy function:

[0226]

[0227] Among them, i,j represents pixel position, ψ i and ψ ij are the contrast between the target itself and its local area, and the preliminary calculated saliency map is obtained, such as Figure 4 As shown in the second figure, for ψ ij The mixed Gaussian model is used to replace the histogram model for calculation, and a constant term is added to the covariance matrix to avoid non-convergence. The color model uses the GMM model. When the pixel labels are used as disordered variables to form a Markov field and global observations can be obtained, these labels are modeled. ij (x i ,x j ) in the complex GMM global model prediction, and finally obtain a more accurate saliency map and the final result. Figure 4The second picture in the first row is the result after the initial extraction, the first picture in the second row is the saliency map, and the fourth picture in the second row is the final saliency extraction result.

[0228] This application uses a dense disordered regular model. If a sparse disordered regular model with an associated background and subject color model is highly correlated with the dense disordered regular model, a more efficient saliency detection can be performed using the dense disordered regular model.

[0229] 5. Retrieval Preprocessing Module

[0230] When conducting a forward evaluation of a hierarchical iterative network, the input image must first be preprocessed, including de-meaning compression and adaptive adjustment of the image data matrix; de-meaning compression uses the trained hierarchical iterative network to remember the mean characteristics of the original training set, so that when retrieving an image, even if the input image does not come from the original training library, it is necessary to subtract the mean of the image training library to ensure the reliability of the algorithm.

[0231] The data matrix of the image is adaptively adjusted. The initial input image size is 224×224×3. The resolution adaptive adjustment method is as follows: Figure 5 , suppose you need to get the value of the unknown function f at point (x, y), and the known condition G 11 =(x1,y1),G 12 =(x1,y2),G 21 =(x2,y1),G 22 =(x2,y2) The f value of the four points is calculated by linear interpolation in the x and y directions. First, interpolation is performed in the x-axis direction:

[0232]

[0233] Similarly, we can get f(x,y2) and interpolate in the y-axis direction. Finally, we can calculate f(x,y):

[0234]

[0235] Convert the image to any resolution so that it can meet the input requirements of the hierarchical iterative network algorithm.

[0236] 6. Precision Matching Module

[0237] After the hash rough search results are obtained, the images with higher scores in the hash search results are found and further compared with the original search images, which is called exact matching.

[0238] For two n-dimensional vectors, we have the following formula:

[0239]

[0240] The actual encoding is calculated using the vector mode, namely:

[0241]

[0242] All the distances obtained by comparison are sorted by size. The smaller the d value, the higher the similarity between the two images. The final image retrieval result can be obtained by inputting images with relatively small distances.

[0243] 7. System Testing and Result Analysis

[0244] (1) Retrieval effect in complex environments

[0245] During the experiment, when the photographed target is in a complex layout, the retrieval results will show a large deviation. The reason is that when the hierarchical iterative network extracts features, it mistakenly considers the background as the object to be retrieved itself, thereby extracting the background features, which ultimately leads to inaccurate retrieval results.

[0246] To solve the above problems, this application uses a subject extraction algorithm to distinguish the background from the subject, and performs retrieval while retaining only the subject, so as to improve the accuracy of retrieval. Figure 6 As shown in the figure, the system did not find any results similar to the target image in the complex layout. For example, in the first image, the background includes objects such as an air purifier, a keyboard, a computer monitor, and a cabinet, resulting in poor search results. The second image shows a fan, but the results contain multiple circular objects resembling tea cakes, resulting in poor search results.

[0247] exist Figure 6 Among the 9 images that the system thinks are closest, the system mistakenly thinks the wallet is a computer and outputs an incorrect result. Figure 7 In the process of extracting the subject, the system can retrieve the correct results. Figure 7 The three images above are the original image, the mask, and the salient target obtained through the mask, i.e., the wallet. During salient extraction, the algorithm may not completely segment the salient individuals. If the extraction effect is too poor, it may sometimes lead to search failure. However, through Figure 8 and Figure 6 The comparison with the second figure and repeated experiments show that the saliency extraction algorithm designed in this application can play a greater practical role.

[0248] (2) Comparison with Pailitao

[0249] Pailitao is a physical object retrieval application launched by the Taobao platform. The experiment used 1.3 million images as the image library, and compared the search results with Pailitao in terms of product similarity and search speed.

[0250] like Figure 9 , Figure 10 , as shown in the figure, compares the retrieval performance of this application and Pailitao. Considering the actual application situation, the retrieval result comparison charts under simple background, complex background and handheld shooting scenes are selected.

[0251] It can be found Figure 9 Despite some background interference, both algorithms achieved good results. In particular, when we partially obscured the searched item with our hand, the algorithm was able to effectively remove the interference. In comparison, our algorithm was more accurate, retrieving more identical items in both the first and second images. As can be seen from the first image, our algorithm selectively selected the "subject"—the red women's handbag—as the search target. This allowed us to extract more information about the red handbag, ultimately identifying the identical item in the search results.

[0252] Figure 10 This is still the case where the algorithm of this application performs better. From the first comparison image, we can see that for an ordinary bicycle lock, "Pai Li Tao" failed to search, while the algorithm of this application performed well. From the second and third comparison images, we can see that when searching for the product "U disk", under the complex layout, this application successfully retrieved the target image with a black object and a partially black background. In contrast, with "Pai Li Tao", no matter how the image was taken, it was impossible to obtain the correct search result. This also shows that the hierarchical iterative algorithm of this application has a certain level of search accuracy and strong robustness.

[0253] In addition, the experiment collected 100 images of various types and compared the algorithm in this application with Pailitao. A top-5 search validation method was used, meaning that a search was considered successful if at least one of the five search results met the search criteria. The corresponding MAP values were also calculated, also based on statistics from the 100 search result images.

[0254] The first five images from this application were compared with the first five images from Pailitao. A total of 100 search results were tallied, and the results showed that the first five images from this application met the requirements in over half of the cases (62%), while the results from Pailitao were less than half (46%).

[0255] After many experimental comparisons, it was found that the algorithm has better results when the background is simple and the object features are obvious, and when facing image deformation, lighting, and only shooting part of the object; this application leads in retrieval accuracy. When the object background is relatively complex, the retrieval effect of both sides is reduced to a certain extent. This application adopts a method that combines saliency extraction and hierarchical iterative network. Under low user participation, the accuracy of the retrieval results has been significantly improved. Compared with "Pai Li Tao", this algorithm has demonstrated higher retrieval accuracy when retrieving certain products.

[0256] (3) Algorithm efficiency test

[0257] To ensure a good user experience, algorithm processing time must be strictly controlled. Testing has shown that the mobile app "Pai Li Tao" takes approximately 2.1 to 4 seconds from taking a photo to displaying the result, which is generally within the user's tolerance range. However, the algorithm itself should ideally take less than 2 seconds to ensure a good user experience even in poor network conditions.

[0258] On this device, the algorithm takes approximately 323 milliseconds. Without the subject extraction module, the algorithm takes only 0.24 seconds, far less than the design requirement of 2 seconds. As the database size increases, the hashing module becomes more time-consuming, but its algorithmic complexity is approximately O(log(N)). This is because the hashing classification efficiency increases with the number of sample data, so the time complexity does not increase linearly.

[0259] The search results show that the software designed by this application meets user search needs. In terms of search time, for image databases with less than 10 million images, this algorithm can control the search time to less than 1.6 seconds. In terms of search accuracy, a comparison with other image retrieval algorithms on the CIFAR-10 image database shows that this algorithm has higher accuracy than traditional algorithms and is superior to the best known algorithms. In comparison with the Taobao platform "Pai Li Tao", under complex background search conditions, this application method has a clear advantage.

Claims

1. An accurate retrieval system focusing on the main image in complex layouts, characterized by: include: The first is the feature topic extraction module, which specifically includes: inverse diffusion and disordered hierarchical descent, super factor hierarchical setting, key point collection weight initialization, hierarchical fusion naturalization, local network activation, and feature topic extraction; the second is the key point retrieval module, which specifically includes: subject core dimensionality reduction hashing, subject core dimensionality reduction and its projection matrix calculation, input data generation, hash function generation and retrieval; the third is the subject extraction module; the fourth is the retrieval preprocessing module; and the fifth is the precise matching module; Adopting the hierarchical iterative network forward evaluation and sequenced adaptive compression calculation method, the high-level data expression of the hierarchical iterative network is first extracted. Then, this data is used to perform similar nearest neighbor retrieval using sequenced adaptive compression to obtain a rough retrieval result. Finally, the rough retrieval result is further processed for precise image matching. If the retrieval result is poor due to complex background, a subject extraction module is added to improve the retrieval accuracy. Through the hierarchical design system, each module of the system is made independent of each other. By extracting image features through a convolutional hierarchical iterative network and using sequence-adaptive compression for efficient, precise, and accurate retrieval, we developed an e-commerce image retrieval application software for large-scale, complex image libraries. Furthermore, by designing a series of algorithms for saliency extraction and subject kernel dimensionality reduction, the software is adaptable to complex backgrounds. In terms of feature extraction, the network structure is iterated based on the convolution hierarchy, and the output data of the last full-link layer is extracted as the input of the sequenced adaptive compression; In image retrieval, we use sequenced adaptive compression to reduce retrieval time. We use a core-kernel dimensionality reduction hashing algorithm to reduce the dimensionality of the data's feature matrix and construct a hash function using an optimal gradient rotation method to reduce quantization-induced errors. During retrieval, we use a feature similarity discrimination algorithm to compare the Hamming distance between hash codes to obtain similar images, and then perform further precise matching to obtain the final result. In addition, a saliency extraction algorithm is used to separate the subject and background to improve the retrieval accuracy; Subject kernel dimensionality reduction hashing: Generate hash codes for the extracted features in the image library, then use the projection matrix generated when processing the image library to generate similar hash codes for the images to be retrieved, and perform similarity matching based on the hash codes; Principal kernel dimensionality reduction and projection matrix calculation: First, calculate the covariance matrix of the sample's feature matrix, and then find its eigenvalues and corresponding eigenvectors. The larger the eigenvalue, the higher the discrimination of the corresponding dimension. The one with the largest eigenvalue is the principal component. After sorting the remaining eigenvalues, the corresponding feature matrix is the projection matrix. The projection matrix can be used to analyze new input data with the same initial dimension. A saliency extraction algorithm: First, perform image annotation under the disordered rule model and obtain the corresponding energy function: Among them, i,j represents pixel position, ψ i and ψ ij are the contrast between the target itself and its local area, and the preliminary calculated saliency map is obtained. For ψ ij The mixed Gaussian model is used to replace the histogram model for calculation, and a constant term is added to the covariance matrix to avoid non-convergence. The color model uses the GMM model. When the pixel labels are used as disordered variables to form the Markov field and global observations can be obtained, these labels are modeled; ψ is eliminated. ij (x i ,x j ) and finally obtain a more accurate saliency map to get the final result.

2. The precise retrieval system for focusing on the main image under complex layout according to claim 1 is characterized in that: Inverse diffusion and disordered hierarchical descent: Inverse diffusion uses the chain rule to calculate the factor hierarchy, that is, given a function f(x), calculate the hierarchy of the function with respect to x, that is, f(x), where x is a multidimensional vector representing all weight factors in the hierarchical iterative network and the input data. Inverse diffusion uses a composite function to find the partial derivative, function f(x,y,z)=(x+y)*z, let the intermediate variable q=x+y, set the initial input value of the function to x=l, y=2, z=4, first perform forward calculation, get q=3, f=q*z, which is 12; when performing inverse diffusion, first transfer to f=q*z, and Further as well as Get the partial derivative of f with respect to its independent variable, and then find the level; For a level point, the function f is the polynomial obtained during training, x, y are input data, z is the weight w, and when it gets the input, it first calculates the output value f and the local level of the output value with respect to the input, that is, and and Then, in the inverse diffusion, the level of the final output pair f of the entire network is obtained. The returned level is multiplied by the obtained local level to obtain the level of each input. Finally, the level value after the inverse diffusion is obtained. Then, the disordered level descent method is used to continuously approach the extreme value, and finally the network converges to the set accuracy. Super-factor layered setting: Super-factor is the framework factor of the hierarchical iterative network. In order to obtain the features of multiple attributes of the image, the number of essential collectors is increased, that is, the depth super-factor, and multiple different essential collectors are designed. Each essential collector has its own set of weights. Each set of weights is an iterative layer, which outputs a certain image with specific features after passing the original image through an essential collector. There are n iterative layers to generate n images of interest in different features. These images are regarded as different channels of an image. When doing a positive evaluation, each movement of the iterative layer is only sensitive to and calculates local data, and traverses the width and height. The window size of the essential collector determines the amount of data obtained by the window at one time. The window is an n×n square. In order to obtain the size of the output data body of a certain size, the depth super-factor is set.

3. The precise retrieval system for focusing on the main image under complex layout according to claim 1 is characterized in that: The key is to initialize the weight of the layer point using the following formula: U(w)=2 / (n in +n out ) where n in and n out is the number of input and output data, n is the total amount of data, and the inner product between the weight vector w and the input x is assumed to be Check the variance of s: Assuming that the mean of input and weight is 0, then E[x i ]=E[w i ]=0, the remaining third term U(w i )U(x i ), assuming that all w i , x i All obey the same distribution. The weight w is initialized based on the code segment w = np.random.randn(n) / sqrt(n) to keep the variance of the input data and output data consistent, avoiding the situation where the high-level data tends to 0 and cannot continue training. To solve the problem of disappearing levels, the synonym optimization formula is: n l For the amount of data per layer, layer disappearance is well controlled.

4. The precise retrieval system for focusing on the main image under complex layout according to claim 1 is characterized in that: Hierarchical fusion naturalization: Before training the hierarchical iterative network, the data is naturalized and the hierarchical factors are constrained: Among them, x i is the i-th input data value, μ represents the mean of all inputs in a level, σ 2 Indicates variance, ε is a value close to 0 to prevent the denominator from being 0. After calculation, we get However, after applying Gaussian constraints, the data’s ability to express the input of the previous layer is weakened. Adding a linear transformation to restore the original data information is as follows: The two factors β and γ are learned through a hierarchical iterative network to restore the information expression of the data.

5. The precise retrieval system for focusing on the main image under complex layout according to claim 1 is characterized in that: Local network activation: When training a hierarchical iterative network, only a local network is used in each calculation, and the rest is in an inactive state. By training many different small networks, the total number of factors remains unchanged. Other units are randomly selected under a certain probability criterion to form a new network. In the next training, the hierarchical points are combined with other hierarchical points to form a different type of network, improving the generalization ability of the network, eliminating the weakening of the joint adaptability between the hierarchical point nodes, and enhancing the generalization ability; Each time a forward calculation is performed, the hierarchical points will be closed with a certain probability, that is, disordered inactivation. In the forward evaluation, if the local network is not used, the formula is: That is a typical linear calculation, l represents the lth layer of the hierarchical iterative network, i represents a specific data point, and b is the corresponding parameter. If a local network is added, the formula becomes: In the disordered deactivation, the level points are activated or set to 0 with a probability of a preset super factor p, and r is the Bernoulli disorder number, ranging from 0 to 1; The local network super factor is set to 0.5, and the local network is used to generate a variety of different network expressions to obtain better generalization ability when processing the same data.

6. The precise retrieval system for focusing on the main image under complex layout according to claim 1 is characterized in that: Feature theme extraction: Based on the high representation of image information by hierarchical iteration, the trained network is used to extract the deep expression of each image as its feature. Specifically, in the last full-link layer, this layer outputs 4096 numerical values. This value is used as a vector to obtain 1×4096-dimensional data, which can represent the original image. This process is similar to the forward calculation during training. The difference is that it does not perform hierarchical backpropagation, but extracts its data as image features in the full-link layer; the designed hierarchical iterative network structure calls the above function, and the output of the previous function is used as the input of the next function. Finally, the corresponding 1×4096-dimensional data can be obtained in the full-link layer. This data serves as the final feature description of the image.

7. The precise retrieval system for focusing on the main image under complex layout according to claim 1 is characterized in that: Essential search module: (1) Main core dimensionality reduction hashing First, we use the subject kernel dimensionality reduction hashing to generate hash codes for all the extracted features in the image library. Then, we use the subject kernel dimensionality reduction hashing to generate similar hash codes for the images to be retrieved, and use the projection matrix generated when processing the image library. Finally, we perform similarity matching based on the hash codes. (2) Dimensionality reduction of the main kernel and calculation of its projection matrix Specifically define X={x1,x2,…x N } is a sample matrix composed of N sample features, and the dimension of each sample is m. First, the mean of the entire data is shown as follows: The purpose of finding the mean is to centralize the data and then solve the following formula: The elements on the diagonal of the matrix C are the autovariance of each sample, and the other positions are the covariance between samples. The covariance and autovariance are represented by a matrix respectively. In order to make the samples uncorrelated, their covariance needs to be 0, that is, for the matrix C, let the positions outside its diagonal be 0, and diagonalize C: let the diagonalized matrix be D, and the following formula is obtained: The matrix D is the covariance matrix of Y. If D is a diagonal matrix, the one with the largest value is the main kernel, and P is the projection matrix to be solved. P is obtained by similarly diagonalizing C. The main kernel dimensionality reduction is represented by the following algorithm: Input: Sample data features that need to be reduced in dimension d is a parameter; Output: data after dimensionality reduction Step 1: Calculate the mean of the sample and subtract the mean from all samples; Step 2: Solve the covariance matrix C of the sample; Step 3: Calculate the eigenvalues of the covariance matrix and obtain the eigenvectors from the eigenvalues; Step 4: Arrange the eigenvalues in order, select the eigenvectors corresponding to the corresponding eigenvalues, and compose them into a projection matrix; Step 5: Multiply the new input data by the projection matrix to obtain the reduced dimensionality data; (3) Generating input data First, process the dimensionality-reduced data, assuming that the input features have been zero-mean, and let H l ={h l,1 ,…,h l,K } to represent the first hash table, where h l,K It represents the Kth hash function in the hash table. For the input data of the lth hash table, it is defined as: X l =X-XV l V l T Among them, X is all the data after the main kernel dimension reduction, X l is the input data required by the first hash table, V is the projection matrix, V l = {vl,…,x} is the set of the 1st to 1+K-1th eigenvectors on V, v is the eigenvector corresponding to a certain eigenvalue, and the input data of l hash tables are obtained; (4) Generation and retrieval of hash functions An optimal gradient rotation method is used to construct the hash function. Multi-table hashing requires the construction of multiple sets of hash functions. The construction process of each table is similar. To construct a single hash table, first define the hash map: h k (X)=sgn(Xw k +b k ) b k As a parameter, the data is zero-meaned and input into the hash function h k , then: h k (X)=sgn(Xw k ) w k is a parameter, adding this factor reduces the direct quantization h k The quantization error caused by (X) = sgn(X) can be further expressed as: Will w k Use W to simply express it, and h k (X) is represented by the letter Q; The symbolic function is not differentiable, and the quadratic equation cannot be solved using conventional optimization methods. Therefore, in the optimization process, the F-norm standard measurement is added to solve the above problem. Given two variables, in order to solve the minimum value, one variable is fixed, and the objective function is minimized under this condition. By first fixing W and then fixing Q, the objective function is continuously optimized, and finally a local minimum is obtained. (1) When W is fixed, the original formula can be written as: Where Q is an n×t matrix, a is the dimension of the feature matrix after projection, which is: The value range of Q is {-1,1}. P and W are both fixed values. To get the maximum value, Q must be positive when P is positive and negative when P is negative, so Q = P. (2) When Q is fixed, the original formula can be written as: Because Q T X is an a×a matrix, for Q T X performs singular value decomposition, S, R and U are all orthogonal matrices, then S T WU is also orthogonal, and its maximum value is 1. When S T The WU value is 1, and the corresponding matrix W is obtained: The following is a method for iteratively optimizing the extreme value of the main kernel dimensionality reduction hash by fixing W and then fixing Q: Input: Initialize the unordered rotation matrix The feature matrix X after the main kernel dimensionality reduction; Loop: About W: fix W and set Q = sgn(XW); About Q: Fixed Q, Q T X performs singular value decomposition, Q T X=UΛS T , w=SU T ; Until: The condition for stopping the loop is met. The condition for stopping the loop is the number of loops or a certain quantization error value; Output: orthogonal rotation matrix W; Obtain the rotation matrix W, that is, the hash function, by solving the formula sgn(Xw k ), get the quantized hash binary code, further hash the samples in the entire image library to get the hash code corresponding to each image, and at the same time, put the images with equal hash codes into the same hash bucket to facilitate retrieval. Finally, calculate the hash code of the image to be retrieved and compare it with the hash bucket in the library. By comparing the Hamming distance between the hash code of the image to be retrieved and the hash bucket, find the hash bucket with the closest distance, and take out the image in the bucket, which is the similar image under hash retrieval.

8. The precise retrieval system for focusing on the main image subject under complex layout according to claim 1 is characterized in that: Subject extraction module: Design a fast detection method for salient subjects to distinguish the target from its background while only occupying a small amount of system resources; A dense disordered regular model is selected. If a sparse disordered regular model with an associated background and subject color model is highly correlated with the dense disordered regular model, a more efficient saliency detection can be performed using the dense disordered regular model.

9. The precise retrieval system for focusing on the main image subject under complex layout according to claim 1, characterized in that: Retrieval preprocessing module: When performing forward evaluation of the hierarchical iterative network, the input image must first be preprocessed, including demeaning compression and adaptive adjustment of the image data matrix. Demeaning compression uses the trained hierarchical iterative network to remember the mean characteristics of the original training set. Therefore, when retrieving an image, even if the input image does not come from the original training library, it is necessary to subtract the mean of the image training library. The data matrix of the image is adaptively adjusted. The initial input image size is 224×224×3. The resolution adaptive adjustment method is: Assume that the value of the unknown function f at the point (x, y) needs to be obtained, and the condition G is known. 11 =(x1,y1),G 12 =(x1,y2),G 21 =(x2,y1),G 22 =(x2,y2) The f value of the four points is calculated by linear interpolation in the x and y directions. First, interpolation is performed in the x-axis direction: Similarly, we can get f(x,y2) and interpolate in the y-axis direction, and finally calculate f(x,y): Convert the image to any resolution so that it can meet the input requirements of the hierarchical iterative network algorithm.

10. The precise retrieval system for focusing on the main image under complex layout according to claim 1, characterized in that: Precise matching module: After the hash rough search results are obtained, the image with a higher score in the hash search results is found and further compared with the original search image, which is called precise matching; For two n-dimensional vectors, we have the following formula: The actual encoding is calculated using the vector mode, namely: All the distances obtained by comparison are sorted by size. The smaller the d value, the higher the similarity between the two images. The final image retrieval result can be obtained by inputting images with relatively small distances.

Citation Information

Patent Citations

  • Pressure curve feature extraction method for pressure pipeline

    CN102174992A

  • Multi-task layered image retrieval method based on depth self-coding convolution neural network

    CN107679250A