Data Retrieval Method and Device

By ensuring identical high-order bits in similarity scores, the method reduces computational load and enhances system performance in image retrieval tasks, addressing the inefficiencies of existing GPU-based image retrieval systems.

CN112464011BActive Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910844299.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-06
Publication Date
2025-07-15
Estimated Expiration
2039-09-06

Smart Images

  • Figure CN112464011B_ABST
    Figure CN112464011B_ABST
Patent Text Reader

Abstract

A data retrieval method and apparatus, which relate to the field of computer technology. After obtaining the feature information of the target data and the feature information of each of the m candidate data, the data retrieval apparatus uses a preset algorithm to determine m target similarities according to the feature information of the target data and the feature information of each candidate data. The value of each target similarity includes n bits, and the values of the consecutive x bits starting from the highest bit are the same; according to the values of the n - x bits (the other bits except the x bits) in each target similarity, a first retrieval operation is performed on the m candidate data to obtain a first retrieval result. Without considering the x bits, the data retrieval apparatus performs retrieval, effectively reducing the amount of calculation and improving the system efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a data retrieval method and apparatus. Background Art

[0002] Currently, a dedicated acceleration hardware such as a graphics processing unit (GPU) or a field-programmable gate array (FPGA) is usually used to improve the rate of retrieving an image from a large number of images.

[0003] Taking a computing device including a GPU (such as a server, a server cluster, or a cloud computing device, etc.) as an example, the computing device can obtain the feature vector of each image in the image database and the feature vector of the image to be queried, and according to the obtained feature vectors, use a selection algorithm (such as the top-K algorithm) to retrieve k images from the image database. After that, the computing device retrieves the images with relatively high similarity to the image to be queried from the k images according to the feature vector of the image to be queried.

[0004] The above method includes two image retrieval processes, and the retrieved images have relatively high accuracy. However, for a GPU, the above method has a large amount of calculation, a large retrieval delay, and low system performance. Summary of the Invention

[0005] This application provides a data retrieval method and apparatus for solving the problems of large calculation amount and low system performance.

[0006] To achieve the above object, the embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, a data retrieval method is provided. After a data retrieval apparatus obtains the feature information of a target data and the feature information of each of m (m is a positive integer) candidate data, a preset algorithm is used to determine m target similarities according to the feature information of the target data and the feature information of each candidate data. Here, the value of each target similarity includes n bits, and the values of consecutive x bits starting from the highest bit are the same. Subsequently, the data retrieval apparatus performs a first retrieval operation on the m candidate data according to the values of n-x bits (other bits except x bits) in each target similarity to obtain a first retrieval result. In this application, x is less than n, x and n are positive integers, and the value of n corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data.

[0008] The number of bits n of the target similarity corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data. Therefore, in a scenario where the dimension of the feature information of the target data and the dimension of the feature information of the candidate data are fixed, the value of n remains unchanged. Since the values of the consecutive x bits starting from the highest bit in each target similarity are the same, the data retrieval device can directly perform a first retrieval operation on the m candidate data according to the n - x bits in each target similarity except for the x bits. Compared with the prior art, if the dimension of the feature information of the target data and the dimension of the feature information of the candidate data in the prior art are the same as those in the present application, the data retrieval method provided by the present application effectively reduces the amount of calculation and improves the system performance.

[0009] In a possible implementation manner of the first aspect above, the preset algorithm includes a first algorithm and a second algorithm. Correspondingly, the method of "the data retrieval device uses the preset algorithm to determine m target similarities according to the feature information of the target data and the feature information of each candidate data" includes: the data retrieval device calculates the similarity between the feature information of the target data and the feature information of each candidate data according to the first algorithm to obtain m initial similarities; then, the data retrieval device calculates the m initial similarities according to the second algorithm to obtain m target similarities. Among them, the first algorithm is a similarity algorithm, and the value of each initial similarity includes n bits.

[0010] After the data retrieval device determines the initial similarity, the data retrieval device also processes the initial similarity according to the second algorithm, so that the values of the consecutive x bits starting from the highest bit of the obtained target similarity are the same.

[0011] In another possible implementation manner of the first aspect above, the method of "the data retrieval device performs a first retrieval operation on the m candidate data according to the n - x bits in each target similarity to obtain a first retrieval result" includes: the data retrieval device selects k first similarities according to the n - x bits in each target similarity and obtains the index values of the candidate data corresponding to the k first similarities. k is a positive integer within a preset value range, and the preset value range is from a first threshold to a second threshold. The first threshold is an integer greater than 1, and the second threshold is an integer less than or equal to m. The value of each of the k first similarities is less than the value of the second similarity, and the second similarity is any one of the m target similarities except for the k first similarities.

[0012] The first search result in this application includes the index values of candidate data corresponding to k first similarities, that is, it includes the index values of k candidate data. Here, k is a positive integer within a preset numerical range. In the prior art, k involved is usually a fixed value. Compared with the prior art, the data retrieval device in this application does not need to accurately obtain candidate data of a certain value, and only needs to ensure that the number of candidate data obtained is within a preset range, further reducing the calculation amount and improving the system performance.

[0013] In another possible implementation manner of the foregoing first aspect, for each target similarity, n - x bits are divided into P consecutive bit segments, and the number of bits in at least one bit segment is greater than 1, where P is an integer greater than or equal to 2. Correspondingly, the method of "the data retrieval device selects k first similarities according to the values of n - x bits in each target similarity" includes: the data retrieval device performs a screening operation on each bit segment of the m target similarities in the order of first screening the bit segments located at the high positions and then screening the bit segments located at the low positions; when the number of the screened target similarities is within the preset numerical range, the screened target similarities are used as the k first similarities; where the screened target similarities include: the target similarities screened according to the first bit segment, or include: the target similarities screened according to the 1st to the a-th bit segments; the first bit segment is the bit segment with the highest position among the P bit segments, and a ∈ [2, P].

[0014] In another possible implementation of the above first aspect, for each target similarity, n - x bits are divided into consecutive P sub - keys, where P is an integer greater than or equal to 2. The method of "the data retrieval device selects k first similarities according to the values of n - x bits in each target similarity" includes: starting from the first sub - key among the P sub - keys in the order from the highest bit to the lowest bit, performing the following processing until k first similarities are selected. Specifically, the processing process is as follows: Determine the value of the i - th sub - key in the m target similarities and the number of times each value appears, where i ∈ [1, P]; Judge whether the number of times the first value appears is within a preset value range. The preset value range is from a first threshold to a second threshold. The first value is: the minimum value of the values of the i - th sub - key in the m target similarities; If the number of times the first value appears is within the preset value range, select the target similarities including the first value, and the target similarities including the first value are the k first similarities; If the number of times the first value appears is less than the first threshold, judge whether the sum of the number of times the first value and the second value appears is within the preset value range. The second value is: the second - smallest value of the values of the i - th sub - key in the m target similarities; Thus, repeat the execution until k first similarities are selected; If the number of times the first value appears is greater than the second threshold, determine the value of the (i + 1) - th sub - key in the target similarity to which the first value belongs and the number of times each value of the (i + 1) - th sub - key appears; Thus, repeat the execution until k first similarities are selected.

[0015] The data retrieval device can complete the selection of the first similarity without sorting the values of the m target similarities by size.

[0016] In another possible implementation of the above first aspect, the feature information is initial feature information or target feature information, and the target feature information is obtained by performing feature processing on the initial feature information.

[0017] Among them, the feature processing is at least one of feature dimension processing, feature equalization processing, or feature quantization processing. The feature dimension processing is used to process the data dimension of the initial feature information into a preset dimension. The feature equalization processing is used to equalize the first feature information, and the equalized first feature information is distributed according to a preset rule. The feature quantization processing is used to quantize the second feature information so that the storage space of the quantized second feature information is less than or equal to the storage space of the second feature information. The first feature information is the initial feature information or the feature information obtained by performing feature dimension processing on the initial feature information. The second feature information is the initial feature information, or the feature information obtained after feature dimension processing, or the feature information obtained after feature equalization processing.

[0018] The initial feature information in this application refers to the feature information obtained after using a feature extraction algorithm to extract features from specified data (such as an image).

[0019] In another possible implementation of the first aspect above, the above-mentioned feature information is target feature information. In this case, the data retrieval device can also perform a second retrieval operation on the first retrieval result according to the initial feature information of the target data to obtain a second retrieval result, so as to improve the accuracy of retrieval.

[0020] In a second aspect, a data retrieval device is provided. The data retrieval device can implement the functions in the first aspect or any of the above possible implementations. These functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0021] The data retrieval device may include an acquisition unit, a determination unit, and a retrieval unit. The acquisition unit, the determination unit, and the retrieval unit can perform the corresponding functions in the data retrieval method of the first aspect and any of its possible implementations above. For example: the above-mentioned acquisition unit is used to acquire the feature information of the target data and the feature information of each of the m candidate data, where m is a positive integer. The above-mentioned determination unit is used to use a preset algorithm to determine m target similarities according to the feature information of the target data and the feature information of each candidate data acquired by the above-mentioned acquisition unit. The value of each target similarity includes n bits, and the values of consecutive x bits starting from the highest bit are the same, where x is less than n, and x and n are positive integers. The value of n corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data. The above-mentioned retrieval unit is used to perform a first retrieval operation on the m candidate data according to the values of the n - x bits in each target similarity determined by the above-mentioned determination unit to obtain a first retrieval result, where n - x bits are the other bits except for the x bits.

[0022] In a possible implementation of the second aspect above, the above-mentioned preset algorithm includes a first algorithm and a second algorithm. Correspondingly, the above-mentioned determination unit is specifically used to: calculate the similarity between the feature information of the target data and the feature information of each candidate data according to the first algorithm to obtain m initial similarities; where the first algorithm is a similarity algorithm, and the value of each initial similarity includes n bits; calculate the m initial similarities according to the second algorithm to obtain m target similarities.

[0023] In another possible implementation of the second aspect described above, the retrieval unit includes a selection module and an acquisition module. The selection module is configured to select k first similarities according to the values of n - x bits in each target similarity, where k is a positive integer within a preset value range, the preset value range is from a first threshold to a second threshold, the first threshold is an integer greater than 1, and the second threshold is an integer less than or equal to m; the value of each of the k first similarities is less than the value of a second similarity, and the second similarity is any one of the m target similarities other than the k first similarities. The acquisition module is configured to acquire the index values of the k first similarities selected by the selection module.

[0024] In another possible implementation of the second aspect described above, for each target similarity, n - x bits are divided into P consecutive bit segments, at least one of the bit segments has more than 1 bit, and P is an integer greater than or equal to 2. The selection module is specifically configured to: perform a screening operation on the m target similarities bit by bit segment in the order of first screening the bit segments located at the higher positions and then screening the bit segments located at the lower positions; when the number of the screened target similarities is within the preset value range, use the screened target similarities as the k first similarities; where the screened target similarities include: the target similarities screened according to the first bit segment, or include: the target similarities screened according to the first to the a-th bit segments; the first bit segment is the highest bit segment among the P bit segments, and a ∈ [2, P].

[0025] In another possible implementation manner of the second aspect described above, for each target similarity, n - x bits are divided into consecutive P sub - keys, where P is an integer greater than or equal to 2. The above - mentioned selection module is specifically configured to: starting from the first sub - key among the P sub - keys in the order from the high - order bit to the low - order bit, perform the following processing until k first similarities are selected. Specifically, the processing process is as follows: Determine the value of the i - th sub - key among the m target similarities and the number of times each value appears, where i ∈ [1, P]; Judge whether the number of times the first value appears is within a preset value range, and the preset value range is from a first threshold to a second threshold. The first value is: the minimum value of the value of the i - th sub - key among the m target similarities; If the number of times the first value appears is within the preset value range, then select the target similarities including the first value, and the target similarities including the first value are the k first similarities; If the number of times the first value appears is less than the first threshold, then judge whether the sum of the number of times the first value and the second value appears is within the preset value range. The second value is: the second - smallest value of the value of the i - th sub - key among the m target similarities; In this way, repeat the execution until k first similarities are selected; If the number of times the first value appears is greater than the second threshold, then determine the value of the (i + 1) - th sub - key among the target similarities to which the first value belongs and the number of times each value of the (i + 1) - th sub - key appears; In this way, repeat the execution until k first similarities are determined.

[0026] In another possible implementation manner of the second aspect described above, the above - mentioned feature information is initial feature information or target feature information, and the target feature information is obtained by performing feature processing on the initial feature information. Here, the feature processing is at least one of feature - dimension processing, feature - equalization processing, or feature - quantization processing. Feature - dimension processing is used to process the data dimension of the initial feature information into a preset dimension. Feature - equalization processing is used to equalize the first feature information, and the equalized first feature information is distributed according to a preset rule. Feature - quantization processing is used to quantize the second feature information so that the storage space of the quantized second feature information is less than or equal to the storage space of the second feature information. The first feature information is the initial feature information or the feature information obtained by performing feature - dimension processing on the initial feature information. The second feature information is the initial feature information, or the feature information obtained after feature - dimension processing, or the feature information obtained after feature - equalization processing.

[0027] In another possible implementation manner of the second aspect described above, the above - mentioned feature information is target feature information. In this case, the above - mentioned retrieval unit is further configured to perform a second retrieval operation on the first retrieval result according to the initial feature information of the target data to obtain a second retrieval result.

[0028] In a third aspect, a data retrieval device is provided, which includes one or more processors and a communication interface; the communication interface is coupled to the one or more processors and is configured to provide data for the one or more processors. When the one or more processors execute program instructions, the data retrieval device implements the data retrieval method as described in the first aspect or various possible implementation manners above.

[0029] In a fourth aspect, a computer-readable storage medium is further provided, in which instructions are stored; when the instructions run on a data retrieval device, the data retrieval device executes the data retrieval method as described in the first aspect or various possible implementation manners above.

[0030] In a fifth aspect, a computer program product is further provided, which includes instructions; when the instructions run on a data retrieval device, the data retrieval device executes the data retrieval method as described in the first aspect or various possible implementation manners above.

[0031] In a sixth aspect, a system chip is further provided, which is applied in a data retrieval device. The data retrieval device includes at least one processor, and the instructions involved are executed in the at least one processor, so that the data retrieval device executes the data retrieval method as described in the first aspect or various possible implementation manners above.

[0032] Optionally, in any of the above aspects or any of the possible implementation manners above, the data retrieval device may obtain the feature information of candidate data from a locally stored feature database (initial feature database / target feature database), or may calculate the feature information of candidate data in real time. This application does not make any limitation in this regard.

[0033] It should be noted that the above computer instructions may be stored in whole or in part on a first computer storage medium. Among them, the first computer storage medium may be packaged together with the processor of the data retrieval device, or may be packaged separately from the processor of the data retrieval device. This application does not make any limitation in this regard.

[0034] The descriptions of the third aspect, the fourth aspect, the fifth aspect, the sixth aspect and their various implementation manners in this application may refer to the detailed descriptions in the first aspect or various implementation manners; and the beneficial effects of the third aspect, the fourth aspect, the fifth aspect, the sixth aspect and their various implementation manners may refer to the beneficial effect analysis in the first aspect or various implementation manners, which will not be elaborated here.

[0035] In this application, the name of the above data retrieval device does not constitute a limitation on the device or functional module itself. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those of this application and fall within the scope of the claims of this application and their equivalent technologies.

[0036] These aspects or other aspects of this application will be made more concise and understandable in the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a schematic structural diagram of an existing image query system;

[0038] Figure 2A is a schematic diagram of the hardware structure of the terminal in an embodiment of the present invention Figure 1 ;

[0039] Figure 2B is the second schematic diagram of the hardware structure of the terminal in an embodiment of the present invention;

[0040] Figure 3 is a schematic flowchart of the data retrieval method in an embodiment of the present invention Figure 1 ;

[0041] Figure 4 is a schematic diagram of the distribution of sub - keys in an embodiment of the present invention;

[0042] Figure 5 is the second schematic flowchart of the data retrieval method in an embodiment of the present invention;

[0043] Figure 6 is a schematic diagram of the structure of the data retrieval device provided in an embodiment of the present invention Figure 1 ;

[0044] Figure 7 is the second schematic diagram of the structure of the data retrieval device provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] For ease of understanding, relevant technical terms related to the embodiments of the present invention are first described.

[0046] Initial feature information: Also known as deep feature information, it refers to the feature information obtained after using a feature extraction algorithm to extract features from specified data.

[0047] Generally, the initial feature information is represented by a vector of dimension n0, where n0 is a positive integer.

[0048] The above-specified data includes, but is not limited to, image data, text data, voice data, etc. The feature extraction algorithm can be an algorithm custom-set by the user side or the system side, such as deep learning algorithms like convolutional neural network (CNN). Among them, the data dimensions of the deep feature information obtained by using different feature extraction algorithms may be different.

[0049] Matched feature information: refers to the information obtained after performing feature dimension processing (feature matching) on the initial feature information. Since the data dimensions of the initial feature information obtained by different feature extraction algorithms are different, it may lead to poor generality in information processing. Therefore, the initial feature information can be subjected to feature dimension processing to obtain matched feature information with a preset data dimension. That is, by inputting initial feature information with different data dimensions, matched feature information with the same corresponding data dimension can be obtained.

[0050] Taking the initial feature information of images A and B as an example, the data dimension of the initial feature information A is greater than that of the initial feature information B. After respectively performing feature dimension processing on the initial feature information A and the initial feature information B, the matched feature information corresponding to the initial feature information A and the matched feature information corresponding to the initial feature information B can be obtained, and the matched feature information of the two images has the same data dimension.

[0051] Balanced feature information: refers to the information obtained after performing feature balancing processing on the feature information (specifically, it can be the initial feature information or the matched feature information). Since the data distribution of the feature information obtained after feature extraction or feature matching is uneven in each dimension. It can be understood that the information included / indicated in each dimension is different, with larger data in some dimensions and smaller data in some dimensions, that is, the feature distribution is uneven. This will lead to low precision in information processing. Therefore, the present application can perform feature balancing processing on the feature information for subsequent related processing using the balanced feature information.

[0052] Quantized feature information: refers to the information obtained after performing feature quantization processing on the feature information (specifically, it can be any one of the initial feature information, the matched feature information, and the balanced feature information).

[0053] During the process of processing massive data, due to the large amount of data resulting in a large occupation of memory space, in order to achieve the high efficiency of information processing and reduce the memory space occupation rate, the feature information can be quantized for subsequent related information processing using the quantized feature information.

[0054] Hash code: The value recreated by scrambling and mixing data using a hash function is called a hash code. The hash function can compress information or data into a digest, making the data volume smaller while fixing the data format.

[0055] Hamming distance: used to represent the number of bits with different values in the same bit position of two codewords.

[0056] For example: taking the hash codes 10101 and 00110 as an example, from left to right, the bits in the first, fourth, and fifth positions are different, so the Hamming distance between these two codewords is 3.

[0057] Cosine distance: The cosine distance d between two pieces of information x and y can be calculated using the following formula (1). Among them, the information x and the information y can be represented by vectors.

[0058]

[0059] Sub key: represents several consecutive bits in a certain type of data (such as floating-point data) in binary form.

[0060] In the same type of data, each data includes the same number of bits. Taking the example of including n (n is a positive integer) bits, the n bits are divided into consecutive P (P is an integer greater than or equal to 2) sub keys, and each sub key includes several consecutive bits. For the same data, the number of bits included in different sub keys can be the same or different.

[0061] Exemplarily, as Figure 4 shown, for m floating-point data: data (1) to data (m), the length of each data is 32 bits. From the high bit to the low bit (taking the highest bit as the 1st bit as an example), the bits included in each data are divided into 5 sub keys: sub key S1 to sub key S5. Sub key S1 includes bits 10 to 14 in data (1) to data (m), sub key S2 includes bits 15 to 18 in data (1) to data (m), sub key S3 includes bits 19 to 23 in data (1) to data (m), sub key S4 includes bits 24 to 28 in data (1) to data (m), and sub key S5 includes bits 29 to 32 in data (1) to data (m), where each bit is one bit.

[0062] Radix symbol: the value of the sub key, that is, the value of several consecutive bits included in the sub key (specifically binary code).

[0063] For example, if sub key S0 includes two consecutive bits, the radix symbol of sub key S0 can be 00, 01, 10, or 11. In the embodiments of the present invention, "the radix symbol of the sub key" and "the value of the sub key" represent the same meaning.

[0064] Exemplarily, as Figure 4As shown, the radix symbol of sub-key S1 in data (1) and data (2) is both 01010, the radix symbol in data (3) is 01110, and the radix symbol in data (m) is 11010; the radix symbol of sub-key S4 in data (1) to data (3) is 1010.

[0065] Radix symbol table: It is used to represent the distribution of the radix symbols included in a certain sub-key in multiple data, and this distribution can be represented by the number of times the radix symbol appears in the multiple data.

[0066] Exemplarily, as Figure 4 shown, in n floating-point data, for sub-key S1, there is no data with a radix symbol of 00000. The data with a radix symbol of 01010 only includes data (1), data (2), and data (m), and the data with a radix symbol of 01110 only includes data (3). Then, Table 1 below shows the radix symbol table corresponding to sub-key S1 in this case. NULL in Table 1 represents empty.

[0067] Optionally, the radix symbol table in the embodiments of the present invention can also be only used to represent the number of times each radix symbol included in a certain sub-key appears in multiple data. In this case, the information of the radix symbol 00000 is not included in the above Table 1.

[0068] Hardware such as GPU / FPGA can effectively improve the efficiency of data retrieval. Taking image retrieval as an example, after the user inputs a query image, a computing device (such as a server, a server cluster, or a cloud computing device, etc.) usually obtains the feature vectors of each image in the image database and the feature vector of the query image to be retrieved, and according to the obtained feature vectors, uses a selection algorithm (such as the top-K algorithm) to retrieve k images (k is a preset value, and the first retrieval operation is performed) from the image database. Then, the computing device retrieves the images with a higher similarity to the query image from the k images (the second retrieval operation is performed) according to the feature vector of the query image to be retrieved.

[0069] Table 1

[0070]

[0071] Such as Figure 1As shown, when it is necessary to query for images similar to Image A, Image A is input into an image query system (taking the image query system in the terminal as an example). The feature extractor 11 extracts features from Image A to obtain the initial feature information 12 of Image A. An image database 13 is pre-configured in the image query system, and the image database 13 includes multiple candidate images. The candidate images in the image database 13 are input into the feature extractor 11 for feature extraction to obtain an initial feature database 14, and the initial feature database 14 includes the initial feature information of each candidate image. To reduce storage space, the initial feature information 12 of Image A and each initial feature information in the initial feature database 14 can also be input into a feature quantizer 15 for feature quantization to obtain a quantized feature database 16 and the quantized feature information 17 of Image A. Subsequently, a retrieval operation for the quantized features (i.e., the first retrieval operation mentioned above) is performed: using the top-K algorithm, the quantized features of the candidate images whose similarity to the quantized feature information of Image A ranks among the top k in the quantized feature database 16 are determined. Further, for the k candidate images, a retrieval operation for the initial features (i.e., the second retrieval operation mentioned above) is performed: the initial features of the candidate images whose similarity to the initial feature information of Image A ranks among the top N in the initial feature database 14 are determined. In this way, the index values (or identifiers) of the top N candidate images can be output. The top N candidate images are the final query results, which are output and displayed for the user to view. Among them, N and k can be positive integers custom-set on the user side or the system side.

[0072] The above calculation of similarity can be achieved by calculating the Hamming distance or cosine distance between the feature information, etc.

[0073] The number of images in the image database 13 is usually very large. Therefore, the number of feature information in the initial feature database 14 and the quantized feature database 16 is also large. The above method needs to perform two retrieval operations, and each retrieval operation needs to calculate the similarity according to the feature information in the feature database (the initial feature database 14 or the quantized feature database 16). For the GPU, the computational complexity of the above method is large, which will increase the retrieval latency.

[0074] In addition, when using the top-K algorithm to obtain k candidate images, the value of k will be very large, and k is a preset value. The value of k is usually in the thousands or tens of thousands. In this case, when performing the above retrieval operation, the computational complexity will be very large. Correspondingly, the retrieval latency will increase and the system performance will be reduced.

[0075] To solve the above problems, an embodiment of the present invention provides a data retrieval method and apparatus. The data retrieval apparatus pre-processes the target similarity between the feature information of the target data (corresponding to the above-mentioned Image A) and the feature information of each candidate data (corresponding to the candidate images in the above-mentioned image database), so that each target similarity is within a preset numerical range, effectively reducing the number of bits to be traversed during the retrieval process. Thus, the amount of calculation and the complexity of calculation are reduced, and the performance of the system is improved. The feature information here is the initial feature information or the quantized feature information. Further, the data retrieval apparatus in the embodiment of the present invention does not need to accurately determine k candidate data, and only needs to determine the number of candidate data within a certain range. Compared with the prior art, the retrieval efficiency of the terminal in the embodiment of the present invention is faster, the amount of calculation and the complexity of calculation are both reduced to a certain extent, and the performance of the system is improved.

[0076] The number of bits included in the target similarity is determined by the dimensions of the feature information of the target data and the feature information of the candidate data, and the dimensions of the feature information of the target data and the feature information of the candidate data are in turn determined by the algorithm used to obtain the feature information. Generally, the algorithm used by the data retrieval apparatus is pre-set and will not change. Therefore, for the same data retrieval apparatus, the number of bits included in the target similarity in the embodiment of the present invention is the same as the number of bits included in the similarity calculated using the prior art.

[0077] Compared with the prior art, the data retrieval method provided by the embodiment of the present invention is applicable to the same data retrieval apparatus.

[0078] In the embodiment of the present invention, the first x bits (bit) of the target similarity are the same. Taking the value of the target similarity as 32 bits as an example, since the first x bits are the same, in the first retrieval, only the remaining 32 - x bits need to be retrieved to obtain the first retrieval result. In the prior art, there is no such rule for the similarity, and all 32 bits need to be retrieved to obtain the retrieval result.

[0079] The data retrieval apparatus in the embodiment of the present application can use a specific similarity algorithm to directly obtain m target similarities with the first x bits being the same, or can determine the target similarity by the following method: first, calculate using any similarity algorithm (corresponding to the first algorithm in the embodiment of the present application) to obtain m initial similarities; then, process each of the obtained initial similarities using a second algorithm (different from the similarity algorithm) to obtain the target similarity with the first x bits being the same. Of course, the embodiment of the present application does not limit the second algorithm.

[0080] Figure 2A Shows a hardware structure of the data retrieval apparatus in the embodiment of the present invention. As Figure 2AAs shown in the figure, the data retrieval device includes a processor 21, a memory 22, a communication interface 23, and a bus 24. The processor 21, the memory 22, and the communication interface 23 can be connected through the bus 24.

[0081] The processor 21 may include one or more processing units. For example, the processor 21 may include an application processor (AP), a modem processor, a GPU, an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0082] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the data retrieval device can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.

[0083] A memory (such as a cache) can be set in the processor 21 for storing instructions and data. In some embodiments, the memory in the processor 21 is a cache memory. This memory can save the instructions or data that the processor 21 has just used or recycled. If the processor 21 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 21, and thus improves the efficiency of the system.

[0084] In some embodiments, the processor 21 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0085] In the embodiments of the present invention, the processor 21 can be used to extract features from the obtained first data (such as target data and candidate data, etc.) to obtain the initial feature information of the first data, such as the initial feature information of the target data and the initial feature information of the candidate data, etc. After obtaining the initial feature information of the first data, the initial feature information of the first data can be stored in the memory 22 (specifically, the hard disk). The processor 21 can also be used to perform feature processing on the initial feature information of the first data to obtain the target feature information of the first data (such as matching feature information, equalization feature information, or quantization feature information). Further, the target feature information of the first data is stored in the memory 22 (specifically, the memory).

[0086] The processor 21 can also be used to determine the target similarity between the target data and each candidate data by using a preset algorithm, and perform a retrieval operation on the candidate data according to the determined target similarity. Among them, the values of the consecutive x bits starting from the highest bit of each determined target similarity are the same. The determination of the target similarity will be described in detail later in this application.

[0087] The memory 22 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 21 executes various functional applications and data processing of the data retrieval device by running the instructions stored in the memory 22. For example, in an embodiment of the present invention, when receiving the target data input by the user, the processor 21 can obtain the initial feature information of the target data by executing the instructions stored in the memory 22. The memory 22 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as an image query function, etc.). The data storage area can store the data created during the use of the data retrieval device (such as the initial feature information of the target data, etc.).

[0088] The memory 22 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0089] The memory 22 includes at least one of a memory and a video memory, and a hard disk.

[0090] In an embodiment of the present invention, the hard disk is used to store the first data (such as target data and candidate data, etc.) and / or the initial feature information of the first data. Exemplarily, when the number of the first data is multiple, the initial feature information of each of the multiple first data can be stored in the initial feature database and stored in the hard disk.

[0091] The memory is used to store the target feature information of the first data, etc. When the number of the first data is multiple, the target feature information of each of the multiple first data can be stored in the target feature database and stored in the memory, which is convenient for high-speed parallel computing during information retrieval and can improve the efficiency of information retrieval.

[0092] In a possible implementation manner, the memory 22 can exist independently of the processor 21. The memory 22 can be connected to the processor 21 through the bus 25 and is used to store instructions or program code. When the processor 21 calls and executes the instructions or program code stored in the memory 22, the data retrieval method provided by the embodiments of the present invention can be implemented.

[0093] In another possible implementation, the memory 22 can also be integrated with the processor 21.

[0094] The communication interface 23 is used to connect the data retrieval device to other devices (such as a server) through a communication network. The communication network can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 23 can include a receiving unit for receiving data and a transmitting unit for transmitting data.

[0095] In an embodiment of the present invention, the communication interface 23 is used to obtain first data (such as target data and candidate data, etc.). The first data can specifically be image data, text data, voice data, etc., and the embodiments of the present invention do not limit this. Taking the first data as image data as an example, the communication interface 23 can be an acquisition channel or component for obtaining the image data.

[0096] The bus 24 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 2A only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0097] Optionally, as Figure 2A shown, the data retrieval device can further include a display screen 25.

[0098] The display screen 25 is used to display images, videos, etc. The display screen 24 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.

[0099] In some embodiments, the data retrieval device may include one or N display screens 25, where N is a positive integer greater than 1. For example, in the embodiments of the present invention, the display screen 25 can be used to display retrieval results.

[0100] The data retrieval device realizes the display function through a GPU, the display screen 25, etc. The GPU is a microprocessor for image processing and is connected to the display screen 25. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 21 may include one or more GPUs, which execute program instructions to generate or change display information.

[0101] Figure 2B Another hardware structure of the data retrieval device in the embodiments of the present invention is shown. As Figure 2B shown, the data retrieval device may include a processor 31 and a communication interface 32. The processor 31 is coupled to the communication interface 32.

[0102] The functions of the processor 31 can refer to the description of the above-mentioned processor 21. In addition, the processor 31 also has a storage function and can perform the functions of the above-mentioned memory 22.

[0103] The communication interface 32 is used to provide data for the processor 31. The communication interface 32 can be an internal interface of the data retrieval device or an external interface of the data retrieval device (equivalent to the communication interface 23).

[0104] It should be noted that Figure 2A (or Figure 2B ) The structure shown does not constitute a limitation on the data retrieval device. Except Figure 2A (or Figure 2B ) The components shown, the data retrieval device may include more or fewer components than shown in the figure, or combine certain components, or arrange different components.

[0105] The data retrieval method provided by the embodiments of the present invention will be described below in conjunction with the accompanying drawings. The data retrieval method provided by the embodiments of the present invention can be applied in Figure 2A or Figure 2B the data retrieval device shown. Unless otherwise specified, the following steps are all executed by the processor 21 or the processor 31 (specifically, it can be a GPU) in the data retrieval device.

[0106] Figure 3 FIG. is a schematic flowchart of a data retrieval method provided by an embodiment of the present invention. As Figure 3 shown, the data retrieval method includes:

[0107] S301. The data retrieval device acquires the feature information of the target data and the feature information of each of the m candidate data, where m is a positive integer.

[0108] The target data can be image data, text data, voice data, or other forms of data, which is not limited in this embodiment.

[0109] The data retrieval device can adopt the following first implementation manner or the second implementation manner to acquire the feature information of the target data and the feature information of each of the m candidate data. Of course, the data retrieval device can also adopt other implementation manners to acquire the feature information of the target data and the feature information of each of the m candidate data, which is not limited in this embodiment.

[0110] The first implementation manner: The data retrieval device first acquires the target data, and then acquires the initial feature information of the target data according to the acquired target data. In this case, the above feature information is the initial feature information.

[0111] Among them, the data retrieval device can acquire the target data from the local database, or receive the target data sent by other devices (such as receiving the image data sent by other data retrieval devices through Bluetooth, receiving the image data transmitted by social applications), or can also adopt other ways to acquire the target data, which is not limited in this embodiment.

[0112] After the data retrieval device obtains the target data, it can use a feature extraction algorithm to extract the features of the target data, so as to obtain the initial feature information of the target data. Among them, the feature extraction algorithms adopted by different manufacturers may be different, and correspondingly, the data dimensions of the initial feature information of the target data obtained may also be different.

[0113] The feature extraction algorithm is an algorithm used to implement feature extraction, such as deep residual network (resnet), Google convolutional neural network (GoogLeNet), perceptron neural network (PNN), Hopfield network, radial basis function (RBF) neural network, feedback neural network and other deep learning algorithms.

[0114] The m candidate data are the information pre-stored in the database. The data retrieval device can obtain the initial feature information of each candidate data among the m candidate data by means of "obtaining the initial feature information of the target data", or the initial feature information of each candidate data is pre-stored in the local information feature database, and the data retrieval device directly obtains the initial feature information of each candidate data among the m candidate data from the information feature database. This embodiment does not limit the manner in which the data retrieval device obtains the initial feature information of each candidate data among the m candidate data.

[0115] The second implementation manner: To adapt to the data retrieval systems of different manufacturers, after obtaining the initial feature information of the target data, the data retrieval device performs feature processing (at least one of feature dimension processing, feature equalization processing, or feature quantization processing) on the initial feature information of the target data to obtain the target feature information of the target data. In this case, the above-mentioned feature information is the initial feature information.

[0116] Optionally, the target feature information of each candidate data among the m candidate data can be the feature information pre-stored in the information feature database. The initial feature information of the candidate data and the target feature information of the candidate data correspond one by one, and both can be distinguished by the identifier / index value of the same candidate data. This embodiment will not elaborate on this in detail.

[0117] The feature dimension processing involved in the second implementation manner refers to: performing feature dimension processing on the initial feature information so that the data dimension of the initial feature information is processed into a preset dimension. That is, the initial feature information extracted by different manufacturers is processed into feature information of a preset dimension. Among them, the preset dimension can be custom-set by the user side or the system side, and it can be a positive integer. This embodiment does not limit this.

[0118] Optionally, the feature information obtained after feature dimension processing can be used as the target feature information.

[0119] Specifically, the data retrieval device can perform feature dimension processing on the initial feature information (of the target data or candidate data) using a pre-stored feature matching matrix, so as to process the data dimension of the initial feature information into a preset dimension. Among them, the feature matching matrix is pre-trained by the system side using the first sample information. The first sample information can be small sample information. For example, the number of the first sample information is less than or equal to a preset number, etc. Regarding how to train the feature matching matrix using the first sample information, reference can be made to the description of the prior art, and this embodiment will not elaborate in detail.

[0120] The feature dimension processing includes any one of the following: dimension increase processing, dimension reduction processing, and equal dimension processing. When the data dimension of the initial feature information is greater than the preset dimension, the feature dimension processing is dimension reduction processing; when the data dimension of the initial feature information is equal to the preset dimension, the feature dimension processing is equal dimension processing; when the data dimension of the initial feature information is less than the preset dimension, the feature dimension processing is dimension increase processing. For the specific descriptions of dimension increase processing, dimension reduction processing, and equal dimension processing, reference can be made to the description of the prior art, and this embodiment will not elaborate in detail.

[0121] The feature balance processing involved in the above second implementation manner refers to: performing feature balancing on the first feature information so that the balanced feature information is distributed according to a preset rule. Here, the first feature information is the initial feature information or the feature information obtained after performing feature dimension processing on the initial feature information, and this feature information can also be called matching feature information.

[0122] Optionally, the feature information obtained after feature balance processing can be used as the target feature information.

[0123] Specifically, the data retrieval device can use a feature balance matrix to balance the first feature information, so that the balanced first feature information (which can also be called balanced feature information in this application) is distributed according to a preset rule. Among them, the preset rule is associated with the method / means of feature balancing, and it can be custom-set by the user side or the system side. For example, the data retrieval device can adopt techniques such as random projection, variance balance, Gaussian balance, etc. to construct the feature balance matrix. Taking Gaussian balance as an example, the balanced first feature information can be in a Gaussian distribution.

[0124] The feature balance matrix can be pre-trained by the data retrieval device using the second sample information. The second sample information can be small sample information. The second sample information and the above first sample information can be the same or different. Regarding how to train the feature balance matrix using the first sample information, reference can be made to the description of the prior art, and this embodiment will not elaborate in detail.

[0125] Exemplarily, the data retrieval device may generate D using the Marsaglia polar method of Gaussian distribution random sequence coding n × D n mutually independent random numbers that follow the standard normal distribution, and arrange them into a D n × D n random matrix. Further, the random matrix is decomposed using the orthogonal triangular QR decomposition method to obtain a random orthogonal matrix as the feature equalization matrix R. Further, the data retrieval device can multiply the matching feature information obtained after feature dimension processing by the feature equalization matrix R to obtain the corresponding equalized feature information.

[0126] The feature quantization processing involved in the above second implementation manner refers to: performing feature quantization processing on the second feature information, such that the storage space of the quantized second feature information (which can also be referred to as the quantized feature information in this application) is less than or equal to the storage space of the second feature information, or such that the data dimension of the quantized second feature information is less than or equal to the data dimension of the second feature information. Among them, the second feature information is the initial feature information, or the feature information obtained after feature dimension processing (i.e., the matching feature information mentioned above), or the feature information obtained after feature equalization processing (i.e., the equalized feature information mentioned above).

[0127] Optionally, the feature information obtained after feature quantization processing can be used as the target feature information.

[0128] Specifically, the data retrieval device can use a preset feature quantization model to perform feature quantization on the second feature information to obtain the corresponding quantized feature information. The data dimension of the quantized feature information is significantly less than or equal to the data dimension of the feature information before quantization, which can reduce the storage space of the memory. That is to say, the storage space of the quantized feature information is less than or equal to the storage space of the feature information before quantization. The feature quantization model can specifically be custom - set on the user side or the system side, or can be obtained by the data retrieval device pre - training using the third sample information. The third sample information can also be small - sample information, which can be the same as or different from the above - mentioned first sample information and the above - mentioned second sample information. This embodiment does not make a limitation on this.

[0129] Exemplarily, taking the quantization of the feature quantization model as the hash - coding criterion as an example, the following formula shows a kind of hash - coding criterion:

[0130]

[0131] Among them, x is the second feature information, which can specifically be the feature values of each dimension in the feature information. y is the quantized second feature information. Taking the second feature information as 8-dimensional feature information as an example, specifically x = (-0.1, 0.2, 0.5, -1, 1.7, 0.8, -0.9, 0.4). Correspondingly, y = (0, 1, 1, 0, 1, 1, 0, 1) is obtained after quantization using the above hash coding criterion.

[0132] Optionally, the feature processing involved in the above second implementation manner can be encapsulated into a feature processing model. That is, the initial feature information of the target data is input into the feature processing model for corresponding feature processing to obtain the target feature information of the target data.

[0133] In this embodiment, for each piece of data (such as target data / candidate data), whether it is the initial feature information or the target feature information, it is a multi-dimensional vector.

[0134] S302. The data retrieval device uses a preset algorithm to determine m target similarities according to the feature information of the target data and the feature information of each candidate data. The value of each target similarity includes n bits, and the values of the consecutive x bits starting from the highest bit are the same.

[0135] The target similarity in this embodiment can be represented by the distance between feature information, specifically in float type or int type.

[0136] If the feature information obtained by the data retrieval device is the initial feature information, after obtaining the initial feature information of the target data and the initial feature information of each candidate data, the data retrieval device calculates the preset distance (such as Hamming distance, cosine distance, Euler distance or asymmetric distance) between the initial feature information of the target data and the initial feature information of each candidate data. Each calculated preset distance represents an initial similarity. The algorithm used to calculate the preset distance here corresponds to the first algorithm in the embodiments of the present invention; then, the data retrieval device uses the "second algorithm corresponding to the feature extraction algorithm" to process each initial similarity so that the values of the consecutive x bits starting from the highest bit of each processed initial similarity (i.e., the target similarity) are the same.

[0137] If the feature information obtained by the data retrieval device is the target feature information, then after obtaining the target feature information of the target data and the target feature information of each candidate data, the data retrieval device calculates the preset distance (such as Hamming distance, cosine distance, Euler distance or asymmetric distance) between the target feature information of the target data and the target feature information of each candidate data. Each preset distance represents an initial similarity. The algorithm used to calculate the preset distance here corresponds to the first algorithm in the embodiments of the present invention; then, the data retrieval device processes each initial similarity using the "second algorithm corresponding to the feature processing algorithm", so that the continuous x bits of the processed each initial similarity (i.e., the target similarity) from the highest bit have the same value.

[0138] Whether it is the initial feature information or the above-mentioned target feature information, the data retrieval device first calculates the initial similarity between the target data and each candidate data according to the first algorithm. Then, the data retrieval device processes each initial similarity using the second algorithm corresponding to the manner of determining the initial feature information / target feature information to determine m target similarities. Among them, both the first algorithm and the second algorithm are custom-set by the user side or the system side, and this embodiment does not limit this.

[0139] Each target similarity among the m target similarities determined by the data retrieval device includes n (n is a positive integer) bits. The value of n corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data. Since the dimension of the initial feature information is related to the feature extraction algorithm, and the dimension of the target feature information is related to the algorithm used for feature processing, it can also be considered that the value of n corresponds to the feature extraction algorithm and the algorithm used for feature processing.

[0140] The continuous x bits of each target similarity among the m target similarities from the highest bit have the same value, and x is a positive integer.

[0141] Exemplarily, as Figure 4 shown, Figure 4 each of the data (1) to data (m) in represents a target similarity, and the length of each data is 32 bits. Among these m data, the continuous 9 bits from the high bit have the same value, which is "010000000".

[0142] S303. The data retrieval device performs a first retrieval operation on the m candidate data according to the value of n - x bits in each target similarity to obtain a first retrieval result.

[0143] Among them, n - x bits are the other bits except x bits in n bits.

[0144] Since the values of the consecutive x bits starting from the highest bit in each target similarity are the same, the data retrieval device can ignore / skip the values of the x bits and directly select k first similarities from the m target similarities according to the values of the n - x bits in each target similarity, and obtain the index values of the candidate data corresponding to the k first similarities (the index values of the candidate data corresponding to the k first similarities are the first retrieval results). In this way, the computational amount of the data retrieval device is effectively reduced. Based on the reduction of the computational amount, the computational complexity is also reduced to a certain extent, improving the performance of the system.

[0145] The value of each of the k first similarities is less than the value of the second similarity, and the second similarity is any one of the m target similarities other than the k first similarities. That is to say, the data retrieval device needs to select k target similarities with smaller values from the m target similarities.

[0146] Specifically, the data retrieval device can adopt the following implementation method I, implementation method II or implementation method III to select k first similarities from the m target similarities. Of course, the data retrieval device can also adopt other implementation methods to select k first similarities from the m target similarities, and this embodiment does not limit this.

[0147] Implementation method I: The data retrieval device arranges the m target similarities in ascending order of the values of the n - x bits in sequence and obtains the first k target similarities in the arrangement, or the data retrieval device arranges the m target similarities in descending order of the values of the n - x bits in sequence and obtains the last k target similarities in the arrangement. Here, the k target similarities obtained are the k first similarities.

[0148] Implementation method II: In this embodiment, for each target similarity, the n - x bits are divided into consecutive P (P is an integer greater than or equal to 2) bit segments, and the number of bits of at least one bit segment is greater than 1. The data retrieval device performs a screening operation on the m target similarities bit segment by bit segment in the order of first screening the bit segments located at the high positions and then screening the bit segments located at the low positions; when the number of the screened target similarities is within the preset value range, the screened target similarities are used as the k first similarities. Among them, the screened target similarities include: the target similarities screened according to the first bit segment, or include: the target similarities screened according to the 1st to the a-th bit segments; the first bit segment is the bit segment with the highest position among the P bit segments, and a ∈ [2, P].

[0149] Implementation III: In this embodiment, for each target similarity, n - x bits are divided into consecutive P (P is an integer greater than or equal to 2) sub - keys (the sub - keys can correspond to the above - mentioned bit segments), and each sub - key includes several consecutive bits. The base symbols of the same sub - key in different target similarities can be the same or different. After determining the m target similarities, the data retrieval device traverses the P sub - keys in the order from the highest bit to the lowest bit until k first similarities are selected from the m target similarities.

[0150] Compared with the above - mentioned Implementation I, when the data retrieval device in the above - mentioned Implementation II and Implementation III performs the first retrieval operation, there is no need to sort the m target similarities, and the performance of the data retrieval device is higher.

[0151] In this embodiment, k can be a preset value or a value within a preset value range. The preset value range is a range customized by the user side or the system side. For example, the preset value range is: [j, j(1 + θ)], where θ is a number greater than 0 and less than 1. If j = 100 and θ = 0.5, then the preset value range is [100, 150].

[0152] Next, Implementation III will be described with k being a preset value.

[0153] Specifically, for the i - th (i is a positive integer) sub - key among the P sub - keys, the data retrieval device determines the base symbols included in the i - th sub - key in the m target similarities and the number of times each base symbol appears (the base symbols included in the i - th sub - key in the m target similarities and the number of times each base symbol appears can be stored in the data retrieval device in the form of a base symbol table). Then, the data retrieval device determines whether the number of times the base symbol ranked first in ascending order of base symbols (referred to as the first base symbol, corresponding to the first value involved in the embodiments of the present invention) appears is equal to k.

[0154] If the number of times the first base symbol appears is equal to k, the data retrieval device terminates the first retrieval operation, determines the target similarity to which the first base symbol belongs as the first similarity, and obtains the index value of the candidate data corresponding to the first similarity, that is, determines the first retrieval result.

[0155] If the number of times the first base symbol appears is less than k, the data retrieval device determines whether the sum of the number of times the first base symbol and the base symbol ranked second in ascending order of base symbols (referred to as the second base symbol, corresponding to the second value involved in the embodiments of the present invention) appears is equal to k. In this way, it is repeatedly executed until k first similarities are determined according to the i - th sub - key.

[0156] If the number of occurrences of the first radix symbol is greater than k, the data retrieval device determines the radix symbols included in the (i + 1)-th sub-key in the target similarity to which the first radix symbol belongs, and the number of occurrences of each radix symbol. Thereafter, the data retrieval device processes the (i + 1)-th sub-key in the same way as the i-th sub-key, and repeats this process until k first similarities are determined.

[0157] Exemplarily, taking m as 100, the first sub-key includes 2 consecutive bits. The radix symbols included in 100 target similarities are "00", "01", and "11". The number of occurrences of the radix symbol "00" in 100 target similarities is 50, the number of occurrences of the radix symbol "01" in 100 target similarities is 30, and the number of occurrences of the radix symbol "11" in 100 target similarities is 20 as an example.

[0158] When k is 50, since the number of occurrences of the radix symbol "00" is equal to 50, the data retrieval device can obtain the index value of the candidate data corresponding to the target similarity to which the radix symbol "00" belongs.

[0159] When k is 70, the number of occurrences of the radix symbol "00" is less than 70. The data retrieval device obtains the index value of the candidate data corresponding to the target similarity to which the radix symbol "00" belongs. In addition, the data retrieval device also needs to consider the number of occurrences of the radix symbol "01". Since the sum of the number of occurrences of the radix symbol "00" and the radix symbol "01" is greater than 70, the data retrieval device further determines: for the target similarities including the radix symbol "01", the radix symbols included in the second sub-key, and the number of occurrences of each radix symbol. Thereafter, the data retrieval device processes the second sub-key in the same way as the first sub-key, and repeats this process until 20 target similarities including the radix symbol "01" are screened out. Subsequently, the data retrieval device obtains the index values of the candidate data corresponding to the 20 screened target similarities.

[0160] When k is 30, since the number of occurrences of the radix symbol "00" is greater than 30, the data retrieval device further determines: for the target similarities including the radix symbol "00", the radix symbols included in the second sub-key, and the number of occurrences of each radix symbol. Thereafter, the data retrieval device processes the second sub-key in the same way as the first sub-key, and repeats this process until 30 target similarities including the radix symbol "00" are screened out. Subsequently, the data retrieval device obtains the index values of the candidate data corresponding to the 30 screened target similarities.

[0161] Next, taking k as a value within a preset numerical range, the implementation case of Implementation III is described.

[0162] Specifically, for the i-th (where i is a positive integer) sub-key among the P sub-keys, the data retrieval device determines the base symbols included in the m target similarities for the i-th sub-key, and the number of occurrences of each base symbol (the base symbols included in the m target similarities for the i-th sub-key and the number of occurrences of each base symbol can be stored in the data retrieval device in the form of a base symbol table). Then, the data retrieval device determines whether the number of occurrences of the first base symbol (refer to the above description) is within a preset value range.

[0163] If the number of occurrences of the first base symbol is within the second preset value range, the data retrieval device terminates the first retrieval operation, determines the target similarity to which the first base symbol belongs as the first similarity, and obtains the index value of the candidate data corresponding to the first similarity, that is, determines the first retrieval result.

[0164] If the number of occurrences of the first base symbol is less than the first threshold (the minimum value within the preset value range), the data retrieval device determines whether the sum of the number of occurrences of the first base symbol and the second base symbol (refer to the above description) is within the second preset value range. In this way, it is repeatedly executed until k (within the preset value range) first similarities are determined according to the i-th sub-key.

[0165] If the number of occurrences of the first base symbol is greater than the second threshold (the maximum value within the preset value range), the data retrieval device determines the base symbols included in the target similarity to which the first base symbol belongs for the (i + 1)-th sub-key, and the number of occurrences of each base symbol. Then, the data retrieval device processes the (i + 1)-th sub-key in the same way as the i-th sub-key. In this way, it is repeatedly executed until k (within the second preset value range) first similarities are determined.

[0166] Exemplarily, taking m as 100, the first sub-key includes 2 consecutive bits, the base symbols included in 100 target similarities are "00", "01", and "11", the number of occurrences of the base symbol "00" in 100 target similarities is 50, the number of occurrences of the base symbol "01" in 100 target similarities is 30, and the number of occurrences of the base symbol "11" in 100 target similarities is 20 as an example.

[0167] When the preset value range is [30, 70], since the number of occurrences of the base symbol "00" is within [30, 70], the data retrieval device only needs to obtain the index value of the candidate data corresponding to the target similarity to which the base symbol "00" belongs.

[0168] When the preset numerical range is [80, 120], the number of occurrences of the base symbol "00" is less than 80, and the data retrieval device also needs to consider the number of occurrences of the base symbol "01". Since the sum of the number of occurrences of the base symbol "00" and the base symbol "01" is equal to 80, and this sum of the number of occurrences is within [80, 120], therefore, the data retrieval device only needs to obtain the index value of the candidate data corresponding to the target similarity to which the base symbol "00" and the base symbol "01" belong.

[0169] When the preset numerical range is [60, 70], the number of occurrences of the base symbol "00" is less than 60, and the data retrieval device also needs to consider the number of occurrences of the base symbol "01". Since the sum of the number of occurrences of the base symbol "00" and the base symbol "01" is equal to 80, and this sum of the number of occurrences is greater than 70, therefore, the data retrieval device further determines: for the target similarity including the base symbol "01", the base symbols included in the second sub-key and the number of occurrences of each base symbol. Then, the data retrieval device processes the second sub-key in the same way as the first sub-key, and repeats this process until 10 to 20 are selected from the target similarities including the base symbol "01". Subsequently, the data retrieval device obtains the index value of the candidate data corresponding to the selected target similarity.

[0170] When the preset numerical range is [20, 45], since the number of occurrences of the base symbol "00" is greater than 45, therefore, the data retrieval device further determines: for the target similarity including the base symbol "00", the base symbols included in the second sub-key and the number of occurrences of each base symbol. Then, the data retrieval device processes the second sub-key in the same way as the first sub-key, and repeats this process until k (within [20, 45]) first similarities are selected from the target similarities including the base symbol "00". Subsequently, the data retrieval device obtains the index value of the candidate data corresponding to the k first similarities.

[0171] It can be seen that compared with the case where k is a preset value, when k is within the preset numerical range, the data retrieval device only needs to determine that the number of first similarities is within the preset numerical range, without needing to be accurate to the preset value, effectively reducing the amount of calculation and the complexity of the calculation, and improving the performance of the system.

[0172] In summary, for the same data retrieval device in the prior art, the data retrieval method provided by the embodiments of the present invention effectively reduces the number of bits that need to be traversed during the retrieval process, thereby reducing the amount of calculation and the complexity of the calculation, and improving the performance of the system.

[0173] In the embodiment of the present invention, the data retrieval device may no longer accurately determine k candidate data, and only needs to determine the number of candidate data within a certain range of k. Compared with the prior art, the retrieval efficiency of the data retrieval device in the embodiment of the present invention is faster, and the calculation amount and calculation complexity are both reduced to a certain extent, improving the performance of the system.

[0174] Of course, in order to improve the retrieval accuracy, in the scenario where the data retrieval device determines the target similarity according to the target feature information of the target data and the target feature information of each candidate data, that is, in the scenario where the target similarity is determined by adopting the above second implementation manner, the data retrieval device may further perform a second retrieval operation on the first retrieval result based on the initial feature information of the target data to obtain a second retrieval result.

[0175] Combined with Figure 3 , such as Figure 5 shown, after S303, the data retrieval method provided by the embodiment of the present invention further includes:

[0176] S501. The data retrieval device performs a second retrieval operation on the first retrieval result according to the initial feature information of the target data to obtain a second retrieval result.

[0177] The data retrieval device retrieves the initial feature information of each of the k candidate data in the first retrieval result according to the initial feature information of the target data to obtain the result information corresponding to the target data.

[0178] Specifically, the data retrieval device obtains the initial feature information of each of the k candidate data according to the index values of the k candidate data. Furthermore, the cosine distances between the initial feature information of each of the k candidate data and the initial feature information of the target data are calculated respectively, and then the obtained k cosine distances are arranged in descending order, and the index values of the candidate data corresponding to the top s cosine distances are selected as the index values of the result information. Furthermore, according to the obtained s index values of the result information, s result information corresponding to the index value is obtained from the database.

[0179] Optionally, the data retrieval device may display the s result information to the user for viewing. Wherein, s is proportional to k (or m). And, s is less than or equal to k and much less than m. s is a positive integer.

[0180] In practical applications, the data retrieval method provided by the embodiments of the present invention can support concurrent retrieval of multiple pieces of information. The initial feature information involved in the embodiments of the present invention can be stored in an initial feature database, which can be specifically stored in the hard disk of the data retrieval device. The target feature information involved in the embodiments of the present invention (specifically, it can be matching feature information, balanced feature information, and quantization feature information) can be stored in a target feature database, which can be specifically stored in the memory of the data retrieval device, such as the video memory of a graphics processing unit (GPU), the memory of a central processing unit (CPU), and dedicated chips, etc.

[0181] By implementing the embodiments of the present invention, the amount of calculation and the complexity of calculation are effectively reduced, and the performance of the system is improved.

[0182] The above mainly introduces the solution provided by the embodiments of the present invention from the perspective of the method. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0183] The embodiments of the present invention can perform function module division on the above data retrieval device, etc. according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software function module. It should be noted that the division of modules in the embodiments of the present invention is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0184] As Figure 6 shown, it is a schematic structural diagram of a data retrieval device provided by the embodiments of the present invention. Figure 6 The data retrieval device 60 shown can be a computing device, or a chip in a computing device, or a system on a chip in a computing device. The data retrieval device 60 can be used to execute the steps performed by the data retrieval device in any of the data retrieval methods provided above.

[0185] The data retrieval device 60 may include: an acquisition unit 601, a determination unit 602, and a retrieval unit 603. Among them, the acquisition unit 601 is configured to acquire the feature information of the target data and the feature information of each of the m candidate data, where m is a positive integer. Exemplarily, the acquisition unit 601 may be configured to execute S301. The determination unit 602 is configured to use a preset algorithm to determine m target similarities according to the feature information of the target data and the feature information of each candidate data acquired by the acquisition unit 601. The value of each target similarity includes n bits, and the values of the consecutive x bits starting from the highest bit are the same, where x is less than n, and x and n are positive integers, and the value of n corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data. Exemplarily, the determination unit 602 may be configured to execute S302. The retrieval unit 603 is configured to perform a first retrieval operation on the m candidate data according to the values of the n - x bits in each similarity determined by the determination unit 602 to obtain a first retrieval result, where n - x bits are the other bits except for the x bits. Exemplarily, the retrieval unit 603 may be configured to execute S303.

[0186] Optionally, the above - mentioned feature information is initial feature information or target feature information, and the target feature information is obtained by performing feature processing on the initial feature information. Among them, the feature processing is at least one of feature dimension processing, feature equalization processing, or feature quantization processing; the feature dimension processing is used to process the data dimension of the initial feature information into a preset dimension; the feature equalization processing is used to equalize the first feature information, and the equalized first feature information is distributed according to a preset rule; the feature quantization processing is used to quantize the second feature information so that the storage space of the quantized second feature information is less than or equal to the storage space of the second feature information; the first feature information is the initial feature information or the feature information obtained by performing feature dimension processing on the initial feature information; the second feature information is the initial feature information, or the feature information obtained by feature dimension processing, or the feature information obtained by feature equalization processing.

[0187] Optionally, the above - mentioned preset algorithm includes a first algorithm and a second algorithm. The determination unit 602 is specifically configured to: according to the first algorithm, calculate the similarity between the feature information of the target data and the feature information of each candidate data to obtain m initial similarities; where the first algorithm is a similarity algorithm, and the value of each initial similarity includes n bits; according to the second algorithm, calculate the m initial similarities to obtain m target similarities.

[0188] Optionally, in combination Figure 6 , as Figure 7 shown, the retrieval unit 603 includes a selection module 6031 and an acquisition module 6032.

[0189] A selection module 6031 is configured to select k first similarities according to the values of n - x bits in each target similarity, where k is a positive integer within a preset value range. The preset value range is from a first threshold to a second threshold. The first threshold is an integer greater than 1, and the second threshold is an integer less than or equal to m. The value of each of the k first similarities is less than the value of a second similarity, and the second similarity is any one of the m target similarities other than the k first similarities.

[0190] An acquisition module 6032 is configured to acquire the index values of the candidate data corresponding to the k first similarities selected by the selection module 6031.

[0191] Optionally, for each target similarity, the n - x bits are divided into consecutive P (P is an integer greater than or equal to 2) bit segments, and the number of bits in at least one bit segment is greater than 1. The selection module 6031 is specifically configured to: perform a screening operation on the m target similarities bit segment by bit segment in the order of first screening the bit segments at the high positions and then screening the bit segments at the low positions; when the number of the screened target similarities is within the preset value range, use the screened target similarities as the k first similarities; where the screened target similarities include: the target similarities screened according to the first bit segment, or include: the target similarities screened according to the first to the a-th bit segments; the first bit segment is the bit segment at the highest position among the P bit segments, and a ∈ [2, P].

[0192] Optionally, for each target similarity, n - x bits are divided into consecutive P sub - keys, where P is an integer greater than or equal to 2. The selection module 6031 is specifically configured to: starting from the first sub - key among the P sub - keys in the order from the high - order bit to the low - order bit, perform the following processing until k first similarities are selected: determine the value of the i - th sub - key among the m target similarities and the number of times each value appears, where i ∈ [1, P]; determine whether the number of times the first value appears is within a preset value range, the preset value range is from a first threshold to a second threshold, and the first value is: the minimum value of the value of the i - th sub - key among the m target similarities; if the number of times the first value appears is within the preset value range, select the target similarity including the first value, and the target similarity including the first value is one of the k first similarities; if the number of times the first value appears is less than the first threshold, determine whether the sum of the number of times the first value and the second value appears is within the preset value range, and the second value is: the second - smallest value of the value of the i - th sub - key among the m target similarities; and so on, repeat the execution until k first similarities are selected; if the number of times the first value appears is greater than the second threshold, determine the value of the (i + 1) - th sub - key in the target similarity to which the first value belongs and the number of times each value of the (i + 1) - th sub - key appears; and so on, repeat the execution until k first similarities are determined.

[0193] Optionally, when the feature information is the target feature information, the retrieval unit 603 is further configured to perform a second retrieval operation on the first retrieval result according to the initial feature information of the target data to obtain a second retrieval result. Exemplarily, the retrieval unit 603 may be configured to execute S501.

[0194] Of course, the data retrieval device 60 provided in the embodiments of the present invention includes but is not limited to the above - mentioned modules. For example, the data retrieval device 60 may further include a storage unit 604. The storage unit 604 may be configured to store the program code of the data retrieval device 60, and may also be configured to store the data obtained during the operation of the data retrieval device 60, such as the initial feature information of the first data, etc.

[0195] In actual implementation, the obtaining unit 601, the determining unit 602, and the retrieval unit 603 may all be implemented by Figure 2A the processor 21 shown in calling the program code in the memory 22. The specific execution process may refer to Figure 3 or Figure 5 the description of the data retrieval method part shown, which will not be elaborated here.

[0196] Or,

[0197] In actual implementation, the obtaining unit 601, the determining unit 602, and the retrieval unit 603 may all be implemented by Figure 2BThe processor 31 shown executes program code to implement. For the specific execution process, reference can be made to Figure 3 or Figure 5 the description in the data retrieval method section shown, which will not be elaborated here.

[0198] For the explanation of the relevant content in this embodiment, reference can be made to the above method embodiment, which will not be elaborated here.

[0199] Another embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored. When the instructions run on the data retrieval device, the data retrieval device executes each step executed by the data retrieval device in the method flow shown in the above method embodiment.

[0200] In another embodiment of the present invention, a computer program product is further provided. The computer program product includes computer execution instructions, and the computer execution instructions are stored in a computer-readable storage medium; at least one processor of the data retrieval device can read the computer execution instructions from the computer-readable storage medium, and at least one processor executes the computer execution instructions to enable the data retrieval device to execute each step executed by the data retrieval device in the method flow shown in the above method embodiment.

[0201] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can appear in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0202] The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data terminal such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0203] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0204] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0205] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0206] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0207] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disc and other various media that can store program codes.

[0208] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described.

Claims

1. A data retrieval method for use in the field of image processing, characterized in that, Including: Obtaining the feature information of the target data and the feature information of each of the m candidate data, where m is a positive integer, and the target data includes image data; Using a preset algorithm, based on the feature information of the target data and the feature information of each candidate data, to determine m target similarities, the value of each target similarity includes n bits, and the values of the consecutive x bits starting from the highest bit are the same, x is less than n, and x and n are positive integers, and the value of n corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data; According to the values of the n - x bits in each target similarity, performing a first retrieval operation on the m candidate data to obtain a first retrieval result, where the n - x bits are the other bits except the x bits; The preset algorithm includes a first algorithm and a second algorithm; The step of using a preset algorithm, based on the feature information of the target data and the feature information of each candidate data, to determine m target similarities includes: According to the first algorithm, calculating the similarity between the feature information of the target data and the feature information of each candidate data to obtain m initial similarities; where the first algorithm is a similarity algorithm, and the value of each initial similarity includes n bits; according to the second algorithm, calculating the m initial similarities to obtain the m target similarities; The step of performing a first retrieval operation on the m candidate data according to the values of the n - x bits in each target similarity to obtain a first retrieval result includes: According to the values of the n - x bits in each target similarity, selecting k first similarities; k is a positive integer within a preset value range, the preset value range is from a first threshold to a second threshold, the first threshold is an integer greater than 1, and the second threshold is an integer less than or equal to m; the value of each of the k first similarities is less than the value of a second similarity, and the second similarity is any one of the m target similarities except the k first similarities; obtaining the index values of the candidate data corresponding to the k first similarities.

2. The data retrieval method according to claim 1, wherein For each target similarity, the n - x bits are divided into consecutive P bit segments, and the number of bits in at least one bit segment is greater than 1, and P is an integer greater than or equal to 2; The step of selecting k first similarities according to the values of the n - x bits in each target similarity includes: Performing a screening operation on each bit segment of the m target similarities in the order of first screening the bit segments located at the high positions and then screening the bit segments located at the low positions; When the number of the screened target similarities is within the preset value range, taking the screened target similarities as the k first similarities; Wherein, the screened target similarities include: the target similarities screened according to the first bit segment, or include: the target similarities screened according to the 1st to the a - th bit segments; the first bit segment is the bit segment with the highest position among the P bit segments, and a ∈ [2, P].

3. The data retrieval method according to claim 1, wherein For each of the target similarities, the n - x bits are divided into consecutive P sub - keys, where P is an integer greater than or equal to 2; Selecting k first similarities according to the values of the n - x bits in each of the target similarities, including: Starting from the first sub - key among the P sub - keys in the order from the highest bit to the lowest bit, perform the following processing until the k first similarities are selected: Determine the value of the i - th sub - key among the m target similarities and the number of times each value appears, where i ∈ [1, P]; Judge whether the number of times the first value appears is within the preset value range, where the first value is the minimum value of the values of the i - th sub - key among the m target similarities; If the number of times the first value appears is within the preset value range, then select the target similarity including the first value, and the target similarity including the first value is one of the k first similarities; If the number of times the first value appears is less than the first threshold, then judge whether the sum of the number of times the first value and the second value appears is within the preset value range, where the second value is the second - smallest value of the values of the i - th sub - key among the m target similarities; and so on, repeat the execution until the k first similarities are selected; If the number of times the first value appears is greater than the second threshold, then determine the value of the (i + 1) - th sub - key in the target similarity to which the first value belongs and the number of times each value of the (i + 1) - th sub - key appears; and so on, repeat the execution until the k first similarities are selected.

4. The data retrieval method according to any one of claims 1-3, characterized in that The feature information is the initial feature information or the target feature information, where the target feature information is obtained by performing feature processing on the initial feature information; Among them, the feature processing is at least one of feature dimension processing, feature equalization processing, or feature quantization processing; the feature dimension processing is used to process the data dimension of the initial feature information into a preset dimension; the feature equalization processing is used to equalize the first feature information, and the equalized first feature information is distributed according to a preset rule; the feature quantization processing is used to quantize the second feature information so that the storage space of the quantized second feature information is less than or equal to the storage space of the second feature information; the first feature information is the initial feature information or the feature information obtained by performing feature dimension processing on the initial feature information; the second feature information is the initial feature information, or the feature information obtained after feature dimension processing, or the feature information obtained after feature equalization processing.

5. The data retrieval method according to claim 4, wherein When the feature information is the target feature information, the data retrieval method further includes: Performing a second retrieval operation on the first retrieval result according to the initial feature information of the target data to obtain a second retrieval result.

6. A data retrieval device for use in the field of image processing, characterized in that, Including: An acquisition unit for acquiring the feature information of the target data and the feature information of each candidate data among m candidate data, where m is a positive integer, and the target data includes image data; A determination unit, configured to use a preset algorithm to determine m target similarities according to the feature information of the target data and the feature information of each candidate data obtained by the obtaining unit, where the value of each target similarity includes n bits, and the values of consecutive x bits starting from the highest bit are the same, x is less than n, and x and n are positive integers, and the value of n corresponds to the dimension of the feature information of the target data and the dimension of the feature information of the candidate data; A retrieval unit, configured to perform a first retrieval operation on the m candidate data according to the values of the n - x bits in each target similarity determined by the determination unit, to obtain a first retrieval result, where the n - x bits are the other bits except the x bits; The preset algorithm includes a first algorithm and a second algorithm; specifically, the determination unit is configured to: According to the first algorithm, calculate the similarity between the feature information of the target data and the feature information of each candidate data, to obtain m initial similarities; where the first algorithm is a similarity algorithm, and the value of each initial similarity includes n bits; according to the second algorithm, calculate the m initial similarities to obtain the m target similarities; The retrieval unit includes a selection module and an acquisition module; The selection module is configured to select k first similarities according to the values of the n - x bits in each target similarity, where k is a positive integer within a preset value range, the preset value range is from a first threshold to a second threshold, the first threshold is an integer greater than 1, and the second threshold is an integer less than or equal to m; the value of each of the k first similarities is less than the value of a second similarity, and the second similarity is any one of the m target similarities except the k first similarities; the acquisition module is configured to acquire the index values of the candidate data corresponding to the k first similarities selected by the selection module.

7. The data retrieval device according to claim 6, wherein For each target similarity, the n - x bits are divided into consecutive P bit segments, at least one of the bit segments has a number of bits greater than 1, and P is an integer greater than or equal to 2; specifically, the selection module is configured to: Perform a screening operation on each bit segment of the m target similarities in the order of first screening the bit segments located at the high position and then screening the bit segments located at the low position; When the number of the screened target similarities is within the preset value range, use the screened target similarities as the k first similarities; Wherein, the screened target similarities include: the target similarities screened according to the first bit segment, or include: the target similarities screened according to the 1st to the a - th bit segments; the first bit segment is the bit segment with the highest position among the P bit segments, and a ∈ [2, P].

8. The data retrieval device according to claim 6, wherein For each target similarity, the n - x bits are divided into consecutive P sub - keys, and P is an integer greater than or equal to 2; specifically, the selection module is configured to: Starting from the first sub-key among the P sub-keys in the order from the highest bit to the lowest bit, perform the following processing until the k first similarities are selected: Determine the value of the i-th sub-key among the m target similarities and the number of times each value appears, where i ∈ [1, P]; Judge whether the number of times the first value appears is within the preset value range, where the preset value range is from the first threshold to the second threshold, and the first value is: the minimum value of the value of the i-th sub-key among the m target similarities; If the number of times the first value appears is within the preset value range, then select the target similarity including the first value, and the target similarity including the first value is the k first similarities; If the number of times the first value appears is less than the first threshold, then judge whether the sum of the number of times the first value and the second value appears is within the preset value range, and the second value is: the second smallest value of the value of the i-th sub-key among the m target similarities; In this way, repeat the execution until the k first similarities are selected; If the number of times the first value appears is greater than the second threshold, then determine the value of the (i + 1)-th sub-key among the target similarities to which the first value belongs and the number of times each value of the (i + 1)-th sub-key appears; In this way, repeat the execution until the k first similarities are determined; 9. The data retrieval device according to any one of claims 6-8, characterized in that The feature information is the initial feature information or the target feature information, and the target feature information is obtained by performing feature processing on the initial feature information; Among them, the feature processing is at least one of feature dimension processing, feature balancing processing, or feature quantization processing; the feature dimension processing is used to process the data dimension of the initial feature information into a preset dimension; the feature balancing processing is used to balance the first feature information, and the balanced first feature information is distributed according to a preset rule; the feature quantization processing is used to quantize the second feature information so that the storage space of the quantized second feature information is less than or equal to the storage space of the second feature information; the first feature information is the initial feature information or the feature information obtained by performing feature dimension processing on the initial feature information; the second feature information is the initial feature information, or the feature information obtained after feature dimension processing, or the feature information obtained after feature balancing processing.

10. The data retrieval device according to claim 9, characterized in that, The feature information is the target feature information; The retrieval unit is further configured to perform a second retrieval operation on the first retrieval result according to the initial feature information of the target data to obtain a second retrieval result.

11. A data retrieval device, characterized in that, The data retrieval device includes: one or more processors and a communication interface; The communication interface is coupled to the one or more processors and is used to provide data for the one or more processors; when the one or more processors execute program instructions, the data retrieval device executes the data retrieval method according to any one of claims 1-5.

12. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the data retrieval device, the data retrieval device is caused to execute the data retrieval method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image retrieval method and device

    CN106886599A

  • Image retrieval method and device, graphic processor and storage medium

    CN109614510A