Information Retrieval Method and Related Device

By selecting the target clustering center and its candidate coding set in the HNSW algorithm and quantizing the search information, the problem of low retrieval efficiency in the prior art is solved, and more efficient and accurate information retrieval is achieved.

CN116226313BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111467761.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-06-20
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The existing HNSW algorithm based on quantitative encoding requires a large amount of decoding and floating-point calculations in information retrieval, resulting in low retrieval efficiency, especially when facing massive and high-dimensional data.

Method used

By obtaining the similarity between multiple candidate cluster centers and the information to be retrieved, select the target cluster center, and obtain its corresponding candidate code set. Quantitatively encode the search information, calculate its similarity to the candidate encoding, and select the target encoding to obtain the search results.

Benefits of technology

Reduce the number of decoding and floating-point calculations, reduce the calculation cost, and improve the efficiency and accuracy of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116226313B_ABST
    Figure CN116226313B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technologies, and provides an information retrieval method and related apparatus to improve the retrieval efficiency. The method includes: selecting at least one target clustering center from a plurality of candidate clustering centers according to first configuration information including the plurality of candidate clustering centers; then, after quantizing and encoding the information to be retrieved to obtain a to-be-retrieved code according to second configuration information including candidate code sets respectively corresponding to the respective target clustering centers, selecting at least one target code from the respective candidate codes; and further obtaining a retrieval result according to the selected target codes. In this way, not only the amount of data for similarity comparison is reduced, the retrieval efficiency is improved, but also the retrieval accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of big data and natural language processing technologies, in the face of massive and high-dimensional data, in order to achieve information retrieval quickly and accurately, the Approximate Nearest Neighbor (ANN) technology has gradually become the focus of attention.

[0003] In related technologies, ANN usually uses the Hierarchical Navigable Small World graphs (HNSW) algorithm based on quantization coding to cluster data.

[0004] The so-called HNSW based on quantization coding means that: through quantization coding, encoded data with a relatively short storage length is used to represent the original data, and by using the characteristic of the clustered distribution between data, according to the distance between each encoded data, the encoded data is classified. In this way, when performing information retrieval, after decoding each encoded data to obtain the original data, the data category to which the retrieval statement belongs can be determined according to the distance between the original data and the retrieval statement, and some or all of the data included in this data category can be returned as the retrieval result.

[0005] However, if the above-mentioned HNSW based on quantization coding is used for retrieval, when determining the data category to which the retrieval statement belongs, a large number of decodings and multiple floating-point calculations need to be performed for each encoded data, which affects the retrieval efficiency. Especially in the face of massive and high-dimensional data, the retrieval efficiency is extremely low. Summary of the Invention

[0006] Embodiments of the present application provide an information retrieval method and related devices to reduce the computational cost and improve the information retrieval efficiency.

[0007] In a first aspect, embodiments of the present application provide an information retrieval method, including:

[0008] Obtain the information to be retrieved;

[0009] Obtain first configuration information including a plurality of candidate clustering centers, and select at least one target clustering center based on the first similarity between each of the plurality of candidate clustering centers and the information to be retrieved; wherein each candidate clustering center corresponds to a data category;

[0010] Obtain second configuration information including candidate code sets respectively corresponding to the at least one target clustering center; wherein each candidate code set is encoded based on each information to be recalled belonging to the corresponding data category;

[0011] Quantize and encode the information to be retrieved to obtain a retrieval code to be retrieved, and select at least one target code based on the second similarity between each candidate code included in each candidate code set and the retrieval code to be retrieved;

[0012] Obtain a retrieval result based on the at least one target code.

[0013] In a second aspect, an embodiment of the present application provides an information retrieval device, including:

[0014] An encoding unit, configured to obtain information to be retrieved;

[0015] A first matching unit, configured to obtain first configuration information including a plurality of candidate clustering centers, and select at least one target clustering center based on the first similarity between each of the plurality of candidate clustering centers and the information to be retrieved; wherein, each candidate clustering center corresponds to a data category;

[0016] An obtaining unit, configured to obtain second configuration information including candidate code sets respectively corresponding to the at least one target clustering center; wherein, each candidate code set is encoded based on each information to be recalled belonging to the corresponding data category;

[0017] A second matching unit, configured to quantize and encode the information to be retrieved to obtain a retrieval code to be retrieved, and select at least one target code based on the second similarity between each candidate code included in each candidate code set and the retrieval code to be retrieved;

[0018] A recall unit, configured to obtain a retrieval result based on the at least one target code.

[0019] As a possible implementation manner, the first matching unit is further configured to determine the plurality of candidate clustering centers in the following manner:

[0020] Obtain each information to be recalled;

[0021] Select a plurality of first clustering centers based on a specified number of data categories and based on the third similarity between each of the information to be recalled;

[0022] Cluster each of the information to be recalled based on the plurality of first clustering centers to obtain a set of information to be recalled respectively corresponding to each of the plurality of first clustering centers;

[0023] Determine second clustering centers respectively corresponding to the plurality of obtained sets of information to be recalled based on each of the information to be recalled included in the plurality of obtained sets of information to be recalled;

[0024] Use the determined plurality of second clustering centers as the plurality of candidate clustering centers.

[0025] As a possible implementation manner, when selecting multiple first clustering centers based on the specified number of data categories and based on the third similarity between each of the information to be recalled, the first matching unit is specifically configured to:

[0026] Iteratively perform the following operations based on the specified number of data categories until the multiple first clustering centers are obtained:

[0027] Select one piece of information to be recalled from each of the information to be recalled as the initial clustering center point;

[0028] Determine the third similarity between each of the other information to be recalled and the initial clustering center point;

[0029] Based on the determined third similarities, select one piece of information to be recalled that meets the similarity condition from each of the information to be recalled, and use the one piece of information to be recalled as a first clustering center.

[0030] As a possible implementation manner, when determining the third similarity between each of the other information to be recalled and the initial clustering center point, the first matching unit is specifically configured to:

[0031] Calculate the Euclidean distance between each of the other information to be recalled and the initial clustering center point, and use the sum of the obtained Euclidean distances as the total distance;

[0032] Based on the Euclidean distances, the total distance, and the total number of information of each of the information to be recalled, determine the third similarity between each of the other information to be recalled and the initial clustering center point.

[0033] As a possible implementation manner, when selecting one piece of information to be recalled that meets the similarity condition from each of the information to be recalled based on the determined third similarities, the first matching unit is specifically configured to:

[0034] If there is one piece of information to be recalled whose similarity accumulation value is greater than the similarity threshold among each of the information to be recalled, determine that the one piece of information to be recalled meets the similarity condition;

[0035] Wherein, the similarity accumulation value is used to represent the accumulation from the third similarity corresponding to the first piece of information to be recalled to the third similarity corresponding to the one piece of information to be recalled, the similarity threshold is obtained based on the total similarity accumulation value and the specified weight, and the total similarity accumulation value is used to represent the accumulation of the third similarities corresponding to each of the information to be recalled.

[0036] As a possible implementation manner, when using the determined multiple second cluster centers as the multiple candidate cluster centers, the first matching unit specifically is configured to:

[0037] Based on the first cluster center set corresponding to each information to be recalled, and based on the second cluster center set corresponding to each information to be recalled, determine the first measurement value and the second measurement value corresponding to each information to be recalled;

[0038] If the first measurement value corresponding to each information to be recalled is greater than the corresponding second measurement value, then use the multiple second cluster centers as the multiple candidate cluster centers; otherwise, based on the multiple second cluster centers, determine new cluster centers, and based on the new cluster centers, obtain the multiple candidate cluster centers;

[0039] Wherein, each first measurement value is used to represent the distance between each second cluster center in the corresponding second cluster center set, and each second measurement value is used to represent the sum of the distances between the corresponding first cluster center set and the second cluster center set.

[0040] As a possible implementation manner, when quantizing and encoding the information to be retrieved to obtain the retrieved code, the second matching unit specifically is configured to:

[0041] Obtain the sub-information to be retrieved corresponding to each feature dimension in the information to be retrieved;

[0042] Based on the specified value range corresponding to each feature dimension, respectively perform quantization encoding on the corresponding sub-information to be retrieved, and obtain the encoded sub-information to be retrieved;

[0043] Based on the obtained encoded sub-information to be retrieved, obtain the retrieved code.

[0044] As a possible implementation manner, when obtaining the retrieval result based on the at least one target code, the recall unit specifically is configured to:

[0045] Decode the at least one target code to obtain the information to be recalled corresponding to each of the at least one target code;

[0046] Based on the fourth similarity between each of the at least one information to be recalled and the information to be retrieved, determine at least one target recall information from the at least one information to be recalled.

[0047] As a possible implementation manner, when decoding the at least one target code to obtain at least one information to be recalled, the recall unit specifically is configured to:

[0048] Obtain the sub-codes corresponding to each feature dimension in the at least one target code;

[0049] Based on the specified value ranges corresponding to each feature dimension, decode the corresponding sub-codes respectively to obtain the corresponding information to be recalled.

[0050] As a possible implementation manner, when determining at least one target recall information from the at least one information to be recalled based on the fourth similarity between each of the at least one information to be recalled and the information to be retrieved, the recall unit is specifically configured to:

[0051] Determine the fourth similarity between each of the at least one information to be recalled and the information to be retrieved;

[0052] Based on the specified recall number, determine at least one target recall information from the at least one information to be recalled according to the values of each fourth similarity.

[0053] As a possible implementation manner, the first configuration information includes the connection relationships between the multiple candidate cluster centers, where the number of other candidate cluster centers connected to each candidate cluster center does not exceed the specified connection number.

[0054] As a possible implementation manner, the second configuration information includes multiple inverted lists, and each inverted list includes a set of candidate codes corresponding to a corresponding candidate cluster center.

[0055] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above information retrieval method.

[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above information retrieval method.

[0057] In a fifth aspect, an embodiment of the present application provides a computer program product, the program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the above information retrieval method.

[0058] In the embodiments of the present application, first configuration information including multiple candidate clustering centers is obtained. Based on the first similarity between each of the multiple candidate clustering centers and the information to be retrieved, at least one target clustering center is selected. Then, second configuration information including the candidate coding sets respectively corresponding to the target clustering centers is obtained. After quantifying and coding the information to be retrieved to obtain the retrieved code, at least one target code is selected according to the second similarity between each candidate code and the retrieved code. Furthermore, according to the selected target codes, a retrieval result is obtained.

[0059] In this way, on the one hand, through the first configuration information including multiple candidate clustering centers, the data category to which the information to be retrieved belongs can be determined, thereby reducing the number of candidate codes for similarity comparison with the retrieved code, improving the retrieval efficiency. At the same time, information retrieval based on clustering improves the retrieval accuracy. On the other hand, compared with comparing the decoded candidate codes with the information to be retrieved, in the embodiments of the present application, by quantifying and coding the information to be retrieved and comparing the retrieved code with the candidate codes, the time consumption caused by decoding a large amount of data is avoided, the retrieval time is reduced, and the retrieval efficiency is further improved.

[0060] Other features and advantages of the present application will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0062] Figure 1 It is a schematic diagram of an application scenario provided in the embodiments of the present application;

[0063] Figure 2 It is a schematic diagram of the interaction between a terminal device and a server provided in the embodiments of the present application;

[0064] Figure 3 It is a schematic flowchart of an index construction method provided in the embodiments of the present application;

[0065] Figure 4 It is a schematic flowchart of a method for determining candidate clustering centers provided in the embodiments of the present application;

[0066] Figure 5 It is a schematic diagram of the vector representation of the information to be recalled provided in the embodiments of the present application;

[0067] Figure 6 It is a schematic flow chart for determining a first clustering center provided in an embodiment of the present application;

[0068] Figure 7 It is a schematic diagram of a third similarity between pieces of information to be recalled provided in an embodiment of the present application;

[0069] Figure 8 It is a schematic diagram of a first clustering center provided in an embodiment of the present application;

[0070] Figure 9 It is a schematic diagram of a second clustering center provided in an embodiment of the present application;

[0071] Figure 10 It is a schematic diagram of HNSW provided in an embodiment of the present application;

[0072] Figure 11 It is a schematic diagram of an inverted index provided in an embodiment of the present application;

[0073] Figure 12 It is a schematic diagram of another inverted index provided in an embodiment of the present application;

[0074] Figure 13 It is a schematic flow chart of an information retrieval method provided in an embodiment of the present application;

[0075] Figure 14 It is a schematic diagram of a second similarity between a code to be recalled and a code to be retrieved provided in an embodiment of the present application;

[0076] Figure 15 It is a schematic flow chart for determining target recalled information provided in an embodiment of the present application;

[0077] Figure 16 It is a schematic logic diagram for determining target recalled information provided in an embodiment of the present application;

[0078] Figure 17 It is a schematic structural diagram of an information retrieval device provided in an embodiment of the present application;

[0079] Figure 18 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0080] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts belong to the scope protected by the technical solutions of this application.

[0081] With the rapid development of big data and natural language processing technologies, in the face of massive and high-dimensional feature vectors, how to quickly and accurately perform retrieval is a problem that must be solved in processing large-scale data. k-Nearest Neighbor (kNN) retrieval is a commonly used information retrieval technology. kNN retrieval selects the k most similar pieces of information from the information set to be recalled based on the distance between feature vectors. However, in the face of massive and high-dimensional data, since ideal retrieval effects and acceptable retrieval times cannot be obtained through kNN retrieval, ANN retrieval has gradually become the focus of attention. The core idea of ANN retrieval is to improve retrieval efficiency at the expense of precision within an acceptable range, rather than being limited to returning the k most similar pieces of data.

[0082] In related technologies, ANN usually uses HNSW based on quantization coding to cluster data. The so-called HNSW based on quantization coding means that through quantization coding, coding data with a relatively short storage length is used to represent the original data, and the cluster-like aggregation distribution characteristics between data are utilized to classify the coding data according to the distance between each coding data. In this way, when performing information retrieval, after decoding each coding data to obtain the original data, the data category to which the retrieval statement belongs can be determined based on the distance between the original data and the retrieval statement, and some or all of the data included in this data category can be returned as the retrieval result.

[0083] However, if the above-mentioned HNSW based on quantization coding is used for retrieval, when determining the data category to which the retrieval statement belongs, a large number of decodings and multiple floating-point calculations need to be performed for each coding data, which affects the retrieval efficiency. Especially in the face of massive and high-dimensional data, the retrieval efficiency is extremely low.

[0084] In order to improve the retrieval efficiency and retrieval accuracy, in the embodiments of the present application, after obtaining the information to be retrieved, first configuration information is obtained. The first configuration information contains a plurality of candidate clustering centers. Based on the first similarity between each of the plurality of candidate clustering centers and the information to be retrieved, at least one target clustering center can be selected. Then, second configuration information is obtained. The second configuration information contains a candidate coding set corresponding to each target clustering center. Then, the information to be retrieved is quantized and coded to obtain a retrieved code. Based on the second similarity between each candidate code and the retrieved code, at least one target code can be selected. Furthermore, based on the selected target code, a retrieval result is obtained.

[0085] In this way, on the one hand, through the first configuration information containing a plurality of candidate clustering centers, the data category to which the information to be retrieved belongs can be determined, thereby reducing the number of candidate codes for similarity comparison with the retrieved code, improving the retrieval efficiency. At the same time, information retrieval based on clustering improves the retrieval accuracy. On the other hand, compared with comparing the candidate code with the information to be retrieved after decoding, in the embodiments of the present application, by quantizing and coding the information to be retrieved and comparing the retrieved code with the candidate code, the time consumption caused by decoding a large amount of data is avoided, the retrieval time is reduced, and the retrieval efficiency is further improved.

[0086] The preferred embodiments of the present application will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0087] Refer to Figure 1 As shown, it is a schematic diagram of an application scenario provided in the embodiments of the present application. This application scenario includes at least a terminal device 101 and a server 102. The number of terminal devices 101 can be one or more, and the number of servers 102 can also be one or more. The present application does not make a specific limitation on the number of terminal devices 101 and servers 102.

[0088] A target application can be installed in the terminal device 101. Among them, the target application can be a client application, a web version application, a mini-program application, etc. In practical applications, the target application can be any application with an information retrieval function. The terminal device 101 can be a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The embodiments of the present application can be applied to various information retrieval scenarios, including but not limited to the navigation field, vehicle-mounted scenarios, cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0089] Server 102 may be the background server of the target application, providing corresponding services for the target application. Server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 101 and the server 102 may be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application.

[0090] In the embodiments of the present application, the information retrieval scenario may specifically be a recommendation scenario or a search scenario. In the recommendation scenario or the search scenario, the information to be retrieved may be one or more retrieval keywords input by the user, and the information to be recalled may be videos, audios, pictures, articles, expressions, etc., but is not limited thereto.

[0091] Refer to Figure 2 As shown, it is an interaction schematic diagram between a terminal device and a server provided in the embodiments of the present application.

[0092] S201. In response to the information to be retrieved input by the target object, the terminal device sends a retrieval request carrying the information to be retrieved to the server. The target object may refer to the user, or an account used by the user, but is not limited thereto. The target object may input one or more retrieval keywords through the operation interface of the terminal device.

[0093] S202. After receiving the retrieval request, the server determines at least one target recall information corresponding to the information to be retrieved from the set of information to be recalled based on the pre-stored first configuration information and second configuration information. The specific information retrieval process is described below.

[0094] S203. The server returns at least one target recall information to the terminal device.

[0095] It should be noted that the information retrieval method in the embodiments of the present application may be executed by the server or the terminal device, which is not limited in this regard.

[0096] The information retrieval method provided in the embodiments of the present application may be divided into two stages: the index construction stage and the information retrieval stage.

[0097] The index construction stage is used to construct the first configuration information including multiple candidate clustering centers, and the first configuration information may also be referred to as the HNSW index hereinafter, and to construct the second configuration information including the candidate coding sets corresponding to the multiple candidate clustering centers, and the second configuration information may also be referred to as the inverted index hereinafter.

[0098] The information retrieval stage is used to determine at least one target recalled information from each recalled information to be retrieved according to the first configuration information and the second configuration information after obtaining the information to be retrieved.

[0099] For ease of understanding, the index construction stage and the information retrieval stage will be described separately below.

[0100] Refer to Figure 3 shown, which is a schematic flowchart of an index construction method provided in an embodiment of the present application. The specific process is as follows:

[0101] S301. Determine k candidate cluster centers based on each recalled information to be retrieved.

[0102] It should be noted that in the embodiment of the present application, k represents the specified number of data categories. For example, the value of k is 6, that is, the specified number of data categories is 6.

[0103] As a possible implementation manner, a preset clustering algorithm can be used to cluster each recalled information to be retrieved to obtain k candidate cluster centers. The preset clustering algorithm can be but is not limited to K-Means clustering, mean shift clustering, agglomerative hierarchical clustering, etc.

[0104] As another possible implementation manner, to improve the clustering efficiency, refer to Figure 4 shown, when executing S301, the following steps can be adopted:

[0105] S3011. Obtain each recalled information to be retrieved.

[0106] In the embodiment of the present application, each recalled information to be retrieved can be represented in vector form. Use the vector data set D to represent the data set composed of each recalled information to be retrieved, D = {o1, o2, o3,..., o n}, where o i is the i-th recalled information to be retrieved, and the total number of information of the recalled information included in D is |D|.

[0107] In the vector data set D, the dimension of each recalled information to be retrieved is n, and each recalled information to be retrieved can be represented as (v1, v2, v3,..., v n ), where the value of n is a positive integer, and v1, v2, v3,..., v n are respectively used to represent the values of the recalled information to be retrieved in n dimensions. Hereinafter, the value of each dimension can also be called a sub-recalled information to be retrieved. Exemplarily, the value of each dimension can be a floating point number, stored in floating point (float) type data, and occupying 4 bytes (B) of storage space.

[0108] It should be noted that in the embodiments of the present application, after obtaining each information to be recalled, feature extraction needs to be performed on the information to be recalled to obtain corresponding feature vectors, so as to implement information retrieval based on the feature vectors. The information to be recalled in the following text is the corresponding feature vector.

[0109] For example, refer to Figure 5 As shown, assume that the vector dataset D contains information to be recalled 1, information to be recalled 2, and information to be recalled 3. The vector representations of information to be recalled 1, information to be recalled 2, and information to be recalled 3 are o1, o2, and o3 respectively. That is, the vector dataset D contains o1, o2, and o3. Among them, information to be recalled 1 is a beauty video, information to be recalled 2 is a food video, and information to be recalled 3 is an article. o1 is (v 11 , v 12 , v 13 , ……, v 1n ), o2 is (v 21 , v 22 , v 23 , ……, v 2n ), and o3 is (v 31 , v 32 , v 33 , ……, v 3n ).

[0110] S3012. Select k first cluster centers based on the specified number of data categories and based on the third similarity between each information to be recalled.

[0111] Specifically, refer to Figure 6 As shown, when executing S3012, the following steps can be adopted:

[0112] S30121. Select one information to be recalled from each information to be recalled as the initial cluster center point.

[0113] For the convenience of description, take the vector dataset D containing o1, o2, o3, ……, o |D| as an example. Among them, o1, o2, o3, ……, o |D| are used to represent each information to be recalled respectively.

[0114] In the embodiments of the present application, one information to be recalled can be randomly selected from each information to be recalled as the initial cluster center.

[0115] For example, randomly select o2 as the initial cluster center from o1, o2, o3, ……, o |D| .

[0116] S30122. Determine the third similarity between each other information to be recalled and the initial cluster center point.

[0117] In the embodiments of the present application, the third similarity between each other information to be recalled and the initial clustering center point can be determined by, but not limited to, the following methods:

[0118] Calculate the Euclidean distance between each other information to be recalled and the initial clustering center point, and take the sum of the obtained Euclidean distances as the total distance;

[0119] Based on each Euclidean distance, the total distance, and the total number of information of each information to be recalled, determine the third similarity between each other information to be recalled and the initial clustering center point.

[0120] Specifically, the Euclidean distance between each other information to be recalled and the initial clustering center point can be calculated by the following formula (1):

[0121] d(o c ,o a )=(v c1 -v a1 ) 2 +(v c2 -v a2 ) 2 +…+(v cn -v an ) 2 Formula (1)

[0122] Wherein, d(o c ,o a ) represents the Euclidean distance between any other information to be recalled and the initial clustering center point, o a represents any other information to be recalled, and o c represents the initial clustering center point.

[0123] The total distance d 总 can be calculated by the following formula (2):

[0124]

[0125] The third similarity between an other information to be recalled and the initial clustering center point can be calculated by the following formula (3):

[0126]

[0127] Wherein, P(o i ) represents the third similarity between any other information to be recalled and the initial clustering center point, and o i represents the i-th other information to be recalled in the vector dataset.

[0128] It should be noted that in the embodiments of the present application, the third similarity corresponding to the initial clustering center can also be calculated using the above formula (3).

[0129] For example, referring to Figure 7 as shown, assume that the vector data set D contains o1, o2, o3, o4, and the initial clustering center is o2. Then, the Euclidean distances between o1, o3, o4 and o2 are calculated respectively using the above formula (1). Among them, the Euclidean distance d(o2, o1) between o1 and o2 is 4, the Euclidean distance d(o2, o3) between o3 and o2 is 6, and the Euclidean distance d(o2, o4) between o4 and o2 is 10. Based on the Euclidean distances between o1, o3, o4 and o2, the total distance d 总 has a value of 20. Then, based on the Euclidean distance d(o2, o1) between o1 and o2, the total distance d 总 and the total number of information items 4 of each information to be recalled, the third similarities between o1, o3, o4 and o2 are obtained using the above formula (3). Among them, the third similarity between o1 and o2 the third similarity between o3 and o2 the third similarity between o4 and o2 the third similarity corresponding to o2 itself is

[0130] S30123. Based on the determined third similarities, select one information to be recalled that meets the similarity condition from each information to be recalled, and use the information to be recalled as a first clustering center.

[0131] In some embodiments, in order to improve the clustering accuracy, after randomly selecting the initial clustering center, a new clustering center, that is, the first clustering center, can be selected according to the similarities between each other information to be recalled and the initial clustering center. Hereinafter, only taking the information to be recalled m as an example, the determination process of the first clustering center will be described. The information to be recalled m is any one of the information to be recalled.

[0132] In the embodiments of the present application, if the information to be recalled m meets the following conditions, it is determined that the information to be recalled m meets the similarity condition: the cumulative similarity value of the information to be recalled m is greater than the similarity threshold, and the similarity threshold is obtained based on the total cumulative similarity value and the specified weight.

[0133] Among them, the cumulative similarity value of the information to be recalled m is used to represent the accumulation from the third similarity corresponding to the first information to be recalled to the third similarity corresponding to the information to be recalled m, and the total cumulative similarity value is used to represent the accumulation of the third similarities corresponding to each information to be recalled.

[0134] Specifically, the similarity accumulation value s1 of the information m to be recalled can be calculated using the following formula (4):

[0135]

[0136] Among them, m represents the position of the information m to be recalled in the vector dataset.

[0137] The similarity threshold s2 can be calculated using the following formula (5):

[0138]

[0139] Among them, rnd represents the specified weight, rnd is a random coefficient, and the value range of rnd is (0, 1).

[0140] It should be noted that in the embodiments of the present application, in order to ensure the clustering efficiency, the information to be recalled that meets the similarity condition can be directly used as the first clustering center, and no further screening is performed on the remaining information to be recalled.

[0141] Still taking the vector dataset D containing o1, o2, o3, o4 as an example, according to formula (4), the similarity accumulation value corresponding to o1 is P(o1) = 0.225, the similarity accumulation value corresponding to o2 is P(o1) + P(o2) = 0.225 + 0.125 = 0.35, the similarity accumulation value corresponding to o3 is P(o1) + P(o2) + P(o3) = 0.225 + 0.125 + 0.275 = 0.625, and the similarity accumulation value corresponding to o4 is P(o1) + P(o2) + P(o3) + P(o4) = 0.225 + 0.125 + 0.275 + 0.375 = 1. Assuming that the value of rnd is 0.5, according to formula (5), the similarity threshold is 0.5. Since the similarity accumulation value corresponding to o3 is greater than the similarity threshold, o3 is selected as a first clustering center.

[0142] S30124. Determine whether the number of the selected first clustering centers reaches k. If so, execute S30125; otherwise, return to execute S30121.

[0143] S30125. Output k first clustering centers.

[0144] For example, assuming that the value of k is 8, among o1, o2, o3, ……, o |D| o1, o2, o3, o4, o5, o6, o7, o8 are all first clustering centers.

[0145] S3013. Based on the k first clustering centers, cluster each piece of information to be recalled to obtain the sets of information to be recalled corresponding to the k first clustering centers respectively.

[0146] It should be noted that in the embodiments of the present application, among the set of information to be recalled corresponding to the first clustering center is the first clustering center.

[0147] For example, referring to Figure 8 as shown, assume that the first clustering centers are o1, o2, o3, o4. Among them, the set of information to be recalled corresponding to o1 includes o1, o5, o6, o7, the set of information to be recalled corresponding to o2 includes o2, o 10 , the set of information to be recalled corresponding to o3 includes o3, o8, o9, and the set of information to be recalled corresponding to o4 includes o4, o 11 , o 12 .

[0148] In the embodiments of the present application, when executing S3013, after obtaining k first clustering centers, an HNSW can be constructed for the k first clustering centers. Furthermore, based on the HNSW constructed for the k first clustering centers, according to the distances between each piece of information to be recalled and the k first clustering centers, the set of information to be recalled corresponding to each of the k first clustering centers can be obtained. For the construction process of the HNSW, refer to S302 below, and for the query process of the HNSW, refer to S1302 below.

[0149] S3014. Based on each piece of information to be recalled included in the obtained k sets of information to be recalled, determine the second clustering center corresponding to each of the k sets of information to be recalled.

[0150] In the embodiments of the present application, the average value of each piece of information to be recalled included in a set of information to be recalled can be used as the second clustering center corresponding to the set of information to be recalled.

[0151] Still taking the first clustering centers as o1, o2, o3, o4 as an example, referring to Figure 9 as shown, the set of information to be recalled corresponding to o1 includes o1, o5, o6, o7. The average value of o1, o5, o6, o7 is o1, and the second clustering center corresponding to the set of information to be recalled corresponding to o1 is o1. The set of information to be recalled corresponding to o2 includes o2, o 10 , o2, o 10 's average value is o 13 , and the second clustering center corresponding to the set of information to be recalled corresponding to o2 is o 13 , the set of information to be recalled corresponding to o3 includes o3, o8, o9. The average value of o3, o8, o9 is o 14 , and the second clustering center corresponding to the set of information to be recalled corresponding to o3 is o 14 , the set of information to be recalled corresponding to o4 includes o4, o 11 , o 12, o4, o 11 , o 12 The average value of o4 and o is o4, and the second clustering center corresponding to the set of information to be recalled corresponding to o4 is o4.

[0152] S3015. Take the determined k second clustering centers as k candidate clustering centers.

[0153] As a possible implementation, the determined k second clustering centers can be directly taken as k candidate clustering centers.

[0154] As another possible implementation, in order to improve the clustering effect and thus improve the retrieval accuracy, it is also possible to determine the first measurement value and the second measurement value corresponding to each piece of information to be recalled based on the set of first clustering centers corresponding to each piece of information to be recalled and the set of second clustering centers corresponding to each piece of information to be recalled. If the first measurement value corresponding to each piece of information to be recalled is greater than the corresponding second measurement value, then take the k second clustering centers as k candidate clustering centers; otherwise, determine new clustering centers based on the k second clustering centers, and obtain k candidate clustering centers based on the new clustering centers.

[0155] Among them, each first measurement value is used to represent the distance between the second clustering centers in the corresponding set of second clustering centers, and each second measurement value is used to represent the sum of the distances between the corresponding set of first clustering centers and the set of second clustering centers.

[0156] The set of first clustering centers and the set of second clustering centers corresponding to each piece of information to be recalled can both be the x clustering centers closest to the information to be recalled, and the number of x can be set according to the actual application scenario. In this article, the set of first clustering centers can also be understood as the set of clustering centers corresponding to the information to be recalled in the previous time, and the set of second clustering centers can also be understood as the set of clustering centers corresponding to the information to be recalled currently.

[0157] Below, only an example where both the set of first clustering centers and the set of second clustering centers contain two clustering centers will be used for illustration.

[0158] Specifically, d(u, v) can be used to represent the first measurement value, and d(u, u1) + d(v, v1) can be used to represent the second measurement value, where u1 and v1 represent the two clustering centers closest to the previous time, and u and v represent the two clustering centers closest to the current time.

[0159] When the value of d(u, v) corresponding to each piece of information to be recalled is greater than d(u, u1) + d(v, v1), take the k second clustering centers as k candidate clustering centers.

[0160] If, among the information to be recalled, there is information to be recalled x1 for which the value of d(u, v) is not greater than d(u, u1) + d(v, v1), then adjust the cluster center corresponding to the information to be recalled x1, and based on the adjusted set of k pieces of information to be recalled, determine new k cluster centers until the value of d(u, v) corresponding to each piece of information to be recalled is greater than d(u, u1) + d(v, v1). The process of determining the new k cluster centers is similar to S3013 - S3014 and will not be elaborated here. That is to say, in the implementation of this application, k candidate cluster centers can be obtained by repeatedly executing the process of S3013 - S3014.

[0161] S302. Obtain first configuration information including k candidate cluster centers according to the k candidate cluster centers.

[0162] In the embodiment of this application, the first configuration information includes the connection relationships between multiple candidate cluster centers, and the number of other candidate cluster centers connected to each candidate cluster center does not exceed a specified connection number.

[0163] The first configuration information can be HNSW, and HNSW can be constructed in the following way: For any one of the k candidate cluster centers, regard it as a node to be inserted. Then determine which layer the node to be inserted can fall into, and in each layer of the graph from this layer downwards, insert the node, and in each layer of the graph, search for m neighbor nodes of the node through a heuristic search algorithm and connect them.

[0164] Refer to Figure 10 As shown, assume that the value of k is 8. It can be understood that after inserting all k candidate cluster centers as nodes into the graph, a multi - layer NSW arranged from top to bottom can be formed. That is to say, the HNSW graph constructed based on the HNSW algorithm can include multiple layers of NSW arranged from top to bottom.

[0165] Figure 10 In, the HNSW graph includes 3 layers of NSW arranged from top to bottom with the number of nodes increasing in turn. Among them, the first layer of NSW includes nodes 1 - node 8, and these 8 nodes are the nodes corresponding to the k second cluster centers respectively. The second layer of NSW includes nodes 1, node 2, node 6, and node 8. The third layer of NSW includes nodes 1 and node 6.

[0166] In Figure 10 each layer of NSW, each node is connected to no more than m neighbor nodes. Exemplarily, the value of m can be 4. Taking node 1 in the first layer of NSW as an example, node 1 is respectively connected to nodes 3, node 4, and node 8. That is to say, in the first layer of NSW, node 1 is connected to 3 neighbor nodes, and nodes 3, node 4, and node 8 can be called the neighbor nodes of node 1.

[0167] Taking node 1 in the second - layer NSW as an example, it is respectively connected to node 2 and node 8. That is to say, in the second - layer NSW, node 1 is connected to two neighbor nodes. In other words, there are two neighbor nodes for node 1 in the second - layer NSW: node 2 and node 8.

[0168] In addition, in the HNSW graph, the NSW in the upper layer has fewer nodes and greater distances between nodes compared to the NSW in the lower layer. Therefore, when searching for a set number of target clustering centers in the top - down order, the target nodes can be roughly located first, and then a fine - grained search can be performed within the range of the rough location, thereby reducing the computational cost.

[0169] S303: Obtain the to - be - recalled information sets corresponding to each of the k candidate clustering centers, and encode the to - be - recalled information sets corresponding to each of the k candidate clustering centers to obtain a second configuration information including candidate encoding sets corresponding to each of the k candidate clustering centers.

[0170] In the embodiments of the present application, the second configuration information includes multiple inverted lists, and each inverted list includes a candidate encoding set corresponding to a corresponding candidate clustering center. The second configuration information can adopt, but is not limited to, an inverted index.

[0171] Refer to Figure 11 As shown, the inverted index includes k inverted lists, and each inverted list corresponds to a cluster, that is, the k inverted lists respectively correspond to K candidate clustering centers. Each inverted list can be represented by a corresponding list identifier. Exemplarily, the list identifier can be represented by a serial number (ID).

[0172] Among them, each inverted list includes each to - be - recalled information in the to - be - recalled information set corresponding to the corresponding candidate clustering center.

[0173] To reduce the system memory overhead, in the embodiments of the present application, after obtaining the to - be - recalled information sets corresponding to each of the k candidate clustering centers, the to - be - recalled information sets corresponding to each of the k candidate clustering centers can be encoded to compress the data set vectors and greatly reduce the memory occupancy.

[0174] When encoding the to - be - recalled information sets corresponding to each of the k candidate clustering centers, Product Quantization (PQ) encoding or Scalar Quantization (SQ) encoding can be used. Since the compression ratio of PQ encoding is usually 4 times that of SQ encoding, the error of PQ encoding is larger than that of SQ encoding during decoding.

[0175] In some embodiments, to ensure the retrieval accuracy, SQ coding can be used to encode the recall information sets corresponding to each of the k candidate clustering centers. Below, only the recall information k1 is taken as an example for illustration. The recall information k1 is any one of the recall information in the recall information set k, and the recall information set k is any one of the k recall information sets.

[0176] Specifically, obtain the recall sub-information corresponding to each feature dimension in the recall information k1, and based on the specified value range corresponding to each feature dimension, respectively perform quantization coding on the corresponding recall sub-information to obtain the encoded recall sub-information, and based on each of the obtained encoded retrieval sub-information, obtain the recall code.

[0177] Among them, the specified value range corresponding to each feature dimension is determined according to the values of all the recall information included in the recall information set k in this feature dimension, and the specified value range includes the maximum value and the minimum value.

[0178] For each dimension v i , record the maximum value vmaxi and the minimum value vmini of all the recall information included in the recall information set k in this dimension; assume that the encoding length in the SQ coding is 8 bits, then an encoded recall sub-information s i can be calculated using formula (6):

[0179]

[0180] For example, assume that v i is 20, vmaxi is 100, vmini is 0, and the encoded recall sub-information calculated using formula (6) is 51.

[0181] Refer to Figure 12 As shown, in the embodiments of the present application, the recall information sets corresponding to each of the k candidate clustering centers can be obtained first to obtain the second configuration information including the recall sets corresponding to each of the k candidate clustering centers. Then, the recall information sets corresponding to each of the k candidate clustering centers are encoded to obtain the recall code sets corresponding to each of the k candidate clustering centers. After that, the recall code sets corresponding to each of the k candidate clustering centers are used to replace the recall sets corresponding to each of the k candidate clustering centers respectively to obtain the second configuration information including the recall code sets corresponding to each of the k candidate clustering centers.

[0182] Next, based on the constructed first configuration information and the second configuration information, the information retrieval stage is described.

[0183] Refer to Figure 13As shown in the figure, it is a schematic flowchart of an information retrieval method provided in an embodiment of the present application. The specific process is as follows:

[0184] S1301. Obtain the information to be retrieved.

[0185] Specifically, when performing quantization encoding on the information to be retrieved, the sub-information to be retrieved corresponding to each feature dimension in the information to be retrieved can be obtained, and based on the specified value range corresponding to each feature dimension, the corresponding sub-information to be retrieved is respectively quantized and encoded to obtain the encoded sub-information to be retrieved, and based on the obtained encoded sub-information to be retrieved, the retrieval code is obtained.

[0186] S1302. Obtain the first configuration information including multiple candidate cluster centers, and select at least one target cluster center based on the first similarity between each of the multiple candidate cluster centers and the information to be retrieved; wherein, each candidate cluster center corresponds to a data category.

[0187] Among them, the first similarity can be represented by, but not limited to, Euclidean distance and cosine distance. The number of selected target cluster centers can be set according to the actual application scenario.

[0188] In some embodiments, after obtaining the first configuration information including multiple candidate cluster centers, at least one target cluster center can be determined by, but not limited to, the following method:

[0189] Starting from the top layer of HNSW, perform layer-by-layer search on each layer in the order from top to bottom until in the bottom layer, Q nearest neighbor nodes are searched from N nodes, and the Q candidate cluster centers corresponding to the Q nearest neighbor nodes are used as the target cluster centers.

[0190] The above layer-by-layer search can specifically include: using the starting node of the current layer as the initial current node, searching for the node closest to the information to be retrieved from the current node and the neighbor nodes having a connection relationship with the current node as the updated current node, determining the current node when the search end condition is reached as the first node, and entering the next layer via the first node. Using the determined first node as the starting node of the next layer.

[0191] It should be noted that when the current layer is the top layer, the above starting node can be any randomly selected node. The above search end condition can include: all neighbor nodes of the starting node have been searched, or the distance between the current node and the information to be retrieved is less than the distance between the neighbor nodes of the current node and the information to be retrieved.

[0192] Taking the search for one target cluster center as an example, refer to Figure 10As shown in the figure, first, a layer search is performed on the third layer. The starting node of the third layer is node 6, and the neighbor nodes of node 6 are node 1. Node 1 is taken as the updated current node. At this time, all the neighbor nodes of node 6 have been searched, and node 1 is taken as the starting node of the next layer. Then, a layer search is performed on the second layer. The starting node of the second layer is node 1, and the neighbor nodes of node 1 are node 2 and node 8. Assume that among node 2 and node 8, the node closest to the information to be retrieved is node 2. Then node 2 is taken as the updated current node. At this time, all the neighbor nodes of node 1 have been searched, and node 2 is taken as the starting node of the next layer. Finally, a layer search is performed on the first layer. The starting node of the second layer is node 2, and the neighbor nodes of node 2 are node 3, node 4, and node 5. Assume that among node 3, node 4, and node 5, the node closest to the information to be retrieved is node 3. At this time, all the neighbor nodes of node 2 have been searched, and the candidate clustering center corresponding to node 3 is taken as the target distance center.

[0193] It should be noted that in the embodiments of the present application, since each node in the HNSW graph corresponds to a candidate clustering center, and the data of the candidate clustering center can specifically be a vector, the above search for the node closest to the information to be retrieved can also be understood as a process of calculating the distance between two vectors.

[0194] S1303. Obtain second configuration information including candidate coding sets respectively corresponding to at least one target clustering center; wherein each candidate coding set is encoded based on each piece of information to be recalled belonging to the corresponding data category.

[0195] In this article, the set of information to be recalled corresponding to the target clustering center is called the candidate coding set.

[0196] S1304. Quantize and encode the information to be retrieved to obtain a retrieved coding, and based on the second similarities between each candidate coding included in each candidate coding set and the retrieved coding, select at least one target coding.

[0197] In the embodiments of the present application, the number of selected target codings can be preset, and the second similarity can be calculated using formula (1).

[0198] For example, as shown in Figure 14 the figure, assume that the number of selected target codings is 2, and the candidate codings include: to-be-recalled coding 11, to-be-recalled coding 12, to-be-recalled coding 13. The second similarities between to-be-recalled coding 11, to-be-recalled coding 12, to-be-recalled coding 13 and the retrieved coding are 85%, 90%, and 95% respectively. From to-be-recalled coding 11, to-be-recalled coding 12, to-be-recalled coding 13, select to-be-recalled coding 12 and to-be-recalled coding 13 as the target codings.

[0199] It should be noted that when performing quantization encoding on the retrieval information, the specified value ranges of the respective feature dimensions corresponding to each target cluster center can be used to perform quantization encoding on the retrieval information to obtain respective retrieval encodings to be retrieved. Then, based on each candidate encoding included in each candidate encoding set, the second similarity with the corresponding retrieval encoding to be retrieved is used to select at least one target encoding. Among them, the specified value ranges of the respective feature dimensions corresponding to each target cluster center are determined according to the information set to be recalled corresponding to the target cluster center. Since the quantization encoding method for the retrieval information to be retrieved is similar to the quantization encoding method in the index construction stage, it will not be elaborated here.

[0200] S1305. Obtain a retrieval result based on at least one target encoding.

[0201] Specifically, referring to Figure 15 as shown, when executing S1305, the following steps can be adopted:

[0202] S13051. Decode at least one target encoding to obtain the information to be recalled corresponding to each of the at least one target encoding.

[0203] To improve the retrieval accuracy, in some embodiments, after selecting at least one target encoding, the at least one target encoding can be decoded to obtain the information to be recalled corresponding to each of the at least one target encoding, and based on the fourth similarity between each of the at least one information to be recalled and the retrieval information to be retrieved, at least one target recall information is determined from the at least one information to be recalled.

[0204] Specifically, when decoding at least one target encoding, the following method can be adopted: obtain the sub-encodings corresponding to the respective feature dimensions in the at least one target encoding, and based on the specified value ranges corresponding to the respective feature dimensions, decode the corresponding sub-encodings respectively to obtain the corresponding information to be recalled.

[0205] The decoded sub-encoding can be calculated using the following formula (7):

[0206]

[0207] where v i ′ represents the decoded sub-encoding, and s i represents the sub-encoding corresponding to a feature dimension.

[0208] For example, assume that s i is 51, vmaxi is 100, and vmini is 0. Using formula (7), the retrieved sub-information after encoding is calculated to be 20.

[0209] S13052. Determine at least one target recall information from at least one recall information candidate based on the fourth similarity between each of the at least one recall information candidate and the information to be retrieved.

[0210] Specifically, when executing S13052, the following steps may be adopted:

[0211] S130521. Determine the fourth similarity between each of the at least one recall information candidate and the information to be retrieved.

[0212] It should be noted that in the embodiments of the present application, the fourth similarity can be calculated using formula (1).

[0213] S130522. Based on the specified recall number, determine at least one target recall information from the at least one recall information candidate according to the values of the respective fourth similarities.

[0214] Specifically, the at least one recall information candidate can be sorted according to the values of the respective fourth similarities, and then, based on the sorting result and the specified recall number, at least one target recall information is determined from the at least one recall information candidate.

[0215] For example, as shown in Figure 16 Suppose the specified recall number is 1,000, and the recall information candidates corresponding to the target encoding include: recall information 11, recall information 12, recall information 13, recall information 14,..., recall information 31, recall information 32, etc. The recall information candidates are sorted in descending order according to the values of the fourth similarities, and the sorting result is in turn: recall information 13, recall information 12, recall information 31, recall information 32,..., recall information 13, recall information 14, etc. Based on the specified recall number, the first 1,000 recall information candidates starting from recall information 13 are used as the target recall information.

[0216] Based on the same inventive concept, an information retrieval device is provided in the embodiments of the present application. As shown in Figure 17 The structural schematic diagram of the information retrieval device 1700 may include:

[0217] An encoding unit 1701, configured to obtain the information to be retrieved;

[0218] A first matching unit 1702, configured to obtain first configuration information including a plurality of candidate cluster centers, and select at least one target cluster center based on the first similarity between each of the plurality of candidate cluster centers and the information to be retrieved; where each candidate cluster center corresponds to a data category;

[0219] An acquisition unit 1703, configured to acquire second configuration information including a candidate coding set corresponding to each of the at least one target clustering center; wherein each candidate coding set is obtained by coding each recall information belonging to the corresponding data category;

[0220] A second matching unit 1704, configured to perform quantization coding on the information to be retrieved to obtain a to-be-retrieved code, and select at least one target code based on second similarities between each candidate code included in each candidate coding set and the to-be-retrieved code;

[0221] A recall unit 1705, configured to obtain a retrieval result based on the at least one target code.

[0222] As a possible implementation manner, the first matching unit 1702 is further configured to determine the multiple candidate clustering centers in the following manner:

[0223] Acquire each recall information;

[0224] Select multiple first clustering centers based on a specified number of data categories and a third similarity between each of the recall information;

[0225] Cluster each of the recall information based on the multiple first clustering centers to obtain a recall information set corresponding to each of the multiple first clustering centers;

[0226] Determine a second clustering center corresponding to each of the multiple obtained recall information sets based on each recall information included in each of the multiple obtained recall information sets;

[0227] Use the determined multiple second clustering centers as the multiple candidate clustering centers.

[0228] As a possible implementation manner, when selecting multiple first clustering centers based on a specified number of data categories and a third similarity between each of the recall information, the first matching unit 1702 is specifically configured to:

[0229] Based on the specified number of data categories, iteratively perform the following operations until the multiple first clustering centers are obtained:

[0230] Select one recall information from each of the recall information as an initial clustering center point;

[0231] Determine a third similarity between each of the other recall information and the initial clustering center point;

[0232] Based on the determined third similarities, select one piece of information to be recalled that meets the similarity condition from the pieces of information to be recalled, and use the one piece of information to be recalled as a first clustering center.

[0233] As a possible implementation manner, when determining the third similarity between each other piece of information to be recalled and the initial clustering center point, the first matching unit 1702 is specifically configured to:

[0234] Calculate the Euclidean distance between each other piece of information to be recalled and the initial clustering center point, and use the sum of the obtained Euclidean distances as the total distance;

[0235] Based on the Euclidean distances, the total distance, and the total number of pieces of information of the pieces of information to be recalled, determine the third similarity between each other piece of information to be recalled and the initial clustering center point.

[0236] As a possible implementation manner, when selecting one piece of information to be recalled that meets the similarity condition from the pieces of information to be recalled based on the determined third similarities, the first matching unit 1702 is specifically configured to:

[0237] If there is one piece of information to be recalled among the pieces of information to be recalled whose similarity accumulation value is greater than the similarity threshold, determine that the one piece of information to be recalled meets the similarity condition;

[0238] Wherein, the similarity accumulation value is used to represent the accumulation from the third similarity corresponding to the first piece of information to be recalled to the third similarity corresponding to the one piece of information to be recalled, the similarity threshold is obtained based on the total similarity accumulation value and a specified weight, and the total similarity accumulation value is used to represent the accumulation of the third similarities corresponding to the pieces of information to be recalled.

[0239] As a possible implementation manner, when using the determined multiple second clustering centers as the multiple candidate clustering centers, the first matching unit 1702 is specifically configured to:

[0240] Based on the set of first clustering centers corresponding to each piece of information to be recalled, and based on the set of second clustering centers corresponding to each piece of information to be recalled, determine the first measurement value and the second measurement value corresponding to each piece of information to be recalled;

[0241] If the first measurement values corresponding to the pieces of information to be recalled are all greater than the corresponding second measurement values, use the multiple second clustering centers as the multiple candidate clustering centers; otherwise, based on the multiple second clustering centers, determine new clustering centers, and based on the new clustering centers, obtain the multiple candidate clustering centers;

[0242] Among them, each first measurement value is used to characterize the distance between each second clustering center in the corresponding second clustering center set, and each second measurement value is used to characterize the sum of the distances between the corresponding first clustering center set and the second clustering center set.

[0243] As a possible implementation manner, when quantizing and encoding the information to be retrieved to obtain the encoded information to be retrieved, the second matching unit 1704 is specifically configured to:

[0244] Obtain the sub-information to be retrieved corresponding to each feature dimension in the information to be retrieved;

[0245] Based on the specified value range corresponding to each feature dimension, respectively perform quantization encoding on the corresponding sub-information to be retrieved to obtain the encoded sub-information to be retrieved;

[0246] Based on the obtained encoded sub-information to be retrieved, obtain the encoded information to be retrieved.

[0247] As a possible implementation manner, when obtaining the retrieval result based on the at least one target code, the recall unit 1705 is specifically configured to:

[0248] Decode the at least one target code to obtain the information to be recalled corresponding to each of the at least one target code;

[0249] Based on the fourth similarity between each of the at least one information to be recalled and the information to be retrieved, determine at least one target recalled information from the at least one information to be recalled.

[0250] As a possible implementation manner, when decoding the at least one target code to obtain at least one information to be recalled, the recall unit 1705 is specifically configured to:

[0251] Obtain the sub-code corresponding to each feature dimension in the at least one target code;

[0252] Based on the specified value range corresponding to each feature dimension, respectively decode the corresponding sub-code to obtain the corresponding information to be recalled.

[0253] As a possible implementation manner, when determining at least one target recalled information from the at least one information to be recalled based on the fourth similarity between each of the at least one information to be recalled and the information to be retrieved, the recall unit 1705 is specifically configured to:

[0254] Determine the fourth similarity between each of the at least one information to be recalled and the information to be retrieved;

[0255] Based on the specified recall number, at least one target recall information is determined from the at least one information to be recalled according to the values of the respective fourth similarities.

[0256] As a possible implementation manner, the first configuration information includes the connection relationships between the multiple candidate cluster centers, where the number of other candidate cluster centers connected to each candidate cluster center does not exceed a specified connection number.

[0257] As a possible implementation manner, the second configuration information includes multiple inverted lists, and each inverted list includes a candidate coding set corresponding to a corresponding candidate cluster center.

[0258] For the convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.

[0259] Regarding the device in the above embodiments, the specific manners in which each unit executes requests have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0260] Those skilled in the art to which the present application pertains can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0261] After introducing the information retrieval method and device of the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.

[0262] Figure 18 is a block diagram of an electronic device 1800 shown according to an exemplary embodiment. The device includes:

[0263] A processor 1810;

[0264] A memory 1820 for storing executable instructions of the processor 1810;

[0265] Wherein, the processor 1810 is configured to execute instructions to implement the information retrieval method in the embodiments of the present application, such as Figure 3 、 Figure 4 、 Figure 6 、 Figure 13 or Figure 15 the steps shown in.

[0266] In an exemplary embodiment, a storage medium including operations is also provided, such as a memory 1820 including operations, and the operations can be executed by a processor 1810 of an electronic device 1800 to complete the above method. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Portable Compact Disk Read Only Memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0267] Based on the same inventive concept, the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the information retrieval method provided in various optional implementation manners in the above embodiments.

[0268] In some possible implementation manners, various aspects of the information retrieval method provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on a computer device, the computer program is used to cause the computer device to execute the steps in the information retrieval method according to various exemplary embodiments described in this specification. For example, the computer device can execute as Figure 3 , Figure 4 , Figure 6 , Figure 13 or Figure 15 shown in the steps.

[0269] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0270] The program product of an embodiment of the present application may be in the form of a CD-ROM and include program code, and can run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.

[0271] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device. Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0272] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. An information retrieval method, characterized in that, The method includes: Obtaining the information to be retrieved; Obtaining first configuration information including a plurality of candidate clustering centers, and selecting at least one target clustering center based on the first similarity between each of the plurality of candidate clustering centers and the information to be retrieved; wherein each candidate clustering center corresponds to a data category; Obtaining second configuration information including candidate code sets corresponding to each of the at least one target clustering center; wherein each candidate code set is encoded based on each information to be recalled belonging to the corresponding data category; Quantifying and encoding the information to be retrieved to obtain a retrieved code, and selecting at least one target code based on the second similarity between each candidate code included in each candidate code set and the retrieved code; Obtaining a retrieval result based on the at least one target code.

2. The method according to claim 1, characterized in that, The plurality of candidate clustering centers are determined by the following method: Obtaining each information to be recalled; Selecting a plurality of first clustering centers based on a specified number of data categories and a third similarity between each of the information to be recalled; Clustering each of the information to be recalled based on the plurality of first clustering centers to obtain sets of information to be recalled corresponding to each of the plurality of first clustering centers; Determining second clustering centers corresponding to each of the obtained sets of information to be recalled based on each of the information to be recalled included in the plurality of obtained sets of information to be recalled; Taking the determined plurality of second clustering centers as the plurality of candidate clustering centers.

3. The method according to claim 2, characterized in that, The selecting a plurality of first clustering centers based on a specified number of data categories and a third similarity between each of the information to be recalled includes: Based on the specified number of data categories, iteratively perform the following operations until the plurality of first clustering centers are obtained: Selecting one of the information to be recalled as an initial clustering center point from each of the information to be recalled; Determining the third similarity between each of the other information to be recalled and the initial clustering center point; Based on the determined third similarities, selecting one of the information to be recalled that meets the similarity condition from each of the information to be recalled, and taking the one of the information to be recalled as a first clustering center.

4. The method according to claim 3, characterized in that, The determining the third similarity between each of the other information to be recalled and the initial clustering center point includes: Calculating the Euclidean distance between each of the other information to be recalled and the initial clustering center point, and taking the sum of the obtained Euclidean distances as the total distance; Based on the Euclidean distances, the total distance, and the total number of information of each of the information to be recalled, determining the third similarity between each of the other information to be recalled and the initial clustering center point.

5. The method according to claim 3, characterized in that, The selecting one of the information to be recalled that meets the similarity condition from each of the information to be recalled based on the determined third similarities includes: If there is one of the information to be recalled whose similarity accumulation value is greater than the similarity threshold among each of the information to be recalled, determining that the one of the information to be recalled meets the similarity condition; Among them, the similarity accumulation value is used to represent the accumulation from the third similarity corresponding to the first information to be recalled to the third similarity corresponding to the information to be recalled. The similarity threshold is obtained based on the total similarity accumulation value and a specified weight, and the total similarity accumulation value is used to represent the accumulation of the third similarities corresponding to each of the information to be recalled.

6. The method according to any one of claims 2-5, characterized in that, The step of using the determined multiple second cluster centers as the multiple candidate cluster centers includes: Determining the first measurement value and the second measurement value corresponding to each of the information to be recalled based on the first cluster center set corresponding to each of the information to be recalled and the second cluster center set corresponding to each of the information to be recalled; If the first measurement value corresponding to each of the information to be recalled is greater than the corresponding second measurement value, then using the multiple second cluster centers as the multiple candidate cluster centers; otherwise, determining new cluster centers based on the multiple second cluster centers, and obtaining the multiple candidate cluster centers based on the new cluster centers; Among them, each first measurement value is used to represent the distance between the second cluster centers in the corresponding second cluster center set, and each second measurement value is used to represent the sum of the distances between the corresponding first cluster center set and the second cluster center set.

7. The method according to any one of claims 1-5, characterized in that, The step of quantifying and encoding the information to be retrieved to obtain a retrieved code includes: Obtaining the sub-information to be retrieved corresponding to each feature dimension in the information to be retrieved; Quantifying and encoding the corresponding sub-information to be retrieved respectively based on the specified value range corresponding to each feature dimension to obtain the encoded sub-information to be retrieved; Obtaining a retrieved code based on the obtained encoded sub-information to be retrieved.

8. The method according to any one of claims 1-5, characterized in that, The step of obtaining a retrieval result based on the at least one target code includes: Decoding the at least one target code to obtain the information to be recalled corresponding to each of the at least one target code; Determining at least one target recalled information from the at least one information to be recalled based on the fourth similarity between each of the at least one information to be recalled and the information to be retrieved.

9. The method according to claim 8, wherein The step of decoding the at least one target code to obtain at least one information to be recalled includes: Obtaining the sub-codes corresponding to each feature dimension in the at least one target code; Decoding the corresponding sub-codes respectively based on the specified value range corresponding to each feature dimension to obtain the corresponding information to be recalled.

10. The method according to claim 8, wherein The step of determining at least one target recalled information from the at least one information to be recalled based on the fourth similarity between each of the at least one information to be recalled and the information to be retrieved includes: Determining the fourth similarity between each of the at least one information to be recalled and the information to be retrieved; Determining at least one target recalled information from the at least one information to be recalled according to the values of each fourth similarity based on the specified number of recalls.

11. The method according to any one of claims 1 - 5, wherein The first configuration information includes the connection relationship between the multiple candidate cluster centers, where the number of other candidate cluster centers connected to each candidate cluster center does not exceed the specified connection number.

12. The method according to any one of claims 1 - 5, wherein The second configuration information includes a plurality of inverted lists, and each inverted list includes a corresponding candidate code set corresponding to a candidate clustering center.

13. An information retrieval device, wherein It includes: A coding unit, configured to obtain information to be retrieved; A first matching unit, configured to obtain first configuration information including a plurality of candidate clustering centers, and select at least one target clustering center based on the first similarity between each of the plurality of candidate clustering centers and the information to be retrieved; wherein, each candidate clustering center corresponds to a data category; An obtaining unit, configured to obtain second configuration information including the candidate code sets respectively corresponding to the at least one target clustering center; wherein, each candidate code set is encoded based on each piece of recall information belonging to the corresponding data category; A second matching unit, configured to perform quantization encoding on the information to be retrieved to obtain a to-be-retrieved code, and select at least one target code based on the second similarity between each candidate code included in each candidate code set and the to-be-retrieved code; A recall unit, configured to obtain a retrieval result based on the at least one target code.

14. An electronic device, wherein It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 12.

15. A computer-readable storage medium, wherein It includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 12.

16. A computer program product, wherein It includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Information retrieval method, device and equipment and computer readable storage medium

    CN111753060A

  • Data retrieval method and device and computer readable storage medium

    CN112418298A