Method, apparatus and device for retrieving library optimization

By mapping the similarity between the retrieval database and the query database in risk control scenarios, the contribution value and retention probability are determined, the retrieval database is optimized, the problem of introducing noisy samples is solved, and the data quality and recognition accuracy are improved.

CN117093608BActive Publication Date: 2026-05-01ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2023-08-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In risk control scenarios, the introduction of noisy samples during the dynamic updating of the retrieval database leads to a decline in data quality, affecting the recognition accuracy and performance of vector retrieval applications.

Method used

By mapping the similarity between search elements in the search database and query elements in the query database, a contribution value is determined. Based on the mapping relationship between the contribution value and the retention probability, the loss function is minimized to determine whether the search element is retained, thereby optimizing the search database.

Benefits of technology

It improves the data quality of the retrieval database, reduces noisy samples, and enhances the recognition accuracy and performance of vector retrieval applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117093608B_ABST
    Figure CN117093608B_ABST
Patent Text Reader

Abstract

The one or more embodiments of the specification disclose a method, device and equipment for retrieval library optimization. The method comprises: mapping the similarity between retrieval elements in a retrieval library and query elements in a query library to obtain a contribution value of each retrieval element to the retrieval score corresponding to the query element pair; determining a mapping relationship between the retrieval score representing the query element and the retention probability of each retrieval element according to the contribution value; determining the retention probability corresponding to each retrieval element in the process of minimizing a first loss function based on the mapping relationship, wherein the first loss function is used to represent the residual between the retrieval score corresponding to the query element and the classification label; and determining a keep or discard indication of whether each retrieval element in the retrieval library is retained according to the retention probability, so as to update the retrieval library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of data processing technology, and in particular to a method, apparatus and equipment for optimizing a search database. Background Technology

[0002] In risk control scenarios, risks are constantly changing. By promptly introducing samples with new risk characteristics into the search database, the database's generalization ability during vector retrieval can be improved. Simultaneously, removing "outdated" samples from the database also helps improve its performance in vector retrieval. This process of updating the search database is known as search database maintenance or optimization.

[0003] However, during the dynamic updating of the search database, such as when new risk samples are added, noisy samples are often introduced. These noisy samples degrade the data quality of the search database and severely impact the performance of downstream vector retrieval applications, thereby reducing the recognition accuracy and increasing the disturbance rate of vector retrieval applications. Therefore, there is an urgent need to provide a better search database optimization solution. Summary of the Invention

[0004] This specification provides a method, apparatus, and device for optimizing a search database, in order to provide a search database optimization scheme that meets the expectations of those involved in the search database.

[0005] In a first aspect, one or more embodiments of this specification provide a method for optimizing a retrieval database, comprising: mapping the similarity between retrieval elements in the retrieval database and query elements in the query database to obtain a contribution value of each retrieval element to the retrieval score corresponding to the query element; determining a mapping relationship characterizing the retrieval score of the query element and the retention probability of each retrieval element based on the contribution value; determining the retention probability corresponding to each retrieval element in the process of minimizing a first loss function based on the mapping relationship, wherein the first loss function is used to characterize the residual between the retrieval score corresponding to the query element and the classification label; and determining a retention / retention indicator for each retrieval element in the retrieval database based on the retention probability, so as to update the retrieval database.

[0006] Secondly, embodiments of this specification provide an apparatus for optimizing a retrieval database, comprising: mapping the similarity between retrieval elements in the retrieval database and query elements in the query database to obtain a contribution value of each retrieval element to the retrieval score corresponding to the query element; determining a mapping relationship characterizing the retrieval score of the query element and the retention probability of each retrieval element based on the contribution value; determining the retention probability corresponding to each retrieval element in the process of minimizing a first loss function based on the mapping relationship, wherein the first loss function is used to characterize the residual between the retrieval score corresponding to the query element and the classification label; and determining a retention / retention indication for each retrieval element in the retrieval database based on the retention probability, so as to update the retrieval database.

[0007] Thirdly, embodiments of this specification provide an electronic device comprising: mapping the similarity between search elements in a search library and query elements in a query library to obtain a contribution value of each search element to the search score corresponding to the query element; determining a mapping relationship characterizing the search score of the query element and the retention probability of each search element based on the contribution value; determining the retention probability corresponding to each search element in the process of minimizing a first loss function based on the mapping relationship, wherein the first loss function is used to characterize the residual between the search score corresponding to the query element and the classification label; and determining a retention / retention indicator for each search element in the search library based on the retention probability, so as to update the search library.

[0008] Fourthly, embodiments of this specification provide a storage medium for storing a computer program that can be executed by a processor to implement the following process: mapping the similarity between search elements in a search library and query elements in a query library to obtain the contribution value of each search element to the search score corresponding to the query element; determining a mapping relationship characterizing the search score of the query element and the retention probability of each search element based on the contribution value; determining the retention probability corresponding to each search element in the process of minimizing a first loss function based on the mapping relationship, wherein the first loss function is used to characterize the residual between the search score corresponding to the query element and the classification label; and determining a retention / deletion indication for each search element in the search library based on the retention probability to update the search library. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic flowchart of a method for optimizing a search database according to an embodiment of this specification.

[0011] Figure 2 This is a schematic diagram illustrating an application scenario of a method for optimizing a search database according to an embodiment of this specification.

[0012] Figure 3 This is a schematic flowchart of a method for optimizing a search database according to an embodiment of this specification.

[0013] Figure 4 This is a schematic diagram of a retrieval database optimization device according to an embodiment of this specification.

[0014] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this specification. Detailed Implementation

[0015] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments in this specification. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this application.

[0016] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this specification can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0017] The following description, in conjunction with the accompanying drawings, details the methods, apparatus, and devices for optimizing search libraries provided in the embodiments of this specification through specific examples and application scenarios.

[0018] Vector retrieval is the process of finding, given an input sample (i.e., the query vector) and a given vector dataset, retrieving K nearest vectors (K-Nearest Neighbors, KNN) based on a metric such as Euclidean distance, cosine distance, inner product, or Hamming distance. These K vectors are then used to determine the category of the input sample. This dataset, composed of historical data and requiring maintenance, is typically called the base database or retrieval database.

[0019] In the application of risk vector retrieval, there are usually two mechanisms for updating the retrieval database: (1) filtering samples through a time window, dynamically adding the latest risk samples, and removing early risk samples; (2) only adding samples without removing them. The first mechanism cannot avoid the online coverage decline caused by the erroneous removal of early high-quality samples, and will also introduce noisy samples. The second mechanism will cause the size of the retrieval database to expand indefinitely and cannot avoid the introduction of noisy samples.

[0020] Therefore, it is necessary to select samples appropriately for inclusion in the search database and optimize the search performance of the database.

[0021] Figure 1 This illustration shows a method for optimizing a search database according to an embodiment of the present invention. This method can be executed by an electronic device, which may include a server and / or a terminal device, wherein the terminal device may be, for example, an in-vehicle terminal or a mobile phone terminal. In other words, the method can be executed by software or hardware installed in the aforementioned electronic device, and the method includes the following steps:

[0022] S102: Map the similarity between the search elements in the search database and the query elements in the query database to obtain the contribution value of each search element to the search score corresponding to the query element.

[0023] The search database, also known as the base database, is the dataset used for optimization in this specification. The search database can be an existing dataset or a dataset obtained by adding new data to an existing search database. This specification does not limit the source of the search database; it can be determined based on the actual situation. Specifically, the data in the search database can be called search elements, and the data in the query database can be called query elements. The query database is a dataset constructed based on the search database, containing the same data as the search database. In one example, the query database can be the search database itself.

[0024] The tags for search elements in the retrieval database can be divided into two categories: black samples and white samples, typically assigned 0 and 1 labels respectively. The specific sample types referred to as black samples and white samples can be determined based on the application scenario of the retrieval database. In a risk control scenario, black samples refer to samples with risk, while white samples refer to samples without risk.

[0025] The retrieval score is a rating of the query elements in the query database. Specifically, the theoretical retrieval score of a retrieval element with a black sample label is consistent with the category label corresponding to the black sample, and the theoretical retrieval score of a retrieval element with a white sample label is consistent with the category label corresponding to the white sample. The closer the retrieval score of a retrieval element with a black sample label is to 0, the higher the quality of the retrieval element; the further the retrieval score of a retrieval element with a black sample label deviates from 0, the lower the quality of the retrieval element. Similarly, the closer the retrieval score of a retrieval element with a white sample label is to 1, the higher the quality of the retrieval element; the further the retrieval score of a retrieval element with a white sample label deviates from 1, the lower the quality of the retrieval element. In one example, a predetermined number of retrieval elements (referred to as recall elements) that are highly relevant to the query elements in the query database can be used to determine the retrieval score.

[0026] In one example, the retrieval score of a query element can be determined based on the pairwise similarity between a retrieval element in the retrieval database and a query element in the query database. The similarity can be measured using methods such as Euclidean distance, cosine similarity, inner product, or Hamming distance. To better illustrate the invention and highlight its main points, the specific embodiments in this specification use vector distance (i.e., Euclidean distance) as the similarity metric. Those skilled in the art should understand that other similarity metrics can also be implemented using similar methods.

[0027] The contribution value is the numerical value by which the search element contributes to the search score of the query element. The vector distance between the search element and the query element is inversely proportional to the contribution value of the search element to the search score of the query element. Generally, the closer the vector distance between the search element and the query element, the higher the similarity between them; the farther the vector distance, the lower the similarity. Therefore, the distance between the search element and the query element can be mapped to obtain the contribution value of the search element to the search score of the query element. This specification does not specify a particular mapping method, which can be determined according to the actual situation. In one example, the mapping method can be 1 / d. β And so on, where d is the vector distance, and the value of parameter β can be determined according to the actual situation.

[0028] Furthermore, when the retrieved element is a white sample, the retrieved element can provide a positive contribution value to the query element; when the retrieved element is a black sample, the retrieved element can provide a negative contribution value to the query element.

[0029] For example, for the retrieval database (base database) S base :

[0030] S base ={(x i ,y i )|i=1,2,3,…,N} (Formula 1)

[0031] Where, x i It is the retrieval database S base The search element in y i It is the category tag corresponding to the search element, y i = 0 or 1.

[0032] Build query library S query And S base =S query We can first calculate the query database S. query All query elements and search database S base The pairwise vector distances of all retrieved elements in the matrix are used to obtain matrix D. N×N , where matrix D N×N Each element d in ij for:

[0033] d ij =||x i -x j ||2 (Formula 2)

[0034] Where, d ij It is the vector distance, x i ∈S query x j ∈S base , i=1,2,3,…,N, j=1,2,3,…,N.

[0035] Furthermore, formulas 3 and 4 can be used to map the vector distance to obtain the contribution value:

[0036] W = F(D) (Formula 3)

[0037] s = 1 / d ij (Formula 4)

[0038] Where D is the query database S query Any query element and search database S base Let F be the matrix formed by the vector distances between any retrieved elements in D, where F is the mapping function, W is the contribution matrix obtained by mapping the vector distances in D, and d ij s is an element in matrix D, and s is a pair of elements in matrix D. ij The contribution values ​​obtained by mapping are i = 1, 2, 3, ..., N, j = 1, 2, 3, ..., N.

[0039] S104: Based on the contribution value, determine the mapping relationship between the retrieval score representing the query element and the retention probability of each retrieval element.

[0040] The retention probability is the probability that a search element is retained in the search database, indicating the quality of the search element. Specifically, high-quality search elements have a higher retention probability, while low-quality search elements have a lower retention probability. By weighting the contribution value of the search element based on the retention probability, the accuracy of the search score can be improved. Therefore, the mapping relationship between the search score representing the query element and the retention probability of the search element can be determined through the contribution value.

[0041] In one example, the mapping relationship could be a linear mapping relationship as follows:

[0042] z N×1 =F(D N×N )p N×1 (Formula 5)

[0043] Among them, z N×1 To represent the query database S query The vector formed by the retrieval scores corresponding to each query element, where F is the mapping function and p N×1 To represent the search database S base The vector formed by the retention probabilities corresponding to each retrieved element, D N×N To represent the query database S query Any query element and search database S base The matrix formed by the vector distances between any retrieved elements.

[0044] In one example, based on the range of the retrieval score (0,1), this mapping relationship can be a non-linear mapping relationship:

[0045] z N×1 =sigmoid(F(D) N×N )p N×1 ) (Formula 6)

[0046] Among them, z N×1 To represent the query database S query The vector formed by the retrieval scores corresponding to each query element, where F is the mapping function and p N×1 To represent the search database S base The vector formed by the retention probabilities corresponding to each retrieved element, D N×N To represent the query database S query Any query element and search database S base The matrix formed by the vector distances between any retrieved elements.

[0047] S106: Based on the mapping relationship, in the process of minimizing the first loss function, the retention probability corresponding to each search element is determined. The first loss function is used to characterize the residual between the search score and the classification label corresponding to the query element.

[0048] Specifically, after obtaining the mapping relationship between the retrieval score representing the query element and the retention probability of the retrieval element in step S104, the first loss function corresponding to this mapping relationship can be determined. This first loss function can be used to represent the residual between the retrieval score corresponding to the query element and the classification label. Furthermore, the retrieval score in the first loss function can be replaced with the retention probability using the mapping relationship, thereby determining the retention probability of each retrieval element while minimizing the loss value of the first loss function.

[0049] Typically, optimizing a search database is equivalent to solving an integer programming problem, where 0 and 1 can represent the selection or rejection of each sample in the database, thereby optimizing the data quality of the base database. This is a computationally complex problem with a time complexity of O(2^3). n The problem of finding the retention probability is time-consuming and difficult. Step S106 solves the retention probability through the first loss function, transforming the optimization process of the search library into a linear integer programming problem and reducing the computational complexity to O(1), making it easy and fast to solve.

[0050] S108: Based on the retention probability, determine whether each search element in the search database should be retained or not, so as to update the search database.

[0051] The retention / deletion instructions can include deletion instructions and / or retention instructions. After determining the retention probability of each search element in step S106, the retention / deletion instructions corresponding to the retention probabilities can be determined. Furthermore, a retention probability threshold can be set. When the retention probability is greater than the threshold, a retention instruction can be selected for the search element so that the search library retains the search element; when the retention probability is less than the threshold, a deletion instruction can be selected for the search element so that the search library deletes the search element.

[0052] In the embodiments of this specification, a first loss function is established based on the mapping relationship constructed according to the contribution value. Based on this mapping relationship, the retention probability of each search element is determined while minimizing the loss value of the first loss function. Then, the retention probability is used to determine the retention or removal indication of each search element in the search database, thereby updating the search database. This process, through the first loss function, transforms the discrete variable of whether an index element in the search database is retained into the continuous variable of the retention probability of the index element. This enables rapid determination of whether each search element in the search database is retained, thereby quickly eliminating noisy samples in the search database, further improving the speed of search database optimization, and ensuring the data quality of the search database.

[0053] Figure 2 A schematic diagram illustrating an application scenario for a search database optimization method is provided, such as... Figure 2 As shown, the retrieval database server 201 sends a retrieval database optimization command to the optimization server 202. After receiving the retrieval database optimization command, the optimization server 202 reads the retrieval database in the first database 203, generates a query database, obtains the optimized retrieval database using the retrieval database optimization method in this specification, and puts the optimized retrieval database into the second database 204. At the same time, it feeds back the optimization results to the retrieval database server 201 so that the retrieval database server 201 can use the optimized retrieval database to perform vector retrieval.

[0054] The retrieval server 201 and the optimization server 202 can be the same server, and the first database 203 and the second database 204 can be the same database.

[0055] In one implementation, step S102 can be executed as follows: steps A1-A3:

[0056] Step A1: Determine the vector distance between the vector of the retrieved element and the vector of the query element;

[0057] Step A2: From the vector distance, determine the preset number of target distances related to each query element. The retrieved element corresponding to the target distance is the recall element corresponding to the query element.

[0058] Step A3: Map the target distance to obtain the contribution value of each recalled element to the retrieval score corresponding to the query element. The absolute value of the contribution value is negatively correlated with the target distance.

[0059] The recalled elements are a predetermined number of search vectors retrieved from the search database that are similar to the vector of the query element, based on vector distance. Specifically, the vector distances between the query element's vector and the vectors of all search elements in the search database are sorted in ascending order. Starting with the smallest non-zero vector distance (excluding the vector distances between the query element and search vectors at the same position in the search database), a predetermined number of vector distances are sequentially obtained as the target distance. The search elements corresponding to the target distances are the recalled elements. This manual does not specify a particular value for the predetermined number; it can be selected according to the actual situation.

[0060] Furthermore, the target distance can be mapped to obtain the contribution value of each recalled element to the retrieval score corresponding to the query element. The absolute value of the contribution is negatively correlated with the target distance. Additionally, recalled elements with white sample labels can provide positive contribution values, while recalled elements with black sample labels can provide negative contribution values.

[0061] For the contribution value matrix W obtained above, the retrieval elements other than the recall elements corresponding to each query element can be set to a specific value (such as 0) to obtain a contribution value matrix that only contains the recall elements.

[0062] In the embodiments of this specification, the retrieval score corresponding to the query element is determined by using only the contribution value of a preset number of recalled elements. This reduces the computational load of determining the contribution value and thus improves the speed of updating the retrieval database.

[0063] In one implementation, step S106 can be executed as follows: steps B1-B2:

[0064] Step B1: Based on the mapping relationship, replace the retrieval score corresponding to the query element in each loss term of the first loss function with the retention probability of the retrieval element corresponding to the query element to obtain the second loss function;

[0065] Step B2: In the process of minimizing the second loss function, determine the retention probability in the loss term corresponding to each search element.

[0066] Specifically, after determining the first loss function, the retrieval score in the first loss function can be replaced with the retention probability in the mapping relationship to obtain the second loss function. Then, in the process of minimizing the second loss function, the value of the retention probability in each loss term is determined, thereby optimizing the retrieval database.

[0067] In one implementation, when the mapping relationship is linear, step B2 can be executed as follows: steps C1-C3:

[0068] Step C1: In the case that the classification label of the search element corresponding to the loss term in the first loss function is a white sample, a first symbol is set for the loss term. The first symbol is used to ensure that the value of the loss term is negative.

[0069] Step C2: In the case that the classification label of the search element corresponding to the loss term in the first loss function is black sample, a second symbol is set for the loss term. The second symbol is used to ensure that the value of the loss term is negative.

[0070] Step C3: Based on the mapping relationship, determine the retention probability corresponding to each search element while minimizing the first loss function with the set symbols.

[0071] Since white samples provide positive search scores and black samples provide negative search scores, to minimize the loss value of the first loss function, different signs can be assigned to the loss terms corresponding to search elements with the classification label of black samples and those with the classification label of white samples. Specifically, a positive sign can be assigned to the loss terms corresponding to search elements with the classification label of white samples, and a negative sign can be assigned to the loss terms corresponding to search elements with the classification label of black samples, so as to minimize the loss value of the first loss function.

[0072] For example, for the linear mapping relationship in Equation 5, its first loss function can be shown in Equation 7:

[0073]

[0074] Where L is the loss value of the first loss function, y i To query the category tags of elements, z i This is the search score corresponding to the queried element.

[0075] In Formula 7, the z-coordinates of the search elements with classification labels for black samples and the search elements with classification labels for white samples are... i The signs are different, through 2y i -1 provides different signs for query elements with different classification labels, enabling each loss term in the first loss function to be negative, thereby minimizing the loss value.

[0076] In the embodiments of this specification, different symbols are provided for the search elements with the classification label of white samples and the search elements with the classification label of black samples, so that the value of each loss term in the first loss function is negative, thereby minimizing the loss value of the first loss function. In this way, the retention probability value corresponding to each loss term in the first loss function can be determined, thereby optimizing the search library.

[0077] In one implementation, step C3 can be executed as follows: steps D1-D3:

[0078] Step D1: In the second loss function, sum the coefficients of the retention probabilities in the loss terms corresponding to the same search element to obtain the total coefficient of the retention probabilities in the loss terms corresponding to the search element.

[0079] Step D2: When the total coefficient is negative, set the retention probability of the search element corresponding to the total coefficient to a first preset value. The first preset value is used to indicate the loss term corresponding to the retained search element.

[0080] Step D3: If the total coefficient is positive, set the retention probability of the search element corresponding to the total coefficient to a second preset value. The second preset value is used to indicate the deletion of the loss item corresponding to the search element.

[0081] Specifically, since the mapping relationship is a mapping between the retrieval scores of all query elements and the retention probabilities of the corresponding retrieval elements, the second loss function, obtained by replacing the retrieval scores in the first loss function with retention probabilities through the mapping relationship, can include all retrieval elements in the search database. Therefore, the coefficients of the retention probabilities corresponding to each retrieval element can be combined. In the process of minimizing the second loss function, the specific value of the retention probability of the retrieval element is determined by the total coefficient of the combined retrieval elements.

[0082] For example, after replacing the first loss function in Formula 7 with a mapping relationship, the resulting second loss function is processed as shown in Formula 8:

[0083]

[0084] Where L is the loss value, N is the number of query elements in the query database, and y i To query the category tag corresponding to an element, z i Let y be the retrieval score corresponding to the query element, i = 1, 2, 3, ..., N, y be the vector composed of the retrieval elements, z be the vector composed of the retrieval scores, p be the vector representing the retention probability of the retrieval element, and W be the contribution value matrix. j The probability of retaining the retrieved element. For p j The total coefficient.

[0085] As can be seen from Formula 8, q ij =(2y i -1)·w ij ,again Therefore, the following formula can be obtained:

[0086]

[0087] Where, q ij It is the coefficient of the retention probability, y i To query the category tag corresponding to an element, y j To retrieve the category tags corresponding to the elements, vector distance The contribution value of the retrieved element to the retrieval score of the query element obtained through mapping, w ij It is the contribution value of the retrieved element to the retrieval score of the query element.

[0088] As can be seen from Formula 9, when the sample labels of the retrieved element and the query element are the same in the same loss term, q ijIt is a positive number; when the sample labels of the retrieved element and the query element in the same loss term are different, q ij It is a negative number. Therefore, in using formula 8... When calculating the loss function, you can directly determine whether the value of each loss term is positive or negative based on whether the sample labels of the retrieved element and the query element are the same.

[0089] In the embodiments of this specification, after determining the coefficient of the retention probability of the loss term corresponding to the search element in the second loss function, the retention probability corresponding to each search element can be determined based on the total coefficient during the process of minimizing the loss value of the second loss function. This process achieves the determination of the retention probability through the total coefficient of the retention probability in the loss term, thereby optimizing the search database.

[0090] In one implementation, step S108 can be executed as follows: steps E1-E2:

[0091] Step E1: When the retention probability of the searched element is a first preset value, the retention indication of the searched element is determined as a retention indication;

[0092] Step E2: When the retention probability of the searched element is the second preset value, the retention / deletion indication of the searched element is determined as a deletion indication.

[0093] As shown in Equation 10, the optimization objective of the second loss function in Equation 8 is:

[0094]

[0095] To minimize the value of the second loss function, for each loss term in Equation 8, when the coefficient of the retention probability is... When it is positive, we can take p. j The coefficient of the retention probability is 0 (first preset value). When the value is negative, p can be taken. j It is 1 (second preset value).

[0096] In the embodiments of this specification, the retention probability of the search element is determined by minimizing the second loss function, thereby directly determining whether each search element in the search database is retained.

[0097] The search database often exhibits an imbalance between white and black samples. For instance, in risk control scenarios, white samples typically have a significant numerical advantage, which can negatively impact the overall coefficient of the searched elements in the second loss function. It is very easy to obtain a positive value, and therefore, when minimizing the second loss function, the p value in the loss term corresponding to the retrieved element will be affected. jThe value is determined to be 0, and therefore, according to Formula 8, black samples in the search library can be easily excluded. In one implementation, step S106 can be executed as follows: steps F1-F3:

[0098] Step F1: In the first loss function, set the first weight for the loss term corresponding to the query element with the classification label of white sample;

[0099] Step F2: Set a second weight for the loss term corresponding to the query element with the classification label "black sample". The first weight is greater than the second weight.

[0100] Step F3: Based on the first loss function with set weights, determine the retention probability of each search element.

[0101] Specifically, to address the issue of black samples being easily excluded, different weights can be assigned to the loss terms corresponding to query elements with different classification labels in the first loss function. In one example, a smaller weight can be assigned to the loss term corresponding to the query element with the classification label "white sample," while a larger weight can be assigned to the loss term corresponding to the query element with the classification label "black sample."

[0102] In one example, weights can be determined based on the sample distribution. Specifically, as shown in Equations 11 and 12, the second weight corresponding to the loss term of the query element with the classification label "black sample" can be set to 1, and the first weight corresponding to the loss term of the query element with the classification label "white sample" can be determined proportionally.

[0103] w negative =1 (Formula 11)

[0104] w positive =(C negative / C positive ) β (Formula 12)

[0105] Among them, W negative It is the second weight corresponding to the loss term of the query element with the classification label "black" in the first loss function, w positive C is the first weight of the loss term corresponding to the query element with the classification label "white sample" in the first loss function. negative C is the number of query elements with the category label "black sample". positive β is the number of query elements with the category label "black sample", and β is a weighting coefficient determined based on the actual situation.

[0106] In the embodiments of this specification, different weights are set for different classification labels in the retrieval library, which solves the problem that a small number of high-quality samples are easily excluded due to the imbalance of samples in the retrieval library, thereby improving the generalization ability of the retrieval library.

[0107] In one implementation, step S106 can be executed as follows: steps G1-G3:

[0108] Step G1: Based on the range of retrieval scores, the retrieval scores in the mapping relationship are processed using an activation function;

[0109] Step G2: Based on the range of values ​​for the retention probability, the retention probability in the first loss function is processed using an activation function;

[0110] Step G3: Based on the mapping relationship processed by the activation function, determine the retention probability corresponding to each search element while minimizing the first loss function processed by the activation function.

[0111] Specifically, since the label for white samples in the retrieval database is 1 and the label for black samples is 0, an activation function can be used to process the retrieval scores in the mapping relationship, mapping the retrieval scores to the range (0,1). For example, after processing the retrieval scores in Formula 5 with an activation function, the resulting mapping relationship is shown in Formula 6, and the loss function corresponding to the processed mapping relationship can be the cross-entropy loss function shown in Formula 13.

[0112]

[0113] Where L is the loss value of the first loss function, N is the number of query elements in the query database, and y i To retrieve the category tag corresponding to an element, z i Let i = 1, 2, 3, ..., N, be the retrieval score for the element being retrieved.

[0114] Based on the mapping relationship after activation function processing, replacing the retrieval score in Formula 13 transforms the optimization of the function value of the first loss function into... This problem can be solved using the Stochastic Gradient Descent (SGD) method.

[0115] Furthermore, in order to ensure that the retention probability corresponding to the retrieved element is in the range of [0,1], the retention probability in the first loss function can be processed by an activation function, as shown in Formula 14.

[0116]

[0117] in, θ is the probability of retaining the retrieved element in the search database, and θ is the transformed variable.

[0118] Formula 14 further transforms the optimization of the first loss function into... Determined based on θ After that, it can be based on Determine whether to keep or remove each black sample in the search database. Furthermore, a retention probability threshold can be set, and samples exceeding this threshold will be excluded. The corresponding samples will be retained if their probability is below the retention threshold. The corresponding sample was deleted.

[0119] In the embodiments of this specification, the retrieval score in the mapping relationship and the retention probability in the first loss function are processed by activation functions to determine the retention or rejection of samples in the retrieval database using the processed loss function. This process transforms the discrete variable of whether an index element in the retrieval database is retained—the quantifiable variable of the retention probability—into a continuous variable. This enables rapid determination of whether each retrieval element in the database is retained, thereby quickly eliminating noisy samples and further improving the speed of retrieval database optimization while ensuring the data quality of the database.

[0120] Figure 3 A schematic flowchart illustrating a method for optimizing a search database is presented. For example... Figure 3 As shown, the following steps H1-H5 can be performed:

[0121] Step H1: Self-retrieval. Specifically, determine the vector distance between each pair of search elements in the search database and construct a distance matrix. The columns of the matrix are the search vectors in the search database, and the rows of the matrix are the query vectors in the same query database. From the vector distances in each row of the matrix, select a preset number (topk) of vector distances with smaller distances. The search vectors corresponding to these preset number of vector distances are the recall vectors corresponding to the query vectors.

[0122] Step H2: Distance Mapping. Specifically, use Formula 15 to map the vector distances in the distance matrix to contribution values, while setting the diagonal lines of the vector matrix to zero.

[0123]

[0124] Where topk is the preset number in step H1, and d topk+1 It is the (+1)th vector distance after sorting the vector distances corresponding to the query vectors in ascending order, d i It is the vector distance corresponding to the query vector, and N is the number of query elements in the query database, i = 1, 2, 3, ..., N.

[0125] Step H3: Weighting. Specifically, it consists of three parts, including:

[0126] (1) Obtaining the symbolic weighted matrix. When the labels of the search vector and the query vector corresponding to the vector distance in the vector matrix are the same, a positive symbolic weight (value of 1) is set for the vector distance; when the labels of the search vector and the query vector corresponding to the vector distance in the vector matrix are different, a negative symbolic weight (value of -1) is set for the vector distance. Based on the symbolic weight corresponding to each vector distance in the vector matrix, the symbolic weighted matrix is ​​obtained.

[0127] (2) Obtaining the Balanced Weighting Matrix. The query vectors in each row of the matrix are weighted: when the label of the query vector in the vector matrix is ​​a white sample, the balance weight corresponding to the vector distance in that row is set to 1; when the label of the query vector is a black sample, the balance weight corresponding to the vector distance in that row is set to C1 / C2, where C1 is the number of white samples in the search database and C2 is the number of black samples in the search database. Based on the balance weight corresponding to each vector distance in the vector matrix, the sign-weighted matrix is ​​obtained.

[0128] (3) Weighting. After applying a sign-weighted matrix and a balanced-weighted matrix to the distance matrix, the total coefficient of each search element in the search database is obtained by summing the columns.

[0129] Step H4: Delete the retrieved elements. Specifically, delete the retrieved vectors corresponding to the negative values ​​obtained by summing the columns in step H3.

[0130] The retrieval classification performance of different retrieval libraries on out-of-time (OOT) datasets is shown in the table below.

[0131] Table 1. Search Classification Results

[0132] Search library used AUC The search database obtained by the scheme in this specification 99.53 Search library that retains all black samples 99.51 Search database with a time window of 90 days 99.14 Search database with a time window of 30 days 94.78

[0133] As shown in Table 1, the retrieval database obtained by the method provided in this specification still has a better retrieval classification effect compared with the retrieval database that retains all black samples. This fully demonstrates that the method provided in this specification can reasonably select black samples to be added to the retrieval database, thereby optimizing the retrieval effect. Furthermore, the retrieval classification effects of the retrieval database with a 90-day time window and the retrieval database with a 30-day time window indicate that the method provided in this specification can fully utilize data for continuous optimization and improve retrieval performance.

[0134] The search database optimization method provided in the embodiments of this specification can be executed by a search database optimization apparatus or a control module within that apparatus for executing the search database optimization method. This specification describes the search database optimization apparatus provided in the embodiments of this specification using an apparatus for executing the search database optimization method as an example.

[0135] Figure 4This is a schematic diagram of the structure of a retrieval database optimization device according to an embodiment of the present invention. Figure 4 As shown, the retrieval database optimization apparatus 400 includes:

[0136] The contribution value determination module 410 is used to map the similarity between the search elements in the search database and the query elements in the query database to obtain the contribution value of each search element to the search score corresponding to the query element.

[0137] The mapping relationship determination module 420 is used to determine the mapping relationship between the retrieval score representing the query element and the retention probability of each retrieval element based on the contribution value.

[0138] The retention probability determination module 430 is used to determine the retention probability of each search element based on the mapping relationship while minimizing the first loss function. The first loss function is used to characterize the residual between the search score and the classification label of the query element.

[0139] The update module 440 is used to determine the retention or rejection indication of each search element in the search database based on the retention probability, so as to update the search database.

[0140] In one embodiment, the contribution value determination module 410 includes:

[0141] The vector distance determination unit is used to determine the vector distance between the vector of the retrieved element and the vector of the query element.

[0142] The target distance determination unit is used to determine a preset number of target distances related to each query element from the vector distances. The retrieval element corresponding to the target distance is the recall element corresponding to the query element.

[0143] The contribution value determination unit is used to map the target distance and obtain the contribution value of each recalled element to the retrieval score corresponding to the query element. The absolute value of the contribution value is negatively correlated with the target distance.

[0144] In one embodiment, the retention probability determination module 430 includes:

[0145] The loss function determination unit is used to replace the retrieval score corresponding to the query element in each loss term of the first loss function with the retention probability of the retrieval element corresponding to the query element based on the mapping relationship, so as to obtain the second loss function;

[0146] The first probability determination unit is used to determine the retention probability in the loss term corresponding to each search element during the process of minimizing the second loss function.

[0147] In one embodiment, when the mapping relationship is a linear mapping relationship, the first probability determination unit is used to:

[0148] In the case that the classification label of the search element corresponding to the loss term in the first loss function is white sample, a first sign is set for the loss term. The first sign is used to ensure that the value of the loss term is negative.

[0149] In the case that the classification label of the search element corresponding to the loss term in the first loss function is black sample, a second sign is set for the loss term. The second sign is used to ensure that the value of the loss term is negative.

[0150] Based on the mapping relationship, the retention probability corresponding to each search element is determined in the process of minimizing the first loss function with the set symbols.

[0151] In one embodiment, based on the mapping relationship, the retention probability corresponding to each retrieved element is determined during the process of minimizing the first loss function with predefined symbols, including:

[0152] In the second loss function, the coefficients of the retention probability in the loss terms corresponding to the same search element are summed to obtain the total coefficient of the retention probability in the loss terms corresponding to the search element.

[0153] When the total coefficient is negative, the retention probability of the search element corresponding to the total coefficient is set to the first preset value. The first preset value is used to indicate the loss term corresponding to the retention of the search element.

[0154] When the total coefficient is positive, the retention probability of the search element corresponding to the total coefficient is set to the second preset value. The second preset value is used to indicate the deletion of the loss term corresponding to the search element.

[0155] In one embodiment, the update module 440 includes:

[0156] The first indication determining unit is used to determine the retention indication of the search element as a retention indication when the retention probability of the search element is a first preset value.

[0157] The second indication determination unit is used to determine the retention indication of the search element as a deletion indication when the retention probability of the search element is a second preset value.

[0158] In one embodiment, the retention probability determination module 430 includes:

[0159] The first weight determination unit is used to set the first weight for the loss term corresponding to the retrieval element with the classification label of white sample in the first loss function;

[0160] The second weight setting unit is used to set a second weight for the loss term corresponding to the retrieval element with the classification label "black sample". The first weight is less than the second weight.

[0161] The second probability determination unit is used to determine the retention probability of each search element in the process of minimizing the first loss function with set weights.

[0162] In one embodiment, the retention probability determination module 430 includes:

[0163] The first processing unit is used to process the retrieval scores in the mapping relationship based on the range of retrieval scores using an activation function;

[0164] The second processing unit is used to process the retention probability in the first loss function through an activation function based on the range of values ​​of the retention probability.

[0165] The third probability determination unit is used to determine the retention probability of each search element based on the mapping relationship processed by the activation function, while minimizing the first loss function processed by the activation function.

[0166] The apparatus for optimizing the search database in the embodiments of this specification can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. The embodiments of this specification do not impose specific limitations.

[0167] The apparatus for optimizing the search database in the embodiments of this specification can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; the embodiments of this specification do not specifically limit it.

[0168] The retrieval database optimization apparatus provided in the embodiments of this specification can achieve... Figure 1 The various processes implemented in the method embodiments are not described in detail here to avoid repetition.

[0169] Based on the same idea, one or more embodiments of this specification also provide an electronic device, such as... Figure 5As shown. Electronic devices can vary considerably due to differences in configuration or performance, and may include one or more processors 501 and memory 502. Memory 502 may store one or more application programs or data. Memory 502 may be temporary or persistent storage. The application programs stored in memory 502 may include one or more modules (not shown), each module may include a series of computer-executable instructions for the electronic device. Furthermore, processor 501 may be configured to communicate with memory 502 and execute the series of computer-executable instructions in memory 502 on the electronic device. The electronic device may also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input / output interfaces 505, and one or more keyboards 506.

[0170] Specifically, in this embodiment, the electronic device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0171] The similarity between the search elements in the search database and the query elements in the query database is mapped to obtain the contribution value of each search element to the search score corresponding to the query element.

[0172] Based on the contribution value, determine the mapping relationship between the retrieval score representing the query element and the retention probability of each retrieval element;

[0173] Based on the mapping relationship, the retention probability corresponding to each retrieval element is determined in the process of minimizing the first loss function. The first loss function is used to characterize the residual between the retrieval score and the classification label corresponding to the query element.

[0174] Based on the retention probability, determine whether each search element in the search database should be retained or not, so as to update the search database.

[0175] One or more embodiments of this specification also provide a storage medium storing one or more computer programs, the one or more computer programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform various processes of the above-described method embodiments for optimizing the search library, and specifically for performing:

[0176] The similarity between the search elements in the search database and the query elements in the query database is mapped to obtain the contribution value of each search element to the search score corresponding to the query element.

[0177] Based on the contribution value, determine the mapping relationship between the retrieval score representing the query element and the retention probability of each retrieval element;

[0178] Based on the mapping relationship, the retention probability corresponding to each retrieval element is determined in the process of minimizing the first loss function. The first loss function is used to characterize the residual between the retrieval score and the classification label corresponding to the query element.

[0179] Based on the retention probability, determine whether each search element in the search database should be retained or not, so as to update the search database.

[0180] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the above-described storage medium embodiment is basically similar to the method embodiment, so the description is relatively simple; relevant parts can be referred to the description of the method embodiment.

[0181] The methods, apparatus, modules, or units described in the above embodiments can be implemented by a computer chip or entity, or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0182] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0183] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0184] This specification describes one or more embodiments of methods, apparatus (systems), and computer program products according to embodiments of this specification with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0185] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0186] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0187] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0188] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0189] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0190] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0191] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. This specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0192] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0193] The above description is merely one or more embodiments of this specification and is not intended to limit this application. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of one or more embodiments of this specification.

Claims

1. A method for optimizing a search database, comprising: The similarity between search elements in the search database and query elements in the query database is mapped to obtain the contribution value of each search element to the search score corresponding to the query element. The contribution value is the value contributed by the search element to the search score of the query element. The similarity between the search element and the query element is positively correlated with the contribution value of the search element to the search score corresponding to the query element. Based on the contribution value, a mapping relationship is determined between the retrieval score representing the query element and the retention probability of each query element; Based on the mapping relationship, in the process of minimizing the first loss function, the retention probability corresponding to each search element is determined. The first loss function is used to characterize the residual between the search score and the classification label corresponding to the query element. Based on the retention probability, a retention / detention indicator is determined for each search element in the search database, so as to update the search database.

2. The method according to claim 1, wherein mapping the similarity between search elements in the search database and query elements in the query database to obtain the contribution value of each search element to the search score corresponding to the query element includes: Determine the vector distance between the vector of the retrieved element and the vector of the query element; From the vector distances, a preset number of target distances related to each query element are determined, and the retrieval element corresponding to the target distance is the recall element corresponding to the query element; The target distance is mapped to obtain the contribution value of each recalled element to the retrieval score corresponding to the query element, and the absolute value of the contribution value is negatively correlated with the target distance.

3. The method according to claim 1, wherein determining the retention probability corresponding to each search element during the process of minimizing the first loss function based on the mapping relationship includes: Based on the mapping relationship, the retrieval score corresponding to the query element in each loss term of the first loss function is replaced with the retention probability of the retrieval element corresponding to the query element to obtain the second loss function; In the process of minimizing the second loss function, the retention probability in the loss term corresponding to each of the retrieved elements is determined.

4. The method according to claim 3, wherein when the mapping relationship is a linear mapping relationship, determining the retention probability in the loss term corresponding to each of the retrieved elements during the process of minimizing the second loss function includes: When the classification label of the search element corresponding to the loss term in the first loss function is a white sample, a first symbol is set for the loss term, and the first symbol is used to ensure that the value of the loss term is negative. When the classification label of the search element corresponding to the loss term in the first loss function is black sample, a second symbol is set for the loss term, and the second symbol is used to ensure that the value of the loss term is negative. Based on the mapping relationship, the retention probability corresponding to each search element is determined during the process of minimizing the first loss function with the set symbol.

5. The method according to claim 4, wherein determining the retention probability corresponding to each of the retrieved elements during the process of minimizing the first loss function with pre-set symbols based on the mapping relationship includes: In the second loss function, the coefficients of the retention probability in the loss terms corresponding to the same search element are summed to obtain the total coefficient of the retention probability in the loss terms corresponding to the search element; When the total coefficient is negative, the retention probability of the search element corresponding to the total coefficient is set to a first preset value, and the first preset value is used to indicate the retention of the loss item corresponding to the search element; When the total coefficient is positive, the retention probability of the search element corresponding to the total coefficient is set to a second preset value, which is used to indicate the deletion of the loss item corresponding to the search element.

6. The method according to claim 5, wherein determining the retention indication for each search element in the search database based on the retention probability comprises: When the retention probability of the search element is a first preset value, the retention indication of the search element is determined as a retention indication; When the retention probability of the search element is a second preset value, the retention / deletion indication of the search element is determined as a deletion indication.

7. The method according to claim 1, wherein determining the retention probability corresponding to each retrieved element during the process of minimizing the first loss function includes: In the first loss function, a first weight is set for the loss term corresponding to the search element whose classification label is white sample; A second weight is set for the loss term corresponding to the search element with the classification label "black sample", where the first weight is less than the second weight; In the process of minimizing the first loss function with set weights, the retention probability corresponding to each retrieved element is determined.

8. The method according to claim 1, wherein determining the retention probability corresponding to each retrieved element during the process of minimizing the first loss function based on the mapping relationship includes: Based on the range of the search scores, the search scores in the mapping relationship are processed by an activation function; Based on the range of values ​​for the retention probability, the retention probability in the first loss function is processed by an activation function; Based on the mapping relationship processed by the activation function, the retention probability corresponding to each search element is determined in the process of minimizing the first loss function processed by the activation function.

9. An apparatus for updating a search database, comprising: The contribution value determination module is used to map the similarity between the search elements in the search database and the query elements in the query database to obtain the contribution value of each search element to the search score corresponding to the query element. The contribution value is the numerical value contributed by the search element to the search score of the query element. The similarity between the search element and the query element is positively correlated with the contribution value of the search element to the search score corresponding to the query element. The mapping relationship determination module is used to determine the mapping relationship between the retrieval score representing the query element and the retention probability of each retrieval element based on the contribution value. The retention probability determination module is used to determine the retention probability corresponding to each of the search elements based on the mapping relationship, while minimizing the first loss function. The first loss function is used to characterize the residual between the search score and the classification label corresponding to the query element. An update module is used to determine, based on the retention probability, whether each search element in the search database should be retained or not, so as to update the search database.

10. An electronic device, comprising: processor, and A memory configured to store computer-executable instructions, which, when executed, enable the processor to: The similarity between search elements in the search database and query elements in the query database is mapped to obtain the contribution value of each search element to the search score corresponding to the query element. The contribution value is the value contributed by the search element to the search score of the query element. The similarity between the search element and the query element is positively correlated with the contribution value of the search element to the search score corresponding to the query element. Based on the contribution value, a mapping relationship is determined between the retrieval score representing the query element and the retention probability of each query element; Based on the mapping relationship, in the process of minimizing the first loss function, the retention probability corresponding to each search element is determined. The first loss function is used to characterize the residual between the search score and the classification label corresponding to the query element. Based on the retention probability, a retention / detention indicator is determined for each search element in the search database, so as to update the search database.

Citation Information

Patent Citations

  • Image target positioning correction method and related equipment

    CN110232713A

  • Blacklist query method and system, electronic equipment and storage medium

    CN113986921A